Class FileUtils

java.lang.Object
org.gbif.utils.file.FileUtils

public final class FileUtils extends Object
Collection of file utils.
This class has only been tested for use with a UTF-8 system encoding.
  • Field Details

  • Constructor Details

  • Method Details

    • classpath2Filepath

      public static String classpath2Filepath(String path)
    • classpathStream

      public static InputStream classpathStream(String path) throws IOException
      Throws:
      IOException
    • columnsToSet

      public static Set<String> columnsToSet(InputStream source, int... column) throws IOException
      Throws:
      IOException
    • columnsToSet

      public static Set<String> columnsToSet(InputStream source, Set<String> resultSet, int... column) throws IOException
      Reads a file and returns a unique set of multiple columns from lines which are no comments (starting with #) and trims whitespace.
      Parameters:
      source - the UTF-8 encoded text file with tab delimited columns
      resultSet - the set implementation to be used. Will not be cleared before reading!
      column - variable length argument of column indices to process
      Returns:
      set of column rows
      Throws:
      IOException
    • copyStreams

      public static void copyStreams(InputStream in, OutputStream out) throws IOException
      Throws:
      IOException
    • copyStreamToFile

      public static void copyStreamToFile(InputStream in, File out) throws IOException
      Throws:
      IOException
    • createTempDir

      public static File createTempDir() throws IOException
      Throws:
      IOException
    • createTempDir

      public static File createTempDir(String prefix, String suffix) throws IOException
      Parameters:
      prefix - The prefix string to be used in generating the file's name; must be at least three characters long
      suffix - The suffix string to be used in generating the file's name; may be null, in which case the suffix ".tmp" will be used
      Throws:
      IOException
    • deleteDirectoryRecursively

      public static void deleteDirectoryRecursively(File directory)
      Delete directory recursively, including all its files, sub-folders, and sub-folder's files.
      Parameters:
      directory - directory to delete recursively
    • escapeFilename

      public static String escapeFilename(String filename)
      Escapes a filename so it is a valid filename on all systems, replacing /. .. \t\r\n.
      Parameters:
      filename - to be escaped
    • getClasspathFile

      public static File getClasspathFile(String path)
    • getInputStream

      public static InputStream getInputStream(File source) throws FileNotFoundException
      Throws:
      FileNotFoundException
    • getInputStreamReader

      Throws:
      FileNotFoundException
    • getInputStreamReader

      Throws:
      FileNotFoundException
    • getLineIterator

      public static org.apache.commons.io.LineIterator getLineIterator(InputStream source)
      Parameters:
      source - the source input stream encoded in UTF-8
    • getLineIterator

      public static org.apache.commons.io.LineIterator getLineIterator(InputStream source, String encoding)
      Parameters:
      source - the source input stream
      encoding - the encoding used by the input stream
    • getUtf8Reader

      Throws:
      FileNotFoundException
    • humanReadableByteCount

      public static String humanReadableByteCount(long bytes, boolean si)
      Converts the byte size into human-readable format. Support both SI and byte format.
    • isCompressedFile

      public static boolean isCompressedFile(File source)
    • readByteBuffer

      public static ByteBuffer readByteBuffer(File file) throws IOException
      Reads a complete file into a byte buffer.
      Throws:
      IOException
    • readByteBuffer

      public static ByteBuffer readByteBuffer(File file, int bufferSize) throws IOException
      Reads the first bytes of a file into a byte buffer.
      Parameters:
      bufferSize - the number of bytes to read from the file
      Throws:
      IOException
    • setLinesPerMemorySort

      public static void setLinesPerMemorySort(int linesPerMemorySort)
      Parameters:
      linesPerMemorySort - are the number of lines that should be sorted in memory, determining the number of file segments to be sorted when doing a Java file sort. Defaults to 100000, if you have memory available a higher value increases performance.
    • startNewUtf8File

      public static Writer startNewUtf8File(File file) throws IOException
      Throws:
      IOException
    • startNewUtf8XmlFile

      public static Writer startNewUtf8XmlFile(File file) throws IOException
      Throws:
      IOException
    • streamToList

      public static LinkedList<String> streamToList(InputStream source) throws IOException
      Takes a utf8 encoded input stream and reads in every line/row into a list.
      Returns:
      list of rows
      Throws:
      IOException
    • streamToList

      public static List<String> streamToList(InputStream source, List<String> resultList) throws IOException
      Reads a file and returns a list of all lines which are no comments (starting with #) and trims whitespace.
      Parameters:
      source - the UTF-8 encoded text file to read
      resultList - the list implementation to be used. Will not be cleared before reading!
      Returns:
      list of lines
      Throws:
      IOException
    • streamToList

      public static LinkedList<String> streamToList(InputStream source, String encoding) throws IOException
      Throws:
      IOException
    • streamToMap

      public static Map<String,String> streamToMap(InputStream source) throws IOException
      Reads a utf8 encoded inut stream, splits
      Throws:
      IOException
    • streamToMap

      public static Map<String,String> streamToMap(InputStream source, int key, int value, boolean trimToNull) throws IOException
      Throws:
      IOException
    • streamToMap

      public static Map<String,String> streamToMap(InputStream source, Map<String,String> result) throws IOException
      Read a hashmap from a tab delimited utf8 input stream using the row number as an integer value and the entire row as the value. Ignores commented rows starting with #.
      Parameters:
      source - tab delimited text file to read
      Throws:
      IOException
    • streamToMap

      public static Map<String,String> streamToMap(InputStream source, Map<String,String> result, int key, int value, boolean trimToNull) throws IOException
      Read a hashmap from a tab delimited utf8 file, ignoring commented rows starting with #.
      Parameters:
      source - tab delimited input stream to read
      key - column number to use as key
      value - column number to use as value
      trimToNull - if true trims map entries to null
      Throws:
      IOException
    • streamToSet

      public static Set<String> streamToSet(InputStream source) throws IOException
      Throws:
      IOException
    • streamToSet

      public static Set<String> streamToSet(InputStream source, Set<String> resultSet) throws IOException
      Reads a file and returns a unique set of all lines which are no comments (starting with #) and trims whitespace.
      Parameters:
      source - the UTF-8 encoded text file to read
      resultSet - the set implementation to be used. Will not be cleared before reading!
      Returns:
      set of unique lines
      Throws:
      IOException
    • toFilePath

      public static String toFilePath(URL url)
    • url2file

      public static File url2file(URL url)
    • getLinesPerMemorySort

      public int getLinesPerMemorySort()
    • mergeSortedFiles

      public void mergeSortedFiles(List<File> sortFiles, Writer sortedFileWriter, Comparator<String> lineComparator) throws IOException
      Merges a list of intermediary sort chunk files into a single sorted file. On completion, the intermediary sort chunk files are deleted.
      Parameters:
      sortFiles - sort chunk files to merge
      sortedFileWriter - writer to merge to. Can already be open and contain data
      lineComparator - To use when determining the order (reuse the one that was used to sort the individual files)
      Throws:
      IOException
    • sort

      public void sort(File input, File sorted, String encoding, int column, String columnDelimiter, Character enclosedBy, String newlineDelimiter, int ignoreHeaderLines) throws IOException
      Sorts the input file into the output file using the supplied delimited line parameters. This method is not reliable when the sort field may contain Unicode codepoints outside the Basic Multilingual Plane, i.e. above ￿. In that case, the sort order differs from Java's String sort order. This should not be a problem for most usage; the Supplementary Multilingual Planes contain ancient scripts, emojis, arrows and so on.
      Parameters:
      input - To sort
      sorted - The sorted version of the input excluding ignored header lines (see ignoreHeaderLines)
      column - the column that keeps the values to sort on
      columnDelimiter - the delimiter that separates columns in a row
      enclosedBy - optional column enclosing character, e.g. a double quote for CSVs
      newlineDelimiter - the chars used for new lines, usually \n, \n\r or \r
      ignoreHeaderLines - number of beginning lines to ignore, e.g. headers
      Throws:
      IOException
    • sort

      public void sort(List<File> inputs, File sorted, String encoding, int column, String columnDelimiter, Character enclosedBy, String newlineDelimiter, int ignoreHeaderLines) throws IOException
      Sorts the input file into the output file using the supplied delimited line parameters. This method is not reliable when the sort field may contain Unicode codepoints outside the Basic Multilingual Plane, i.e. above ￿. In that case, the sort order differs from Java's String sort order. This should not be a problem for most usage; the Supplementary Multilingual Planes contain ancient scripts, emojis, arrows and so on.
      Parameters:
      inputs - To sort
      sorted - The sorted version of the input excluding ignored header lines (see ignoreHeaderLines)
      column - the column that keeps the values to sort on
      columnDelimiter - the delimiter that separates columns in a row
      enclosedBy - optional column enclosing character, e.g. a double quote for CSVs
      newlineDelimiter - the chars used for new lines, usually \n, \n\r or \r
      ignoreHeaderLines - number of beginning lines to ignore, e.g. headers
      Throws:
      IOException
    • sort

      public void sort(File input, File sorted, String encoding, int column, String columnDelimiter, Character enclosedBy, String newlineDelimiter, int ignoreHeaderLines, Comparator<String> lineComparator, boolean ignoreCase) throws IOException
      Sorts the input file into the output file using the supplied delimited line parameters. This method is not reliable when the sort field may contain Unicode codepoints outside the Basic Multilingual Plane, i.e. above ￿. In that case, the sort order differs from Java's String sort order. This should not be a problem for most usage; the Supplementary Multilingual Planes contain ancient scripts, emojis, arrows and so on. This method is globally synchronized, in case multiple sorts are attempted to the same file simultaneously. This could be improved to allow synchronizing against the destination file, rather than for all sorts.
      Parameters:
      input - To sort
      sorted - The sorted version of the input excluding ignored header lines (see ignoreHeaderLines)
      column - the column that keeps the values to sort on
      columnDelimiter - the delimiter that separates columns in a row
      enclosedBy - optional column enclosing character, e.g. a double quote for CSVs
      newlineDelimiter - the chars used for new lines, usually \n, \r\n or \r
      ignoreHeaderLines - number of beginning lines to ignore, e.g. headers
      lineComparator - used to sort the output
      ignoreCase - ignore case order, this parameter couldn't have any effect if the LineComparator is used
      Throws:
      IOException
    • sort

      public void sort(List<File> inputs, File sorted, String encoding, int column, String columnDelimiter, Character enclosedBy, String newlineDelimiter, int ignoreHeaderLines, Comparator<String> lineComparator, boolean ignoreCase) throws IOException
      Sorts the input file into the output file using the supplied delimited line parameters. This method is not reliable when the sort field may contain Unicode codepoints outside the Basic Multilingual Plane, i.e. above ￿. In that case, the sort order differs from Java's String sort order. This should not be a problem for most usage; the Supplementary Multilingual Planes contain ancient scripts, emojis, arrows and so on. This method is globally synchronized, in case multiple sorts are attempted to the same file simultaneously. This could be improved to allow synchronizing against the destination file, rather than for all sorts.
      Parameters:
      inputs - To sort
      sorted - The sorted version of the input excluding ignored header lines (see ignoreHeaderLines)
      column - the column that keeps the values to sort on
      columnDelimiter - the delimiter that separates columns in a row
      enclosedBy - optional column enclosing character, e.g. a double quote for CSVs
      newlineDelimiter - the chars used for new lines, usually \n, \r\n or \r
      ignoreHeaderLines - number of beginning lines to ignore, e.g. headers
      lineComparator - used to sort the output
      ignoreCase - ignore case order, this parameter couldn't have any effect if the LineComparator is used
      Throws:
      IOException
    • sortInJava

      public void sortInJava(File input, File sorted, String encoding, Comparator<String> lineComparator, int ignoreHeaderLines) throws IOException
      Sorts the input file into the output file using the supplied lineComparator.
      Parameters:
      input - To sort
      sorted - The sorted version of the input excluding ignored header lines (see ignoreHeaderLines)
      lineComparator - To use during comparison
      ignoreHeaderLines - number of beginning lines to ignore, e.g. headers
      Throws:
      IOException
    • sortInJava

      public void sortInJava(List<File> inputs, File sorted, String encoding, Comparator<String> lineComparator, int ignoreHeaderLines) throws IOException
      Sorts the input file into the output file using the supplied lineComparator.
      Parameters:
      inputs - To sort
      sorted - The sorted version of the input excluding ignored header lines (see ignoreHeaderLines)
      lineComparator - To use during comparison
      ignoreHeaderLines - number of beginning lines to ignore, e.g. headers
      Throws:
      IOException
    • split

      public List<File> split(File input, int linesPerOutput, String extension) throws IOException
      Splits the supplied file into files of set line size and with a suffix.
      Parameters:
      input - To split up
      linesPerOutput - Lines per split file
      extension - The file extension to use - e.g. ".txt"
      Returns:
      The split files
      Throws:
      IOException
    • touch

      public static void touch(File file) throws IOException
      Creates an empty file or updates the last updated timestamp on the same as the unix command of the same name.

      From Guava.

      Parameters:
      file - the file to create or update
      Throws:
      IOException - if an I/O error occurs
    • getFileExtension

      public static String getFileExtension(String fullName)
      Returns the file extension for the given file name, or the empty string if the file has no extension. The result does not include the '.'.

      Note: This method simply returns everything after the last '.' in the file's name as determined by File.getName(). It does not account for any filesystem-specific behavior that the File API does not already account for. For example, on NTFS it will report "txt" as the extension for the filename "foo.exe:.txt" even though NTFS will drop the ":.txt" part of the name when the file is actually created on the filesystem due to NTFS's Alternate Data Streams.

      From Guava.

    • createParentDirs

      public static void createParentDirs(File file) throws IOException
      Creates any necessary but nonexistent parent directories of the specified file. Note that if this operation fails it may have succeeded in creating some (but not all) of the necessary parent directories.

      From Guava.

      Throws:
      IOException - if an I/O error occurs, or if any necessary but nonexistent parent directories of the specified file could not be created.