For the complete documentation index, see llms.txt. This page is also available as Markdown.

CSV Loader

A license is required to access Spark functionality on the DNAnexus Platform. Contact DNAnexus Sales for more information.

Overview

The CSV Loader ingests CSV files into a database. The input CSV files are loaded into a Parquet-format database and tables that can be queried using Spark SQL.

You can load a single CSV file or many CSV files. In the case of many files, all files must be syntactically equal.

For example:

  • All files must have the same separator. This can be a comma, tab, or another consistent delimiter.

  • All files must include a header line, or all files must exclude it

Each CSV file is loaded into its own table within the specified database.

How to Run CSV Loader

Input:

  • csv: (array) CSV files to load into the database.

Required Parameters:

  • database_name: name of the database to load the CSV files into.

  • create_mode: strict mode creates a database and tables from scratch, and optimistic mode creates a database and tables if they do not already exist.

  • insert_mode: append appends data to the end of tables, and overwrite is equivalent to truncating the tables and then appending to them.

  • table_name: array of table names, one for each corresponding CSV file by array index.

  • type: cluster type. Use spark for Spark apps.

Other Options:

  • spark_read_csv_header: (boolean) default false -- whether the first line of each CSV is used as column names for the corresponding table.

  • spark_read_csv_sep: (string) default , -- separator character used by each CSV.

  • spark_read_csv_infer_schema: (boolean) default false -- whether the input schema is inferred from the data.

Basic Run

The following case creates a brand new database and loads data into two new tables:

Last updated

Was this helpful?