SeqFu Install
Recipes / Prepare and validate a sample sheet

Prepare and validate a sample sheet

Generate a metadata or manifest file from a directory of reads and validate it before running a pipeline.

Commands used
Input
Directory, TSV-CSV
Level
Beginner

Most pipelines (QIIME 2, nf-core, Dadaist2…) start from a sample sheet listing the samples and their files. Writing it by hand is tedious and error-prone.

1. Make sure the reads are sound

seqfu check --dir reads/

2. Generate the sample sheet

seqfu metadata extracts the sample ID from the file names and pairs R1 and R2. The default format is a QIIME 2 import manifest:

seqfu metadata reads/ > manifest.tsv
sample-id	forward-absolute-filepath	reverse-absolute-filepath
sample1	/data/reads/sample1_R1.fq.gz	/data/reads/sample1_R2.fq.gz
sample2	/data/reads/sample2_R1.fq.gz	/data/reads/sample2_R2.fq.gz

Pick another format with -f; run seqfu metadata formats to list them (ampliseq, mag, rnaseq, bactopia, irida, dadaist, lotus, qiime1, qiime2…):

seqfu metadata -f ampliseq reads/ > samplesheet.tsv

If the sample ID is not the first _-separated part of the file name, change the separator with --split and the part(s) to keep with --pos.

3. Validate the table

Tables edited in a spreadsheet often end up with a missing or extra column in some rows. seqfu tabcheck checks that every row has the same number of fields, and exits with an error otherwise:

seqfu tabcheck --header manifest.tsv samplesheet.tsv
File	PassQC	Columns	Rows	Separator
manifest.tsv	Pass	3	3	[tab]
samplesheet.tsv	Error[row=2;expected=2;observed=3;reason=inconsistent-column-count]

Use --inspect on a valid table to see the inferred type of each column.