Amplicon datasets contain many identical reads. Dereplicating them (keeping one copy of each unique sequence plus its abundance) makes clustering and denoising much faster.
1. Dereplicate
seqfu derep prints unique sequences sorted by
abundance, with the abundance in the name in the ;size=N format understood by USEARCH and
VSEARCH:
seqfu derep reads.fa.gz > uniques.fa
head -n 1 uniques.fa
>seq.1;size=18335
Use --json derep.json to save the mapping between input and output sequences, or -c to
write the size as a comment instead of in the name.
2. Drop singletons
Sequences seen only once are often errors. --min-size keeps only the abundant ones:
seqfu derep -m 2 reads.fa.gz > uniques.min2.fa
Skipped 13575 clusters having less than 2 sequences.
3. Dereplicate across samples
derep reads the size= annotations of its input, so files that were dereplicated
separately can be combined without losing the original counts:
for f in samples/*.fa.gz; do
seqfu derep "$f" > "derep/$(basename "$f" .fa.gz).fa"
done
seqfu derep derep/*.fa > all_uniques.fa
Add --ignore-size if you want each input record to count as one.
4. Inspect the result
# How many unique sequences, and how long?
seqfu stats -n uniques.min2.fa
# Longest first (sort also removes duplicates)
seqfu sort uniques.min2.fa | head -n 2
# Which uniques contain the forward primer?
seqfu grep -o CCTACGGGNGGCWGCAG uniques.min2.fa