DNA Metabarcoding Infrastructure

Metabarcoding outgrew custom pipelines. spcfy is what comes next.

Shared infrastructure for labs, bioinformaticians, and researchers who need DNA Metabarcoding data to be reproducible, comparable, and decision-ready. One standard, every project, every lab. Not in months. In hours.

4h
avg. turnaround
6+
databases queried
1000+
samples per run

Not ready for a demo? Read the workflow documentation →

Covering CO1, 16S, ITS2, 12S, PITS2. Animals, plants, fungi, bacteria.

Find your entry point

Creating a global ecosystem for biodiversity data.

spcfy connects data producers who generate and process biodiversity data with data users who need reliable results for research, monitoring and decisions. Pick your role, or follow the shared logic from FASTQ to insight.

Working across more than one role? Many research institutes, consultancies and lab teams do. spcfy is built to connect these roles on one shared foundation, not to lock them into separate tools.

The scale problem

DNA Metabarcoding has outgrown spreadsheets.

A single project generates thousands of OTUs across hundreds of samples, each with ecological metadata. Species lists do not carry that complexity. Lab-specific pipeline stacks do not scale with it.

You need infrastructure built for this volume.

Without it, every new project starts from scratch. Every comparison needs a caveat. Every dataset ages the moment the next reference database version drops.

Project size
Typical
Large program
Samples
50
1,000
Paired-end FASTQ files
100
2,000
Data volume
~5 GB
~200 GB
OTUs detected
3,000+
12,000+
spcfy processing time
~2 hours
~8 hours

Raw sequence output has no memory.
spcfy gives it one.

Every OTU gets a stable identity. Every result carries its pipeline version, its reference database, its QC threshold. Your 2024 data is still comparable in 2026.

How it works

Results that hold up. Across labs, databases, and time.

Every order goes through a multi-database consensus pipeline, built by bioinformaticians who run their own labs.

Algorithm

SPARK · Consensus Taxonomy

SPARK evaluates all BLAST hits and selects the most reliable match using a weighted score that combines identity, taxonomic resolution, and consistency across hits. Do not rely on a single top hit, identify the overall best. Robust. Precise. Built for reproducibility.

Algorithm

LCA · Taxonomic Annotation

spcfy applies a rich set of taxonomic annotations, worldwide or regional BOLD, NCBI GenBank, SILVA, UNITE, RDP, tailored to different amplicons. It then condenses them into a Least Common Ancestor taxonomic consensus. Built for easy interpretability.

Algorithm

SPIN · Stable OTU Identity

SPIN assigns each sequence a stable spcfy identification number, analogous to a BIN but designed for metabarcoding. OTU sequences are matched against a reference SPIN database, allowing sequences from different runs to be merged instantly without re-clustering. Consistent tracking. Direct comparability of taxa across samples, projects, and time.

spcfAI, alpha

Speak to your data

Ask questions across your biodiversity datasets in plain language. Surface patterns, flag anomalies, generate reports. Powered by spcfAI, built on the same standardized data layer. Available today via an API, using spcfy's built-in LLM or your own connected Claude or ChatGPT model.

spcfAI is live in alpha. Deeper predictive layers arrive with spcfyPREDICT, still in development.
Full workflow documentation →
Data ownership

Your data stays yours. Full provenance, full control.

spcfy is built on data ownership as a first principle. Every dataset, raw and processed, is available for download at any time, in standard formats. No lock-in.

You own your data

All data is hosted in Germany, GDPR-compliant, with clear access controls. You decide who sees what, when, and in what format.

Every project carries a locked audit trail: pipeline version, reference database version, QC thresholds. Your results stay reproducible and citable.

Raw and processed datasets are available for download in standard formats at any time. Full data ownership is given from day one.

FAIR-aligned publishing, when you want it

In development

When publication is the goal, spcfy is designed to make it straightforward. Standardized exports, Darwin Core compatibility, DOI-ready metadata, and full parameter provenance will travel with your dataset.

Direct connections to public repositories are currently in development. Once available, all exports will be structured for direct submission to:

GBIF ENA BOLD Custom
Ecosystem

Labs analyze, customers explore, and everything stays comparable.

spcfy is not a tool for one lab. It is shared infrastructure where every lab's output follows the same standard, so customers can actually compare results across labs and over time.

From sequencing run to client-ready delivery, without custom scripts or manual QC.

Step 1 · Labs
Upload FASTQ
Step 2 · spcfy
Process & standardize
Step 3 · spcfy
Visualize & compare
Step 4 · Customers
Explore results

Four tiers. One ecosystem.

Start with the data foundation. Add analytical depth as your needs grow.

BASE
spcfyBASE

The data foundation. From raw FASTQ to standardized, comparable output.

INSIGHT
spcfyINSIGHT

The analytical layer. Explore, compare, and share biodiversity intelligence.

PREDICT
spcfyPREDICT

Predictive layers and Nature KPIs, built on top of INSIGHT.

Coming soon
spcfAI
spcfAI

A natural-language interface for your biodiversity data, with API access to spcfy's built-in LLM or your own Claude or ChatGPT model.

Live, alpha
Full pricing and feature comparison →
iBOL 2026 · Silver Sponsor
10th International Barcode of Life Conference (iBOL 2026)

spcfy is a Silver Sponsor of iBOL 2026.

The 10th International Barcode of Life Conference brings around 500 scientists working in DNA barcoding and metabarcoding to Bangkok, this year under the theme "Building on Barcodes: Impacting Science and Society."

Come find us on site for a direct conversation about your workflow, your data, and whether spcfy fits.

When November 2–6, 2026 Where Bangkok, Thailand
Planning to attend iBOL 2026?

Let's set up a time to talk.

Book an online demo before the conference, or find us on site in Bangkok. New accounts also get 4 weeks of freemium access with 15 processing credits, no commitment required.

From field sample to biodiversity signal.

In hours. On infrastructure built for this volume and complexity.

Questions first? Write to us at hello@spcfy.io