The Benchmark Matters: Evaluating TruPath Genome Beyond Individual Loci

TruPath Genome combines Illumina paired-end sequencing with long-range proximity information and advanced informatics to provide long-range genomic insight in a streamlined workflow on the NovaSeq™ X Sequencing System. The important question for researchers is not which technology can be used to tell the best story, but which technology can be trusted; which approach produces accurate, comprehensive, and reproducible results when evaluated against established benchmarks.

TruPath Genome brings long-range genomic insight to the NovaSeq™ X Sequencing System by identifying reads originating from the same input molecule based on their relative location on the flow cell surface. Advanced informatics combine high-quality Illumina paired-end reads with long-range proximity information to analyze genomic variation.

Read more about TruPath Genome here.

Recent customer research illustrates the potential for TruPath Genome to consolidate analyses that can otherwise require multiple methods in short-read and long-read sequencing research workflows. More broadly, TruPath expands the genomic regions and variant classes that can be evaluated with Illumina sequencing, including regions that have historically been difficult to characterize accurately. Those capabilities should be evaluated on evidence, not category assumptions.

Pacific Biosciences recently published a blog challenging TruPath Genome performance using selected loci and IGV screenshots. Questions like these deserve a careful technical response, so we examined the examples and compared the broader data against established benchmark truth sets.

Our analysis identified several methodological issues that materially affect the conclusions. More importantly, the exercise illustrates a broader principle: individual screenshots can raise useful questions, but platform-level performance should be judged through comprehensive, reproducible benchmarking.

Why individual loci are insufficient for technology-level conclusions

HG002 is the most thoroughly characterized human genome in existence. The Genome in a Bottle v4.2.1 small-variant benchmark, the Challenging Medically Relevant Genes (CMRG) benchmark, and the T2T-HG002 Q100 diploid assembly benchmark exist precisely so that platform comparisons look at a wide swath of the genome and don’t rely on individual loci.

PacBio's blog claims that some variants detected by PacBio HiFi are not called by both standard Illumina and TruPath sequencing. The author includes a single IGV screenshot showing PacBio read support for three variants that lack TruPath read support. 

Individual loci can surface important questions, but a single discordant variant viewed in IGV is not representative of a technology’s overall accuracy. Platform-level comparisons require appropriately matched benchmark datasets and metrics. 

To assess small-variant performance beyond individual loci, we measured false negatives using the full NIST HG002 v5.0q truth set. In this analysis, PacBio HiFi had at least 25% more small-variant false negatives and at least 22% more total small-variant errors than TruPath (Figure 1). These results apply to the evaluated small-variant benchmark and should not be generalized to every variant class. The broader methodological point is that technology-level conclusions require benchmark data matched to the variant class being evaluated.

Individual loci still warrant investigation. The next step, however, is to determine the extent of the issue using an industry-standard benchmark. That progression from isolated observation to comprehensive measurement is essential for meaningful scientific conclusions.

Why pileup visualization alone can mislead

Pacbio’s second example in FMR1 shows a PacBio pileup with a 33 bp CGG insertion and a TruPath pileup with “not a single read alignment containing this insertion,” and concludes that the call must have been inferred from the pangenome without read support. That interpretation depends on which reads the interpreter has selected to display in the pileup and which they have missed or ignored, not on the full underlying read evidence.

The screenshot omits the soft-clipped portions of each read. When those bases are displayed, it is clear that the supporting signal is present in the TruPath reads, as shown here: 

Soft-clipped alignments commonly occur when reads containing a repeat expansion or inserted sequence are aligned to a reference genome because the inserted sequence has no corresponding matching sequence in the reference.

The conclusion therefore reflects the visualization settings, not an absence of TruPath read support. The variant is called correctly in the VCF and supported by the underlying reads. This is why a default pileup view should not be treated as definitive evidence of variant-calling performance.

In these examples, the role of the pangenome reference in DRAGEN is mischaracterized. As described in our publications and online documentation, the pangenome helps reduce mapping ambiguity for reads containing population-specific variants and informs variant calling. DRAGEN Germline does not emit variants in the absence of explicit supporting read evidence.

What the NCS1 example actually shows

The PacBio post reports a heterozygous 8 bp deletion and homozygous 129 bp insertion in the TruPath calls, versus a phased 368 bp / 129 bp insertion pair in PacBio that matches the benchmark.

This example identifies a real limitation in the TruPath software version used in the analysis: the locus was not called correctly. The underlying reads supporting both insertions were nevertheless present in the TruPath data, indicating a software-calling limitation rather than a failure to capture the sequence evidence. Analyzing with our most up-to-date DRAGEN build in which this limitation has already been addressed, the 368 bp and 129 bp insertions and the phased A deletion are called concordantly with the benchmark.

Modern SV callers, including DRAGEN-SV and Sawfish, do not call variants by reading pileups. They reassemble candidate variant sequences locally and score the read evidence against those candidates. Reads supporting a large insertion are routinely split, clipped, or assigned MAPQ0 in the reference-projected alignment, which is why a pileup alone is not an appropriate basis for evaluating structural-variant calls.

Across successive releases over the past five years, DRAGEN has reduced genome-wide SNV errors on HG002 by roughly 88% (version 3.6 to 4.5, GIAB v4.2.1). The NCS1 example illustrates how software enhancements materially affect analysis, and we remain committed to continually improving TruPath Genome with future updates like these. 

Summary

PacBio’s title calls for measuring over guessing. On that principle, we agree. Selected screenshots can identify loci worth investigating, but they are not a substitute for benchmarked analysis across the relevant truth sets and variant classes.

The appropriate standard is transparent, reproducible benchmarking that makes both strengths and limitations visible. We welcome scrutiny of TruPath Genome and will continue to provide data that enable researchers to evaluate its performance on that basis.