Skip to content

Benchmarks

F1 against NAM, all isoforms:

LevelHelixerHelixForge
Exon54.961.4
Intron62.772.1
Intron chain24.642.5
Locus37.555.5

Transcripts whose entire intron chain exactly matches a NAM transcript rise from 12,308 to 27,096. The gain holds when both annotations are reduced to one transcript per gene (so it is not an artifact of emitting more transcripts) and at the coding (CDS) level (so it is not confined to UTRs). Improvement is uniform across all ten chromosomes (intron-chain sensitivity SD 0.6), with no regressing chromosome.

GenesTranscriptsTranscripts per gene
Helixer44,48044,4801.00
HelixForge44,30771,6391.62
NAM (reference)39,03571,7911.84

HelixForge recovers most of the reference's isoform multiplicity from a single-isoform prior.

RNA-seq validates the added splice junctions

Section titled “RNA-seq validates the added splice junctions”

Splice junctions absent from the reference, adjudicated against the RNA-seq evidence:

Novel junctions with RNA supportReference RNA-supported junctions recovered
Helixer5.0%77.9% (32,267 missed)
HelixForge22.4%93.6% (9,353 missed)

The junctions HelixForge adds are RNA-supported 4.5× more often than the input predictor's, and it simultaneously recovers more of the reference's real junctions rather than trading one for the other. The result is robust to the read-depth threshold, and 98.5% of supported novel junctions carry canonical GT–AG motifs.

Median CDS length is 1,026 nt for HelixForge and 1,023 nt for NAM; raw Helixer under-calls at 855 nt. Reconciliation restores reference-like coding length, not only UTR extent.