Benchmarks
Structural agreement with the reference
Section titled “Structural agreement with the reference”F1 against NAM, all isoforms:
| Level | Helixer | HelixForge |
|---|---|---|
| Exon | 54.9 | 61.4 |
| Intron | 62.7 | 72.1 |
| Intron chain | 24.6 | 42.5 |
| Locus | 37.5 | 55.5 |
Transcripts whose entire intron chain exactly matches a NAM transcript rise from 12,308 to 27,096. The gain holds when both annotations are reduced to one transcript per gene (so it is not an artifact of emitting more transcripts) and at the coding (CDS) level (so it is not confined to UTRs). Improvement is uniform across all ten chromosomes (intron-chain sensitivity SD 0.6), with no regressing chromosome.
Isoform reconstruction
Section titled “Isoform reconstruction”| Genes | Transcripts | Transcripts per gene | |
|---|---|---|---|
| Helixer | 44,480 | 44,480 | 1.00 |
| HelixForge | 44,307 | 71,639 | 1.62 |
| NAM (reference) | 39,035 | 71,791 | 1.84 |
HelixForge recovers most of the reference's isoform multiplicity from a single-isoform prior.
RNA-seq validates the added splice junctions
Section titled “RNA-seq validates the added splice junctions”Splice junctions absent from the reference, adjudicated against the RNA-seq evidence:
| Novel junctions with RNA support | Reference RNA-supported junctions recovered | |
|---|---|---|
| Helixer | 5.0% | 77.9% (32,267 missed) |
| HelixForge | 22.4% | 93.6% (9,353 missed) |
The junctions HelixForge adds are RNA-supported 4.5× more often than the input predictor's, and it simultaneously recovers more of the reference's real junctions rather than trading one for the other. The result is robust to the read-depth threshold, and 98.5% of supported novel junctions carry canonical GT–AG motifs.
Coding structure
Section titled “Coding structure”Median CDS length is 1,026 nt for HelixForge and 1,023 nt for NAM; raw Helixer under-calls at 855 nt. Reconciliation restores reference-like coding length, not only UTR extent.