TRANSFORMER PORTFOLIO · PROJECT 04
Long-Document Question Answering
An evidence-grounded Document AI system combining a QASPER-fine-tuned Longformer evaluation pipeline with a free, serverless browser QA demonstration.
LongformerQASPERTransformers.jsONNX RuntimeEvidence Grounding
REAL BENCHMARK
200 evaluation examples
QASPER evaluation results
Exact Match12.50%Fine-tuned Longformer
Token F126.66%vs. 16.16% base Longformer
Evidence Recovery49.00%vs. 30.00% base Longformer
Evidence Token Recall60.14%vs. 45.88% base Longformer
| Model | Exact Match | Token F1 | Evidence Recovery | Evidence Token Recall |
|---|---|---|---|---|
| BERT truncated to 512 | 1.50% | 7.37% | 26.00% | 41.34% |
| Base Longformer + windows | 6.00% | 16.16% | 30.00% | 45.88% |
| QASPER-fine-tuned Longformer | 12.50% | 26.66% | 49.00% | 60.14% |
STEP 1
Provide a document
Sourcepasted-text
Characters0
Words0
PrivacyBrowser only
STEP 2
Ask a focused question
ModelNot loaded
InferenceNot started
Choose a sample, upload a document, or paste text to begin.
STEP 3
Inspect grounded output
Total chunks—
QA candidates—
Latency—
Runtime—
Answer
No answer generated yet.
Confidence proxy
—
Supporting paragraph
No supporting paragraph selected yet.
Highlighted evidence
Highlighted evidence will appear here.
Diagnostics and candidate answers
{}SYSTEM DESIGN
Long-document processing flow
Document input
↓Overlapping chunks
↓Lexical candidate retrieval
↓Browser Transformer QA
↓Answer + supporting evidence
MODEL DISCLOSURE
Core model versus live browser model
| Component | Evaluated Python project | Live Static Space |
|---|---|---|
| Model | anmol-unitmole/longformer-qasper-document-qa | Xenova/distilbert-base-cased-distilled-squad |
| Context strategy | Longformer sparse attention + sliding windows | Retrieval over overlapping short chunks |
| Inference location | Python / PyTorch | Visitor browser / ONNX Runtime |
| Purpose | Training, benchmarking, long-context evaluation | Free interactive portfolio demonstration |