This tool analyzes your uploaded clinical text, finds the best supporting evidence per field, and helps you complete the EBMT Leukemia Registry form.
Workflow
1→2
or
→3→4→5→6→7→8
After indexing, evidence is retrieved per field; confirm the chosen snippet (or extract value) before printing the EBMT form.
Work is saved in localStorage, so refresh won't lose progress.
NOTE: This is a beta version and may not reflect the entire EBMT form.
2) Evidence load
Tip: keep evidence as plain text. (Avoid Personal Information in public demos.)
Show files
3) Form (Pages 1–3 scope)
Focus: Header, Disease, Haematological values, Chromosome analysis (with one field per abnormality).
About
Motivation: EBMT leukemia registry entries require careful mapping from clinical text to structured fields.
This prototype keeps the user in control: it retrieves likely relevant evidence snippets, then you confirm or extract values.
How it works (technical)
Upload plain-text evidence and tag each file (e.g. doctor letter, lab results, cytogenetics).
Evidence text is chunked, then each chunk is embedded into a vector using Xenova/all-MiniLM-L6-v2.
Building the index enables fast retrieval: for each form field, the tool embeds the field’s query and finds the most similar evidence chunks (cosine similarity), constrained by the field’s allowed tags.
Date fields use a dedicated keyword/context scorer over pre-extracted date hits, then fall back to embedding retrieval if needed.
Everything runs in your browser; embeddings may be cached in-memory depending on localStorage size limits.
Why evidence-first?
The tool does not attempt to be clinically authoritative. It surfaces multiple candidate snippets per field (“Suggest”) so you can review context, copy it, and ensure the final registry output matches the underlying evidence.
Filling the EBMT Leukemia Registry form
After indexing, click Suggest per field, then use Use snippet / Extract value and optionally Confirm before printing the form.
Runtime note
CORS note: Do not open as file://. Use GitHub Pages or run a local server.
Contact
Carlos Vega: write me at carlos.vega[at]lih.lu
Embeddings Index
What you’re seeing: during indexing, each uploaded text chunk is converted into an embedding (a vector) using Xenova/all-MiniLM-L6-v2.
Retrieval later compares the embedded form-field query to these chunk vectors (cosine similarity) to rank the most relevant evidence snippets.
Column meanings: File = original evidence file, Tags = tags assigned to that file, Snippet = the chunk text used for the embedding, Embedding = whether the vector exists (and its dimension), Storage = whether embeddings were kept in-memory or are missing due to localStorage size limits.