The science
What CureCrunch computes, where the data comes from, how results are checked, and what they cannot tell you.
Data sources
CureCrunch currently analyzes open-access data from The Cancer Genome Atlas (TCGA) PanCancer studies, obtained through cBioPortal. We have prepared 20 studies, including cancers of the brain (glioblastoma), prostate, lung, ovary, breast, stomach, liver, bladder, head and neck, skin, uterus, kidney and colon and rectum. For mutation analyses we also use the PanCanAtlas MC3 mutation calls (Ellrott et al., Cell Systems, 2018).
We use only open-access data. Volunteer devices do not receive controlled-access or patient-level restricted data.
Analyses
Mutation frequency
For each gene, the share of tumors in a study with a coding mutation. Verified units are pooled across a study into one frequency per gene.
Tumor mutational burden (TMB)
The number of coding mutations per megabase of sequenced genome in each tumor. A tumor above 10 mutations per megabase is commonly labeled high. We follow the convention of Chalmers et al. (Genome Medicine, 2017) for the MC3 data and report the median and the share of high-burden tumors.
Copy-number variation
Whether a gene has extra copies (amplified) or lost copies (deleted) in each tumor. A gene is reported as recurrently amplified or deleted when it is changed in at least 10% of the tumors in a work unit.
Copy-number coverage is limited today. Work units prepared so far cover only part of the genome, and we are expanding to a curated set of cancer genes. Results should be read as results for the genes covered, not for the whole genome.
In development
Kaplan-Meier survival analysis is in development and is not yet open to volunteers.
How results are checked
- Every work unit is computed on at least two devices, and the summaries must agree (see How it works).
- The analysis program is built reproducibly and tested against more than 500 test cases, matched byte for byte against a second, independently written implementation.
- For mutation frequency, survival correlation and tumor mutational burden, our pipeline matches published reference values on 29 of 29 active benchmark checks. Eight further checks are still under review or blocked, and they are not counted in that figure.
- As an early sanity check on copy-number results, prostate tumor results recover the TMPRSS2–ERG deletion on chromosome 21q22, which is well documented in the literature.
What results can and cannot tell you
CureCrunch results describe groups of tumors in public research datasets. They are not medical advice, they say nothing about any one person's health, and they do not guide treatment.
- Results are early. Volunteer results are still being verified, and nothing is published as a finding until it has been reviewed by a statistician and a cancer researcher.
- TCGA studies are not a random sample of all patients. A frequency in TCGA may differ from the frequency in a different population.
- Summaries pooled from many small units can differ from a single analysis run on the whole dataset. We are writing up each pooling method for outside review.
Acknowledgments
The TCGA Research Network, the cBioPortal team and the patients who contributed samples made this data available. The Cancer Genome Atlas is a joint effort of the National Cancer Institute and the National Human Genome Research Institute.