README
======

Dataset title:
Principles of subclonal gene dosage across human cancers

Recommended Citation:  
Kolbeinsdottir, S., & Enge, M. (2026). Data for: Principles of subclonal gene dosage across human cancers (Version 1) [Data set]. Karolinska Institutet. https://doi.org/10.48723/gr42-ja30

See the Researchdata.se catalogue entry for this dataset - accessible via the DOI link above - for compete metadata, including the dataset description, contact information and information about how to obtain the dataset. 


Responsible organization:
Karolinska Institutet 

Description: The data shared here is joint single-cell Whole genome sequencing (WGS) and single-cell mRNA sequencing from primary patient material from different cancer types. Primary samples from pediatric acute lymphoblastic leukemias, acute myeloid leukemias, breast cancers, melanomas and sarcomas were processed using Direct nuclear tagmentation and RNA-sequencing (DNTR-seq) (doi: 10.1016/j.molcel.2020.09.025). Data shared here are the fastq files of all single cells sequenced in this study.

License: [State any license you are imposing and explain any exceptions for particular files or data.]

This dataset contains personal data and is subject to legal protections. This dataset can be requested from the Researchdata.se research data catalog using the link above.

Essential information for dataset reuse.
----------------------------------------

The data shared here are paired-end fasts files obtained by single-cell WGS and sincle-cell mRNA sequencing. The samples are primary samples from patients diagnosed with the cancers described above. Solid tumor samples were mechanically and enzymatically dissociated, and all samples were enriched for live cells prior to flow cytometry isolation of single cells. Single cells isolated from the samples were processed using DNTR-seq, in which the nuclei undergoes direct tagmentation, enabling even coverage of the genome at ultra-low coverage, while the mRNA fraction is prepared with a smartseq2 based protocol, enabling whole transcript sequencing. For the study the data was processed using the ASCENT pipeline (https://github.com/EngeLab/ASCENT)


Datafile descriptions
----------------------

Datafiles are paired-end fastq files (21826 read 1 and 21826 read 2 files), combined 2 TB of data. 

Reference information
---------------------

In order to analyse the data run the pipeline available on GitHub (https://github.com/EngeLab/ASCENT), further information found in the original publication (doi:10.1093/nar/gkaf919). 


Data Quality:
The raw files included here are sequencing data for cells that passed the QC of the ASCENT pipeline from all patients included. 

