This dataset was created by DNAstack with data from NCBI's Sequence Read Archive (SRA) in order to serve the research community during the COVID-19 pandemic, providing up-to-date SARS-CoV-2 data that has been processed alongside other public data sources using a standardized workflow. The use of a standardized workflow to produce this harmonized dataset allows public data that were generated using different methodologies to be combined and compared for a more powerful global analysis of available SARS-CoV-2 data, allowing researchers rapid access to aggregated downstream results for quick and accurate insight generation. The full bioinformatics processing pipeline written in WDL can be found on Dockstore (https://dockstore.org/workflows/github.com/DNAstack/covid-processing-pipeline/covid-19-varcal-from-assembly:master) and on Github (https://github.com/DNAstack/covid-processing-pipeline). Methodology: SRA-formatted data was converted to the standard FASTQ format using the sra-toolkit (https://github.com/ncbi/sra-tools). FASTQ files were aligned to the SARS-CoV-2 reference genome (https://www.ncbi.nlm.nih.gov/nuccore/MN908947) to produce alignment files (BAM format), which were...
Access is restricted by the data custodian. The collection's description and structure are public; the data is not available for querying through this network.
NCBI SRA SARS-CoV-2 Genomes is published on Viral AI.