forked from awslabs/open-data-registry
-
Notifications
You must be signed in to change notification settings - Fork 0
/
1000-genomes-data-lakehouse-ready.yaml
33 lines (32 loc) · 2.24 KB
/
1000-genomes-data-lakehouse-ready.yaml
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
Name: 1000 Genomes Phase 3 Reanalysis with DRAGEN 3.5 - Data Lakehouse Ready
Description: "The 1000 Genomes Project is an international collaboration which has established the most detailed catalogue of human genetic variation, including SNPs, structural variants, and their haplotype context. There were a total of 3202 individuals sequenced as part of Phase 3 of this project. The high coverage samples were processed using the Illumina DRAGEN v3.5.7b pipeline and are available at s3://1000genomes-dragen/. This dataset contains the VCFs transformed to Parquet/ORC in 3 different schemas - partitioned by samples, partitioned by chromosome and a nested data format. These representations of the 1000 Genomes DRAGEN data are stored in Parquet/ORC format and can be queried through [Amazon Athena](https://aws.amazon.com/athena/?whats-new-cards.sort-by=item.additionalFields.postDateTime&whats-new-cards.sort-order=desc). To add these tables to your Glue Data Catalog and for sample queries on this dataset, please refer to the link in our Documentation."
Documentation: https://github.com/aws-samples/data-lake-as-code/tree/roda#readme
Contact: https://github.com/aws-samples/data-lake-as-code/issues
ManagedBy: "[Amazon Web Services](https://aws.amazon.com/)"
UpdateFrequency: Not updated
Tags:
- biology
- bioinformatics
- genetic
- genomic
- Homo sapiens
- life sciences
- parquet
- population genetics
- vcf
License: "Data from the 1000 Genomes Project is now available without embargo, following the final publication from the project. Use of the data should be cited in the usual way, with current details available at http://www.internationalgenome.org/faq/how-do-i-cite-1000-genomes-project."
Resources:
- Description: Parquet representations of 1000 Genomes VCF outputs from DRAGEN, ready for enrollment into Data Lake as Code.
ARN: arn:aws:s3:::aws-roda-hcls-datalake/thousandgenomes_dragen
Region: us-east-1
Type: S3 Bucket
DataAtWork:
Tutorials:
- Title: Sample Queries on the 1000 Genomes, gnomAD and ClinVar data Lake
URL: https://github.com/aws-samples/aws-genomics-datalake/blob/main/1000Genomes.ipynb
AuthorName: Sujaya Srinivasan
Services:
- Athena
- Glue
Tools & Applications:
Publications: