Person:

Sunyaev, Shamil

Loading...
Profile Picture

Email Address

AA Acceptance Date

Birth Date

Research Projects

Organizational Units

Job Title

Last Name

Sunyaev

First Name

Shamil

Name

Sunyaev, Shamil

Search Results

Now showing 1 - 10 of 37
  • Publication

    Biocomputing Enters its Adolescence

    (BioMed Central, 2005) Sunyaev, Shamil
  • Publication

    Balancing Selection on a Regulatory Region Exhibiting Ancient Variation That Predates Human–Neandertal Divergence

    (Public Library of Science, 2013) Gokcumen, Omer; Zhu, Qihui; Mulder, Lubbertus C. F.; Iskow, Rebecca C.; Austermann, Christian; Scharer, Christopher D.; Raj, Towfique; Boss, Jeremy M.; Sunyaev, Shamil; Price, Alkes; Stranger, Barbara; Simon, Viviana; Lee, Charles

    Ancient population structure shaping contemporary genetic variation has been recently appreciated and has important implications regarding our understanding of the structure of modern human genomes. We identified a ∼36-kb DNA segment in the human genome that displays an ancient substructure. The variation at this locus exists primarily as two highly divergent haplogroups. One of these haplogroups (the NE1 haplogroup) aligns with the Neandertal haplotype and contains a 4.6-kb deletion polymorphism in perfect linkage disequilibrium with 12 single nucleotide polymorphisms (SNPs) across diverse populations. The other haplogroup, which does not contain the 4.6-kb deletion, aligns with the chimpanzee haplotype and is likely ancestral. Africans have higher overall pairwise differences with the Neandertal haplotype than Eurasians do for this NE1 locus (p<10−15). Moreover, the nucleotide diversity at this locus is higher in Eurasians than in Africans. These results mimic signatures of recent Neandertal admixture contributing to this locus. However, an in-depth assessment of the variation in this region across multiple populations reveals that African NE1 haplotypes, albeit rare, harbor more sequence variation than NE1 haplotypes found in Europeans, indicating an ancient African origin of this haplogroup and refuting recent Neandertal admixture. Population genetic analyses of the SNPs within each of these haplogroups, along with genome-wide comparisons revealed significant FST (p = 0.00003) and positive Tajima's D (p = 0.00285) statistics, pointing to non-neutral evolution of this locus. The NE1 locus harbors no protein-coding genes, but contains transcribed sequences as well as sequences with putative regulatory function based on bioinformatic predictions and in vitro experiments. We postulate that the variation observed at this locus predates Human–Neandertal divergence and is evolving under balancing selection, especially among European populations.

  • Publication

    Patterns and rates of exonic de novo mutations in autism spectrum disorders

    (2013) Neale, Benjamin; Kou, Yan; Liu, Li; Ma'ayan, Avi; Samocha, Kaitlin E.; Sabo, Aniko; Lin, Chiao-Feng; Stevens, Christine; Wang, Li-San; Makarov, Vladimir; Polak, Paz; Yoon, Seungtai; Maguire, Jared; Crawford, Emily L.; Campbell, Nicholas G.; Geller, Evan T.; Valladares, Otto; Shafer, Chad; Liu, Han; Zhao, Tuo; Cai, Guiqing; Lihm, Jayon; Dannenfelser, Ruth; Jabado, Omar; Peralta, Zuleyma; Nagaswamy, Uma; Muzny, Donna; Reid, Jeffrey G.; Newsham, Irene; Wu, Yuanqing; Lewis, Lora; Han, Yi; Voight, Benjamin F.; Lim, Elaine; Rossin, Elizabeth; Kirby, Andrew; Flannick, Jason; Fromer, Menachem; Shakir, Khalid; Fennell, Tim; Garimella, Kiran; Banks, Eric; Poplin, Ryan; Gabriel, Stacey; DePristo, Mark; Wimbish, Jack R.; Boone, Braden E.; Levy, Shawn E.; Betancur, Catalina; Sunyaev, Shamil; Boerwinkle, Eric; Buxbaum, Joseph D.; Cook, Edwin H.; Devlin, Bernie; Gibbs, Richard A.; Roeder, Kathryn; Schellenberg, Gerard D.; Sutcliffe, James S.; Daly, Mark

    Autism spectrum disorders (ASD) are believed to have genetic and environmental origins, yet in only a modest fraction of individuals can specific causes be identified1,2. To identify further genetic risk factors, we assess the role of de novo mutations in ASD by sequencing the exomes of ASD cases and their parents (n= 175 trios). Fewer than half of the cases (46.3%) carry a missense or nonsense de novo variant and the overall rate of mutation is only modestly higher than the expected rate. In contrast, there is significantly enriched connectivity among the proteins encoded by genes harboring de novo missense or nonsense mutations, and excess connectivity to prior ASD genes of major effect, suggesting a subset of observed events are relevant to ASD risk. The small increase in rate of de novo events, when taken together with the connections among the proteins themselves and to ASD, are consistent with an important but limited role for de novo point mutations, similar to that documented for de novo copy number variants. Genetic models incorporating these data suggest that the majority of observed de novo events are unconnected to ASD, those that do confer risk are distributed across many genes and are incompletely penetrant (i.e., not necessarily causal). Our results support polygenic models in which spontaneous coding mutations in any of a large number of genes increases risk by 5 to 20-fold. Despite the challenge posed by such models, results from de novo events and a large parallel case-control study provide strong evidence in favor of CHD8 and KATNAL2 as genuine autism risk factors.

  • Publication

    Deleterious Alleles in the Human Genome Are on Average Younger Than Neutral Alleles of the Same Frequency

    (Public Library of Science, 2013) Kiezun, Adam; Pulit, Sara L.; Francioli, Laurent C.; van Dijk, Freerk; Swertz, Morris; Boomsma, Dorret I.; van Duijn, Cornelia M.; Slagboom, P. Eline; van Ommen, G. J. B.; Wijmenga, Cisca; de Bakker, Paul; Sunyaev, Shamil

    Large-scale population sequencing studies provide a complete picture of human genetic variation within the studied populations. A key challenge is to identify, among the myriad alleles, those variants that have an effect on molecular function, phenotypes, and reproductive fitness. Most non-neutral variation consists of deleterious alleles segregating at low population frequency due to incessant mutation. To date, studies characterizing selection against deleterious alleles have been based on allele frequency (testing for a relative excess of rare alleles) or ratio of polymorphism to divergence (testing for a relative increase in the number of polymorphic alleles). Here, starting from Maruyama's theoretical prediction (Maruyama T (1974), Am J Hum Genet USA 6:669–673) that a (slightly) deleterious allele is, on average, younger than a neutral allele segregating at the same frequency, we devised an approach to characterize selection based on allelic age. Unlike existing methods, it compares sets of neutral and deleterious sequence variants at the same allele frequency. When applied to human sequence data from the Genome of the Netherlands Project, our approach distinguishes low-frequency coding non-synonymous variants from synonymous and non-coding variants at the same allele frequency and discriminates between sets of variants independently predicted to be benign or damaging for protein structure and function. The results confirm the abundance of slightly deleterious coding variation in humans.

  • Publication

    Analysis of Sequence Conservation at Nucleotide Resolution

    (Public Library of Science, 2007) Asthana, Saurabh; Roytberg, Mikhail; Stamatoyannopoulos, John; Sunyaev, Shamil

    One of the major goals of comparative genomics is to understand the evolutionary history of each nucleotide in the human genome sequence, and the degree to which it is under selective pressure. Ascertainment of selective constraint at nucleotide resolution is particularly important for predicting the functional significance of human genetic variation and for analyzing the sequence substructure of cis-regulatory sequences and other functional elements. Current methods for analysis of sequence conservation are focused on delineation of conserved regions comprising tens or even hundreds of consecutive nucleotides. We therefore developed a novel computational approach designed specifically for scoring evolutionary conservation at individual base-pair resolution. Our approach estimates the rate at which each nucleotide position is evolving, computes the probability of neutrality given this rate estimate, and summarizes the result in a Sequence CONservation Evaluation (SCONE) score. We computed SCONE scores in a continuous fashion across 1% of the human genome for which high-quality sequence information from up to 23 genomes are available. We show that SCONE scores are clearly correlated with the allele frequency of human polymorphisms in both coding and noncoding regions. We find that the majority of noncoding conserved nucleotides lie outside of longer conserved elements predicted by other conservation analyses, and are experiencing ongoing selection in modern humans as evident from the allele frequency spectrum of human polymorphism. We also applied SCONE to analyze the distribution of conserved nucleotides within functional regions. These regions are markedly enriched in individually conserved positions and short (<15 bp) conserved “chunks.” Our results collectively suggest that the majority of functionally important noncoding conserved positions are highly fragmented and reside outside of canonically defined long conserved noncoding sequences. A small subset of these fragmented positions may be identified with high confidence.

  • Publication

    Computational and Statistical Approaches to Analyzing Variants Identified by Exome Sequencing

    (BioMed Central, 2011) Stitziel, Nathan O.; Kiezun, Adam; Sunyaev, Shamil

    New sequencing technology has enabled the identification of thousands of single nucleotide polymorphisms in the exome, and many computational and statistical approaches to identify disease-association signals have emerged.

  • Publication

    Cell-of-origin chromatin organization shapes the mutational landscape of cancer

    (2015) Polak, Paz; Karlić, Rosa; Koren, Amnon; Thurman, Robert; Sandstrom, Richard; Lawrence, Michael; Reynolds, Alex; Rynes, Eric; Vlahoviček, Kristian; Stamatoyannopoulos, John A.; Sunyaev, Shamil

    Cancer is a disease potentiated by mutations in somatic cells. Cancer mutations are not distributed uniformly along the genome. Instead, different genomic regions vary by up to 5-fold in the local density of somatic mutations1, posing a fundamental problem for statistical methods of cancer genomics. Epigenomic organization has been proposed as a major determinant of the cancer mutational landscape1-5. However, both somatic mutagenesis and epigenomic features are highly cell-type-specific6,7. We investigated the distribution of mutations in multiple samples of diverse cancer types and compared them to cell-type-specific epigenomic features. Here, we show that chromatin accessibility and modification, together with replication timing, explain up to 86% of the variance in mutation rates along cancer genomes. Overwhelmingly, the best predictors of local somatic mutation density are epigenomic features derived from the most likely cell type of origin of the corresponding malignancy. Moreover, we find that cell-of-origin chromatin features are much stronger determinants of cancer mutation profiles than chromatin features of cognate cancer cell lines. We show further that the cell type of origin of a cancer can be accurately determined based on the distribution of mutations along its genome. Thus, DNA sequence of a cancer genome encompasses a wealth of information about the identity and epigenomic features of its cell of origin.

  • Publication

    Integrative analysis of 111 reference human epigenomes

    (2015) Kundaje, Anshul; Meuleman, Wouter; Ernst, Jason; Bilenky, Misha; Yen, Angela; Kheradpour, Pouya; Zhang, Zhizhuo; Heravi-Moussavi, Alireza; Liu, Yaping; Amin, Viren; Ziller, Michael; Whitaker, John W; Schultz, Matthew D; Sandstrom, Richard S; Eaton, Matthew L; Wu, Yi-Chieh; Wang, Jianrong; Ward, Lucas D; Sarkar, Abhishek; Quon, Gerald; Pfenning, Andreas; Wang, Xinchen; Claussnitzer, Melina; Coarfa, Cristian; Harris, R Alan; Shoresh, Noam; Epstein, Charles B; Gjoneska, Elizabeta; Leung, Danny; Xie, Wei; Hawkins, R David; Lister, Ryan; Hong, Chibo; Gascard, Philippe; Mungall, Andrew J; Moore, Richard; Chuah, Eric; Tam, Angela; Canfield, Theresa K; Hansen, R Scott; Kaul, Rajinder; Sabo, Peter J; Bansal, Mukul S; Carles, Annaick; Dixon, Jesse R; Farh, Kai-How; Feizi, Soheil; Karlic, Rosa; Kim, Ah-Ram; Kulkarni, Ashwinikumar; Li, Daofeng; Lowdon, Rebecca; Mercer, Tim R; Neph, Shane J; Onuchic, Vitor; Polak, Paz; Rajagopal, Nisha; Ray, Pradipta; Sallari, Richard C; Siebenthall, Kyle T; Sinnott-Armstrong, Nicholas; Stevens, Michael; Thurman, Robert E; Wu, Jie; Zhang, Bo; Zhou, Xin; Beaudet, Arthur E; Boyer, Laurie A; De Jager, Philip; Farnham, Peggy J; Fisher, Susan J; Haussler, David; Jones, Steven; Li, Wei; Marra, Marco; McManus, Michael T; Sunyaev, Shamil; Thomson, James A; Tlsty, Thea D; Tsai, Li-Huei; Wang, Wei; Waterland, Robert A; Zhang, Michael; Chadwick, Lisa H; Bernstein, Bradley; Costello, Joseph F; Ecker, Joseph R; Hirst, Martin; Meissner, Alexander; Milosavljevic, Aleksandar; Ren, Bing; Stamatoyannopoulos, John A; Wang, Ting; Kellis, Manolis

    The reference human genome sequence set the stage for studies of genetic variation and its association with human disease, but a similar reference has lacked for epigenomic studies. To address this need, the NIH Roadmap Epigenomics Consortium generated the largest collection to-date of human epigenomes for primary cells and tissues. Here, we describe the integrative analysis of 111 reference human epigenomes generated as part of the program, profiled for histone modification patterns, DNA accessibility, DNA methylation, and RNA expression. We establish global maps of regulatory elements, define regulatory modules of coordinated activity, and their likely activators and repressors. We show that disease and trait-associated genetic variants are enriched in tissue-specific epigenomic marks, revealing biologically-relevant cell types for diverse human traits, and providing a resource for interpreting the molecular basis of human disease. Our results demonstrate the central role of epigenomic information for understanding gene regulation, cellular differentiation, and human disease.

  • Publication

    Multiple rare alleles at LDLR and APOA5 confer risk for early-onset myocardial infarction

    (2014) Do, Ron; Stitziel, Nathan O.; Won, Hong-Hee; Jørgensen, Anders Berg; Duga, Stefano; Merlini, Pier Angelica; Kiezun, Adam; Farrall, Martin; Goel, Anuj; Zuk, Or; Guella, Illaria; Asselta, Rosanna; Lange, Leslie A.; Peloso, Gina M; Auer, Paul L.; Girelli, Domenico; Martinelli, Nicola; Farlow, Deborah N.; DePristo, Mark A.; Roberts, Robert; Stewart, Alexander F.R.; Saleheen, Danish; Danesh, John; Epstein, Stephen E.; Sivapalaratnam, Suthesh; Hovingh, G. Kees; Kastelein, John J.; Samani, Nilesh J.; Schunkert, Heribert; Erdmann, Jeanette; Shah, Svati H.; Kraus, William E.; Davies, Robert; Nikpay, Majid; Johansen, Christopher T.; Wang, Jian; Hegele, Robert A.; Hechter, Eliana; Marz, Winfried; Kleber, Marcus E.; Huang, Jie; Johnson, Andrew D.; Li, Mingyao; Burke, Greg L.; Gross, Myron; Liu, Yongmei; Assimes, Themistocles L.; Heiss, Gerardo; Lange, Ethan M.; Folsom, Aaron R.; Taylor, Herman A.; Olivieri, Oliviero; Hamsten, Anders; Clarke, Robert; Reilly, Dermot F.; Yin, Wu; Rivas, Manuel A.; Donnelly, Peter; Rossouw, Jacques E.; Psaty, Bruce M.; Herrington, David M.; Wilson, James G.; Rich, Stephen S.; Bamshad, Michael J.; Tracy, Russell P.; Cupples, L. Adrienne; Rader, Daniel J.; Reilly, Muredach P.; Spertus, John A.; Cresci, Sharon; Hartiala, Jaana; Tang, W.H. Wilson; Hazen, Stanley L.; Allayee, Hooman; Reiner, Alex P.; Carlson, Christopher S.; Kooperberg, Charles; Jackson, Rebecca D.; Boerwinkle, Eric; Lander, Eric S.; Schwartz, Stephen M.; Siscovick, David S.; McPherson, Ruth; Tybjaerg-Hansen, Anne; Abecasis, Goncalo R.; Watkins, Hugh; Nickerson, Deborah A.; Ardissino, Diego; Sunyaev, Shamil; O’Donnell, Christopher J.; Altshuler, David; Gabriel, Stacey; Kathiresan, Sekar

    Summary Myocardial infarction (MI), a leading cause of death around the world, displays a complex pattern of inheritance1,2. When MI occurs early in life, the role of inheritance is substantially greater1. Previously, rare mutations in low-density lipoprotein (LDL) genes have been shown to contribute to MI risk in individual families3–8 whereas common variants at more than 45 loci have been associated with MI risk in the population9–15. Here, we evaluate the contribution of rare mutations to MI risk in the population. We sequenced the protein-coding regions of 9,793 genomes from patients with MI at an early age (≤50 years in males and ≤60 years in females) along with MI-free controls. We identified two genes where rare coding-sequence mutations were more frequent in cases versus controls at exome-wide significance. At low-density lipoprotein receptor (LDLR), carriers of rare, damaging mutations (3.1% of cases versus 1.3% of controls) were at 2.4-fold increased risk for MI; carriers of null alleles at LDLR were at even higher risk (13-fold difference). This sequence-based estimate of the proportion of early MI cases due to LDLR mutations is remarkably similar to an estimate made more than 40 years ago using total cholesterol16. At apolipoprotein A-V (APOA5), carriers of rare nonsynonymous mutations (1.4% of cases versus 0.6% of controls) were at 2.2-fold increased risk for MI. When compared with non-carriers, LDLR mutation carriers had higher plasma LDL cholesterol whereas APOA5 mutation carriers had higher plasma triglycerides. Recent evidence has connected MI risk with coding sequence mutations at two genes functionally related to APOA5, namely lipoprotein lipase15,17 and apolipoprotein C318,19. When combined, these observations suggest that, beyond LDL cholesterol, disordered metabolism of triglyceride-rich lipoproteins contributes to MI risk.

  • Publication

    No evidence that selection has been less effective at removing deleterious mutations in Europeans than in Africans

    (2014) Do, Ron; Balick, Daniel; Li, Heng; Adzhubei, Ivan; Sunyaev, Shamil; Reich, David

    Non-African populations have experienced size reductions in the time since their split from West Africans, leading to the hypothesis that natural selection to remove weakly deleterious mutations has been less effective in the history of non-Africans. To test this hypothesis, we measured the per-genome accumulation of non-synonymous substitutions across diverse pairs of populations. We find no evidence for a higher load of deleterious mutations in non-Africans. However, we detect significant differences among more divergent populations, as archaic Denisovans have accumulated non-synonymous mutations faster than either modern humans or Neanderthals. To reconcile these findings with patterns that have been interpreted as evidence of less effective removal of deleterious mutations in non-Africans than in West Africans, we use simulations to show that the observed patterns are not likely to reflect changes in the effectiveness of selection after the populations split, and instead are likely to be driven by other population genetic factors.