Person: Lange, Christoph
Email Address
AA Acceptance Date
Birth Date
Research Projects
Organizational Units
Job Title
Last Name
First Name
Name
Search Results
Publication Asthma-susceptibility variants identified using probands in case-control and family-based analyses
(BioMed Central, 2010) Murphy, Amy J; Soto-Quiros, Manuel E; Avila, Lydiana; Celedón, Juan C; O'Connor, George T; Himes, Blanca; Lasky-Su, Jessica; Wu, Ann; Wilk, Jemma; Hunninghake, Gary; Klanderman, Barbara; Lazarus, Ross; Lange, Christoph; Raby, Benjamin; Silverman, Edwin; Weiss, ScottBackground: Asthma is a chronic respiratory disease whose genetic basis has been explored for over two decades, most recently via genome-wide association studies. We sought to find asthma-susceptibility variants by using probands from a single population in both family-based and case-control association designs. Methods: We used probands from the Childhood Asthma Management Program (CAMP) in two primary genome-wide association study designs: (1) probands were combined with publicly available population controls in a case-control design, and (2) probands and their parents were used in a family-based design. We followed a two-stage replication process utilizing three independent populations to validate our primary findings. Results: We found that single nucleotide polymorphisms with similar case-control and family-based association results were more likely to replicate in the independent populations, than those with the smallest p-values in either the case-control or family-based design alone. The single nucleotide polymorphism that showed the strongest evidence for association to asthma was rs17572584, which replicated in 2/3 independent populations with an overall p-value among replication populations of 3.5E-05. This variant is near a gene that encodes an enzyme that has been implicated to act coordinately with modulators of Th2 cell differentiation and is expressed in human lung. Conclusions: Our results suggest that using probands from family-based studies in case-control designs, and combining results of both family-based and case-control approaches, may be a way to augment our ability to find SNPs associated with asthma and other complex diseases.
Publication The Association of a SNP Upstream of INSIG2 with Body Mass Index is Reproduced in Several but Not All Cohorts
(Public Library of Science, 2007) Emilsson, Valur; Hinney, Anke; Heid, Iris M; Zhu, Xiaofeng; Thorleifsson, Gudmar; Gunnarsdottir, Steinunn; Walters, G. Bragi; Thorsteinsdottir, Unnur; Kong, Augustine; Gulcher, Jeffrey; Nguyen, Thuy Trang; Scherag, André; Pfeufer, Arne; Meitinger, Thomas; Brönner, Günter; Rief, Winfried; Soto-Quiros, Manuel E; Avila, Lydiana; Groop, Leif; Tuomi, Tiinamaija; Isomaa, Bo; Bengtsson, Kristina; Butler, Johannah L; Vollmert, Caren; Celedón, Juan C; Wichmann, H. Erich; Hebebrand, Johannes; Stefansson, Kari; Abecasis, Gonçalo; Lyon, Helen N.; Lasky-Su, Jessica; Klanderman, Barbara; Raby, Benjamin; Silverman, Edwin; Weiss, Scott; Laird, Nan; Ding, Xiao; Cooper, Richard S; Fox, Caroline; O'Donnell, Christopher; Lange, Christoph; Hirschhorn, JoelA SNP upstream of the INSIG2 gene, rs7566605, was recently found to be associated with obesity as measured by body mass index (BMI) by Herbert and colleagues. The association between increased BMI and homozygosity for the minor allele was first observed in data from a genome-wide association scan of 86,604 SNPs in 923 related individuals from the Framingham Heart Study offspring cohort. The association was reproduced in four additional cohorts, but was not seen in a fifth cohort. To further assess the general reproducibility of this association, we genotyped rs7566605 in nine large cohorts from eight populations across multiple ethnicities (total n = 16,969). We tested this variant for association with BMI in each sample under a recessive model using family-based, population-based, and case-control designs. We observed a significant (p < 0.05) association in five cohorts but saw no association in three other cohorts. There was variability in the strength of association evidence across examination cycles in longitudinal data from unrelated individuals in the Framingham Heart Study Offspring cohort. A combined analysis revealed significant independent validation of this association in both unrelated (p = 0.046) and family-based (p = 0.004) samples. The estimated risk conferred by this allele is small, and could easily be masked by small sample size, population stratification, or other confounders. These validation studies suggest that the original association is less likely to be spurious, but the failure to observe an association in every data set suggests that the effect of SNP rs7566605 on BMI may be heterogeneous across population samples.
Publication ‘Location, Location, Location’: a spatial approach for rare variant analysis and an application to a study on non-syndromic cleft lip with or without cleft palate
(Oxford University Press, 2012) Fier, Heide; Won, Sungho; Prokopenko, Dmitry; AlChawa, Taofik; Ludwig, Kerstin U.; Fimmers, Rolf; Silverman, Edwin; Pagano, Marcello; Mangold, Elisabeth; Lange, ChristophMotivation: For the analysis of rare variants in sequence data, numerous approaches have been suggested. Fixed and flexible threshold approaches collapse the rare variant information of a genomic region into a test statistic with reduced dimensionality. Alternatively, the rare variant information can be combined in statistical frameworks that are based on suitable regression models, machine learning, etc. Although the existing approaches provide powerful tests that can incorporate information on allele frequencies and prior biological knowledge, differences in the spatial clustering of rare variants between cases and controls cannot be incorporated. Based on the assumption that deleterious variants and protective variants cluster or occur in different parts of the genomic region of interest, we propose a testing strategy for rare variants that builds on spatial cluster methodology and that guides the identification of the biological relevant segments of the region. Our approach does not require any assumption about the directions of the genetic effects. Results: In simulation studies, we assess the power of the clustering approach and compare it with existing methodology. Our simulation results suggest that the clustering approach for rare variants is well powered, even in situations that are ideal for standard methods. The efficiency of our spatial clustering approach is not affected by the presence of rare variants that have opposite effect size directions. An application to a sequencing study for non-syndromic cleft lip with or without cleft palate (NSCL/P) demonstrates its practical relevance. The proposed testing strategy is applied to a genomic region on chromosome 15q13.3 that was implicated in NSCL/P etiology in a previous genome-wide association study, and its results are compared with standard approaches. Availability: Source code and documentation for the implementation in R will be provided online. Currently, the R-implementation only supports genotype data. We currently are working on an extension for VCF files. Contact: heide.fier@googlemail.com
Publication PBAT: A comprehensive software package for genome-wide association analysis of complex family-based studies
(BioMed Central, 2005) Van Steen, Kristel; Lange, ChristophThe PBAT software package (v2.5) provides a unique set of tools for complex family-based association analysis at a genome-wide level. PBAT can handle nuclear families with missing parental genotypes, extended pedigrees with missing genotypic information, analysis of single nucleotide polymorphisms (SNPs), haplotype analysis, quantitative traits, multivariate/longitudinal data and time to onset phenotypes. The data analysis can be adjusted for covariates and gene/environment interactions. Haplotype-based features include sliding windows and the reconstruction of the haplotypes of the probands. PBAT's screening tools allow the user successfully to handle the multiple comparisons problem at a genome-wide level, even for 100,000 SNPs and more. Moreover, PBAT is computationally fast. A genome scan of 300,000 SNPs in 2,000 trios takes 4 central processing unit (CPU)-days. PBAT is available for Linux, Sun Solaris and Windows XP.
Publication Screening and Replication using the Same Data Set: Testing Strategies for Family-Based Studies in which All Probands Are Affected
(Public Library of Science, 2008) Murphy, Amy; Weiss, Scott; Lange, ChristophFor genome-wide association studies in family-based designs, we propose a powerful two-stage testing strategy that can be applied in situations in which parent-offspring trio data are available and all offspring are affected with the trait or disease under study. In the first step of the testing strategy, we construct estimators of genetic effect size in the completely ascertained sample of affected offspring and their parents that are statistically independent of the family-based association/transmission disequilibrium tests (FBATs/TDTs) that are calculated in the second step of the testing strategy. For each marker, the genetic effect is estimated (without requiring an estimate of the SNP allele frequency) and the conditional power of the corresponding FBAT/TDT is computed. Based on the power estimates, a weighted Bonferroni procedure assigns an individually adjusted significance level to each SNP. In the second stage, the SNPs are tested with the FBAT/TDT statistic at the individually adjusted significance levels. Using simulation studies for scenarios with up to 1,000,000 SNPs, varying allele frequencies and genetic effect sizes, the power of the strategy is compared with standard methodology (e.g., FBATs/TDTs with Bonferroni correction). In all considered situations, the proposed testing strategy demonstrates substantial power increases over the standard approach, even when the true genetic model is unknown and must be selected based on the conditional power estimates. The practical relevance of our methodology is illustrated by an application to a genome-wide association study for childhood asthma, in which we detect two markers meeting genome-wide significance that would not have been detected using standard methodology.
Publication Meta-Analysis of the INSIG2 Association with Obesity Including 74,345 Individuals: Does Heterogeneity of Estimates Relate to Study Design?
(Public Library of Science, 2009) Heid, Iris M.; Huth, Cornelia; Loos, Ruth J. F.; Kronenberg, Florian; Adamkova, Vera; Anand, Sonia S.; Ardlie, Kristin; Biebermann, Heike; Bjerregaard, Peter; Boeing, Heiner; Bouchard, Claude; Ciullo, Marina; Cooper, Jackie A.; Corella, Dolores; Dina, Christian; Engert, James C.; Fisher, Eva; Francès, Francesc; Froguel, Philippe; Hebebrand, Johannes; Hegele, Robert A.; Hinney, Anke; Hoehe, Margret R.; Hubacek, Jaroslav A.; Humphries, Steve E.; Hunt, Steven C.; Illig, Thomas; Järvelin, Marjo-Riita; Kaakinen, Marika; Kollerits, Barbara; Krude, Heiko; Kumar, Jitender; Lange, Leslie A.; Langer, Birgit; Li, Shengxu; Luchner, Andreas; Meyre, David; Mohlke, Karen L.; Mooser, Vincent; Nebel, Almut; Nguyen, Thuy Trang; Paulweber, Bernhard; Perusse, Louis; Rankinen, Tuomo; Rosskopf, Dieter; Schreiber, Stefan; Sengupta, Shantanu; Sorice, Rossella; Suk, Anita; Thorleifsson, Gudmar; Thorsteinsdottir, Unnur; Völzke, Henry; Vimaleswaran, Karani S.; Wareham, Nicholas J.; Waterworth, Dawn; Yusuf, Salim; Lindgren, Cecilia; McCarthy, Mark I.; Wichmann, H.-Erich; Allison, David B.; Hu, Frank; Qi, Lu; Lyon, Helen N.; Lange, Christoph; Hirschhorn, Joel; Laird, NanThe INSIG2 rs7566605 polymorphism was identified for obesity (BMI≥30 kg/m2) in one of the first genome-wide association studies, but replications were inconsistent. We collected statistics from 34 studies (n = 74,345), including general population (GP) studies, population-based studies with subjects selected for conditions related to a better health status (‘healthy population’, HP), and obesity studies (OB). We tested five hypotheses to explore potential sources of heterogeneity. The meta-analysis of 27 studies on Caucasian adults (n = 66,213) combining the different study designs did not support overall association of the CC-genotype with obesity, yielding an odds ratio (OR) of 1.05 (p-value = 0.27). The I2 measure of 41% (p-value = 0.015) indicated between-study heterogeneity. Restricting to GP studies resulted in a declined I2 measure of 11% (p-value = 0.33) and an OR of 1.10 (p-value = 0.015). Regarding the five hypotheses, our data showed (a) some difference between GP and HP studies (p-value = 0.012) and (b) an association in extreme comparisons (BMI≥32.5, 35.0, 37.5, 40.0 kg/m2 versus BMI less than;25 kg/m2) yielding ORs of 1.16, 1.18, 1.22, or 1.27 (p-values 0.001 to 0.003), which was also underscored by significantly increased CC-genotype frequencies across BMI categories (10.4% to 12.5%, p-value for trend = 0.0002). We did not find evidence for differential ORs (c) among studies with higher than average obesity prevalence compared to lower, (d) among studies with BMI assessment after the year 2000 compared to those before, or (e) among studies from older populations compared to younger. Analysis of non-Caucasian adults (n = 4889) or children (n = 3243) yielded ORs of 1.01 (p-value = 0.94) or 1.15 (p-value = 0.22), respectively. There was no evidence for overall association of the rs7566605 polymorphism with obesity. Our data suggested an association with extreme degrees of obesity, and consequently heterogeneous effects from different study designs may mask an underlying association when unaccounted for. The importance of study design might be under-recognized in gene discovery and association replication so far.
Publication Using Canonical Correlation Analysis to Discover Genetic Regulatory Variants
(Public Library of Science, 2010) Naylor, Melissa G.; Lin, Xihong; Weiss, Scott; Raby, Benjamin; Lange, ChristophBackground: Discovering genetic associations between genetic markers and gene expression levels can provide insight into gene regulation and, potentially, mechanisms of disease. Such analyses typically involve a linkage or association analysis in which expression data are used as phenotypes. This approach leads to a large number of multiple comparisons and may therefore lack power. We assess the potential of applying canonical correlation analysis to partitioned genomewide data as a method for discovering regulatory variants. Methodology/Principal Findings: Simulations suggest that canonical correlation analysis has higher power than standard pairwise univariate regression to detect single nucleotide polymorphisms when the expression trait has low heritability. The increase in power is even greater under the recessive model. We demonstrate this approach using the Childhood Asthma Management Program data. Conclusions/Significance: Our approach reduces multiple comparisons and may provide insight into the complex relationships between genotype and gene expression.
Publication On the Analysis of Genome-Wide Association Studies in Family-Based Designs: A Universal, Robust Analysis Approach and an Application to Four Genome-Wide Association Studies
(Public Library of Science, 2009) Won, Sungho; Wilk, Jemma; Mathias, Rasika A.; O'Donnell, Christopher; Silverman, Edwin; Barnes, Kathleen; O'Connor, George T.; Weiss, Scott; Lange, ChristophFor genome-wide association studies in family-based designs, we propose a new, universally applicable approach. The new test statistic exploits all available information about the association, while, by virtue of its design, it maintains the same robustness against population admixture as traditional family-based approaches that are based exclusively on the within-family information. The approach is suitable for the analysis of almost any trait type, e.g. binary, continuous, time-to-onset, multivariate, etc., and combinations of those. We use simulation studies to verify all theoretically derived properties of the approach, estimate its power, and compare it with other standard approaches. We illustrate the practical implications of the new analysis method by an application to a lung-function phenotype, forced expiratory volume in one second (FEV1) in 4 genome-wide association studies.
Publication Genomic Screening in Family-Based Association Testing
(BioMed Central, 2005) Murphy, Amy; McQueen, Matthew B; Lasky-Su, Jessica; Kraft, Peter; Lazarus, Ross; Laird, Nan; Lange, Christoph; Van Steen, KristelDue to the recent gains in the availability of single-nucleotide polymorphism data, genome-wide association testing has become feasible. It is hoped that this additional data may confirm the presence of disease susceptibility loci, and identify new genetic determinants of disease. However, the problem of multiple comparisons threatens to diminish any potential gains from this newly available data. To circumvent the multiple comparisons issue, we utilize a recently developed screening technique using family-based association testing. This screening methodology allows for the identification of the most promising single-nucleotide polymorphisms for testing without biasing the nominal significance level of our test statistic. We compare the results of our screening technique across univariate and multivariate family-based association tests. From our analyses, we observe that the screening technique, applied to different settings, is fairly consistent in identifying optimal markers for testing. One of the identified markers, TSC0047225, was significantly associated with both the ttth1 (p = 0.004) and ttth1-ttth4 (p = 0.004) phenotype(s). We find that both univariate- and multivariate-based screening techniques are powerful tools for detecting an association.
Publication Handling the data management needs of high-throughput sequencing data: SpeedGene, a compression algorithm for the efficient storage of genetic data
(BioMed Central, 2012) Qiao, Dandi; Yip, Wai-Ki; Lange, ChristophBackground: As Next-Generation Sequencing data becomes available, existing hardware environments do not provide sufficient storage space and computational power to store and process the data due to their enormous size. This is and will be a frequent problem that is encountered everyday by researchers who are working on genetic data. There are some options available for compressing and storing such data, such as general-purpose compression software, PBAT/PLINK binary format, etc. However, these currently available methods either do not offer sufficient compression rates, or require a great amount of CPU time for decompression and loading every time the data is accessed. Results: Here, we propose a novel and simple algorithm for storing such sequencing data. We show that, the compression factor of the algorithm ranges from 16 to several hundreds, which potentially allows SNP data of hundreds of Gigabytes to be stored in hundreds of Megabytes. We provide a C++ implementation of the algorithm, which supports direct loading and parallel loading of the compressed format without requiring extra time for decompression. By applying the algorithm to simulated and real datasets, we show that the algorithm gives greater compression rate than the commonly used compression methods, and the data-loading process takes less time. Also, The C++ library provides direct-data-retrieving functions, which allows the compressed information to be easily accessed by other C++ programs. Conclusions: The SpeedGene algorithm enables the storage and the analysis of next generation sequencing data in current hardware environment, making system upgrades unnecessary.
- «
- 1 (current)
- 2
- 3
- »