Person: Philippakis, Anthony Andrew
Email Address
AA Acceptance Date
Birth Date
Research Projects
Organizational Units
Job Title
Last Name
First Name
Name
Search Results
Publication Predicting the Binding Preference of Transcription Factors to Individual DNA (\kappa)-mers
(Oxford University Press, 2008) Alleyne, Trevis M.; Peña-Castillo, Lourdes; Badis, Gwenael; Talukder, Shaheynoor; Berger, Michael F.; Gehrke, Andrew R.; Philippakis, Anthony Andrew; Bulyk, Martha; Morris, Quaid D.; Hughes, Timothy R.Motivation: Recognition of specific DNA sequences is a central mechanism by which transcription factors (TFs) control gene expression. Many TF-binding preferences, however, are unknown or poorly characterized, in part due to the difficulty associated with determining their specificity experimentally, and an incomplete understanding of the mechanisms governing sequence specificity. New techniques that estimate the affinity of TFs to all possible (\kappa)-mers provide a new opportunity to study DNA–protein interaction mechanisms, and may facilitate inference of binding preferences for members of a given TF family when such information is available for other family members. Results: We employed a new dataset consisting of the relative preferences of mouse homeodomains for all eight-base DNA sequences in order to ask how well we can predict the binding profiles of homeodomains when only their protein sequences are given. We evaluated a panel of standard statistical inference techniques, as well as variations of the protein features considered. Nearest neighbour among functionally important residues emerged among the most effective methods. Our results underscore the complexity of TF–DNA recognition, and suggest a rational approach for future analyses of TF families.
Publication Mapping Copy Number Variation by Population Scale Genome Sequencing
(Nature Publishing Group, 2011) Mills, Ryan Edward; Handsaker, Robert; Korn, Joshua; Nemesh, James; Shi, Xinghua; Lee, Charles; McCarroll, Steven; Altshuler, David; Gabriel, Stacey B.; Lander, Eric; Ambrogio, Lauren; Bloom, Toby; Cibulskis, Kristian; Fennell, Tim J.; Jaffe, David B.; Shefler, Erica; Sougnez, Carrie L.; Daly, Mark; DePristo, Mark A.; Ball, Aaron D.; Banks, Eric; Browning, Brian L.; Garimella, Kiran V.; Grossman, Sharon; Hanna, Matt; Hartl, Chris; Kernytsky, Andrew M.; Li, Heng; Maguire, Jared R.; McKenna, Aaron; Philippakis, Anthony Andrew; Poplin, Ryan E.; Price, Alkes; Rivas, Manuel A.; Sabeti, Pardis; Schaffner, Stephen; Shlyakhter, Ilya; Wilkinson, JaneGenomic structural variants (SVs) are abundant in humans, differing from other forms of variation in extent, origin and functional impact. Despite progress in SV characterization, the nucleotide resolution architecture of most SVs remains unknown. We constructed a map of unbalanced SVs (that is, copy number variants) based on whole genome DNA sequencing data from 185 human genomes, integrating evidence from complementary SV discovery approaches with extensive experimental validations. Our map encompassed 22,025 deletions and 6,000 additional SVs, including insertions and tandem duplications. Most SVs (53%) were mapped to nucleotide resolution, which facilitated analysing their origin and functional impact. We examined numerous whole and partial gene deletions with a genotyping approach and observed a depletion of gene disruptions amongst high frequency deletions. Furthermore, we observed differences in the size spectra of SVs originating from distinct formation mechanisms, and constructed a map of SV hotspots formed by common mechanisms. Our analytical framework and SV map serves as a resource for sequencing-based association studies.
Publication A Map of Human Genome Variation from Population Scale Sequencing
(Nature Publishing Group, 2010) Altshuler, David; Lander, Eric; Ambrogio, Lauren; Bloom, Toby; Cibulskis, Kristian; Fennell, Tim J.; Gabriel, Stacey B.; Jaffe, David B.; Shefler, Erica; Sougnez, Carrie L.; Lee, Charles; Mills, Ryan Edward; Shi, Xinghua; Daly, Mark; DePristo, Mark A.; Ball, Aaron D.; Banks, Eric; Browning, Brian L.; Garimella, Kiran V.; Grossman, Sharon; Handsaker, Robert; Hanna, Matt; Hartl, Chris; Kernytsky, Andrew M.; Korn, Joshua M.; Li, Heng; Maguire, Jared R.; McCarroll, Steven; Nemesh, James C.; McKenna, Aaron; Philippakis, Anthony Andrew; Poplin, Ryan E.; Price, Alkes; Rivas, Manuel A.; Sabeti, Pardis; Schaffner, Stephen; Shlyakhter, IlyaThe 1000 Genomes Project aims to provide a deep characterization of human genome sequence variation as a foundation for investigating the relationship between genotype and phenotype. Here we present results of the pilot phase of the project, designed to develop and compare different strategies for genome-wide sequencing with high-throughput platforms. We undertook three projects: low-coverage whole-genome sequencing of 179 individuals from four populations; high-coverage sequencing of two mother–father–child trios; and exon-targeted sequencing of 697 individuals from seven populations. We describe the location, allele frequency and local haplotype structure of approximately 15 million single nucleotide polymorphisms, 1 million short insertions and deletions, and 20,000 structural variants, most of which were previously undescribed. We show that, because we have catalogued the vast majority of common variation, over 95% of the currently accessible variants found in any individual are present in this data set. On average, each person is found to carry approximately 250 to 300 loss-of-function variants in annotated genes and 50 to 100 variants previously implicated in inherited disorders. We demonstrate how these results can be used to inform association and functional studies. From the two trios, we directly estimate the rate of de novo germline base substitution mutations to be approximately (10^{−8}) per base pair per generation. We explore the data with regard to signatures of natural selection, and identify a marked reduction of genetic variation in the neighbourhood of genes, due to selection at linked sites. These methods and public data will support the next phase of human genetic research.
Publication Inferring Condition-Specific Transcription Factor Function from DNA Binding and Gene Expression Data
(2007) McCord, Rachel Patton; Berger, Michael F; Philippakis, Anthony Andrew; Bulyk, MarthaNumerous genomic and proteomic datasets are permitting the elucidation of transcriptional regulatory networks in the yeast Saccharomyces cerevisiae. However, predicting the condition dependence of regulatory network interactions has been challenging, because most protein–DNA interactions identified in vivo are from assays performed in one or a few cellular states. Here, we present a novel method to predict the condition-specific functions of S. cerevisiae transcription factors (TFs) by integrating 1327 microarray gene expression data sets and either comprehensive TF binding site data from protein binding microarrays (PBMs) or in silico motif data. Importantly, our method does not impose arbitrary thresholds for calling target regions ‘bound' or genes ‘differentially expressed', but rather allows all the information derived from a TF binding or gene expression experiment to be considered. We show that this method can identify environmental, physical, and genetic interactions, as well as distinct sets of genes that might be activated or repressed by a single TF under particular conditions. This approach can be used to suggest conditions for directed in vivo experimentation and to predict TF function.
Publication Expression-Guided in Silico Evaluation of Candidate Cis Regulatory Codes for Drosophila Muscle Founder Cells
(Public Library of Science, 2006) Philippakis, Anthony Andrew; Busser, Brian W; Gisselbrecht, Stephen S; He, Fangxue Sherry; Estrada, Beatriz; Michelson, Alan; Bulyk, MarthaWhile combinatorial models of transcriptional regulation can be inferred for metazoan systems from a priori biological knowledge, validation requires extensive and time-consuming experimental work. Thus, there is a need for computational methods that can evaluate hypothesized cis regulatory codes before the difficult task of experimental verification is undertaken. We have developed a novel computational framework (termed “CodeFinder”) that integrates transcription factor binding site and gene expression information to evaluate whether a hypothesized transcriptional regulatory model (TRM; i.e., a set of co-regulating transcription factors) is likely to target a given set of co-expressed genes. Our basic approach is to simultaneously predict cis regulatory modules (CRMs) associated with a given gene set and quantify the enrichment for combinatorial subsets of transcription factor binding site motifs comprising the hypothesized TRM within these predicted CRMs. As a model system, we have examined a TRM experimentally demonstrated to drive the expression of two genes in a sub-population of cells in the developing Drosophila mesoderm, the somatic muscle founder cells. This TRM was previously hypothesized to be a general mode of regulation for genes expressed in this cell population. In contrast, the present analyses suggest that a modified form of this cis regulatory code applies to only a subset of founder cell genes, those whose gene expression responds to specific genetic perturbations in a similar manner to the gene on which the original model was based. We have confirmed this hypothesis by experimentally discovering six (out of 12 tested) new CRMs driving expression in the embryonic mesoderm, four of which drive expression in founder cells.