Awua, Adolf K; Adanu, Richard M K; Wiredu, Edwin K; Afari, Edwin A; Zubuch, Vanessa A; Asmah, Richard H; Severini, Alberto
2017-04-21
In addition to being useful for classification, sequence variations of human Papillomavirus (HPV) genotypes have been implicated in differential oncogenic potential and a differential association with the different histological forms of invasive cervical cancer. These associations have also been indicated for HPV genotype lineages and sub-lineages. In order to better understand the potential implications of lineage variation in the occurrence of cervical cancers in Ghana, we studied the lineages of the three most prevalent HPV genotypes among women with normal cytology as baseline to further studies. Of previously collected self- and health personnel-collected cervical specimen, 54, which were positive for HPV16, 18 and 45, were selected and the long control region (LCR) of each HPV genotype was separately amplified by a nested PCR. DNA sequences of 41 isolates obtained with the forward and reverse primers by Sanger sequencing were analysed. Nucleotide sequence variations of the HPV16 genotypes were observed at 30 positions within the LCR (7460 - 7840). Of these, 19 were the known variations for the lineages B and C (African lineages), while the other 11 positions had variations unique to the HPV16 isolates of this study. For the HPV18 isolates, the variations were at 35 positions, 22 of which were known variations of Africa lineages and the other 13 were unique variations observed for the isolates obtained in this study (at positions 7799 and 7813). HPV45 isolates had variations at 35 positions and 2 (positions 7114 and 97) were unique to the isolates of this study. This study provides the first data on the lineages of HPV 16, 18 and 45 isolates from Ghana. Although the study did not obtain full genome sequence data for a comprehensive comparison with known lineages, these genotypes were predominately of the Africa lineages and had some unique sequence variations at positions that suggest potential oncogenic implications. These data will be useful for comparison with lineages of these genotypes from women with cervical lesion and all the forms of invasive cervical cancers.
Caporale, Lynn Helena
2012-09-01
This overview of a special issue of Annals of the New York Academy of Sciences discusses uneven distribution of distinct types of variation across the genome, the dependence of specific types of variation upon distinct classes of DNA sequences and/or the induction of specific proteins, the circumstances in which distinct variation-generating systems are activated, and the implications of this work for our understanding of evolution and of cancer. Also discussed is the value of non text-based computational methods for analyzing information carried by DNA, early insights into organizational frameworks that affect genome behavior, and implications of this work for comparative genomics. © 2012 New York Academy of Sciences.
Reverse Transcription Errors and RNA-DNA Differences at Short Tandem Repeats.
Fungtammasan, Arkarachai; Tomaszkiewicz, Marta; Campos-Sánchez, Rebeca; Eckert, Kristin A; DeGiorgio, Michael; Makova, Kateryna D
2016-10-01
Transcript variation has important implications for organismal function in health and disease. Most transcriptome studies focus on assessing variation in gene expression levels and isoform representation. Variation at the level of transcript sequence is caused by RNA editing and transcription errors, and leads to nongenetically encoded transcript variants, or RNA-DNA differences (RDDs). Such variation has been understudied, in part because its detection is obscured by reverse transcription (RT) and sequencing errors. It has only been evaluated for intertranscript base substitution differences. Here, we investigated transcript sequence variation for short tandem repeats (STRs). We developed the first maximum-likelihood estimator (MLE) to infer RT error and RDD rates, taking next generation sequencing error rates into account. Using the MLE, we empirically evaluated RT error and RDD rates for STRs in a large-scale DNA and RNA replicated sequencing experiment conducted in a primate species. The RT error rates increased exponentially with STR length and were biased toward expansions. The RDD rates were approximately 1 order of magnitude lower than the RT error rates. The RT error rates estimated with the MLE from a primate data set were concordant with those estimated with an independent method, barcoded RNA sequencing, from a Caenorhabditis elegans data set. Our results have important implications for medical genomics, as STR allelic variation is associated with >40 diseases. STR nonallelic transcript variation can also contribute to disease phenotype. The MLE and empirical rates presented here can be used to evaluate the probability of disease-associated transcripts arising due to RDD. © The Author 2016. Published by Oxford University Press on behalf of the Society for Molecular Biology and Evolution.
A survey of copy number variation in the porcine genome detected from whole-genome sequence
USDA-ARS?s Scientific Manuscript database
An important challenge to post-genomic biology is relating observed phenotypic variation to the underlying genotypic variation. Genome-wide association studies (GWAS) have made thousands of connections between single nucleotide polymorphisms (SNPs) and phenotypes, implicating regions of the genome t...
Frye, Mark A; Ryu, Euijung; Nassan, Malik; Jenkins, Gregory D; Andreazza, Ana C; Evans, Jared M; McElroy, Susan L; Oglesbee, Devin; Highsmith, W Edward; Biernacka, Joanna M
2017-01-01
Converging genetic, postmortem gene-expression, cellular, and neuroimaging data implicate mitochondrial dysfunction in bipolar disorder. This study was conducted to investigate whether mitochondrial DNA (mtDNA) haplogroups and single nucleotide variants (SNVs) are associated with sub-phenotypes of bipolar disorder. MtDNA from 224 patients with Bipolar I disorder (BPI) was sequenced, and association of sequence variations with 3 sub-phenotypes (psychosis, rapid cycling, and adolescent illness onset) was evaluated. Gene-level tests were performed to evaluate overall burden of minor alleles for each phenotype. The haplogroup U was associated with a higher risk of psychosis. Secondary analyses of SNVs provided nominal evidence for association of psychosis with variants in the tRNA, ND4 and ND5 genes. The association of psychosis with ND4 (gene that encodes NADH dehydrogenase 4) was further supported by gene-level analysis. Preliminary analysis of mtDNA sequence data suggests a higher risk of psychosis with the U haplogroup and variation in the ND4 gene implicated in electron transport chain energy regulation. Further investigation of the functional consequences of this mtDNA variation is encouraged. Copyright © 2016. Published by Elsevier Ltd.
Identification of structural variation in mouse genomes.
Keane, Thomas M; Wong, Kim; Adams, David J; Flint, Jonathan; Reymond, Alexandre; Yalcin, Binnaz
2014-01-01
Structural variation is variation in structure of DNA regions affecting DNA sequence length and/or orientation. It generally includes deletions, insertions, copy-number gains, inversions, and transposable elements. Traditionally, the identification of structural variation in genomes has been challenging. However, with the recent advances in high-throughput DNA sequencing and paired-end mapping (PEM) methods, the ability to identify structural variation and their respective association to human diseases has improved considerably. In this review, we describe our current knowledge of structural variation in the mouse, one of the prime model systems for studying human diseases and mammalian biology. We further present the evolutionary implications of structural variation on transposable elements. We conclude with future directions on the study of structural variation in mouse genomes that will increase our understanding of molecular architecture and functional consequences of structural variation.
Goossens, Dirk; Moens, Lotte N; Nelis, Eva; Lenaerts, An-Sofie; Glassee, Wim; Kalbe, Andreas; Frey, Bruno; Kopal, Guido; De Jonghe, Peter; De Rijk, Peter; Del-Favero, Jurgen
2009-03-01
We evaluated multiplex PCR amplification as a front-end for high-throughput sequencing, to widen the applicability of massive parallel sequencers for the detailed analysis of complex genomes. Using multiplex PCR reactions, we sequenced the complete coding regions of seven genes implicated in peripheral neuropathies in 40 individuals on a GS-FLX genome sequencer (Roche). The resulting dataset showed highly specific and uniform amplification. Comparison of the GS-FLX sequencing data with the dataset generated by Sanger sequencing confirmed the detection of all variants present and proved the sensitivity of the method for mutation detection. In addition, we showed that we could exploit the multiplexed PCR amplicons to determine individual copy number variation (CNV), increasing the spectrum of detected variations to both genetic and genomic variants. We conclude that our straightforward procedure substantially expands the applicability of the massive parallel sequencers for sequencing projects of a moderate number of amplicons (50-500) with typical applications in resequencing exons in positional or functional candidate regions and molecular genetic diagnostics. 2008 Wiley-Liss, Inc.
A global reference for human genetic variation
2016-01-01
The 1000 Genomes Project set out to provide a comprehensive description of common human genetic variation by applying whole-genome sequencing to a diverse set of individuals from multiple populations. Here we report completion of the project, having reconstructed the genomes of 2,504 individuals from 26 populations using a combination of low-coverage whole-genome sequencing, deep exome sequencing, and dense microarray genotyping. We characterized a broad spectrum of genetic variation, in total over 88 million variants (84.7 million single nucleotide polymorphisms (SNPs), 3.6 million short insertions/deletions (indels), and 60,000 structural variants), all phased onto high-quality haplotypes. This resource includes >99% of SNP variants with a frequency of >1% for a variety of ancestries. We describe the distribution of genetic variation across the global sample, and discuss the implications for common disease studies. PMID:26432245
He, Xiao-Lan; Li, Qian; Peng, Wei-Hong; Zhou, Jie; Cao, Xue-Lian; Wang, Di; Huang, Zhong-Qian; Tan, Wei; Li, Yu; Gan, Bing-Cheng
2017-06-26
The internal transcribed spacer (ITS), RNA polymerase II second largest subunit (RPB2), and elongation factor 1-alpha (EF1α) are often used in fungal taxonomy and phylogenetic analysis. As we know, an ideal molecular marker used in molecular identification and phylogenetic studies is homogeneous within species, and interspecific variation exceeds intraspecific variation. However, during our process of performing ITS, RPB2, and EF1α sequencing on the Pleurotus spp., we found that intra-isolate sequence polymorphism might be present in these genes because direct sequencing of PCR products failed in some isolates. Therefore, we detected intra- and inter-isolate variation of the three genes in Pleurotus by polymerase chain reaction amplification and cloning in this study. Results showed that intra-isolate variation of ITS was not uncommon but the polymorphic level in each isolate was relatively low in Pleurotus; intra-isolate variations of EF1α and RPB2 sequences were present in an unexpectedly high amount. The polymorphism level differed significantly between ITS, RPB2, and EF1α in the same individual, and the intra-isolate heterogeneity level of each gene varied between isolates within the same species. Intra-isolate and intraspecific variation of ITS in the tested isolates was less than interspecific variation, and intra-isolate and intraspecific variation of RPB2 was probably equal with interspecific divergence. Meanwhile, intra-isolate and intraspecific variation of EF1α could exceed interspecific divergence. These findings suggested that RPB2 and EF1α are not desirable barcoding candidates for Pleurotus. We also discussed the reason why rDNA and protein-coding genes showed variants within a single isolate in Pleurotus, but must be addressed in further research. Our study demonstrated that intra-isolate variation of ribosomal and protein-coding genes are likely widespread in fungi. This has implications for studies on fungal evolution, taxonomy, phylogenetics, and population genetics. More extensive sampling of these genes and other candidates will be required to ensure reliability as phylogenetic markers and DNA barcodes.
Clan Genomics and the Complex Architecture of Human Disease
Belmont, John W.; Boerwinkle, Eric
2013-01-01
Human diseases are caused by alleles that encompass the full range of variant types, from single-nucleotide changes to copy-number variants, and these variations span a broad frequency spectrum, from the very rare to the common. The picture emerging from analysis of whole-genome sequences, the 1000 Genomes Project pilot studies, and targeted genomic sequencing derived from very large sample sizes reveals an abundance of rare and private variants. One implication of this realization is that recent mutation may have a greater influence on disease susceptibility or protection than is conferred by variations that arose in distant ancestors. PMID:21962505
USDA-ARS?s Scientific Manuscript database
Genotyping-by-sequencing (GBS) was performed on 257 Phytophthora infestans isolates belonging to four clonal lineages to study within-lineage diversity. The four lineages used in the study included US-8 (n=28), US-11 (n=27), US-23 (n=166), and US-24 (n=36), with isolates originating from 23 of the U...
2010-01-01
Background Accurate diagnosis is essential for prompt and appropriate treatment of malaria. While rapid diagnostic tests (RDTs) offer great potential to improve malaria diagnosis, the sensitivity of RDTs has been reported to be highly variable. One possible factor contributing to variable test performance is the diversity of parasite antigens. This is of particular concern for Plasmodium falciparum histidine-rich protein 2 (PfHRP2)-detecting RDTs since PfHRP2 has been reported to be highly variable in isolates of the Asia-Pacific region. Methods The pfhrp2 exon 2 fragment from 458 isolates of P. falciparum collected from 38 countries was amplified and sequenced. For a subset of 80 isolates, the exon 2 fragment of histidine-rich protein 3 (pfhrp3) was also amplified and sequenced. DNA sequence and statistical analysis of the variation observed in these genes was conducted. The potential impact of the pfhrp2 variation on RDT detection rates was examined by analysing the relationship between sequence characteristics of this gene and the results of the WHO product testing of malaria RDTs: Round 1 (2008), for 34 PfHRP2-detecting RDTs. Results Sequence analysis revealed extensive variations in the number and arrangement of various repeats encoded by the genes in parasite populations world-wide. However, no statistically robust correlation between gene structure and RDT detection rate for P. falciparum parasites at 200 parasites per microlitre was identified. Conclusions The results suggest that despite extreme sequence variation, diversity of PfHRP2 does not appear to be a major cause of RDT sensitivity variation. PMID:20470441
Khodakov, Dmitriy; Wang, Chunyan; Zhang, David Yu
2016-10-01
Nucleic acid sequence variations have been implicated in many diseases, and reliable detection and quantitation of DNA/RNA biomarkers can inform effective therapeutic action, enabling precision medicine. Nucleic acid analysis technologies being translated into the clinic can broadly be classified into hybridization, PCR, and sequencing, as well as their combinations. Here we review the molecular mechanisms of popular commercial assays, and their progress in translation into in vitro diagnostics. Copyright © 2016 The Authors. Published by Elsevier B.V. All rights reserved.
GENOMIC BASIS OF AGING AND LIFE HISTORY EVOLUTION IN DROSOPHILA MELANOGASTER
Remolina, Silvia C.; Chang, Peter L.; Leips, Jeff; Nuzhdin, Sergey V.; Hughes, Kimberly A.
2015-01-01
Natural diversity in aging and other life history patterns is a hallmark of organismal variation. Related species, populations, and individuals within populations show genetically based variation in life span and other aspects of age-related performance. Population differences are especially informative because these differences can be large relative to within-population variation and because they occur in organisms with otherwise similar genomes. We used experimental evolution to produce populations divergent for life span and late-age fertility and then used deep genome sequencing to detect sequence variants with nucleotide-level resolution. Several genes and genome regions showed strong signatures of selection, and the same regions were implicated in independent comparisons, suggesting that the same alleles were selected in replicate lines. Genes related to oogenesis, immunity, and protein degradation were implicated as important modifiers of late-life performance. Expression profiling and functional annotation narrowed the list of strong candidate genes to 38, most of which are novel candidates for regulating aging. Life span and early-age fecundity were negatively correlated among populations; therefore the alleles we identified also are candidate regulators of a major life-history trade-off. More generally, we argue that hitchhiking mapping can be a powerful tool for uncovering the molecular bases of quantitative genetic variation. PMID:23106705
Li, Dora A; Walker, Esther; Francki, Michael G
2015-12-01
Carotenoids (especially lutein) are known to be the pigment source for flour b* colour in bread wheat. Flour b* colour variation is controlled by a quantitative trait locus (QTL) on wheat chromosome 7AL and one gene from the carotenoid pathway, phytoene synthase, was functionally associated with the QTL on 7AL in some, but not all, wheat genotypes. A SNP marker within a sequence similar to catalase (Cat3-A1snp) derived from full-length (FL) cDNA (AK332460), however, was consistently associated with the QTL on 7AL and implicated in regulating hydrogen peroxide (H2O2) to control carotenoid accumulation affecting flour b* colour. The number of catalase genes on chromosome 7AL was investigated in this study to identify which gene may be implicated in flour b* variation and two were identified through interrogation of the draft wheat genome survey sequence consisting of five exons and a further two members having eight exons identified through comparative analysis with the single catalase gene on rice chromosome 6, PCR amplification and sequencing. It was evident that the catalase genes on chromosome 7A had duplicated and diverged during evolution relative to its counterpart on rice chromosome 6. The detection of transcripts in seeds, the co-location with Cat3-A1snp marker and maximised alignment of FL-cDNA (AK332460) with cognate genomic sequence indicated that TaCat3-A1 was the member of the catalase gene family associated with flour b* colour variation. Re-sequencing identified three alleles from three wheat varieties, TaCat3-A1a, TaCat3-A1b and TaCat3-A1c, and their predicted protein identified differences in peroxisomal targeting signal tri-peptide domain in the carboxyl terminal end providing new insights into their potential role in regulating cellular H2O2 that contribute to flour b* colour variation.
Hartl, Daniel L.
2008-01-01
Simple models of molecular evolution assume that sequences evolve by a Poisson process in which nucleotide or amino acid substitutions occur as rare independent events. In these models, the expected ratio of the variance to the mean of substitution counts equals 1, and substitution processes with a ratio greater than 1 are called overdispersed. Comparing the genomes of 10 closely related species of Drosophila, we extend earlier evidence for overdispersion in amino acid replacements as well as in four-fold synonymous substitutions. The observed deviation from the Poisson expectation can be described as a linear function of the rate at which substitutions occur on a phylogeny, which implies that deviations from the Poisson expectation arise from gene-specific temporal variation in substitution rates. Amino acid sequences show greater temporal variation in substitution rates than do four-fold synonymous sequences. Our findings provide a general phenomenological framework for understanding overdispersion in the molecular clock. Also, the presence of substantial variation in gene-specific substitution rates has broad implications for work in phylogeny reconstruction and evolutionary rate estimation. PMID:18480070
Zhu, X Q; Gasser, R B
1998-06-01
In this study, we assessed single-strand conformation polymorphism (SSCP)-based approaches for their capacity to fingerprint sequence variation in ribosomal DNA (rDNA) of ascaridoid nematodes of veterinary and/or human health significance. The second internal transcribed spacer region (ITS-2) of rDNA was utilised as the target region because it is known to provide species-specific markers for this group of parasites. ITS-2 was amplified by PCR from genomic DNA derived from individual parasites and subjected to analysis. Direct SSCP analysis of amplicons from seven taxa (Toxocara vitulorum, Toxocara cati, Toxocara canis, Toxascaris leonina, Baylisascaris procyonis, Ascaris suum and Parascaris equorum) showed that the single-strand (ss) ITS-2 patterns produced allowed their unequivocal identification to species. While no variation in SSCP patterns was detected in the ITS-2 within four species for which multiple samples were available, the method allowed the direct display of four distinct sequence types of ITS-2 among individual worms of T. cati. Comparison of SSCP/sequencing with the methods of dideoxy fingerprinting (ddF) and restriction endonuclease fingerprinting (REF) revealed that also ddF allowed the definition of the four sequence types, whereas REF displayed three of four. The findings indicate the usefulness of the SSCP-based approaches for the identification of ascaridoid nematodes to species, the direct display of sequence variation in rDNA and the detection of population variation. The ability to fingerprint microheterogeneity in ITS-2 rDNA using such approaches also has implications for studying fundamental aspects relating to mutational change in rDNA.
Fane, Anne; Sarovich, Derek S.; Price, Erin P.; Rush, Catherine M.; Govan, Brenda L.; Parker, Elizabeth; Mayo, Mark; Currie, Bart J.; Ketheesan, Natkunam
2017-01-01
Neurologic melioidosis is a serious, potentially fatal form of Burkholderia pseudomallei infection. Recently, we reported that a subset of clinical isolates of B. pseudomallei from Australia have heightened virulence and potential for dissemination to the central nervous system. In this study, we demonstrate that this subset has a B. mallei–like sequence variation of the actin-based motility gene, bimA. Compared with B. pseudomallei isolates having typical bimA alleles, isolates that contain the B. mallei–like variation demonstrate increased persistence in phagocytic cells and increased virulence with rapid systemic dissemination and replication within multiple tissues, including the brain and spinal cord, in an experimental model. These findings highlight the implications of bimA variation on disease progression of B. pseudomallei infection and have considerable clinical and public health implications with respect to the degree of neurotropic threat posed to human health. PMID:28418830
Morris, Jodie L; Fane, Anne; Sarovich, Derek S; Price, Erin P; Rush, Catherine M; Govan, Brenda L; Parker, Elizabeth; Mayo, Mark; Currie, Bart J; Ketheesan, Natkunam
2017-05-01
Neurologic melioidosis is a serious, potentially fatal form of Burkholderia pseudomallei infection. Recently, we reported that a subset of clinical isolates of B. pseudomallei from Australia have heightened virulence and potential for dissemination to the central nervous system. In this study, we demonstrate that this subset has a B. mallei-like sequence variation of the actin-based motility gene, bimA. Compared with B. pseudomallei isolates having typical bimA alleles, isolates that contain the B. mallei-like variation demonstrate increased persistence in phagocytic cells and increased virulence with rapid systemic dissemination and replication within multiple tissues, including the brain and spinal cord, in an experimental model. These findings highlight the implications of bimA variation on disease progression of B. pseudomallei infection and have considerable clinical and public health implications with respect to the degree of neurotropic threat posed to human health.
Ovine Reference Materials and Assays for Prion Genetic Testing
USDA-ARS?s Scientific Manuscript database
Background: Genetic predisposition to scrapie in sheep is associated with variation in the peptide sequence of the ovine prion protein encoded by Prnp. Codon variants implicated in scrapie susceptibility or disease progression include those at amino acid positions 112, 136, 141, 154, and 171. Nin...
Wang, Yan; Liu, Guo-Hua; Li, Jia-Yuan; Xu, Min-Jun; Ye, Yong-Gang; Zhou, Dong-Hui; Song, Hui-Qun; Lin, Rui-Qing; Zhu, Xing-Quan
2013-02-01
This study examined sequence variation in three mitochondrial DNA (mtDNA) regions, namely cytochrome c oxidase subunit 1 (cox1), NADH dehydrogenase subunit 5 (nad5) and cytochrome b (cytb), among Trichuris ovis isolates from different hosts in Guangdong Province, China. A portion of the cox1 (pcox1), nad5 (pnad5) and cytb (pcytb) genes was amplified separately from individual whipworms by PCR, and was subjected to sequencing from both directions. The size of the sequences of pcox1, pnad5 and pcytb was 618, 240 and 464 bp, respectively. Although the intra-specific sequence variations within T. ovis were 0-0.8% for pcox1, 0-0.8% for pnad5 and 0-1.9% for pcytb, the inter-specific sequence differences among members of the genus Trichuris were significantly higher, being 24.3-26.5% for pcox1, 33.7-56.4% for pnad5 and 24.8-26.1% for pcytb, respectively. Phylogenetic analyses using combined sequences of pcox1, pnad5 and pcytb, with three different computational algorithms (maximum likelihood, maximum parsimony and Bayesian inference), indicated that all of the T. ovis isolates grouped together with high statistical support. These findings demonstrated the existence of intra-specific variation in mtDNA sequences among T. ovis isolates from different hosts, and have implications for studying molecular epidemiology and population genetics of T. ovis.
Novel rare variations of the oxytocin receptor (OXTR) gene in autism spectrum disorder individuals.
Liu, Xiaoxi; Kawashima, Minae; Miyagawa, Taku; Otowa, Takeshi; Latt, Khun Zaw; Thiri, Myo; Nishida, Hisami; Sugiyama, Toshiro; Tsurusaki, Yoshinori; Matsumoto, Naomichi; Mabuchi, Akihiko; Tokunaga, Katsushi; Sasaki, Tsukasa
2015-01-01
The oxytocin receptor (OXTR) gene has been implicated as a risk gene for autism spectrum disorder (ASD)-a neurodevelopmental disorder with essential features of impairments in social communication and reciprocal interaction. The genetic associations between common variations in OXTR and ASD have been reported in multiple ethnic populations. However, little is known about the distribution of rare variations within OXTR in ASD patients. In this study, we resequenced the full length of OXTR in 105 ASD individuals using an approach that combined the power of next-generation sequencing technology, long-range PCR and DNA pooling. We demonstrated that rare variants with minor allele frequency as low as 0.05% could be reliably detected by our method. We identified 28 novel variants including potential functional variants in the intron region and one rare missense variant (R150S). We subsequently performed Sanger sequencing and validated five novel variants located in previously suggested candidate regions in ASD individuals. Further sequencing of 312 healthy subjects showed that the burden of rare variants is significantly higher in ASDs compared with healthy individuals. Our results support that the rare variation in OXTR gene might be involved in ASD.
A map of human genome variation from population-scale sequencing.
Abecasis, Gonçalo R; Altshuler, David; Auton, Adam; Brooks, Lisa D; Durbin, Richard M; Gibbs, Richard A; Hurles, Matt E; McVean, Gil A
2010-10-28
The 1000 Genomes Project aims to provide a deep characterization of human genome sequence variation as a foundation for investigating the relationship between genotype and phenotype. Here we present results of the pilot phase of the project, designed to develop and compare different strategies for genome-wide sequencing with high-throughput platforms. We undertook three projects: low-coverage whole-genome sequencing of 179 individuals from four populations; high-coverage sequencing of two mother-father-child trios; and exon-targeted sequencing of 697 individuals from seven populations. We describe the location, allele frequency and local haplotype structure of approximately 15 million single nucleotide polymorphisms, 1 million short insertions and deletions, and 20,000 structural variants, most of which were previously undescribed. We show that, because we have catalogued the vast majority of common variation, over 95% of the currently accessible variants found in any individual are present in this data set. On average, each person is found to carry approximately 250 to 300 loss-of-function variants in annotated genes and 50 to 100 variants previously implicated in inherited disorders. We demonstrate how these results can be used to inform association and functional studies. From the two trios, we directly estimate the rate of de novo germline base substitution mutations to be approximately 10(-8) per base pair per generation. We explore the data with regard to signatures of natural selection, and identify a marked reduction of genetic variation in the neighbourhood of genes, due to selection at linked sites. These methods and public data will support the next phase of human genetic research.
Reuter, Miriam S.; Walker, Susan; Thiruvahindrapuram, Bhooma; Whitney, Joe; Cohn, Iris; Sondheimer, Neal; Yuen, Ryan K.C.; Trost, Brett; Paton, Tara A.; Pereira, Sergio L.; Herbrick, Jo-Anne; Wintle, Richard F.; Merico, Daniele; Howe, Jennifer; MacDonald, Jeffrey R.; Lu, Chao; Nalpathamkalam, Thomas; Sung, Wilson W.L.; Wang, Zhuozhi; Patel, Rohan V.; Pellecchia, Giovanna; Wei, John; Strug, Lisa J.; Bell, Sherilyn; Kellam, Barbara; Mahtani, Melanie M.; Bassett, Anne S.; Bombard, Yvonne; Weksberg, Rosanna; Shuman, Cheryl; Cohn, Ronald D.; Stavropoulos, Dimitri J.; Bowdin, Sarah; Hildebrandt, Matthew R.; Wei, Wei; Romm, Asli; Pasceri, Peter; Ellis, James; Ray, Peter; Meyn, M. Stephen; Monfared, Nasim; Hosseini, S. Mohsen; Joseph-George, Ann M.; Keeley, Fred W.; Cook, Ryan A.; Fiume, Marc; Lee, Hin C.; Marshall, Christian R.; Davies, Jill; Hazell, Allison; Buchanan, Janet A.; Szego, Michael J.; Scherer, Stephen W.
2018-01-01
BACKGROUND: The Personal Genome Project Canada is a comprehensive public data resource that integrates whole genome sequencing data and health information. We describe genomic variation identified in the initial recruitment cohort of 56 volunteers. METHODS: Volunteers were screened for eligibility and provided informed consent for open data sharing. Using blood DNA, we performed whole genome sequencing and identified all possible classes of DNA variants. A genetic counsellor explained the implication of the results to each participant. RESULTS: Whole genome sequencing of the first 56 participants identified 207 662 805 sequence variants and 27 494 copy number variations. We analyzed a prioritized disease-associated data set (n = 1606 variants) according to standardized guidelines, and interpreted 19 variants in 14 participants (25%) as having obvious health implications. Six of these variants (e.g., in BRCA1 or mosaic loss of an X chromosome) were pathogenic or likely pathogenic. Seven were risk factors for cancer, cardiovascular or neurobehavioural conditions. Four other variants — associated with cancer, cardiac or neurodegenerative phenotypes — remained of uncertain significance because of discrepancies among databases. We also identified a large structural chromosome aberration and a likely pathogenic mitochondrial variant. There were 172 recessive disease alleles (e.g., 5 individuals carried mutations for cystic fibrosis). Pharmacogenomics analyses revealed another 3.9 potentially relevant genotypes per individual. INTERPRETATION: Our analyses identified a spectrum of genetic variants with potential health impact in 25% of participants. When also considering recessive alleles and variants with potential pharmacologic relevance, all 56 participants had medically relevant findings. Although access is mostly limited to research, whole genome sequencing can provide specific and novel information with the potential of major impact for health care. PMID:29431110
Reuter, Miriam S; Walker, Susan; Thiruvahindrapuram, Bhooma; Whitney, Joe; Cohn, Iris; Sondheimer, Neal; Yuen, Ryan K C; Trost, Brett; Paton, Tara A; Pereira, Sergio L; Herbrick, Jo-Anne; Wintle, Richard F; Merico, Daniele; Howe, Jennifer; MacDonald, Jeffrey R; Lu, Chao; Nalpathamkalam, Thomas; Sung, Wilson W L; Wang, Zhuozhi; Patel, Rohan V; Pellecchia, Giovanna; Wei, John; Strug, Lisa J; Bell, Sherilyn; Kellam, Barbara; Mahtani, Melanie M; Bassett, Anne S; Bombard, Yvonne; Weksberg, Rosanna; Shuman, Cheryl; Cohn, Ronald D; Stavropoulos, Dimitri J; Bowdin, Sarah; Hildebrandt, Matthew R; Wei, Wei; Romm, Asli; Pasceri, Peter; Ellis, James; Ray, Peter; Meyn, M Stephen; Monfared, Nasim; Hosseini, S Mohsen; Joseph-George, Ann M; Keeley, Fred W; Cook, Ryan A; Fiume, Marc; Lee, Hin C; Marshall, Christian R; Davies, Jill; Hazell, Allison; Buchanan, Janet A; Szego, Michael J; Scherer, Stephen W
2018-02-05
The Personal Genome Project Canada is a comprehensive public data resource that integrates whole genome sequencing data and health information. We describe genomic variation identified in the initial recruitment cohort of 56 volunteers. Volunteers were screened for eligibility and provided informed consent for open data sharing. Using blood DNA, we performed whole genome sequencing and identified all possible classes of DNA variants. A genetic counsellor explained the implication of the results to each participant. Whole genome sequencing of the first 56 participants identified 207 662 805 sequence variants and 27 494 copy number variations. We analyzed a prioritized disease-associated data set ( n = 1606 variants) according to standardized guidelines, and interpreted 19 variants in 14 participants (25%) as having obvious health implications. Six of these variants (e.g., in BRCA1 or mosaic loss of an X chromosome) were pathogenic or likely pathogenic. Seven were risk factors for cancer, cardiovascular or neurobehavioural conditions. Four other variants - associated with cancer, cardiac or neurodegenerative phenotypes - remained of uncertain significance because of discrepancies among databases. We also identified a large structural chromosome aberration and a likely pathogenic mitochondrial variant. There were 172 recessive disease alleles (e.g., 5 individuals carried mutations for cystic fibrosis). Pharmacogenomics analyses revealed another 3.9 potentially relevant genotypes per individual. Our analyses identified a spectrum of genetic variants with potential health impact in 25% of participants. When also considering recessive alleles and variants with potential pharmacologic relevance, all 56 participants had medically relevant findings. Although access is mostly limited to research, whole genome sequencing can provide specific and novel information with the potential of major impact for health care. © 2018 Joule Inc. or its licensors.
Ekanayake, Saliya; Ruan, Yang; Schütte, Ursel M. E.; Kaonongbua, Wittaya; Fox, Geoffrey; Ye, Yuzhen; Bever, James D.
2016-01-01
ABSTRACT Arbuscular mycorrhizal (AM) fungi form mutualisms with plant roots that increase plant growth and shape plant communities. Each AM fungal cell contains a large amount of genetic diversity, but it is unclear if this diversity varies across evolutionary lineages. We found that sequence variation in the nuclear large-subunit (LSU) rRNA gene from 29 isolates representing 21 AM fungal species generally assorted into genus- and species-level clades, with the exception of species of the genera Claroideoglomus and Entrophospora. However, there were significant differences in the levels of sequence variation across the phylogeny and between genera, indicating that it is an evolutionarily constrained trait in AM fungi. These consistent patterns of sequence variation across both phylogenetic and taxonomic groups pose challenges to interpreting operational taxonomic units (OTUs) as approximations of species-level groups of AM fungi. We demonstrate that the OTUs produced by five sequence clustering methods using 97% or equivalent sequence similarity thresholds failed to match the expected species of AM fungi, although OTUs from AbundantOTU, CD-HIT-OTU, and CROP corresponded better to species than did OTUs from mothur or UPARSE. This lack of OTU-to-species correspondence resulted both from sequences of one species being split into multiple OTUs and from sequences of multiple species being lumped into the same OTU. The OTU richness therefore will not reliably correspond to the AM fungal species richness in environmental samples. Conservatively, this error can overestimate species richness by 4-fold or underestimate richness by one-half, and the direction of this error will depend on the genera represented in the sample. IMPORTANCE Arbuscular mycorrhizal (AM) fungi form important mutualisms with the roots of most plant species. Individual AM fungi are genetically diverse, but it is unclear whether the level of this diversity differs among evolutionary lineages. We found that the amount of sequence variation in an rRNA gene that is commonly used to identify AM fungal species varied significantly between evolutionary groups that correspond to different genera, with the exception of two genera that are genetically indistinguishable from each other. When we clustered groups of similar sequences into operational taxonomic units (OTUs) using five different clustering methods, these patterns of sequence variation caused the number of OTUs to either over- or underestimate the actual number of AM fungal species, depending on the genus. Our results indicate that OTU-based inferences about AM fungal species composition from environmental sequences can be improved if they take these taxonomically structured patterns of sequence variation into account. PMID:27260357
USDA-ARS?s Scientific Manuscript database
Porcine reproductive and respiratory syndrome virus (PRRSV) is widespread with a high variation in sequence and virulence among the divergent strains and causes an economically destructive disease. A viral ovarian domain protease (vOTU) has been previously identified within the nonstructural protein...
An Exome Sequencing Study to Assess the Role of Rare Genetic Variation in Pulmonary Fibrosis.
Petrovski, Slavé; Todd, Jamie L; Durheim, Michael T; Wang, Quanli; Chien, Jason W; Kelly, Fran L; Frankel, Courtney; Mebane, Caroline M; Ren, Zhong; Bridgers, Joshua; Urban, Thomas J; Malone, Colin D; Finlen Copeland, Ashley; Brinkley, Christie; Allen, Andrew S; O'Riordan, Thomas; McHutchison, John G; Palmer, Scott M; Goldstein, David B
2017-07-01
Idiopathic pulmonary fibrosis (IPF) is an increasingly recognized, often fatal lung disease of unknown etiology. The aim of this study was to use whole-exome sequencing to improve understanding of the genetic architecture of pulmonary fibrosis. We performed a case-control exome-wide collapsing analysis including 262 unrelated individuals with pulmonary fibrosis clinically classified as IPF according to American Thoracic Society/European Respiratory Society/Japanese Respiratory Society/Latin American Thoracic Association guidelines (81.3%), usual interstitial pneumonia secondary to autoimmune conditions (11.5%), or fibrosing nonspecific interstitial pneumonia (7.2%). The majority (87%) of case subjects reported no family history of pulmonary fibrosis. We searched 18,668 protein-coding genes for an excess of rare deleterious genetic variation using whole-exome sequence data from 262 case subjects with pulmonary fibrosis and 4,141 control subjects drawn from among a set of individuals of European ancestry. Comparing genetic variation across 18,668 protein-coding genes, we found a study-wide significant (P < 4.5 × 10 -7 ) case enrichment of qualifying variants in TERT, RTEL1, and PARN. A model qualifying ultrarare, deleterious, nonsynonymous variants implicated TERT and RTEL1, and a model specifically qualifying loss-of-function variants implicated RTEL1 and PARN. A subanalysis of 186 case subjects with sporadic IPF confirmed TERT, RTEL1, and PARN as study-wide significant contributors to sporadic IPF. Collectively, 11.3% of case subjects with sporadic IPF carried a qualifying variant in one of these three genes compared with the 0.3% carrier rate observed among control subjects (odds ratio, 47.7; 95% confidence interval, 21.5-111.6; P = 5.5 × 10 -22 ). We identified TERT, RTEL1, and PARN-three telomere-related genes previously implicated in familial pulmonary fibrosis-as significant contributors to sporadic IPF. These results support the idea that telomere dysfunction is involved in IPF pathogenesis.
The population genomics of rhesus macaques (Macaca mulatta) based on whole-genome sequences
Xue, Cheng; Raveendran, Muthuswamy; Harris, R. Alan; Fawcett, Gloria L.; Liu, Xiaoming; White, Simon; Dahdouli, Mahmoud; Rio Deiros, David; Below, Jennifer E.; Salerno, William; Cox, Laura; Fan, Guoping; Ferguson, Betsy; Horvath, Julie; Johnson, Zach; Kanthaswamy, Sree; Kubisch, H. Michael; Liu, Dahai; Platt, Michael; Smith, David G.; Sun, Binghua; Vallender, Eric J.; Wang, Feng; Wiseman, Roger W.; Chen, Rui; Muzny, Donna M.; Gibbs, Richard A.; Yu, Fuli; Rogers, Jeffrey
2016-01-01
Rhesus macaques (Macaca mulatta) are the most widely used nonhuman primate in biomedical research, have the largest natural geographic distribution of any nonhuman primate, and have been the focus of much evolutionary and behavioral investigation. Consequently, rhesus macaques are one of the most thoroughly studied nonhuman primate species. However, little is known about genome-wide genetic variation in this species. A detailed understanding of extant genomic variation among rhesus macaques has implications for the use of this species as a model for studies of human health and disease, as well as for evolutionary population genomics. Whole-genome sequencing analysis of 133 rhesus macaques revealed more than 43.7 million single-nucleotide variants, including thousands predicted to alter protein sequences, transcript splicing, and transcription factor binding sites. Rhesus macaques exhibit 2.5-fold higher overall nucleotide diversity and slightly elevated putative functional variation compared with humans. This functional variation in macaques provides opportunities for analyses of coding and noncoding variation, and its cellular consequences. Despite modestly higher levels of nonsynonymous variation in the macaques, the estimated distribution of fitness effects and the ratio of nonsynonymous to synonymous variants suggest that purifying selection has had stronger effects in rhesus macaques than in humans. Demographic reconstructions indicate this species has experienced a consistently large but fluctuating population size. Overall, the results presented here provide new insights into the population genomics of nonhuman primates and expand genomic information directly relevant to primate models of human disease. PMID:27934697
ERIC Educational Resources Information Center
Bornkessel-Schlesewsky, Ina; Grewe, Tanja; Schlesewsky, Matthias
2012-01-01
Prior research on the neural bases of syntactic comprehension suggests that activation in the left inferior frontal gyrus (lIFG) correlates with the processing of word order variations. However, there are inconsistencies with respect to the specific subregion within the IFG that is implicated by these findings: the pars opercularis or the pars…
The cancer transcriptome is shaped by genetic changes, variation in gene transcription, mRNA processing, editing and stability, and the cancer microbiome. Deciphering this variation and understanding its implications on tumorigenesis requires sophisticated computational analyses. Most RNA-Seq analyses rely on methods that first map short reads to a reference genome, and then compare them to annotated transcripts or assemble them. However, this strategy can be limited when the cancer genome is substantially different than the reference or for detecting sequences from the cancer microbiome.
Abo, Ryan P; Ducar, Matthew; Garcia, Elizabeth P; Thorner, Aaron R; Rojas-Rudilla, Vanesa; Lin, Ling; Sholl, Lynette M; Hahn, William C; Meyerson, Matthew; Lindeman, Neal I; Van Hummelen, Paul; MacConaill, Laura E
2015-02-18
Genomic structural variation (SV), a common hallmark of cancer, has important predictive and therapeutic implications. However, accurately detecting SV using high-throughput sequencing data remains challenging, especially for 'targeted' resequencing efforts. This is critically important in the clinical setting where targeted resequencing is frequently being applied to rapidly assess clinically actionable mutations in tumor biopsies in a cost-effective manner. We present BreaKmer, a novel approach that uses a 'kmer' strategy to assemble misaligned sequence reads for predicting insertions, deletions, inversions, tandem duplications and translocations at base-pair resolution in targeted resequencing data. Variants are predicted by realigning an assembled consensus sequence created from sequence reads that were abnormally aligned to the reference genome. Using targeted resequencing data from tumor specimens with orthogonally validated SV, non-tumor samples and whole-genome sequencing data, BreaKmer had a 97.4% overall sensitivity for known events and predicted 17 positively validated, novel variants. Relative to four publically available algorithms, BreaKmer detected SV with increased sensitivity and limited calls in non-tumor samples, key features for variant analysis of tumor specimens in both the clinical and research settings. © The Author(s) 2014. Published by Oxford University Press on behalf of Nucleic Acids Research.
Berg, Ingrid L; Neumann, Rita; Lam, Kwan-Wood G; Sarbajna, Shriparna; Odenthal-Hesse, Linda; May, Celia A; Jeffreys, Alec J
2010-10-01
PRDM9 has recently been identified as a likely trans regulator of meiotic recombination hot spots in humans and mice. PRDM9 contains a zinc finger array that, in humans, can recognize a short sequence motif associated with hot spots, with binding to this motif possibly triggering hot-spot activity via chromatin remodeling. We now report that human genetic variation at the PRDM9 locus has a strong effect on sperm hot-spot activity, even at hot spots lacking the sequence motif. Subtle changes within the zinc finger array can create hot-spot nonactivating or enhancing variants and can even trigger the appearance of a new hot spot, suggesting that PRDM9 is a major global regulator of hot spots in humans. Variation at the PRDM9 locus also influences aspects of genome instability-specifically, a megabase-scale rearrangement underlying two genomic disorders as well as minisatellite instability-implicating PRDM9 as a risk factor for some pathological genome rearrangements.
Berg, Ingrid L.; Neumann, Rita; Lam, Kwan-Wood G.; Sarbajna, Shriparna; Odenthal-Hesse, Linda; May, Celia A.; Jeffreys, Alec J.
2011-01-01
PRDM9 has recently been identified as a likely trans-regulator of meiotic recombination hot spots in humans and mice1-3. The protein contains a zinc finger array that in humans can recognise a short sequence motif associated with hot spots4, with binding to this motif possibly triggering hot-spot activity via chromatin remodelling5. We now show that variation in the zinc finger array in humans has a profound effect on sperm hot-spot activity, even at hot spots lacking the sequence motif. Very subtle changes within the array can create hot-spot non-activating and enhancing alleles, and even trigger the appearance of a new hot spot. PRDM9 thus appears to be the preeminent global regulator of hot spots in humans. Variation at this locus also influences aspects of genome instability, specifically a megabase-scale rearrangement underlying two genomic disorders6 as well as minisatellite instability7, implicating PRDM9 as a risk factor for some pathological genome rearrangements. PMID:20818382
Werling, Donna M; Brand, Harrison; An, Joon-Yong; Stone, Matthew R; Zhu, Lingxue; Glessner, Joseph T; Collins, Ryan L; Dong, Shan; Layer, Ryan M; Markenscoff-Papadimitriou, Eirene; Farrell, Andrew; Schwartz, Grace B; Wang, Harold Z; Currall, Benjamin B; Zhao, Xuefang; Dea, Jeanselle; Duhn, Clif; Erdman, Carolyn A; Gilson, Michael C; Yadav, Rachita; Handsaker, Robert E; Kashin, Seva; Klei, Lambertus; Mandell, Jeffrey D; Nowakowski, Tomasz J; Liu, Yuwen; Pochareddy, Sirisha; Smith, Louw; Walker, Michael F; Waterman, Matthew J; He, Xin; Kriegstein, Arnold R; Rubenstein, John L; Sestan, Nenad; McCarroll, Steven A; Neale, Benjamin M; Coon, Hilary; Willsey, A Jeremy; Buxbaum, Joseph D; Daly, Mark J; State, Matthew W; Quinlan, Aaron R; Marth, Gabor T; Roeder, Kathryn; Devlin, Bernie; Talkowski, Michael E; Sanders, Stephan J
2018-05-01
Genomic association studies of common or rare protein-coding variation have established robust statistical approaches to account for multiple testing. Here we present a comparable framework to evaluate rare and de novo noncoding single-nucleotide variants, insertion/deletions, and all classes of structural variation from whole-genome sequencing (WGS). Integrating genomic annotations at the level of nucleotides, genes, and regulatory regions, we define 51,801 annotation categories. Analyses of 519 autism spectrum disorder families did not identify association with any categories after correction for 4,123 effective tests. Without appropriate correction, biologically plausible associations are observed in both cases and controls. Despite excluding previously identified gene-disrupting mutations, coding regions still exhibited the strongest associations. Thus, in autism, the contribution of de novo noncoding variation is probably modest in comparison to that of de novo coding variants. Robust results from future WGS studies will require large cohorts and comprehensive analytical strategies that consider the substantial multiple-testing burden.
Reed, K M; Dorschner, M O; Todd, T N; Phillips, R B
1998-09-01
Sequence variation in the control region (D-loop) of the mitochondrial DNA (mtDNA) was examined to assess the genetic distinctiveness of the shortjaw cisco (Coregonus zenithicus). Individuals from within the Great Lakes Basin as well as inland lakes outside the basin were sampled. DNA fragments containing the entire D-loop were amplified by PCR from specimens of C. zenithicus and the related species C. artedi, C. hoyi, C. kiyi, and C. clupeaformis. DNA sequence analysis revealed high similarity within and among species and shared polymorphism for length variants. Based on this analysis, the shortjaw cisco is not genetically distinct from other cisco species.
USDA-ARS?s Scientific Manuscript database
Maruca vitrata is a polyphagous insect pest on a wide variety of leguminous plants in the tropics and subtropics. The contribution of host-associated genetic variation on population structure was investigated using analysis mitochondrial cox1 sequence and microsatellite marker data from M. vitrata c...
USDA-ARS?s Scientific Manuscript database
The recently cloned blast resistance (R) gene Pi-km protects rice crops against specific races of the fungal pathogen Magnaporthe oryzae in a gene-for-gene manner. The use of blast R genes remains the most cost-effective method for an integrated disease management strategy. To facilitate rice breed...
The Contribution of Mosaic Variants to Autism Spectrum Disorder.
Freed, Donald; Pevsner, Jonathan
2016-09-01
De novo mutation is highly implicated in autism spectrum disorder (ASD). However, the contribution of post-zygotic mutation to ASD is poorly characterized. We performed both exome sequencing of paired samples and analysis of de novo variants from whole-exome sequencing of 2,388 families. While we find little evidence for tissue-specific mosaic mutation, multi-tissue post-zygotic mutation (i.e. mosaicism) is frequent, with detectable mosaic variation comprising 5.4% of all de novo mutations. We identify three mosaic missense and likely-gene disrupting mutations in genes previously implicated in ASD (KMT2C, NCKAP1, and MYH10) in probands but none in siblings. We find a strong ascertainment bias for mosaic mutations in probands relative to their unaffected siblings (p = 0.003). We build a model of de novo variation incorporating mosaic variants and errors in classification of mosaic status and from this model we estimate that 33% of mosaic mutations in probands contribute to 5.1% of simplex ASD diagnoses (95% credible interval 1.3% to 8.9%). Our results indicate a contributory role for multi-tissue mosaic mutation in some individuals with an ASD diagnosis.
Fast T2*-weighted MRI of the prostate at 3 Tesla.
Hardman, Rulon L; El-Merhi, Fadi; Jung, Adam J; Ware, Steve; Thompson, Ian M; Friel, Harry T; Peng, Qi
2011-04-01
To describe a rapid T2*-weighted (T2*W), three-dimensional (3D) echo planar imaging (EPI) sequence and its application in mapping local magnetic susceptibility variations in 3 Tesla (T) prostate MRI. To compare the sensitivity of T2*W EPI with routinely used T1-weighted turbo-spin echo sequence (T1W TSE) in detecting hemorrhage and the implications on sequences sensitive to field inhomogeneities such as MR spectroscopy (MRS). B(0) susceptibility weighted mapping was performed using a 3D EPI sequence featuring a 2D spatial excitation pulse with gradients of spiral k-space trajectory. A series of 11 subjects were imaged using 3T MRI and combination endorectal (ER) and six-channel phased array cardiac coils. T1W TSE and T2*W EPI sequences were analyzed quantitatively for hemorrhage contrast. Point resolved spectroscopy (PRESS MRS) was performed and data quality was analyzed. Two types of susceptibility variation were identified: hemorrhagic and nonhemorrhagic T2*W-positive areas. Post-biopsy hemorrhage lesions showed on average five times greater contrast on the T2*W images than T1W TSE images. Six nonhemorrhage regions of severe susceptibility artifact were apparent on the T2*W images that were not seen on standard T1W or T2W images. All nonhemorrhagic susceptibility artifact regions demonstrated compromised spectral quality on 3D MRS. The fast T2*W EPI sequence identifies hemorrhagic and nonhemorrhagic areas of susceptibility variation that may be helpful in prostate MRI planning at 3.0T. Copyright © 2011 Wiley-Liss, Inc.
Rebelling for a Reason: Protein Structural “Outliers”
Arumugam, Gandhimathi; Nair, Anu G.; Hariharaputran, Sridhar; Ramanathan, Sowdhamini
2013-01-01
Analysis of structural variation in domain superfamilies can reveal constraints in protein evolution which aids protein structure prediction and classification. Structure-based sequence alignment of distantly related proteins, organized in PASS2 database, provides clues about structurally conserved regions among different functional families. Some superfamily members show large structural differences which are functionally relevant. This paper analyses the impact of structural divergence on function for multi-member superfamilies, selected from the PASS2 superfamily alignment database. Functional annotations within superfamilies, with structural outliers or ‘rebels’, are discussed in the context of structural variations. Overall, these data reinforce the idea that functional similarities cannot be extrapolated from mere structural conservation. The implication for fold-function prediction is that the functional annotations can only be inherited with very careful consideration, especially at low sequence identities. PMID:24073209
Dumas, Laura; Dickens, C Michael; Anderson, Nathan; Davis, Jonathan; Bennett, Beth; Radcliffe, Richard A; Sikela, James M
2014-06-01
It has been well documented that genetic factors can influence predisposition to develop alcoholism. While the underlying genomic changes may be of several types, two of the most common and disease associated are copy number variations (CNVs) and sequence alterations of protein coding regions. The goal of this study was to identify CNVs and single-nucleotide polymorphisms that occur in gene coding regions that may play a role in influencing the risk of an individual developing alcoholism. Toward this end, two mouse strains were used that have been selectively bred based on their differential sensitivity to alcohol: the Inbred long sleep (ILS) and Inbred short sleep (ISS) mouse strains. Differences in initial response to alcohol have been linked to risk for alcoholism, and the ILS/ISS strains are used to investigate the genetics of initial sensitivity to alcohol. Array comparative genomic hybridization (arrayCGH) and exome sequencing were conducted to identify CNVs and gene coding sequence differences, respectively, between ILS and ISS mice. Mouse arrayCGH was performed using catalog Agilent 1 × 244 k mouse arrays. Subsequently, exome sequencing was carried out using an Illumina HiSeq 2000 instrument. ArrayCGH detected 74 CNVs that were strain-specific (38 ILS/36 ISS), including several ISS-specific deletions that contained genes implicated in brain function and neurotransmitter release. Among several interesting coding variations detected by exome sequencing was the gain of a premature stop codon in the alpha-amylase 2B (AMY2B) gene specifically in the ILS strain. In total, exome sequencing detected 2,597 and 1,768 strain-specific exonic gene variants in the ILS and ISS mice, respectively. This study represents the most comprehensive and detailed genomic comparison of ILS and ISS mouse strains to date. The two complementary genome-wide approaches identified strain-specific CNVs and gene coding sequence variations that should provide strong candidates to contribute to the alcohol-related phenotypic differences associated with these strains.
Nesbitt, T Clint; Tanksley, Steven D
2002-01-01
Sequence variation was sampled in cultivated and related wild forms of tomato at fw2.2--a fruit weight QTL key to the evolution of domesticated tomatoes. Variation at fw2.2 was contrasted with variation at four other loci not involved in fruit weight determination. Several conclusions could be reached: (1) Fruit weight variation attributable to fw2.2 is not caused by variation in the FW2.2 protein sequence; more likely, it is due to transcriptional variation associated with one or more of eight nucleotide changes unique to the promoter of large-fruit alleles; (2) fw2.2 and loci not involved in fruit weight have not evolved at distinguishably different rates in cultivated and wild tomatoes, despite the fact that fw2.2 was likely a target of selection during domestication; (3) molecular-clock-based estimates suggest that the large-fruit allele of fw2.2, now fixed in most cultivated tomatoes, arose in tomato germplasm long before domestication; (4) extant accessions of L. esculentum var. cerasiforme, the subspecies thought to be the most likely wild ancestor of domesticated tomatoes, appear to be an admixture of wild and cultivated tomatoes rather than a transitional step from wild to domesticated tomatoes; and (5) despite the fact that cerasiforme accessions are polymorphic for large- and small-fruit alleles at fw2.2, no significant association was detected between fruit size and fw2.2 genotypes in the subspecies--as tested by association genetic studies in the relatively small sample studied--suggesting the role of other fruit weight QTL in fruit weight variation in cerasiforme. PMID:12242247
1978-01-01
Three immunologically cross-reactive and non-cross-reactive streptococcal M proteins were analyzed by a chromatographic tryptic peptide mapping system. The results indicate that cross-reactions correlate with the extent of structural similarity among the M protein molecules analyzed. The data also reveal that free lysine is released by the action of trypsin from these three M proteins, suggesting a common lys-lys or arg-lys sequence. In addition, only one peptide has been found to be common within all three M types. This limited structural relatedness among the three M proteins examined indicates that sequence variation plays a major role in the immunological specificity of the M antigens. However, despite sequence variation, all M protein molecules have a common antiphagocytic activity. The fact that no common opsonic antibody has yet been found, even against limited M types, argues against this biological activity being solely the result of a common sequence. Based on these data, it is suggested that the antiphagocytic effect of M protein may be due to a conformationally created environment on the surface of the molecule which is selected by both immunological and biological pressure. PMID:355596
Exome Sequencing Identifies Three Novel Candidate Genes Implicated in Intellectual Disability
Azam, Maleeha; Ayub, Humaira; Vissers, Lisenka E. L. M.; Gilissen, Christian; Ali, Syeda Hafiza Benish; Riaz, Moeen; Veltman, Joris A.; Pfundt, Rolph; van Bokhoven, Hans; Qamar, Raheel
2014-01-01
Intellectual disability (ID) is a major health problem mostly with an unknown etiology. Recently exome sequencing of individuals with ID identified novel genes implicated in the disease. Therefore the purpose of the present study was to identify the genetic cause of ID in one syndromic and two non-syndromic Pakistani families. Whole exome of three ID probands was sequenced. Missense variations in two plausible novel genes implicated in autosomal recessive ID were identified: lysine (K)-specific methyltransferase 2B (KMT2B), zinc finger protein 589 (ZNF589), as well as hedgehog acyltransferase (HHAT) with a de novo mutation with autosomal dominant mode of inheritance. The KMT2B recessive variant is the first report of recessive Kleefstra syndrome-like phenotype. Identification of plausible causative mutations for two recessive and a dominant type of ID, in genes not previously implicated in disease, underscores the large genetic heterogeneity of ID. These results also support the viewpoint that large number of ID genes converge on limited number of common networks i.e. ZNF589 belongs to KRAB-domain zinc-finger proteins previously implicated in ID, HHAT is predicted to affect sonic hedgehog, which is involved in several disorders with ID, KMT2B associated with syndromic ID fits the epigenetic module underlying the Kleefstra syndromic spectrum. The association of these novel genes in three different Pakistani ID families highlights the importance of screening these genes in more families with similar phenotypes from different populations to confirm the involvement of these genes in pathogenesis of ID. PMID:25405613
Simmons, Sheri L; Dibartolo, Genevieve; Denef, Vincent J; Goltsman, Daniela S Aliaga; Thelen, Michael P; Banfield, Jillian F
2008-07-22
Deeply sampled community genomic (metagenomic) datasets enable comprehensive analysis of heterogeneity in natural microbial populations. In this study, we used sequence data obtained from the dominant member of a low-diversity natural chemoautotrophic microbial community to determine how coexisting closely related individuals differ from each other in terms of gene sequence and gene content, and to uncover evidence of evolutionary processes that occur over short timescales. DNA sequence obtained from an acid mine drainage biofilm was reconstructed, taking into account the effects of strain variation, to generate a nearly complete genome tiling path for a Leptospirillum group II species closely related to L. ferriphilum (sampling depth approximately 20x). The population is dominated by one sequence type, yet we detected evidence for relatively abundant variants (>99.5% sequence identity to the dominant type) at multiple loci, and a few rare variants. Blocks of other Leptospirillum group II types ( approximately 94% sequence identity) have recombined into one or more variants. Variant blocks of both types are more numerous near the origin of replication. Heterogeneity in genetic potential within the population arises from localized variation in gene content, typically focused in integrated plasmid/phage-like regions. Some laterally transferred gene blocks encode physiologically important genes, including quorum-sensing genes of the LuxIR system. Overall, results suggest inter- and intrapopulation genetic exchange involving distinct parental genome types and implicate gain and loss of phage and plasmid genes in recent evolution of this Leptospirillum group II population. Population genetic analyses of single nucleotide polymorphisms indicate variation between closely related strains is not maintained by positive selection, suggesting that these regions do not represent adaptive differences between strains. Thus, the most likely explanation for the observed patterns of polymorphism is divergence of ancestral strains due to geographic isolation, followed by mixing and subsequent recombination.
Denef, Vincent J; Goltsman, Daniela S. Aliaga; Thelen, Michael P; Banfield, Jillian F
2008-01-01
Deeply sampled community genomic (metagenomic) datasets enable comprehensive analysis of heterogeneity in natural microbial populations. In this study, we used sequence data obtained from the dominant member of a low-diversity natural chemoautotrophic microbial community to determine how coexisting closely related individuals differ from each other in terms of gene sequence and gene content, and to uncover evidence of evolutionary processes that occur over short timescales. DNA sequence obtained from an acid mine drainage biofilm was reconstructed, taking into account the effects of strain variation, to generate a nearly complete genome tiling path for a Leptospirillum group II species closely related to L. ferriphilum (sampling depth ∼20×). The population is dominated by one sequence type, yet we detected evidence for relatively abundant variants (>99.5% sequence identity to the dominant type) at multiple loci, and a few rare variants. Blocks of other Leptospirillum group II types (∼94% sequence identity) have recombined into one or more variants. Variant blocks of both types are more numerous near the origin of replication. Heterogeneity in genetic potential within the population arises from localized variation in gene content, typically focused in integrated plasmid/phage-like regions. Some laterally transferred gene blocks encode physiologically important genes, including quorum-sensing genes of the LuxIR system. Overall, results suggest inter- and intrapopulation genetic exchange involving distinct parental genome types and implicate gain and loss of phage and plasmid genes in recent evolution of this Leptospirillum group II population. Population genetic analyses of single nucleotide polymorphisms indicate variation between closely related strains is not maintained by positive selection, suggesting that these regions do not represent adaptive differences between strains. Thus, the most likely explanation for the observed patterns of polymorphism is divergence of ancestral strains due to geographic isolation, followed by mixing and subsequent recombination. PMID:18651792
Reed, Kent M.; Dorschner, Michael O.; Todd, Thomas N.; Phillips, Ruth B.
1998-01-01
Sequence variation in the control region (D-loop) of the mitochondrial DNA (mtDNA) was examined to assess the genetic distinctiveness of the shortjaw cisco (Coregonus zenithicus). Individuals from within the Great Lakes Basin as well as inland lakes outside the basin were sampled. DNA fragments containing the entire D-loop were amplified by PCR from specimens ofC. zenithicus and the related species C. artedi, C. hoyi, C. kiyi, and C. clupeaformis. DNA sequence analysis revealed high similarity within and among species and shared polymorphism for length variants. Based on this analysis, the shortjaw cisco is not genetically distinct from other cisco species.
Ackerman, Sara L; Koenig, Barbara A
2018-01-01
Increasingly used for clinical purposes, genome and exome sequencing can generate clinically relevant information that is not directly related to the reason for testing (incidental or secondary findings). Debates about the ethical implications of secondary findings were sparked by the American College of Medical Genetics (ACMG) 2013 policy statement, which recommended that laboratories report pathogenic alterations in 56 genes. Although wide variation in laboratories' secondary findings policies has been reported, little is known about its causes. We interviewed 18 laboratory directors and genetic counselors at 10 U.S. laboratories to investigate the motivations and interests shaping secondary findings reporting policies for clinical exome sequencing. Analysis of interview transcripts and laboratory documents was informed by sociological theories of standardization. Laboratories varied widely in terms of the types of secondary findings reported, consent-form language, and choices offered to patients. In explaining their adaptation of the ACMG report, our participants weighed genetic information's clinical, moral, professional, and commercial value in an attempt to maximize benefits for patients and families, minimize the costs of sequencing and analysis, adhere to professional norms, attract customers, and contend with the uncertain clinical implications of much of the genetic information generated. Nearly all laboratories in our study voluntarily adopted ACMG's recommendations, but their actual practices varied considerably and were informed by laboratory-specific judgments about clinical utility and patient benefit. Our findings offer a compelling example of standardization as a complex process that rarely leads simply to uniformity of practice. As laboratories take on a more prominent role in decisions about the return of genetic information, strategies are needed to inform patients, families, and clinicians about the differences between laboratories' practices and ensure that the consent process prompts a discussion of the value of additional genetic information for patients and their families.
Kim, Daniel Seung; Burt, Amber A; Ranchalis, Jane E; Wilmot, Beth; Smith, Joshua D; Patterson, Karynne E; Coe, Bradley P; Li, Yatong K; Bamshad, Michael J; Nikolas, Molly; Eichler, Evan E; Swanson, James M; Nigg, Joel T; Nickerson, Deborah A; Jarvik, Gail P
2017-06-01
Attention-Deficit Hyperactivity Disorder (ADHD) has high heritability; however, studies of common variation account for <5% of ADHD variance. Using data from affected participants without a family history of ADHD, we sought to identify de novo variants that could account for sporadic ADHD. Considering a total of 128 families, two analyses were conducted in parallel: first, in 11 unaffected parent/affected proband trios (or quads with the addition of an unaffected sibling) we completed exome sequencing. Six de novo missense variants at highly conserved bases were identified and validated from four of the 11 families: the brain-expressed genes TBC1D9, DAGLA, QARS, CSMD2, TRPM2, and WDR83. Separately, in 117 unrelated probands with sporadic ADHD, we sequenced a panel of 26 genes implicated in intellectual disability (ID) and autism spectrum disorder (ASD) to evaluate whether variation in ASD/ID-associated genes were also present in participants with ADHD. Only one putative deleterious variant (Gln600STOP) in CHD1L was identified; this was found in a single proband. Notably, no other nonsense, splice, frameshift, or highly conserved missense variants in the 26 gene panel were identified and validated. These data suggest that de novo variant analysis in families with independently adjudicated sporadic ADHD diagnosis can identify novel genes implicated in ADHD pathogenesis. Moreover, that only one of the 128 cases (0.8%, 11 exome, and 117 MIP sequenced participants) had putative deleterious variants within our data in 26 genes related to ID and ASD suggests significant independence in the genetic pathogenesis of ADHD as compared to ASD and ID phenotypes. © 2017 Wiley Periodicals, Inc. © 2017 Wiley Periodicals, Inc.
Interplay between social experiences and the genome: epigenetic consequences for behavior.
Champagne, Frances A
2012-01-01
Social experiences can have a persistent effect on biological processes leading to phenotypic diversity. Variation in gene regulation has emerged as a mechanism through which the interplay between DNA and environments leads to the biological encoding of these experiences. Epigenetic modifications-molecular pathways through which transcription is altered without altering the underlying DNA sequence-play a critical role in the normal process of development and are being increasingly explored as a mechanism linking environmental experiences to long-term biobehavioral outcomes. In this review, evidence implicating epigenetic factors, such as DNA methylation and histone modifications, in the link between social experiences occurring during the postnatal period and in adulthood and altered neuroendocrine and behavioral outcomes will be highlighted. In addition, the role of epigenetic mechanisms in shaping variation in social behavior and the implications of epigenetics for our understanding of the transmission of traits across generations will be discussed. Copyright © 2012 Elsevier Inc. All rights reserved.
Effect of read-mapping biases on detecting allele-specific expression from RNA-sequencing data
Degner, Jacob F.; Marioni, John C.; Pai, Athma A.; Pickrell, Joseph K.; Nkadori, Everlyne; Gilad, Yoav; Pritchard, Jonathan K.
2009-01-01
Motivation: Next-generation sequencing has become an important tool for genome-wide quantification of DNA and RNA. However, a major technical hurdle lies in the need to map short sequence reads back to their correct locations in a reference genome. Here, we investigate the impact of SNP variation on the reliability of read-mapping in the context of detecting allele-specific expression (ASE). Results: We generated 16 million 35 bp reads from mRNA of each of two HapMap Yoruba individuals. When we mapped these reads to the human genome we found that, at heterozygous SNPs, there was a significant bias toward higher mapping rates of the allele in the reference sequence, compared with the alternative allele. Masking known SNP positions in the genome sequence eliminated the reference bias but, surprisingly, did not lead to more reliable results overall. We find that even after masking, ∼5–10% of SNPs still have an inherent bias toward more effective mapping of one allele. Filtering out inherently biased SNPs removes 40% of the top signals of ASE. The remaining SNPs showing ASE are enriched in genes previously known to harbor cis-regulatory variation or known to show uniparental imprinting. Our results have implications for a variety of applications involving detection of alternate alleles from short-read sequence data. Availability: Scripts, written in Perl and R, for simulating short reads, masking SNP variation in a reference genome and analyzing the simulation output are available upon request from JFD. Raw short read data were deposited in GEO (http://www.ncbi.nlm.nih.gov/geo/) under accession number GSE18156. Contact: jdegner@uchicago.edu; marioni@uchicago.edu; gilad@uchicago.edu; pritch@uchicago.edu Supplementary information: Supplementary data are available at Bioinformatics online. PMID:19808877
Iskow, Rebecca C.; Austermann, Christian; Scharer, Christopher D.; Raj, Towfique; Boss, Jeremy M.; Sunyaev, Shamil; Price, Alkes; Stranger, Barbara; Simon, Viviana; Lee, Charles
2013-01-01
Ancient population structure shaping contemporary genetic variation has been recently appreciated and has important implications regarding our understanding of the structure of modern human genomes. We identified a ∼36-kb DNA segment in the human genome that displays an ancient substructure. The variation at this locus exists primarily as two highly divergent haplogroups. One of these haplogroups (the NE1 haplogroup) aligns with the Neandertal haplotype and contains a 4.6-kb deletion polymorphism in perfect linkage disequilibrium with 12 single nucleotide polymorphisms (SNPs) across diverse populations. The other haplogroup, which does not contain the 4.6-kb deletion, aligns with the chimpanzee haplotype and is likely ancestral. Africans have higher overall pairwise differences with the Neandertal haplotype than Eurasians do for this NE1 locus (p<10−15). Moreover, the nucleotide diversity at this locus is higher in Eurasians than in Africans. These results mimic signatures of recent Neandertal admixture contributing to this locus. However, an in-depth assessment of the variation in this region across multiple populations reveals that African NE1 haplotypes, albeit rare, harbor more sequence variation than NE1 haplotypes found in Europeans, indicating an ancient African origin of this haplogroup and refuting recent Neandertal admixture. Population genetic analyses of the SNPs within each of these haplogroups, along with genome-wide comparisons revealed significant FST (p = 0.00003) and positive Tajima's D (p = 0.00285) statistics, pointing to non-neutral evolution of this locus. The NE1 locus harbors no protein-coding genes, but contains transcribed sequences as well as sequences with putative regulatory function based on bioinformatic predictions and in vitro experiments. We postulate that the variation observed at this locus predates Human–Neandertal divergence and is evolving under balancing selection, especially among European populations. PMID:23593015
Genetic architecture of natural variation in Drosophila melanogaster aggressive behavior
Shorter, John; Couch, Charlene; Huang, Wen; Carbone, Mary Anna; Peiffer, Jason; Anholt, Robert R. H.; Mackay, Trudy F. C.
2015-01-01
Aggression is an evolutionarily conserved complex behavior essential for survival and the organization of social hierarchies. With the exception of genetic variants associated with bioamine signaling, which have been implicated in aggression in many species, the genetic basis of natural variation in aggression is largely unknown. Drosophila melanogaster is a favorable model system for exploring the genetic basis of natural variation in aggression. Here, we performed genome-wide association analyses using the inbred, sequenced lines of the Drosophila melanogaster Genetic Reference Panel (DGRP) and replicate advanced intercross populations derived from the most and least aggressive DGRP lines. We identified genes that have been previously implicated in aggressive behavior as well as many novel loci, including gustatory receptor 63a (Gr63a), which encodes a subunit of the receptor for CO2, and genes associated with development and function of the nervous system. Although genes from the two association analyses were largely nonoverlapping, they mapped onto a genetic interaction network inferred from an analysis of pairwise epistasis in the DGRP. We used mutations and RNAi knock-down alleles to functionally validate 79% of the candidate genes and 75% of the candidate epistatic interactions tested. Epistasis for aggressive behavior causes cryptic genetic variation in the DGRP that is revealed by changing allele frequencies in the outbred populations derived from extreme DGRP lines. This phenomenon may pertain to other fitness traits and species, with implications for evolution, applied breeding, and human genetics. PMID:26100892
Leung, Ross Ka-Kit; Dong, Zhi Qiang; Sa, Fei; Chong, Cheong Meng; Lei, Si Wan; Tsui, Stephen Kwok-Wing; Lee, Simon Ming-Yuen
2014-02-01
Minor variants have significant implications in quasispecies evolution, early cancer detection and non-invasive fetal genotyping but their accurate detection by next-generation sequencing (NGS) is hampered by sequencing errors. We generated sequencing data from mixtures at predetermined ratios in order to provide insight into sequencing errors and variations that can arise for which simulation cannot be performed. The information also enables better parameterization in depth of coverage, read quality and heterogeneity, library preparation techniques, technical repeatability for mathematical modeling, theory development and simulation experimental design. We devised minor variant authentication rules that achieved 100% accuracy in both testing and validation experiments. The rules are free from tedious inspection of alignment accuracy, sequencing read quality or errors introduced by homopolymers. The authentication processes only require minor variants to: (1) have minimum depth of coverage larger than 30; (2) be reported by (a) four or more variant callers, or (b) DiBayes or LoFreq, plus SNVer (or BWA when no results are returned by SNVer), and with the interassay coefficient of variation (CV) no larger than 0.1. Quantification accuracy undermined by sequencing errors could neither be overcome by ultra-deep sequencing, nor recruiting more variant callers to reach a consensus, such that consistent underestimation and overestimation (i.e. low CV) were observed. To accommodate stochastic error and adjust the observed ratio within a specified accuracy, we presented a proof of concept for the use of a double calibration curve for quantification, which provides an important reference towards potential industrial-scale fabrication of calibrants for NGS.
DOE Office of Scientific and Technical Information (OSTI.GOV)
Muchero, Wellington; Labbe, Jessy L; Priya, Ranjan
2014-01-01
To date, Populus ranks among a few plant species with a complete genome sequence and other highly developed genomic resources. With the first genome sequence among all tree species, Populus has been adopted as a suitable model organism for genomic studies in trees. However, far from being just a model species, Populus is a key renewable economic resource that plays a significant role in providing raw materials for the biofuel and pulp and paper industries. Therefore, aside from leading frontiers of basic tree molecular biology and ecological research, Populus leads frontiers in addressing global economic challenges related to fuel andmore » fiber production. The latter fact suggests that research aimed at improving quality and quantity of Populus as a raw material will likely drive the pursuit of more targeted and deeper research in order to unlock the economic potential tied in molecular biology processes that drive this tree species. Advances in genome sequence-driven technologies, such as resequencing individual genotypes, which in turn facilitates large scale SNP discovery and identification of large scale polymorphisms are key determinants of future success in these initiatives. In this treatise we discuss implications of genome sequence-enable technologies on Populus genomic and genetic studies of complex and specialized-traits.« less
Geoffroy, Véronique; Stoetzel, Corinne; Scheidecker, Sophie; Schaefer, Elise; Perrault, Isabelle; Bär, Séverine; Kröll, Ariane; Delbarre, Marion; Antin, Manuela; Leuvrey, Anne-Sophie; Henry, Charline; Blanché, Hélène; Decker, Eva; Kloth, Katja; Klaus, Günter; Mache, Christoph; Martin-Coignard, Dominique; McGinn, Steven; Boland, Anne; Deleuze, Jean-François; Friant, Sylvie; Saunier, Sophie; Rozet, Jean-Michel; Bergmann, Carsten; Dollfus, Hélène; Muller, Jean
2018-04-24
Ciliopathies represent a wide spectrum of rare diseases with overlapping phenotypes and a high genetic heterogeneity. Among those, IFT140 is implicated in a variety of phenotypes ranging from isolated retinis pigmentosa to more syndromic cases. Using whole-genome sequencing in patients with uncharacterized ciliopathies, we identified a novel recurrent tandem duplication of exon 27-30 (6.7 kb) in IFT140, c.3454-488_4182+2588dup p.(Tyr1152_Thr1394dup), missed by whole-exome sequencing. Pathogenicity of the mutation was assessed on the patients' skin fibroblasts. Several hundreds of patients with a ciliopathy phenotype were screened and biallelic mutations were identified in 11 families representing 12 pathogenic variants of which seven are novel. Among those unrelated families especially with a Mainzer-Saldino syndrome, eight carried the same tandem duplication (two at the homozygous state and six at the heterozygous state). In conclusion, we demonstrated the implication of structural variations in IFT140-related diseases expanding its mutation spectrum. We also provide evidences for a unique genomic event mediated by an Alu-Alu recombination occurring on a shared haplotype. We confirm that whole-genome sequencing can be instrumental in the ability to detect structural variants for genomic disorders. © 2018 Wiley Periodicals, Inc.
Chowanadisai, Winyoo; Kelleher, Shannon L; Nemeth, Jennifer F; Yachetti, Stephen; Kuhlman, Charles F; Jackson, Joan G; Davis, Anne M; Lien, Eric L; Lönnerdal, Bo
2005-05-01
Variability in the protein composition of breast milk has been observed in many women and is believed to be due to natural variation of the human population. Single nucleotide polymorphisms (SNPs) are present throughout the entire human genome, but the impact of this variation on human milk composition and biological activity and infant nutrition and health is unclear. The goals of this study were to characterize a variant of human alpha-lactalbumin observed in milk from a Filipino population by determining the location of the polymorphism in the amino acid and genomic sequences of alpha-lactalbumin. Milk and blood samples were collected from 20 Filipino women, and milk samples were collected from an additional 450 women from nine different countries. alpha-Lactalbumin concentration was measured by high-performance liquid chromatography (HPLC), and milk samples containing the variant form of the protein were identified with both HPLC and mass spectrometry (MS). The molecular weight of the variant form was measured by MS, and the location of the polymorphism was narrowed down by protein reduction, alkylation and trypsin digestion. Genomic DNA was isolated from whole blood, and the polymorphism location and subject genotype were determined by amplifying the entire coding sequence of human alpha-lactalbumin by PCR, followed by DNA sequencing. A variant form of alpha-lactalbumin was observed in HPLC chromatograms, and the difference in molecular weight was determined by MS (wild type=14,070 Da, variant=14,056 Da). Protein reduction and digestion narrowed the polymorphism between the 33rd and 77th amino acid of the protein. The genetic polymorphism was identified as adenine to guanine, which translates to a substitution from isoleucine to valine at amino acid 46. The frequency of variation was higher in milk from China, Japan and Philippines, which suggests that this polymorphism is most prevalent in Asia. There are SNPs in the genome for human milk proteins and their implications for protein bioactivity and infant nutrition need to be considered.
Schiessl, Sarah; Samans, Birgit; Hüttel, Bruno; Reinhard, Richard; Snowdon, Rod J.
2014-01-01
Flowering, the transition from the vegetative to the generative phase, is a decisive time point in the lifecycle of a plant. Flowering is controlled by a complex network of transcription factors, photoreceptors, enzymes and miRNAs. In recent years, several studies gave rise to the hypothesis that this network is also strongly involved in the regulation of other important lifecycle processes ranging from germination and seed development through to fundamental developmental and yield-related traits. In the allopolyploid crop species Brassica napus, (genome AACC), homoeologous copies of flowering time regulatory genes are implicated in major phenological variation within the species, however the extent and control of intraspecific and intergenomic variation among flowering-time regulators is still unclear. To investigate differences among B. napus morphotypes in relation to flowering-time gene variation, we performed targeted deep sequencing of 29 regulatory flowering-time genes in four genetically and phenologically diverse B. napus accessions. The genotype panel included a winter-type oilseed rape, a winter fodder rape, a spring-type oilseed rape (all B. napus ssp. napus) and a swede (B. napus ssp. napobrassica), which show extreme differences in winter-hardiness, vernalization requirement and flowering behavior. A broad range of genetic variation was detected in the targeted genes for the different morphotypes, including non-synonymous SNPs, copy number variation and presence-absence variation. The results suggest that this broad variation in vernalization, clock and signaling genes could be a key driver of morphological differentiation for flowering-related traits in this recent allopolyploid crop species. PMID:25202314
Yang, Jian-Rong; Maclean, Calum J; Park, Chungoo; Zhao, Huabin; Zhang, Jianzhi
2017-09-01
It is commonly, although not universally, accepted that most intra and interspecific genome sequence variations are more or less neutral, whereas a large fraction of organism-level phenotypic variations are adaptive. Gene expression levels are molecular phenotypes that bridge the gap between genotypes and corresponding organism-level phenotypes. Yet, it is unknown whether natural variations in gene expression levels are mostly neutral or adaptive. Here we address this fundamental question by genome-wide profiling and comparison of gene expression levels in nine yeast strains belonging to three closely related Saccharomyces species and originating from five different ecological environments. We find that the transcriptome-based clustering of the nine strains approximates the genome sequence-based phylogeny irrespective of their ecological environments. Remarkably, only ∼0.5% of genes exhibit similar expression levels among strains from a common ecological environment, no greater than that among strains with comparable phylogenetic relationships but different environments. These and other observations strongly suggest that most intra and interspecific variations in yeast gene expression levels result from the accumulation of random mutations rather than environmental adaptations. This finding has profound implications for understanding the driving force of gene expression evolution, genetic basis of phenotypic adaptation, and general role of stochasticity in evolution. © The Author 2017. Published by Oxford University Press on behalf of the Society for Molecular Biology and Evolution.
Taylor, E B; Pollard, S; Louie, D
1999-07-01
Bull trout, Salvelinus confluentus (Salmonidae), are distributed in northwestern North America from Nevada to Yukon Territory, largely in interior drainages. The species is of conservation concern owing to declines in abundance, particularly in southern portions of its range. To investigate phylogenetic structure within bull trout that might form the basis for the delineation of major conservation units, we conducted a mitochondrial DNA (mtDNA) survey in bull trout from throughout its range. Restriction fragment length polymorphism (RFLP) analysis of four segments of the mtDNA genome with 11 restriction enzymes resolved 21 composite haplotypes that differed by an average of 0.5% in sequence. One group of haplotypes predominated in 'coastal' areas (west of the coastal mountain ranges) while another predominated in 'interior' regions (east of the coastal mountains). The two putative lineages differed by 0.8% in sequence and were also resolved by sequencing a portion of the ND1 gene in a representative of each RFLP haplotype. Significant variation existed within individual sample sites (12% of total variation) and among sites within major geographical regions (33%), but most variation (55%) was associated with differences between coastal and interior regions. We concluded that: (i) bull trout are subdivided into coastal and interior lineages; (ii) this subdivision reflects recent historical isolation in two refugia south of the Cordilleran ice sheet during the Pleistocene: the Chehalis and Columbia refugia; and (iii) most of the molecular variation resides at the interpopulation and inter-region levels. Conservation efforts, therefore, should focus on maintaining as many populations as possible across as many geographical regions as possible within both coastal and interior lineages.
Ellis, Lisa L.; Huang, Wen; Quinn, Andrew M.; Ahuja, Astha; Alfrejd, Ben; Gomez, Francisco E.; Hjelmen, Carl E.; Moore, Kristi L.; Mackay, Trudy F. C.; Johnston, J. Spencer; Tarone, Aaron M.
2014-01-01
We determined female genome sizes using flow cytometry for 211 Drosophila melanogaster sequenced inbred strains from the Drosophila Genetic Reference Panel, and found significant conspecific and intrapopulation variation in genome size. We also compared several life history traits for 25 lines with large and 25 lines with small genomes in three thermal environments, and found that genome size as well as genome size by temperature interactions significantly correlated with survival to pupation and adulthood, time to pupation, female pupal mass, and female eclosion rates. Genome size accounted for up to 23% of the variation in developmental phenotypes, but the contribution of genome size to variation in life history traits was plastic and varied according to the thermal environment. Expression data implicate differences in metabolism that correspond to genome size variation. These results indicate that significant genome size variation exists within D. melanogaster and this variation may impact the evolutionary ecology of the species. Genome size variation accounts for a significant portion of life history variation in an environmentally dependent manner, suggesting that potential fitness effects associated with genome size variation also depend on environmental conditions. PMID:25057905
Erceg, Jelena; Saunders, Timothy E.; Girardot, Charles; Devos, Damien P.; Hufnagel, Lars; Furlong, Eileen E. M.
2014-01-01
Deciphering the specific contribution of individual motifs within cis-regulatory modules (CRMs) is crucial to understanding how gene expression is regulated and how this process is affected by sequence variation. But despite vast improvements in the ability to identify where transcription factors (TFs) bind throughout the genome, we are limited in our ability to relate information on motif occupancy to function from sequence alone. Here, we engineered 63 synthetic CRMs to systematically assess the relationship between variation in the content and spacing of motifs within CRMs to CRM activity during development using Drosophila transgenic embryos. In over half the cases, very simple elements containing only one or two types of TF binding motifs were capable of driving specific spatio-temporal patterns during development. Different motif organizations provide different degrees of robustness to enhancer activity, ranging from binary on-off responses to more subtle effects including embryo-to-embryo and within-embryo variation. By quantifying the effects of subtle changes in motif organization, we were able to model biophysical rules that explain CRM behavior and may contribute to the spatial positioning of CRM activity in vivo. For the same enhancer, the effects of small differences in motif positions varied in developmentally related tissues, suggesting that gene expression may be more susceptible to sequence variation in one tissue compared to another. This result has important implications for human eQTL studies in which many associated mutations are found in cis-regulatory regions, though the mechanism for how they affect tissue-specific gene expression is often not understood. PMID:24391522
Horizontal gene transfer of chromosomal Type II toxin-antitoxin systems of Escherichia coli.
Ramisetty, Bhaskar Chandra Mohan; Santhosh, Ramachandran Sarojini
2016-02-01
Type II toxin-antitoxin systems (TAs) are small autoregulated bicistronic operons that encode a toxin protein with the potential to inhibit metabolic processes and an antitoxin protein to neutralize the toxin. Most of the bacterial genomes encode multiple TAs. However, the diversity and accumulation of TAs on bacterial genomes and its physiological implications are highly debated. Here we provide evidence that Escherichia coli chromosomal TAs (encoding RNase toxins) are 'acquired' DNA likely originated from heterologous DNA and are the smallest known autoregulated operons with the potential for horizontal propagation. Sequence analyses revealed that integration of TAs into the bacterial genome is unique and contributes to variations in the coding and/or regulatory regions of flanking host genome sequences. Plasmids and genomes encoding identical TAs of natural isolates are mutually exclusive. Chromosomal TAs might play significant roles in the evolution and ecology of bacteria by contributing to host genome variation and by moderation of plasmid maintenance. © FEMS 2015. All rights reserved. For permissions, please e-mail: journals.permissions@oup.com.
Bhatia, Shipra; Gordon, Christopher T.; Foster, Robert G.; Melin, Lucie; Abadie, Véronique; Baujat, Geneviève; Vazquez, Marie-Paule; Amiel, Jeanne; Lyonnet, Stanislas; van Heyningen, Veronica; Kleinjan, Dirk A.
2015-01-01
Disruption of gene regulation by sequence variation in non-coding regions of the genome is now recognised as a significant cause of human disease and disease susceptibility. Sequence variants in cis-regulatory elements (CREs), the primary determinants of spatio-temporal gene regulation, can alter transcription factor binding sites. While technological advances have led to easy identification of disease-associated CRE variants, robust methods for discerning functional CRE variants from background variation are lacking. Here we describe an efficient dual-colour reporter transgenesis approach in zebrafish, simultaneously allowing detailed in vivo comparison of spatio-temporal differences in regulatory activity between putative CRE variants and assessment of altered transcription factor binding potential of the variant. We validate the method on known disease-associated elements regulating SHH, PAX6 and IRF6 and subsequently characterise novel, ultra-long-range SOX9 enhancers implicated in the craniofacial abnormality Pierre Robin Sequence. The method provides a highly cost-effective, fast and robust approach for simultaneously unravelling in a single assay whether, where and when in embryonic development a disease-associated CRE-variant is affecting its regulatory function. PMID:26030420
Skoglund, Pontus; Höglund, Jacob
2010-04-23
Population variation in the degree of seasonal polymorphism is rare in birds, and the genetic basis of this phenomenon remains largely undescribed. Both sexes of Scandinavian and Scottish Willow grouse (Lagopus lagopus) display marked differences in their winter phenotypes, with Scottish grouse retaining a pigmented plumage year-round and Scandinavian Willow grouse molting to a white morph during winter. A widely studied pathway implicated in vertebrate pigmentation is the melanin system, for which functional variation has been characterised in many taxa. We sequenced coding regions from four genes involved in melanin pigmentation (DCT, MC1R, TYR and TYRP1), and an additional control involved in the melanocortin pathway (AGRP), to investigate the genetic basis of winter plumage in Lagopus. Despite the well documented role of the melanin system in animal coloration, we found no plumage-associated polymorphism or evidence for selection in a total of approximately 2.6 kb analysed sequence. Our results indicate that the genetic basis of alternating between pigmented and unpigmented seasonal phenotypes is more likely explained by regulatory changes controlling the expression of these or other loci in the physiological pathway leading to pigmentation.
Genetic background effects in quantitative genetics: gene-by-system interactions.
Sardi, Maria; Gasch, Audrey P
2018-04-11
Proper cell function depends on networks of proteins that interact physically and functionally to carry out physiological processes. Thus, it seems logical that the impact of sequence variation in one protein could be significantly influenced by genetic variants at other loci in a genome. Nonetheless, the importance of such genetic interactions, known as epistasis, in explaining phenotypic variation remains a matter of debate in genetics. Recent work from our lab revealed that genes implicated from an association study of toxin tolerance in Saccharomyces cerevisiae show extensive interactions with the genetic background: most implicated genes, regardless of allele, are important for toxin tolerance in only one of two tested strains. The prevalence of background effects in our study adds to other reports of widespread genetic-background interactions in model organisms. We suggest that these effects represent many-way interactions with myriad features of the cellular system that vary across classes of individuals. Such gene-by-system interactions may influence diverse traits and require new modeling approaches to accurately represent genotype-phenotype relationships across individuals.
Genetics of Inflammatory Bowel Diseases
McGovern, Dermot; Kugathasan, Subra; Cho, Judy H.
2015-01-01
In this Review, we provide an update on genome-wide association studies (GWAS) in inflammatory bowel disease (IBD). In addition, we summarize progress in defining the functional consequences of associated alleles for coding and non-coding genetic variation. In the small minority of loci where major association signals correspond to non-synonymous variation, we summarize studies defining their functional effects and implications for therapeutic targeting. Importantly, the large majority of GWAS-associated loci involve non-coding variation, many of which modulate levels of gene expression. Recent expression quantitative trait loci (eQTL) studies have established that expression of the large majority of human genes is regulated by non-coding genetic variation. Significant advances in defining the epigenetic landscape have demonstrated that IBD GWAS signals are highly enriched within cell-specific active enhancer marks. Studies in European ancestry populations have dominated the landscape of IBD genetics studies, but increasingly, studies in Asian and African-American populations are being reported. Common variation accounts for only a modest fraction of the predicted heritability and the role of rare genetic variation of higher effects (i.e. odds ratios markedly deviating from one) is increasingly being identified through sequencing efforts. These sequencing studies have been particularly productive in very-early onset, more severe cases. A major challenge in IBD genetics will be harnessing the vast array of genetic discovery for clinical utility, through emerging precision medicine initiatives. We discuss the rapidly evolving area of direct to consumer genetic testing, as well as the current utility of clinical exome sequencing, especially in very early onset, severe IBD cases. We summarize recent progress in the pharmacogenetics of IBD with respect of partitioning patient responses to anti-TNF and thiopurine therapies. Highly collaborative studies across research centers and across subspecialties and disciplines will be required to fully realize the promise of genetic discovery in IBD. PMID:26255561
Short-Sequence DNA Repeats in Prokaryotic Genomes
van Belkum, Alex; Scherer, Stewart; van Alphen, Loek; Verbrugh, Henri
1998-01-01
Short-sequence DNA repeat (SSR) loci can be identified in all eukaryotic and many prokaryotic genomes. These loci harbor short or long stretches of repeated nucleotide sequence motifs. DNA sequence motifs in a single locus can be identical and/or heterogeneous. SSRs are encountered in many different branches of the prokaryote kingdom. They are found in genes encoding products as diverse as microbial surface components recognizing adhesive matrix molecules and specific bacterial virulence factors such as lipopolysaccharide-modifying enzymes or adhesins. SSRs enable genetic and consequently phenotypic flexibility. SSRs function at various levels of gene expression regulation. Variations in the number of repeat units per locus or changes in the nature of the individual repeat sequences may result from recombination processes or polymerase inadequacy such as slipped-strand mispairing (SSM), either alone or in combination with DNA repair deficiencies. These rather complex phenomena can occur with relative ease, with SSM approaching a frequency of 10−4 per bacterial cell division and allowing high-frequency genetic switching. Bacteria use this random strategy to adapt their genetic repertoire in response to selective environmental pressure. SSR-mediated variation has important implications for bacterial pathogenesis and evolutionary fitness. Molecular analysis of changes in SSRs allows epidemiological studies on the spread of pathogenic bacteria. The occurrence, evolution and function of SSRs, and the molecular methods used to analyze them are discussed in the context of responsiveness to environmental factors, bacterial pathogenicity, epidemiology, and the availability of full-genome sequences for increasing numbers of microorganisms, especially those that are medically relevant. PMID:9618442
Sabir, Jamal; Schwarz, Erika; Ellison, Nicholas; Zhang, Jin; Baeshen, Nabih A; Mutwakil, Muhammed; Jansen, Robert; Ruhlman, Tracey
2014-08-01
Land plant plastid genomes (plastomes) provide a tractable model for evolutionary study in that they are relatively compact and gene dense. Among the groups that display an appropriate level of variation for structural features, the inverted-repeat-lacking clade (IRLC) of papilionoid legumes presents the potential to advance general understanding of the mechanisms of genomic evolution. Here, are presented six complete plastome sequences from economically important species of the IRLC, a lineage previously represented by only five completed plastomes. A number of characters are compared across the IRLC including gene retention and divergence, synteny, repeat structure and functional gene transfer to the nucleus. The loss of clpP intron 2 was identified in one newly sequenced member of IRLC, Glycyrrhiza glabra. Using deeply sequenced nuclear transcriptomes from two species helped clarify the nature of the functional transfer of accD to the nucleus in Trifolium, which likely occurred in the lineage leading to subgenus Trifolium. Legumes are second only to cereal crops in agricultural importance based on area harvested and total production. Genetic improvement via plastid transformation of IRLC crop species is an appealing proposition. Comparative analyses of intergenic spacer regions emphasize the need for complete genome sequences for developing transformation vectors for plastid genetic engineering of legume crops. © 2014 Society for Experimental Biology, Association of Applied Biologists and John Wiley & Sons Ltd.
Gasser, R B; Rossi, L; Zhu, X
1999-11-01
The sequence of the second internal transcribed spacer of ribosomal DNA was determined for four species of Nematodirus (Nematodirus rupicaprae, Nematodirus oiratianus, Nematodirus davtiani alpinus and Nematodirus europaeus) from roe deer or alpine chamois. The second internal transcribed spacer of the four species varied in length from 228 to 236 bp, and the G + C contents ranged from 41 to 44%. While no intraspecific sequence variation was detected among multiple samples representing three of the taxa, sequence differences of 5.9-9.7% were detected among the four species, Nematodirus davtiani alpinus and N. rupicaprae were genetically most similar (94.1%), followed by N. oiratianus, N. europaeus and N. rupicaprae (91.1-91.5%), whereas N. oiratianus was genetically most different from N. davtiani alpinus. The interspecific sequence differences were exploited for the delineation of the four species by PCR-based restriction fragment length polymorphism (using two enzymes) and single-strand conformation polymorphism. The results have implications for diagnosis, epidemiology and for studying the systematics of the Nematodirinae.
The adaptive evolution of the mammalian mitochondrial genome
da Fonseca, Rute R; Johnson, Warren E; O'Brien, Stephen J; Ramos, Maria João; Antunes, Agostinho
2008-01-01
Background The mitochondria produce up to 95% of a eukaryotic cell's energy through oxidative phosphorylation. The proteins involved in this vital process are under high functional constraints. However, metabolic requirements vary across species, potentially modifying selective pressures. We evaluate the adaptive evolution of 12 protein-coding mitochondrial genes in 41 placental mammalian species by assessing amino acid sequence variation and exploring the functional implications of observed variation in secondary and tertiary protein structures. Results Wide variation in the properties of amino acids were observed at functionally important regions of cytochrome b in species with more-specialized metabolic requirements (such as adaptation to low energy diet or large body size, such as in elephant, dugong, sloth, and pangolin, and adaptation to unusual oxygen requirements, for example diving in cetaceans, flying in bats, and living at high altitudes in alpacas). Signatures of adaptive variation in the NADH dehydrogenase complex were restricted to the loop regions of the transmembrane units which likely function as protons pumps. Evidence of adaptive variation in the cytochrome c oxidase complex was observed mostly at the interface between the mitochondrial and nuclear-encoded subunits, perhaps evidence of co-evolution. The ATP8 subunit, which has an important role in the assembly of F0, exhibited the highest signal of adaptive variation. ATP6, which has an essential role in rotor performance, showed a high adaptive variation in predicted loop areas. Conclusion Our study provides insight into the adaptive evolution of the mtDNA genome in mammals and its implications for the molecular mechanism of oxidative phosphorylation. We present a framework for future experimental characterization of the impact of specific mutations in the function, physiology, and interactions of the mtDNA encoded proteins involved in oxidative phosphorylation. PMID:18318906
Genetic heterogeneity of diffuse large B-cell lymphoma.
Zhang, Jenny; Grubor, Vladimir; Love, Cassandra L; Banerjee, Anjishnu; Richards, Kristy L; Mieczkowski, Piotr A; Dunphy, Cherie; Choi, William; Au, Wing Yan; Srivastava, Gopesh; Lugar, Patricia L; Rizzieri, David A; Lagoo, Anand S; Bernal-Mizrachi, Leon; Mann, Karen P; Flowers, Christopher; Naresh, Kikkeri; Evens, Andrew; Gordon, Leo I; Czader, Magdalena; Gill, Javed I; Hsi, Eric D; Liu, Qingquan; Fan, Alice; Walsh, Katherine; Jima, Dereje; Smith, Lisa L; Johnson, Amy J; Byrd, John C; Luftig, Micah A; Ni, Ting; Zhu, Jun; Chadburn, Amy; Levy, Shawn; Dunson, David; Dave, Sandeep S
2013-01-22
Diffuse large B-cell lymphoma (DLBCL) is the most common form of lymphoma in adults. The disease exhibits a striking heterogeneity in gene expression profiles and clinical outcomes, but its genetic causes remain to be fully defined. Through whole genome and exome sequencing, we characterized the genetic diversity of DLBCL. In all, we sequenced 73 DLBCL primary tumors (34 with matched normal DNA). Separately, we sequenced the exomes of 21 DLBCL cell lines. We identified 322 DLBCL cancer genes that were recurrently mutated in primary DLBCLs. We identified recurrent mutations implicating a number of known and not previously identified genes and pathways in DLBCL including those related to chromatin modification (ARID1A and MEF2B), NF-κB (CARD11 and TNFAIP3), PI3 kinase (PIK3CD, PIK3R1, and MTOR), B-cell lineage (IRF8, POU2F2, and GNA13), and WNT signaling (WIF1). We also experimentally validated a mutation in PIK3CD, a gene not previously implicated in lymphomas. The patterns of mutation demonstrated a classic long tail distribution with substantial variation of mutated genes from patient to patient and also between published studies. Thus, our study reveals the tremendous genetic heterogeneity that underlies lymphomas and highlights the need for personalized medicine approaches to treating these patients.
Variation in recombination rate may bias human genetic disease mapping studies.
Boyle, A Susannah; Noor, Mohamed A F
2004-11-01
The availability of the human genome sequence and variability information (as from the International HapMap project) will enhance our ability to map genetic disorders and choose targets for therapeutic intervention. However, several factors, such as regional variation in recombination rate, can bias conclusions from genetic mapping studies. Here, we examine the impact of regional variation in recombination rate across the human genome. Through computer simulations and literature surveys, we conclude that genetic disorders have been mapped to regions of low recombination more often than expected if such diseases were randomly distributed across the genome. This concentration in low recombination regions may be an artifact, and disorders appearing to be caused by a few genes of large effect may be polygenic. Future genetic mapping studies should be conscious of this potential complication by noting the regional recombination rate of regions implicated in diseases.
Korber, B T; Osmanov, S; Esparza, J; Myers, G
1994-11-01
The World Health Organization Global Programme on AIDS (WHO/GPA) is conducting a large-scale collaborative study of human immunodeficiency virus type 1 (HIV-1) variation, based in four potential vaccine-trial site countries: Brazil, Rwanda, Thailand, and Uganda. Through the course of this study, it was crucial to keep track of certain attributes of the samples from which the viral nucleotide sequences were derived (e.g., country of origin and viral culture characterization), so that meaningful sequence comparisons could be made. Here we describe a system developed in the context of the WHO/GPA study that summarizes such critical attributes by representing them as standardized characters directly incorporated into sequence names. This nomenclature allows linkage of clinical, phenotypic, and geographic information with molecular data. We propose that other investigators involved in human immunodeficiency virus (HIV) nucleotide sequencing efforts adopt a similar standardized sequence nomenclature to facilitate cross-study sequence comparison. HIV sequence data are being generated at an ever-increasing rate; directly coupled to this increase is our deepening understanding of biological parameters that influence or result from sequence variability. A standardized sequence nomenclature that includes relevant biological information would enable researchers to better utilize the growing body of sequence data, and enhance their ability to interpret the biological implications of their own data through facilitating comparisons with previously published work.
Galan, Maxime; Guivier, Emmanuel; Caraux, Gilles; Charbonnel, Nathalie; Cosson, Jean-François
2010-05-11
High-throughput sequencing technologies offer new perspectives for biomedical, agronomical and evolutionary research. Promising progresses now concern the application of these technologies to large-scale studies of genetic variation. Such studies require the genotyping of high numbers of samples. This is theoretically possible using 454 pyrosequencing, which generates billions of base pairs of sequence data. However several challenges arise: first in the attribution of each read produced to its original sample, and second, in bioinformatic analyses to distinguish true from artifactual sequence variation. This pilot study proposes a new application for the 454 GS FLX platform, allowing the individual genotyping of thousands of samples in one run. A probabilistic model has been developed to demonstrate the reliability of this method. DNA amplicons from 1,710 rodent samples were individually barcoded using a combination of tags located in forward and reverse primers. Amplicons consisted in 222 bp fragments corresponding to DRB exon 2, a highly polymorphic gene in mammals. A total of 221,789 reads were obtained, of which 153,349 were finally assigned to original samples. Rules based on a probabilistic model and a four-step procedure, were developed to validate sequences and provide a confidence level for each genotype. The method gave promising results, with the genotyping of DRB exon 2 sequences for 1,407 samples from 24 different rodent species and the sequencing of 392 variants in one half of a 454 run. Using replicates, we estimated that the reproducibility of genotyping reached 95%. This new approach is a promising alternative to classical methods involving electrophoresis-based techniques for variant separation and cloning-sequencing for sequence determination. The 454 system is less costly and time consuming and may enhance the reliability of genotypes obtained when high numbers of samples are studied. It opens up new perspectives for the study of evolutionary and functional genetics of highly polymorphic genes like major histocompatibility complex genes in vertebrates or loci regulating self-compatibility in plants. Important applications in biomedical research will include the detection of individual variation in disease susceptibility. Similarly, agronomy will benefit from this approach, through the study of genes implicated in productivity or disease susceptibility traits.
Functional annotation of HOT regions in the human genome: implications for human disease and cancer
Li, Hao; Chen, Hebing; Liu, Feng; Ren, Chao; Wang, Shengqi; Bo, Xiaochen; Shu, Wenjie
2015-01-01
Advances in genome-wide association studies (GWAS) and large-scale sequencing studies have resulted in an impressive and growing list of disease- and trait-associated genetic variants. Most studies have emphasised the discovery of genetic variation in coding sequences, however, the noncoding regulatory effects responsible for human disease and cancer biology have been substantially understudied. To better characterise the cis-regulatory effects of noncoding variation, we performed a comprehensive analysis of the genetic variants in HOT (high-occupancy target) regions, which are considered to be one of the most intriguing findings of recent large-scale sequencing studies. We observed that GWAS variants that map to HOT regions undergo a substantial net decrease and illustrate development-specific localisation during haematopoiesis. Additionally, genetic risk variants are disproportionally enriched in HOT regions compared with LOT (low-occupancy target) regions in both disease-relevant and cancer cells. Importantly, this enrichment is biased toward disease- or cancer-specific cell types. Furthermore, we observed that cancer cells generally acquire cancer-specific HOT regions at oncogenes through diverse mechanisms of cancer pathogenesis. Collectively, our findings demonstrate the key roles of HOT regions in human disease and cancer and represent a critical step toward further understanding disease biology, diagnosis, and therapy. PMID:26113264
Functional annotation of HOT regions in the human genome: implications for human disease and cancer.
Li, Hao; Chen, Hebing; Liu, Feng; Ren, Chao; Wang, Shengqi; Bo, Xiaochen; Shu, Wenjie
2015-06-26
Advances in genome-wide association studies (GWAS) and large-scale sequencing studies have resulted in an impressive and growing list of disease- and trait-associated genetic variants. Most studies have emphasised the discovery of genetic variation in coding sequences, however, the noncoding regulatory effects responsible for human disease and cancer biology have been substantially understudied. To better characterise the cis-regulatory effects of noncoding variation, we performed a comprehensive analysis of the genetic variants in HOT (high-occupancy target) regions, which are considered to be one of the most intriguing findings of recent large-scale sequencing studies. We observed that GWAS variants that map to HOT regions undergo a substantial net decrease and illustrate development-specific localisation during haematopoiesis. Additionally, genetic risk variants are disproportionally enriched in HOT regions compared with LOT (low-occupancy target) regions in both disease-relevant and cancer cells. Importantly, this enrichment is biased toward disease- or cancer-specific cell types. Furthermore, we observed that cancer cells generally acquire cancer-specific HOT regions at oncogenes through diverse mechanisms of cancer pathogenesis. Collectively, our findings demonstrate the key roles of HOT regions in human disease and cancer and represent a critical step toward further understanding disease biology, diagnosis, and therapy.
Satellite DNA and cytogenetic evolution: molecular aspects and implications for man. [Kangaroo rats
DOE Office of Scientific and Technical Information (OSTI.GOV)
Hatch, F.T.; Mazrimas, J.
1977-02-28
Simple, highly reiterated DNA sequences, often observed in density gradients as satellite DNAs, exist in condensed heterochromatin. This material is predominantly located at chromosomal centromeres, occasionally at telomeres, or intercalated within arms; in a few species it occupies entire chromosome arms. Satellite DNAs are a highly variable component of the genome of most higher eukaryotes, but their functions have remained speculative. The genus of kangaroo rats (Dipodomys) exhibits remarkable interspecies variations in content of three satellite DNAs, consisting of simple sequences 3 to 10 base pairs long, and in species karyotypes. A broad range of diploid-DNA content is correlated withmore » satellite-DNA content. The latter is correlated positively with predominance of biarmed over uniarmed chromosomes (high fundamental number FN) and inversely with two anatomical indices (leg-bone-length ratios) of specialization for the jumping gait. Karyotypic variation is achieved via chromosomal rearrangements, e.g., Robertsonian fusion, C-band heteromorphism, and pericentric inversion. Environmental adaptation is achieved, in part, by reassortment of gene-linkage groups and regulatory controls as a result of the chromosomal rearrangements. The foregoing relationships led to the postulation that highly reiterated DNA sequences play a supragenic, global role in environmental adaptation and the evolution of new species.« less
Hand, Melanie L.; Spangenberg, German C.; Forster, John W.; Cogan, Noel O. I.
2013-01-01
Chloroplast genome sequences are of broad significance in plant biology, due to frequent use in molecular phylogenetics, comparative genomics, population genetics, and genetic modification studies. The present study used a second-generation sequencing approach to determine and assemble the plastid genomes (plastomes) of four representatives from the agriculturally important Lolium-Festuca species complex of pasture grasses (Lolium multiflorum, Festuca pratensis, Festuca altissima, and Festuca ovina). Total cellular DNA was extracted from either roots or leaves, was sequenced, and the output was filtered for plastome-related reads. A comparison between sources revealed fewer plastome-related reads from root-derived template but an increase in incidental bacterium-derived sequences. Plastome assembly and annotation indicated high levels of sequence identity and a conserved organization and gene content between species. However, frequent deletions within the F. ovina plastome appeared to contribute to a smaller plastid genome size. Comparative analysis with complete plastome sequences from other members of the Poaceae confirmed conservation of most grass-specific features. Detailed analysis of the rbcL–psaI intergenic region, however, revealed a “hot-spot” of variation characterized by independent deletion events. The evolutionary implications of this observation are discussed. The complete plastome sequences are anticipated to provide the basis for potential organelle-specific genetic modification of pasture grasses. PMID:23550121
Zhou, Y; Ingelman-Sundberg, M; Lauschke, V M
2017-10-01
Genetic polymorphisms in cytochrome P450 (CYP) genes can result in altered metabolic activity toward a plethora of clinically important medications. Thus, single nucleotide variants and copy number variations in CYP genes are major determinants of drug pharmacokinetics and toxicity and constitute pharmacogenetic biomarkers for drug dosing, efficacy, and safety. Strikingly, the distribution of CYP alleles differs considerably between populations with important implications for personalized drug therapy and healthcare programs. To provide a global distribution map of CYP alleles with clinical importance, we integrated whole-genome and exome sequencing data from 56,945 unrelated individuals of five major human populations. By combining this dataset with population-specific linkage information, we derive the frequencies of 176 CYP haplotypes, providing an extensive resource for major genetic determinants of drug metabolism. Furthermore, we aggregated this dataset into spectra of predicted functional variability in the respective populations and discuss the implications for population-adjusted pharmacological treatment strategies. © 2017 The Authors Clinical Pharmacology & Therapeutics published by Wiley Periodicals, Inc. on behalf of American Society for Clinical Pharmacology and Therapeutics.
Bergman, C M; Kreitman, M
2001-08-01
Comparative genomic approaches to gene and cis-regulatory prediction are based on the principle that differential DNA sequence conservation reflects variation in functional constraint. Using this principle, we analyze noncoding sequence conservation in Drosophila for 40 loci with known or suspected cis-regulatory function encompassing >100 kb of DNA. We estimate the fraction of noncoding DNA conserved in both intergenic and intronic regions and describe the length distribution of ungapped conserved noncoding blocks. On average, 22%-26% of noncoding sequences surveyed are conserved in Drosophila, with median block length approximately 19 bp. We show that point substitution in conserved noncoding blocks exhibits transition bias as well as lineage effects in base composition, and occurs more than an order of magnitude more frequently than insertion/deletion (indel) substitution. Overall, patterns of noncoding DNA structure and evolution differ remarkably little between intergenic and intronic conserved blocks, suggesting that the effects of transcription per se contribute minimally to the constraints operating on these sequences. The results of this study have implications for the development of alignment and prediction algorithms specific to noncoding DNA, as well as for models of cis-regulatory DNA sequence evolution.
Elhassan, Nuha; Gebremeskel, Eyoab Iyasu; Elnour, Mohamed Ali; Isabirye, Dan; Okello, John; Hussien, Ayman; Kwiatksowski, Dominic; Hirbo, Jibril; Tishkoff, Sara; Ibrahim, Muntaser E
2014-01-01
Human genetic variation particularly in Africa is still poorly understood. This is despite a consensus on the large African effective population size compared to populations from other continents. Based on sequencing of the mitochondrial Cytochrome C Oxidase subunit II (MT-CO2), and genome wide microsatellite data we observe evidence suggesting the effective size (Ne) of humans to be larger than the current estimates, with a foci of increased genetic diversity in east Africa, and a population size of east Africans being at least 2-6 fold larger than other populations. Both phylogenetic and network analysis indicate that east Africans possess more ancestral lineages in comparison to various continental populations placing them at the root of the human evolutionary tree. Our results also affirm east Africa as the likely spot from which migration towards Asia has taken place. The study reflects the spectacular level of sequence variation within east Africans in comparison to the global sample, and appeals for further studies that may contribute towards filling the existing gaps in the database. The implication of these data to current genomic research, as well as the need to carry out defined studies of human genetic variation that includes more African populations; particularly east Africans is paramount.
Role of promoter DNA sequence variations on the binding of EGR1 transcription factor.
Mikles, David C; Schuchardt, Brett J; Bhat, Vikas; McDonald, Caleb B; Farooq, Amjad
2014-05-01
In response to a wide variety of stimuli such as growth factors and hormones, EGR1 transcription factor is rapidly induced and immediately exerts downstream effects central to the maintenance of cellular homeostasis. Herein, our biophysical analysis reveals that DNA sequence variations within the target gene promoters tightly modulate the energetics of binding of EGR1 and that nucleotide substitutions at certain positions are much more detrimental to EGR1-DNA interaction than others. Importantly, the reduction in binding affinity poorly correlates with the loss of enthalpy and gain of entropy-a trend indicative of a complex interplay between underlying thermodynamic factors due to the differential role of water solvent upon nucleotide substitution. We also provide a rationale for the physical basis of the effect of nucleotide substitutions on the EGR1-DNA interaction at atomic level. Taken together, our study bears important implications on understanding the molecular determinants of a key protein-DNA interaction at the cross-roads of human health and disease. Copyright © 2014 Elsevier Inc. All rights reserved.
Arlt, Martin F.; Ozdemir, Alev Cagla; Birkeland, Shanda R.; Lyons, Robert H.; Glover, Thomas W.; Wilson, Thomas E.
2011-01-01
Copy-number variants (CNVs) are a major source of genetic variation in human health and disease. Previous studies have implicated replication stress as a causative factor in CNV formation. However, existing data are technically limited in the quality of comparisons that can be made between human CNVs and experimentally induced variants. Here, we used two high-resolution strategies—single nucleotide polymorphism (SNP) arrays and mate-pair sequencing—to compare CNVs that occur constitutionally to those that arise following aphidicolin-induced DNA replication stress in the same human cells. Although the optimized methods provided complementary information, sequencing was more sensitive to small variants and provided superior structural descriptions. The majority of constitutional and all aphidicolin-induced CNVs appear to be formed via homology-independent mechanisms, while aphidicolin-induced CNVs were of a larger median size than constitutional events even when mate-pair data were considered. Aphidicolin thus appears to stimulate formation of CNVs that closely resemble human pathogenic CNVs and the subset of larger nonhomologous constitutional CNVs. PMID:21212237
Abdul-Wajid, Sarah; Veeman, Michael T; Chiba, Shota; Turner, Thomas L; Smith, William C
2014-05-01
Studies in tunicates such as Ciona have revealed new insights into the evolutionary origins of chordate development. Ciona populations are characterized by high levels of natural genetic variation, between 1 and 5%. This variation has provided abundant material for forward genetic studies. In the current study, we make use of deep sequencing and homozygosity mapping to map spontaneous mutations in outbred populations. With this method we have mapped two spontaneous developmental mutants. In Ciona intestinalis we mapped a short-tail mutation with strong phenotypic similarity to a previously identified mutant in the related species Ciona savignyi. Our bioinformatic approach mapped the mutation to a narrow interval containing a single mutated gene, α-laminin3,4,5, which is the gene previously implicated in C. savignyi. In addition, we mapped a novel genetic mutation disrupting neural tube closure in C. savignyi to a T-type Ca(2+) channel gene. The high efficiency and unprecedented mapping resolution of our study is a powerful advantage for developmental genetics in Ciona, and may find application in other outbred species.
Poon, Art F. Y; Kosakovsky Pond, Sergei L.; Bennett, Phil; Richman, Douglas D; Leigh Brown, Andrew J.; Frost, Simon D. W
2007-01-01
CD8+ cytotoxic T-lymphocytes (CTLs) perform a critical role in the immune control of viral infections, including those caused by human immunodeficiency virus type 1 (HIV-1) and hepatitis C virus (HCV). As a result, genetic variation at CTL epitopes is strongly influenced by host-specific selection for either escape from the immune response, or reversion due to the replicative costs of escape mutations in the absence of CTL recognition. Under strong CTL-mediated selection, codon positions within epitopes may immediately “toggle” in response to each host, such that genetic variation in the circulating virus population is shaped by rapid adaptation to immune variation in the host population. However, this hypothesis neglects the substantial genetic variation that accumulates in virus populations within hosts. Here, we evaluate this quantity for a large number of HIV-1– (n ≥ 3,000) and HCV-infected patients (n ≥ 2,600) by screening bulk RT-PCR sequences for sequencing “mixtures” (i.e., ambiguous nucleotides), which act as site-specific markers of genetic variation within each host. We find that nonsynonymous mixtures are abundant and significantly associated with codon positions under host-specific CTL selection, which should deplete within-host variation by driving the fixation of the favored variant. Using a simple model, we demonstrate that this apparently contradictory outcome can be explained by the transmission of unfavorable variants to new hosts before they are removed by selection, which occurs more frequently when selection and transmission occur on similar time scales. Consequently, the circulating virus population is shaped by the transmission rate and the disparity in selection intensities for escape or reversion as much as it is shaped by the immune diversity of the host population, with potentially serious implications for vaccine design. PMID:17397261
Marsden, Clare D; Ortega-Del Vecchyo, Diego; O'Brien, Dennis P; Taylor, Jeremy F; Ramirez, Oscar; Vilà, Carles; Marques-Bonet, Tomas; Schnabel, Robert D; Wayne, Robert K; Lohmueller, Kirk E
2016-01-05
Population bottlenecks, inbreeding, and artificial selection can all, in principle, influence levels of deleterious genetic variation. However, the relative importance of each of these effects on genome-wide patterns of deleterious variation remains controversial. Domestic and wild canids offer a powerful system to address the role of these factors in influencing deleterious variation because their history is dominated by known bottlenecks and intense artificial selection. Here, we assess genome-wide patterns of deleterious variation in 90 whole-genome sequences from breed dogs, village dogs, and gray wolves. We find that the ratio of amino acid changing heterozygosity to silent heterozygosity is higher in dogs than in wolves and, on average, dogs have 2-3% higher genetic load than gray wolves. Multiple lines of evidence indicate this pattern is driven by less efficient natural selection due to bottlenecks associated with domestication and breed formation, rather than recent inbreeding. Further, we find regions of the genome implicated in selective sweeps are enriched for amino acid changing variants and Mendelian disease genes. To our knowledge, these results provide the first quantitative estimates of the increased burden of deleterious variants directly associated with domestication and have important implications for selective breeding programs and the conservation of rare and endangered species. Specifically, they highlight the costs associated with selective breeding and question the practice favoring the breeding of individuals that best fit breed standards. Our results also suggest that maintaining a large population size, rather than just avoiding inbreeding, is a critical factor for preventing the accumulation of deleterious variants.
Marsden, Clare D.; Ortega-Del Vecchyo, Diego; O’Brien, Dennis P.; Taylor, Jeremy F.; Ramirez, Oscar; Vilà, Carles; Marques-Bonet, Tomas; Schnabel, Robert D.; Wayne, Robert K.; Lohmueller, Kirk E.
2016-01-01
Population bottlenecks, inbreeding, and artificial selection can all, in principle, influence levels of deleterious genetic variation. However, the relative importance of each of these effects on genome-wide patterns of deleterious variation remains controversial. Domestic and wild canids offer a powerful system to address the role of these factors in influencing deleterious variation because their history is dominated by known bottlenecks and intense artificial selection. Here, we assess genome-wide patterns of deleterious variation in 90 whole-genome sequences from breed dogs, village dogs, and gray wolves. We find that the ratio of amino acid changing heterozygosity to silent heterozygosity is higher in dogs than in wolves and, on average, dogs have 2–3% higher genetic load than gray wolves. Multiple lines of evidence indicate this pattern is driven by less efficient natural selection due to bottlenecks associated with domestication and breed formation, rather than recent inbreeding. Further, we find regions of the genome implicated in selective sweeps are enriched for amino acid changing variants and Mendelian disease genes. To our knowledge, these results provide the first quantitative estimates of the increased burden of deleterious variants directly associated with domestication and have important implications for selective breeding programs and the conservation of rare and endangered species. Specifically, they highlight the costs associated with selective breeding and question the practice favoring the breeding of individuals that best fit breed standards. Our results also suggest that maintaining a large population size, rather than just avoiding inbreeding, is a critical factor for preventing the accumulation of deleterious variants. PMID:26699508
Agunbiade, Tolulope A.; Coates, Brad S.; Datinon, Benjamin; Djouaka, Rousseau; Sun, Weilin; Tamò, Manuele; Pittendrigh, Barry R.
2014-01-01
Maruca vitrata Fabricius (Lepidoptera: Crambidae) is a polyphagous insect pest that feeds on a variety of leguminous plants in the tropics and subtropics. The contribution of host-associated genetic variation on population structure was investigated using analysis of mitochondrial cytochrome oxidase 1 (cox1) sequence and microsatellite marker data from M. vitrata collected from cultivated cowpea (Vigna unguiculata L. Walp.), and alternative host plants Pueraria phaseoloides (Roxb.) Benth. var. javanica (Benth.) Baker, Loncocarpus sericeus (Poir), and Tephrosia candida (Roxb.). Analyses of microsatellite data revealed a significant global FST estimate of 0.05 (P≤0.001). The program STRUCTURE estimated 2 genotypic clusters (co-ancestries) on the four host plants across 3 geographic locations, but little geographic variation was predicted among genotypes from different geographic locations using analysis of molecular variance (AMOVA; among group variation −0.68%) or F-statistics (F ST Loc = −0.01; P = 0.62). These results were corroborated by mitochondrial haplotype data (φSTLoc = 0.05; P = 0.92). In contrast, genotypes obtained from different host plants showed low but significant levels of genetic variation (F ST Host = 0.04; P = 0.01), which accounted for 4.08% of the total genetic variation, but was not congruent with mitochondrial haplotype analyses (φSTHost = 0.06; P = 0.27). Variation among host plants at a location and host plants among locations showed no consistent evidence for M. vitrata population subdivision. These results suggest that host plants do not significantly influence the genetic structure of M. vitrata, and this has implications for biocontrol agent releases as well as insecticide resistance management (IRM) for M. vitrata in West Africa. PMID:24647356
Blocks of limited haplotype diversity revealed by high-resolution scanning of human chromosome 21.
Patil, N; Berno, A J; Hinds, D A; Barrett, W A; Doshi, J M; Hacker, C R; Kautzer, C R; Lee, D H; Marjoribanks, C; McDonough, D P; Nguyen, B T; Norris, M C; Sheehan, J B; Shen, N; Stern, D; Stokowski, R P; Thomas, D J; Trulson, M O; Vyas, K R; Frazer, K A; Fodor, S P; Cox, D R
2001-11-23
Global patterns of human DNA sequence variation (haplotypes) defined by common single nucleotide polymorphisms (SNPs) have important implications for identifying disease associations and human traits. We have used high-density oligonucleotide arrays, in combination with somatic cell genetics, to identify a large fraction of all common human chromosome 21 SNPs and to directly observe the haplotype structure defined by these SNPs. This structure reveals blocks of limited haplotype diversity in which more than 80% of a global human sample can typically be characterized by only three common haplotypes.
Kelleher, Raymond J; Geigenmüller, Ute; Hovhannisyan, Hayk; Trautman, Edwin; Pinard, Robert; Rathmell, Barbara; Carpenter, Randall; Margulies, David
2012-01-01
Identification of common molecular pathways affected by genetic variation in autism is important for understanding disease pathogenesis and devising effective therapies. Here, we test the hypothesis that rare genetic variation in the metabotropic glutamate-receptor (mGluR) signaling pathway contributes to autism susceptibility. Single-nucleotide variants in genes encoding components of the mGluR signaling pathway were identified by high-throughput multiplex sequencing of pooled samples from 290 non-syndromic autism cases and 300 ethnically matched controls on two independent next-generation platforms. This analysis revealed significant enrichment of rare functional variants in the mGluR pathway in autism cases. Higher burdens of rare, potentially deleterious variants were identified in autism cases for three pathway genes previously implicated in syndromic autism spectrum disorder, TSC1, TSC2, and SHANK3, suggesting that genetic variation in these genes also contributes to risk for non-syndromic autism. In addition, our analysis identified HOMER1, which encodes a postsynaptic density-localized scaffolding protein that interacts with Shank3 to regulate mGluR activity, as a novel autism-risk gene. Rare, potentially deleterious HOMER1 variants identified uniquely in the autism population affected functionally important protein regions or regulatory sequences and co-segregated closely with autism among children of affected families. We also identified rare ASD-associated coding variants predicted to have damaging effects on components of the Ras/MAPK cascade. Collectively, these findings suggest that altered signaling downstream of mGluRs contributes to the pathogenesis of non-syndromic autism.
Hovhannisyan, Hayk; Trautman, Edwin; Pinard, Robert; Rathmell, Barbara; Carpenter, Randall; Margulies, David
2012-01-01
Identification of common molecular pathways affected by genetic variation in autism is important for understanding disease pathogenesis and devising effective therapies. Here, we test the hypothesis that rare genetic variation in the metabotropic glutamate-receptor (mGluR) signaling pathway contributes to autism susceptibility. Single-nucleotide variants in genes encoding components of the mGluR signaling pathway were identified by high-throughput multiplex sequencing of pooled samples from 290 non-syndromic autism cases and 300 ethnically matched controls on two independent next-generation platforms. This analysis revealed significant enrichment of rare functional variants in the mGluR pathway in autism cases. Higher burdens of rare, potentially deleterious variants were identified in autism cases for three pathway genes previously implicated in syndromic autism spectrum disorder, TSC1, TSC2, and SHANK3, suggesting that genetic variation in these genes also contributes to risk for non-syndromic autism. In addition, our analysis identified HOMER1, which encodes a postsynaptic density-localized scaffolding protein that interacts with Shank3 to regulate mGluR activity, as a novel autism-risk gene. Rare, potentially deleterious HOMER1 variants identified uniquely in the autism population affected functionally important protein regions or regulatory sequences and co-segregated closely with autism among children of affected families. We also identified rare ASD-associated coding variants predicted to have damaging effects on components of the Ras/MAPK cascade. Collectively, these findings suggest that altered signaling downstream of mGluRs contributes to the pathogenesis of non-syndromic autism. PMID:22558107
Frueh, Felix W
2010-05-01
The past decade of pharmacogenomics was driven by the sequencing of the human genome to create ever denser maps of genetic variations for studying the diversity across individuals. Today, genotyping technology is available at a fraction of the cost of what it was 10 years ago and many pharmacogenomic variations have been studied in detail. Still, we are only starting to gain an understanding of how pharmacogenomic-guided drug therapy affects clinical outcomes: real-world studies that demonstrate the clinical effectiveness and address the economic implications of pharmacogenomics are needed to help decide when and how to implement pharmacogenomics in clinical practice, how to regulate pharmacogenomic testing and how the healthcare system will integrate this new science into an environment of rapidly increasing cost.
Translation of Nutritional Genomics into Nutrition Practice: The Next Step.
Murgia, Chiara; Adamski, Melissa M
2017-04-06
Genetics is an important piece of every individual health puzzle. The completion of the Human Genome Project sequence has deeply changed the research of life sciences including nutrition. The analysis of the genome is already part of clinical care in oncology, pharmacology, infectious disease and, rare and undiagnosed diseases. The implications of genetic variations in shaping individual nutritional requirements have been recognised and conclusively proven, yet routine use of genetic information in nutrition and dietetics practice is still far from being implemented. This article sets out the path that needs to be taken to build a framework to translate gene-nutrient interaction studies into best-practice guidelines, providing tools that health professionals can use to understand whether genetic variation affects nutritional requirements in their daily clinical practice.
Challenges and opportunities for improving food quality and nutrition through plant biotechnology.
Francis, David; Finer, John J; Grotewold, Erich
2017-04-01
Plant biotechnology has been around since the advent of humankind, resulting in tremendous improvements in plant cultivation through crop domestication, breeding and selection. The emergence of transgenic approaches involving the introduction of defined DNA sequences into plants by humans has rapidly changed the surface of our planet by further expanding the gene pool used by plant breeders for plant improvement. Transgenic approaches in food plants have raised concerns on the merits, social implications, ecological risks and true benefits of plant biotechnology. The recently acquired ability to precisely edit plant genomes by modifying native genes without introducing new genetic material offers new opportunities to rapidly exploit natural variation, create new variation and incorporate changes with the goal to generate more productive and nutritious plants. Copyright © 2016 Elsevier Ltd. All rights reserved.
Translation of Nutritional Genomics into Nutrition Practice: The Next Step
Murgia, Chiara; Adamski, Melissa M.
2017-01-01
Genetics is an important piece of every individual health puzzle. The completion of the Human Genome Project sequence has deeply changed the research of life sciences including nutrition. The analysis of the genome is already part of clinical care in oncology, pharmacology, infectious disease and, rare and undiagnosed diseases. The implications of genetic variations in shaping individual nutritional requirements have been recognised and conclusively proven, yet routine use of genetic information in nutrition and dietetics practice is still far from being implemented. This article sets out the path that needs to be taken to build a framework to translate gene–nutrient interaction studies into best-practice guidelines, providing tools that health professionals can use to understand whether genetic variation affects nutritional requirements in their daily clinical practice. PMID:28383492
Evolutionary and biophysical relationships among the papillomavirus E2 proteins.
Blakaj, Dukagjin M; Fernandez-Fuentes, Narcis; Chen, Zigui; Hegde, Rashmi; Fiser, Andras; Burk, Robert D; Brenowitz, Michael
2009-01-01
Infection by human papillomavirus (HPV) may result in clinical conditions ranging from benign warts to invasive cancer. The HPV E2 protein represses oncoprotein transcription and is required for viral replication. HPV E2 binds to palindromic DNA sequences of highly conserved four base pair sequences flanking an identical length variable 'spacer'. E2 proteins directly contact the conserved but not the spacer DNA. Variation in naturally occurring spacer sequences results in differential protein affinity that is dependent on their sensitivity to the spacer DNA's unique conformational and/or dynamic properties. This article explores the biophysical character of this core viral protein with the goal of identifying characteristics that associated with risk of virally caused malignancy. The amino acid sequence, 3d structure and electrostatic features of the E2 protein DNA binding domain are highly conserved; specific interactions with DNA binding sites have also been conserved. In contrast, the E2 protein's transactivation domain does not have extensive surfaces of highly conserved residues. Rather, regions of high conservation are localized to small surface patches. Implications to cancer biology are discussed.
Piton, Amélie; Redin, Claire; Mandel, Jean-Louis
2013-01-01
Because of the unbalanced sex ratio (1.3–1.4 to 1) observed in intellectual disability (ID) and the identification of large ID-affected families showing X-linked segregation, much attention has been focused on the genetics of X-linked ID (XLID). Mutations causing monogenic XLID have now been reported in over 100 genes, most of which are commonly included in XLID diagnostic gene panels. Nonetheless, the boundary between true mutations and rare non-disease-causing variants often remains elusive. The sequencing of a large number of control X chromosomes, required for avoiding false-positive results, was not systematically possible in the past. Such information is now available thanks to large-scale sequencing projects such as the National Heart, Lung, and Blood (NHLBI) Exome Sequencing Project, which provides variation information on 10,563 X chromosomes from the general population. We used this NHLBI cohort to systematically reassess the implication of 106 genes proposed to be involved in monogenic forms of XLID. We particularly question the implication in XLID of ten of them (AGTR2, MAGT1, ZNF674, SRPX2, ATP6AP2, ARHGEF6, NXF5, ZCCHC12, ZNF41, and ZNF81), in which truncating variants or previously published mutations are observed at a relatively high frequency within this cohort. We also highlight 15 other genes (CCDC22, CLIC2, CNKSR2, FRMPD4, HCFC1, IGBP1, KIAA2022, KLF8, MAOA, NAA10, NLGN3, RPL10, SHROOM4, ZDHHC15, and ZNF261) for which replication studies are warranted. We propose that similar reassessment of reported mutations (and genes) with the use of data from large-scale human exome sequencing would be relevant for a wide range of other genetic diseases. PMID:23871722
Hierarchy and extremes in selections from pools of randomized proteins
Boyer, Sébastien; Biswas, Dipanwita; Kumar Soshee, Ananda; Scaramozzino, Natale; Nizak, Clément; Rivoire, Olivier
2016-01-01
Variation and selection are the core principles of Darwinian evolution, but quantitatively relating the diversity of a population to its capacity to respond to selection is challenging. Here, we examine this problem at a molecular level in the context of populations of partially randomized proteins selected for binding to well-defined targets. We built several minimal protein libraries, screened them in vitro by phage display, and analyzed their response to selection by high-throughput sequencing. A statistical analysis of the results reveals two main findings. First, libraries with the same sequence diversity but built around different “frameworks” typically have vastly different responses; second, the distribution of responses of the best binders in a library follows a simple scaling law. We show how an elementary probabilistic model based on extreme value theory rationalizes the latter finding. Our results have implications for designing synthetic protein libraries, estimating the density of functional biomolecules in sequence space, characterizing diversity in natural populations, and experimentally investigating evolvability (i.e., the potential for future evolution). PMID:26969726
Hierarchy and extremes in selections from pools of randomized proteins.
Boyer, Sébastien; Biswas, Dipanwita; Kumar Soshee, Ananda; Scaramozzino, Natale; Nizak, Clément; Rivoire, Olivier
2016-03-29
Variation and selection are the core principles of Darwinian evolution, but quantitatively relating the diversity of a population to its capacity to respond to selection is challenging. Here, we examine this problem at a molecular level in the context of populations of partially randomized proteins selected for binding to well-defined targets. We built several minimal protein libraries, screened them in vitro by phage display, and analyzed their response to selection by high-throughput sequencing. A statistical analysis of the results reveals two main findings. First, libraries with the same sequence diversity but built around different "frameworks" typically have vastly different responses; second, the distribution of responses of the best binders in a library follows a simple scaling law. We show how an elementary probabilistic model based on extreme value theory rationalizes the latter finding. Our results have implications for designing synthetic protein libraries, estimating the density of functional biomolecules in sequence space, characterizing diversity in natural populations, and experimentally investigating evolvability (i.e., the potential for future evolution).
Arko, B; Prezelj, J; Komel, R; Kocijancic, A; Hudler, P; Marc, J
2002-09-01
Osteoprotegerin (OPG) is a recently discovered member of the TNF receptor superfamily that acts as an important paracrine regulator of bone remodeling. OPG knockout mice develop severe osteoporosis, whereas administration of OPG can prevent ovariectomy-induced bone loss. These findings implicate a role for OPG in the development of osteoporosis. In the present study, we screened the OPG gene promoter for sequence variations and examined their association with bone mineral density (BMD) in 103 osteoporotic postmenopausal women. Single-strand conformation polymorphism analysis followed by DNA sequencing revealed a presence of four nucleotide substitutions: 209 G-->A, 245 T-->G, 889 C-->T, and 950 T-->C. The frequencies of genotypes were as follows: GG (89.3%), GA (10.7%) for 209 G-->A polymorphism; TT (89.3%), TG (10.7%) for 245 T-->G polymorphism; and TT (25.2%), TC (53.4%), CC (21.4%) for 950 T-->C polymorphism. Substitution 889 C-->T was found in only two patients. Statistically significant association of genotypes with BMD at the lumbar spine (P = 0.005) was observed for 209 G-->A and 245 T-->G polymorphisms. Haplotype GATG was associated with lower BMD as compared with GGTT haplotype. Our results suggest that 209 G-->A and 245 T-->G polymorphisms in the OPG gene promoter may contribute to the genetic regulation of BMD.
Population Genomics of Reduced Vancomycin Susceptibility in Staphylococcus aureus
Rishishwar, Lavanya; Kraft, Colleen S.
2016-01-01
ABSTRACT The increased prevalence of vancomycin-intermediate Staphylococcus aureus (VISA) is an emerging health care threat. Genome-based comparative methods hold great promise to uncover the genetic basis of the VISA phenotype, which remains obscure. S. aureus isolates were collected from a single individual that presented with recurrent staphylococcal bacteremia at three time points, and the isolates showed successively reduced levels of vancomycin susceptibility. A population genomic approach was taken to compare patient S. aureus isolates with decreasing vancomycin susceptibility across the three time points. To do this, patient isolates were sequenced to high coverage (~500×), and sequence reads were used to model site-specific allelic variation within and between isolate populations. Population genetic methods were then applied to evaluate the overall levels of variation across the three time points and to identify individual variants that show anomalous levels of allelic change between populations. A successive reduction in the overall levels of population genomic variation was observed across the three time points, consistent with a population bottleneck resulting from antibiotic treatment. Despite this overall reduction in variation, a number of individual mutations were swept to high frequency in the VISA population. These mutations were implicated as potentially involved in the VISA phenotype and interrogated with respect to their functional roles. This approach allowed us to identify a number of mutations previously implicated in VISA along with allelic changes within a novel class of genes, encoding LPXTG motif-containing cell-wall-anchoring proteins, which shed light on a novel mechanistic aspect of vancomycin resistance. IMPORTANCE The emergence and spread of antibiotic resistance among bacterial pathogens are two of the gravest threats to public health facing the world today. We report the development and application of a novel population genomic technique aimed at uncovering the evolutionary dynamics and genetic determinants of antibiotic resistance in Staphylococcus aureus. This method was applied to S. aureus cultures isolated from a single patient who showed decreased susceptibility to the vancomycin antibiotic over time. Our approach relies on the increased resolution afforded by next-generation genome-sequencing technology, and it allowed us to discover a number of S. aureus mutations, in both known and novel gene targets, which appear to have evolved under adaptive pressure to evade vancomycin mechanisms of action. The approach we lay out in this work can be applied to resistance to any number of antibiotics across numerous species of bacterial pathogens. PMID:27446992
A comprehensive analysis of rare genetic variation in amyotrophic lateral sclerosis in the UK.
Morgan, Sarah; Shatunov, Aleksey; Sproviero, William; Jones, Ashley R; Shoai, Maryam; Hughes, Deborah; Al Khleifat, Ahmad; Malaspina, Andrea; Morrison, Karen E; Shaw, Pamela J; Shaw, Christopher E; Sidle, Katie; Orrell, Richard W; Fratta, Pietro; Hardy, John; Pittman, Alan; Al-Chalabi, Ammar
2017-06-01
Amyotrophic lateral sclerosis is a progressive neurodegenerative disease of motor neurons. About 25 genes have been verified as relevant to the disease process, with rare and common variation implicated. We used next generation sequencing and repeat sizing to comprehensively assay genetic variation in a panel of known amyotrophic lateral sclerosis genes in 1126 patient samples and 613 controls. About 10% of patients were predicted to carry a pathological expansion of the C9orf72 gene. We found an increased burden of rare variants in patients within the untranslated regions of known disease-causing genes, driven by SOD1, TARDBP, FUS, VCP, OPTN and UBQLN2. We found 11 patients (1%) carried more than one pathogenic variant (P = 0.001) consistent with an oligogenic basis of amyotrophic lateral sclerosis. These findings show that the genetic architecture of amyotrophic lateral sclerosis is complex and that variation in the regulatory regions of associated genes may be important in disease pathogenesis. © The Author (2017). Published by Oxford University Press on behalf of the Guarantors of Brain.
How important is interannual variability in the climatic interpretation of moraine sequences?
NASA Astrophysics Data System (ADS)
Leonard, E. M.; Laabs, B. J. C.; Plummer, M. A.
2017-12-01
Mountain glaciers respond to both long-term climate and interannual forcing. Anderson et al. (2014) pointed out that kilometer-scale fluctuations in glacier length may result from interannual variability in temperature and precipitation given a "steady" climate with no long-term trends in mean or variability of temperature and precipitation. They cautioned that use of outermost moraines from the Last Glacial Maximum (LGM) as indicators of LGM climate will, because of the role of interannual forcing, result in overestimation of the magnitude of long-term temperature depression and/or precipitation enhancement. Here we assess the implications of these ideas, by examining the effect of interannual variability on glacier length and inferred magnitude of LGM climate change from present under both an assumed steady LGM climate and an LGM climate with low-magnitude, long-period variation in summer temperature and annual precipitation. We employ both the original 1-stage linear glacier model (Roe and O'Neal, 2009) used by Anderson et al. (2014) and a newer 3-stage linear model (Roe and Baker, 2014). We apply the models to two reconstructed LGM glaciers in the Colorado Sangre de Cristo Mountains. Three-stage-model results indicate that, absent long-term variations through a 7500-year-long LGM, interannual variability would result in overestimation of mean LGM temperature depression from the outermost moraine of 0.2-0.6°C. If small long-term cyclic variations of temperature (±0.5°C) and precipitation (±5%) are introduced, the overestimation of LGM temperature depression reduces to less than 0.4°C, and if slightly greater long-term variation (±1.0°C and ±10% precipitation) is introduced, the magnitude of overestimation is 0.3°C or less. Interannual variability may produce a moraine sequence that differs from the sequence that would be expected were glacier length forced only by long-term climate. With small amplitude (±0.5°C and ±5% precipitation) long-term variation, the moraine sequence expected if forced by a combination of interannual variability and long-term climate differs from that expected based on long-term climate forcing alone in 38% of model runs. With the larger amplitude long-term forcing (±1.0°C and ±10% precipitation) this difference occurs in 20% of model runs.
Zheng, Deyou
2008-01-01
Background Sequencing and annotation of several mammalian genomes have revealed that segmental duplications are a common architectural feature of primate genomes; in fact, about 5% of the human genome is composed of large blocks of interspersed segmental duplications. These segmental duplications have been implicated in genomic copy-number variation, gene novelty, and various genomic disorders. However, the molecular processes involved in the evolution and regulation of duplicated sequences remain largely unexplored. Results In this study, the profile of about 20 histone modifications within human segmental duplications was characterized using high-resolution, genome-wide data derived from a ChIP-Seq study. The analysis demonstrates that derivative loci of segmental duplications often differ significantly from the original with respect to many histone methylations. Further investigation showed that genes are present three times more frequently in the original than in the derivative, whereas pseudogenes exhibit the opposite trend. These asymmetries tend to increase with the age of segmental duplications. The uneven distribution of genes and pseudogenes does not, however, fully account for the asymmetry in the profile of histone modifications. Conclusion The first systematic analysis of histone modifications between segmental duplications demonstrates that two seemingly 'identical' genomic copies are distinct in their epigenomic properties. Results here suggest that local chromatin environments may be implicated in the discrimination of derived copies of segmental duplications from their originals, leading to a biased pseudogenization of the new duplicates. The data also indicate that further exploration of the interactions between histone modification and sequence degeneration is necessary in order to understand the divergence of duplicated sequences. PMID:18598352
NASA Astrophysics Data System (ADS)
Pinton, Annamaria; Giordano, Guido; Speranza, Fabio; Þórðarson, Þorvaldur
2018-01-01
The impact of Holocene eruptive events from hot spots like Iceland may have had significant global implications; thus, dating and knowledge of past eruptions chronology is important. However, at high-latitude volcanic islands, the paucity of soils severely limits 14C dating, while the poor K content of basalts strongly restricts the use of K/Ar and Ar/Ar methods. Even tephrochronology, based on 14C age determinations, refers to layers that rarely lie directly above lava flows to be dated. We report on the paleomagnetic dating of 25 sites from the Reykjanes Peninsula and the Tungnaá lava sequence of Iceland. The gathered paleomagnetic directions were compared with the available reference paleosecular variation curves of the Earth magnetic field to obtain the possible emplacement age intervals. To test the method's validity, we sampled the precisely dated Laki (1783-1784 AD) and Eldgjà (934-938 AD) lavas. The age windows obtained for these events encompass the true flow ages. For sites from the Reykjanes peninsula and the Tugnaá lava sequence, we derived multiple possible eruption events and ages. In the Reykjanes peninsula, we propose an older emplacement age (immediately following the 870 AD Iceland Settlement age) for Ogmundarhraun and Kapelluhraun lava fields. For pre-historical (older than the settlement age) Tugnaá eruptions, the method has a dating precision of 300-400 years which allows an increase of the detail in the chronostratigraphy and distribution of lavas in the Tugnaá sequence.
Miller, K.G.; Mountain, Gregory S.; Browning, J.V.; Kominz, M.; Sugarman, P.J.; Christie-Blick, N.; Katz, M.E.; Wright, J.D.
1998-01-01
The New Jersey Sea Level Transect was designed to evaluate the relationships among global sea level (eustatic) change, unconformity-bounded sequences, and variations in subsidence, sediment supply, and climate on a passive continental margin. By sampling and dating Cenozoic strata from coastal plain and continental slope locations, we show that sequence boundaries correlate (within ??0.5 myr) regionally (onshore-offshore) and interregionally (New Jersey-Alabama-Bahamas), implicating a global cause. Sequence boundaries correlate with ??18O increases for at least the past 42 myr, consistent with an ice volume (glacioeustatic) control, although a causal relationship is not required because of uncertainties in ages and correlations. Evidence for a causal connection is provided by preliminary Miocene data from slope Site 904 that directly link ??18O increases with sequence boundaries. We conclude that variation in the size of ice sheets has been a primary control on the formation of sequence boundaries since ~42 Ma. We speculate that prior to this, the growth and decay of small ice sheets caused small-amplitude sea level changes (<20 m) in this supposedly ice-free world because Eocene sequence boundaries also appear to correlate with minor ??18O increases. Subsidence estimates (backstripping) indicate amplitudes of short-term (million-year scale) lowerings that are consistent with estimates derived from ??18O studies (25-50 m in the Oligocene-middle Miocene and 10-20 m in the Eocene) and a long-term lowering of 150-200 m over the past 65 myr, consistent with estimates derived from volume changes on mid-ocean ridges. Although our results are consistent with the general number and timing of Paleocene to middle Miocene sequences published by workers at Exxon Production Research Company, our estimates of sea level amplitudes are substantially lower than theirs. Lithofacies patterns within sequences follow repetitive, predictable patterns: (1) coastal plain sequences consist of basal transgressive sands overlain by regressive highstand silts and quartz sands; and (2) although slope lithofacies variations are subdued, reworked sediments constitute lowstand deposits, causing the strongest, most extensive seismic reflections. Despite a primary eustatic control on sequence boundaries, New Jersey sequences were also influenced by changes in tectonics, sediment supply, and climate. During the early to middle Eocene, low siliciclastic and high pelagic input associated with warm climates resulted in widespread carbonate deposition and thin sequences. Late middle Eocene and earliest Oligocene cooling events curtailed carbonate deposition in the coastal plain and slope, respectively, resulting in a switch to siliciclastic sedimentation. In onshore areas, Oligocene sequences are thin owing to low siliciclastic and pelagic input, and their distribution is patchy, reflecting migration or progradation of depocenters; in contrast, Miocene onshore sequences are thicker, reflecting increased sediment supply, and they are more complete downdip owing to simple tectonics. We conclude that the New Jersey margin provides a natural laboratory for unraveling complex interactions of eustasy, tectonics, changes in sediment supply, and climate change.
Wang, Baosheng; Khalili Mahani, Marjan; Ng, Wei Lun; Kusumi, Junko; Phi, Hai Hong; Inomata, Nobuyuki; Wang, Xiao-Ru; Szmidt, Alfred E
2014-01-01
Pinus krempfii Lecomte is a morphologically and ecologically unique pine, endemic to Vietnam. It is regarded as vulnerable species with distribution limited to just two provinces: Khanh Hoa and Lam Dong. Although a few phylogenetic studies have included this species, almost nothing is known about its genetic features. In particular, there are no studies addressing the levels and patterns of genetic variation in natural populations of P. krempfii. In this study, we sampled 57 individuals from six natural populations of P. krempfii and analyzed their sequence variation in ten nuclear gene regions (approximately 9 kb) and 14 mitochondrial (mt) DNA regions (approximately 10 kb). We also analyzed variation at seven chloroplast (cp) microsatellite (SSR) loci. We found very low haplotype and nucleotide diversity at nuclear loci compared with other pine species. Furthermore, all investigated populations were monomorphic across all mitochondrial DNA (mtDNA) regions included in our study, which are polymorphic in other pine species. Population differentiation at nuclear loci was low (5.2%) but significant. However, structure analysis of nuclear loci did not detect genetically differentiated groups of populations. Approximate Bayesian computation (ABC) using nuclear sequence data and mismatch distribution analysis for cpSSR loci suggested recent expansion of the species. The implications of these findings for the management and conservation of P. krempfii genetic resources were discussed. PMID:25360263
The heterogeneity of human CD127(+) innate lymphoid cells revealed by single-cell RNA sequencing.
Björklund, Åsa K; Forkel, Marianne; Picelli, Simone; Konya, Viktoria; Theorell, Jakob; Friberg, Danielle; Sandberg, Rickard; Mjösberg, Jenny
2016-04-01
Innate lymphoid cells (ILCs) are increasingly appreciated as important participants in homeostasis and inflammation. Substantial plasticity and heterogeneity among ILC populations have been reported. Here we have delineated the heterogeneity of human ILCs through single-cell RNA sequencing of several hundreds of individual tonsil CD127(+) ILCs and natural killer (NK) cells. Unbiased transcriptional clustering revealed four distinct populations, corresponding to ILC1 cells, ILC2 cells, ILC3 cells and NK cells, with their respective transcriptomes recapitulating known as well as unknown transcriptional profiles. The single-cell resolution additionally divulged three transcriptionally and functionally diverse subpopulations of ILC3 cells. Our systematic comparison of single-cell transcriptional variation within and between ILC populations provides new insight into ILC biology during homeostasis, with additional implications for dysregulation of the immune system.
Function and regulation of AUTS2, a gene implicated in autism and human evolution.
Oksenberg, Nir; Stevison, Laurie; Wall, Jeffrey D; Ahituv, Nadav
2013-01-01
Nucleotide changes in the AUTS2 locus, some of which affect only noncoding regions, are associated with autism and other neurological disorders, including attention deficit hyperactivity disorder, epilepsy, dyslexia, motor delay, language delay, visual impairment, microcephaly, and alcohol consumption. In addition, AUTS2 contains the most significantly accelerated genomic region differentiating humans from Neanderthals, which is primarily composed of noncoding variants. However, the function and regulation of this gene remain largely unknown. To characterize auts2 function, we knocked it down in zebrafish, leading to a smaller head size, neuronal reduction, and decreased mobility. To characterize AUTS2 regulatory elements, we tested sequences for enhancer activity in zebrafish and mice. We identified 23 functional zebrafish enhancers, 10 of which were active in the brain. Our mouse enhancer assays characterized three mouse brain enhancers that overlap an ASD-associated deletion and four mouse enhancers that reside in regions implicated in human evolution, two of which are active in the brain. Combined, our results show that AUTS2 is important for neurodevelopment and expose candidate enhancer sequences in which nucleotide variation could lead to neurological disease and human-specific traits.
Mutsaerts, Henri J M M; van Osch, Matthias J P; Zelaya, Fernando O; Wang, Danny J J; Nordhøy, Wibeke; Wang, Yi; Wastling, Stephen; Fernandez-Seara, Maria A; Petersen, E T; Pizzini, Francesca B; Fallatah, Sameeha; Hendrikse, Jeroen; Geier, Oliver; Günther, Matthias; Golay, Xavier; Nederveen, Aart J; Bjørnerud, Atle; Groote, Inge R
2015-06-01
A main obstacle that impedes standardized clinical and research applications of arterial spin labeling (ASL), is the substantial differences between the commercial implementations of ASL from major MRI vendors. In this study, we compare a single identical 2D gradient-echo EPI pseudo-continuous ASL (PCASL) sequence implemented on 3T scanners from three vendors (General Electric Healthcare, Philips Healthcare and Siemens Healthcare) within the same center and with the same subjects. Fourteen healthy volunteers (50% male, age 26.4±4.7years) were scanned twice on each scanner in an interleaved manner within 3h. Because of differences in gradient and coil specifications, two separate studies were performed with slightly different sequence parameters, with one scanner used across both studies for comparison. Reproducibility was evaluated by means of quantitative cerebral blood flow (CBF) agreement and inter-session variation, both on a region-of-interest (ROI) and voxel level. In addition, a qualitative similarity comparison of the CBF maps was performed by three experienced neuro-radiologists. There were no CBF differences between vendors in study 1 (p>0.1), but there were CBF differences of 2-19% between vendors in study 2 (p<0.001 in most gray matter ROIs) and 10-22% difference in CBF values obtained with the same vendor between studies (p<0.001 in most gray matter ROIs). The inter-vendor inter-session variation was not significantly larger than the intra-vendor variation in all (p>0.1) but one of the ROIs (p<0.001). This study demonstrates the possibility to acquire comparable cerebral CBF maps on scanners of different vendors. Small differences in sequence parameters can have a larger effect on the reproducibility of ASL than hardware or software differences between vendors. These results suggest that researchers should strive to employ identical labeling and readout strategies in multi-center ASL studies. Copyright © 2015 Elsevier Inc. All rights reserved.
Clinical Interpretation and Implications of Whole-Genome Sequencing
Dewey, Frederick E.; Grove, Megan E.; Pan, Cuiping; Goldstein, Benjamin A.; Bernstein, Jonathan A.; Chaib, Hassan; Merker, Jason D.; Goldfeder, Rachel L.; Enns, Gregory M.; David, Sean P.; Pakdaman, Neda; Ormond, Kelly E.; Caleshu, Colleen; Kingham, Kerry; Klein, Teri E.; Whirl-Carrillo, Michelle; Sakamoto, Kenneth; Wheeler, Matthew T.; Butte, Atul J.; Ford, James M.; Boxer, Linda; Ioannidis, John P. A.; Yeung, Alan C.; Altman, Russ B.; Assimes, Themistocles L.; Snyder, Michael; Ashley, Euan A.; Quertermous, Thomas
2014-01-01
IMPORTANCE Whole-genome sequencing (WGS) is increasingly applied in clinical medicine and is expected to uncover clinically significant findings regardless of sequencing indication. OBJECTIVES To examine coverage and concordance of clinically relevant genetic variation provided by WGS technologies; to quantitate inherited disease risk and pharmacogenomic findings in WGS data and resources required for their discovery and interpretation; and to evaluate clinical action prompted by WGS findings. DESIGN, SETTING, AND PARTICIPANTS An exploratory study of 12 adult participants recruited at Stanford University Medical Center who underwent WGS between November 2011 and March 2012. A multidisciplinary team reviewed all potentially reportable genetic findings. Five physicians proposed initial clinical follow-up based on the genetic findings. MAIN OUTCOMES AND MEASURES Genome coverage and sequencing platform concordance in different categories of genetic disease risk, person-hours spent curating candidate disease-risk variants, interpretation agreement between trained curators and disease genetics databases, burden of inherited disease risk and pharmacogenomic findings, and burden and interrater agreement of proposed clinical follow-up. RESULTS Depending on sequencing platform, 10% to 19% of inherited disease genes were not covered to accepted standards for single nucleotide variant discovery. Genotype concordance was high for previously described single nucleotide genetic variants (99%-100%) but low for small insertion/deletion variants (53%-59%). Curation of 90 to 127 genetic variants in each participant required a median of 54 minutes (range, 5-223 minutes) per genetic variant, resulted in moderate classification agreement between professionals (Gross κ, 0.52; 95%CI, 0.40-0.64), and reclassified 69%of genetic variants cataloged as disease causing in mutation databases to variants of uncertain or lesser significance. Two to 6 personal disease-risk findings were discovered in each participant, including 1 frameshift deletion in the BRCA1 gene implicated in hereditary breast and ovarian cancer. Physician review of sequencing findings prompted consideration of a median of 1 to 3 initial diagnostic tests and referrals per participant, with fair interrater agreement about the suitability of WGS findings for clinical follow-up (Fleiss κ, 0.24; P < 001). CONCLUSIONS AND RELEVANCE In this exploratory study of 12 volunteer adults, the use of WGS was associated with incomplete coverage of inherited disease genes, low reproducibility of detection of genetic variation with the highest potential clinical effects, and uncertainty about clinically reportable findings. In certain cases, WGS will identify clinically actionable genetic variants warranting early medical intervention. These issues should be considered when determining the role of WGS in clinical medicine. PMID:24618965
Argüello-García, Raúl; Cruz-Soto, Maricela; Romero-Montoya, Lydia; Ortega-Pierres, Guadalupe
2009-12-01
The susceptibility of Giardia duodenalis trophozoites exposed in vitro to sublethal concentrations of metronidazole (MTZ) and albendazole (ABZ) may exhibit inter-culture (variability) and intra-culture (variation) differences in drug susceptibility. It was previously reported that MTZ-resistant trophozoites may display changes in pyruvate:ferredoxin oxidoreductase (PFOR) expression while changes at the beta-tubulin molecule are apparently absent in ABZ-resistant cultures. To assess the levels of gene expression of these molecules, we obtained cloned cultures growing at concentrations up to 23 microM MTZ (WBRM23) and up to 8muM ABZ (WBRA8) and gene sequence and expression of pfor and beta-tubulin loci were compared with these of drug-susceptible clone WB1. Neither the pfor nor the beta-tubulin genes showed changes at sequence level but the MTZ-resistant clones WBRM21 and WBRM23 showed up-regulation of the pfor RNA using the gdh gene as reference. By using WB1 and WBRA8 clones in representational difference analyses of gene expression (RDA) an insert referred to as ARR-VSP was selected and sequenced. It showed the highest homology to one VSP molecule in the Giardia Genome Database (orf GL50803_101765). This isogene was up-regulated in five ABZ-resistant clones and the clone WBRA8 exhibited the highest RNA expression level. When successive progenies of clones WB1, WBRM23 and WBRA8 were analyzed in Northern blot assays to detect pfor and ARR-VSP RNAs respectively, the expression patterns showed variation for both genes but it was much lower in the clone WBRA8. These results suggest that G. duodenalis cultures either susceptible or resistant to MTZ and ABZ may display variability and variation at RNA expression levels albeit these were more marked in the MTZ-resistant parasites. These data might have further implications defining major mechanisms involved in drug resistance of Giardia.
Mutations in the LHX2 gene are not a frequent cause of micro/anophthalmia
Desmaison, Annaïck; Vigouroux, Adeline; Rieubland, Claudine; Peres, Christine; Calvas, Patrick
2010-01-01
Purpose Microphthalmia and anophthalmia are at the severe end of the spectrum of abnormalities in ocular development. A few genes (orthodenticle homeobox 2 [OTX2], retina and anterior neural fold homeobox [RAX], SRY-box 2 [SOX2], CEH10 homeodomain-containing homolog [CHX10], and growth differentiation factor 6 [GDF6]) have been implicated mainly in isolated micro/anophthalmia but causative mutations of these genes explain less than a quarter of these developmental defects. The essential role of the LIM homeobox 2 (LHX2) transcription factor in early eye development has recently been documented. We postulated that mutations in this gene could lead to micro/anophthalmia, and thus performed molecular screening of its sequence in patients having micro/anophthalmia. Methods Seventy patients having non-syndromic forms of colobomatous microphthalmia (n=25), isolated microphthalmia (n=18), or anophthalmia (n=17), and syndromic forms of micro/anophthalmia (n=10) were included in this study after negative molecular screening for OTX2, RAX, SOX2, and CHX10 mutations. Mutation screening of LHX2 was performed by direct sequencing of the coding sequences and intron/exon boundaries. Results Two heterozygous variants of unknown significance (c.128C>G [p.Pro43Arg]; c.776C>A [p.Pro259Gln]) were identified in LHX2 among the 70 patients. These variations were not identified in a panel of 100 control patients of mixed origins. The variation c.776C>A (p.Pro259Gln) was considered as non pathogenic by in silico analysis, while the variation c.128C>G (p.Pro43Arg) considered as deleterious by in silico analysis and was inherited from the asymptomatic father. Conclusions Mutations in LHX2 do not represent a frequent cause of micro/anophthalmia. PMID:21203406
Mutations in the LHX2 gene are not a frequent cause of micro/anophthalmia.
Desmaison, Annaïck; Vigouroux, Adeline; Rieubland, Claudine; Peres, Christine; Calvas, Patrick; Chassaing, Nicolas
2010-12-18
Microphthalmia and anophthalmia are at the severe end of the spectrum of abnormalities in ocular development. A few genes (orthodenticle homeobox 2 [OTX2], retina and anterior neural fold homeobox [RAX], SRY-box 2 [SOX2], CEH10 homeodomain-containing homolog [CHX10], and growth differentiation factor 6 [GDF6]) have been implicated mainly in isolated micro/anophthalmia but causative mutations of these genes explain less than a quarter of these developmental defects. The essential role of the LIM homeobox 2 (LHX2) transcription factor in early eye development has recently been documented. We postulated that mutations in this gene could lead to micro/anophthalmia, and thus performed molecular screening of its sequence in patients having micro/anophthalmia. Seventy patients having non-syndromic forms of colobomatous microphthalmia (n=25), isolated microphthalmia (n=18), or anophthalmia (n=17), and syndromic forms of micro/anophthalmia (n=10) were included in this study after negative molecular screening for OTX2, RAX, SOX2, and CHX10 mutations. Mutation screening of LHX2 was performed by direct sequencing of the coding sequences and intron/exon boundaries. Two heterozygous variants of unknown significance (c.128C>G [p.Pro43Arg]; c.776C>A [p.Pro259Gln]) were identified in LHX2 among the 70 patients. These variations were not identified in a panel of 100 control patients of mixed origins. The variation c.776C>A (p.Pro259Gln) was considered as non pathogenic by in silico analysis, while the variation c.128C>G (p.Pro43Arg) considered as deleterious by in silico analysis and was inherited from the asymptomatic father. Mutations in LHX2 do not represent a frequent cause of micro/anophthalmia.
Garrido-Martín, Diego; Pazos, Florencio
2018-02-27
The exponential accumulation of new sequences in public databases is expected to improve the performance of all the approaches for predicting protein structural and functional features. Nevertheless, this was never assessed or quantified for some widely used methodologies, such as those aimed at detecting functional sites and functional subfamilies in protein multiple sequence alignments. Using raw protein sequences as only input, these approaches can detect fully conserved positions, as well as those with a family-dependent conservation pattern. Both types of residues are routinely used as predictors of functional sites and, consequently, understanding how the sequence content of the databases affects them is relevant and timely. In this work we evaluate how the growth and change with time in the content of sequence databases affect five sequence-based approaches for detecting functional sites and subfamilies. We do that by recreating historical versions of the multiple sequence alignments that would have been obtained in the past based on the database contents at different time points, covering a period of 20 years. Applying the methods to these historical alignments allows quantifying the temporal variation in their performance. Our results show that the number of families to which these methods can be applied sharply increases with time, while their ability to detect potentially functional residues remains almost constant. These results are informative for the methods' developers and final users, and may have implications in the design of new sequencing initiatives.
Bellissimo, Daniel B; Christopherson, Pamela A; Flood, Veronica H; Gill, Joan Cox; Friedman, Kenneth D; Haberichter, Sandra L; Shapiro, Amy D; Abshire, Thomas C; Leissinger, Cindy; Hoots, W Keith; Lusher, Jeanne M; Ragni, Margaret V; Montgomery, Robert R
2012-03-01
Diagnosis and classification of VWD is aided by molecular analysis of the VWF gene. Because VWF polymorphisms have not been fully characterized, we performed VWF laboratory testing and gene sequencing of 184 healthy controls with a negative bleeding history. The controls included 66 (35.9%) African Americans (AAs). We identified 21 new sequence variations, 13 (62%) of which occurred exclusively in AAs and 2 (G967D, T2666M) that were found in 10%-15% of the AA samples, suggesting they are polymorphisms. We identified 14 sequence variations reported previously as VWF mutations, the majority of which were type 1 mutations. These controls had VWF Ag levels within the normal range, suggesting that these sequence variations might not always reduce plasma VWF levels. Eleven mutations were found in AAs, and the frequency of M740I, H817Q, and R2185Q was 15%-18%. Ten AA controls had the 2N mutation H817Q; 1 was homozygous. The average factor VIII level in this group was 99 IU/dL, suggesting that this variation may confer little or no clinical symptoms. This study emphasizes the importance of sequencing healthy controls to understand ethnic-specific sequence variations so that asymptomatic sequence variations are not misidentified as mutations in other ethnic or racial groups.
Wu, Fengnian; Jiang, Hongyan; Beattie, G Andrew C; Holford, Paul; Chen, Jianchi; Wallis, Christopher M; Zheng, Zheng; Deng, Xiaoling; Cen, Yijing
2018-04-24
Diaphorina citri (Asian citrus psyllid; ACP) transmits 'Candidatus Liberibacter asiaticus' associated with citrus Huanglongbing (HLB). ACP has been reported in 11 provinces/regions in China, yet its population diversity remains unclear. In this study, we evaluated ACP population diversity in China using representative whole mitochondrial genome (mitogenome) sequences. Additional mitogenome sequences outside China were also acquired and evaluated. The sizes of the 27 ACP mitogenome sequences ranged from 14 986 to 15 030 bp. Along with three previously published mitogenome sequences, the 30 sequences formed three major mitochondrial groups (MGs): MG1, present in southwestern China and occurring at elevations above 1000 m; MG2, present in southeastern China and Southeast Asia (Cambodia, Indonesia, Malaysia, and Vietnam) and occurring at elevations below 180 m; and MG3, present in the USA and Pakistan. Single nucleotide polymorphisms in five genes (cox2, atp8, nad3, nad1 and rrnL) contributed mostly in the ACP diversity. Among these genes, rrnL had the most variation. Mitogenome sequences analyses revealed two major phylogenetic groups of ACP present in China as well as a possible unique group present currently in Pakistan and the USA. The information could have significant implications for current ACP control and HLB management. © 2018 Society of Chemical Industry. © 2018 Society of Chemical Industry.
Non-codingRNA sequence variations in human chronic lymphocytic leukemia and colorectal cancer.
Wojcik, Sylwia E; Rossi, Simona; Shimizu, Masayoshi; Nicoloso, Milena S; Cimmino, Amelia; Alder, Hansjuerg; Herlea, Vlad; Rassenti, Laura Z; Rai, Kanti R; Kipps, Thomas J; Keating, Michael J; Croce, Carlo M; Calin, George A
2010-02-01
Cancer is a genetic disease in which the interplay between alterations in protein-coding genes and non-coding RNAs (ncRNAs) plays a fundamental role. In recent years, the full coding component of the human genome was sequenced in various cancers, whereas such attempts related to ncRNAs are still fragmentary. We screened genomic DNAs for sequence variations in 148 microRNAs (miRNAs) and ultraconserved regions (UCRs) loci in patients with chronic lymphocytic leukemia (CLL) or colorectal cancer (CRC) by Sanger technique and further tried to elucidate the functional consequences of some of these variations. We found sequence variations in miRNAs in both sporadic and familial CLL cases, mutations of UCRs in CLLs and CRCs and, in certain instances, detected functional effects of these variations. Furthermore, by integrating our data with previously published data on miRNA sequence variations, we have created a catalog of DNA sequence variations in miRNAs/ultraconserved genes in human cancers. These findings argue that ncRNAs are targeted by both germ line and somatic mutations as well as by single-nucleotide polymorphisms with functional significance for human tumorigenesis. Sequence variations in ncRNA loci are frequent and some have functional and biological significance. Such information can be exploited to further investigate on a genome-wide scale the frequency of genetic variations in ncRNAs and their functional meaning, as well as for the development of new diagnostic and prognostic markers for leukemias and carcinomas.
Non-codingRNA sequence variations in human chronic lymphocytic leukemia and colorectal cancer
Wojcik, Sylwia E.; Rossi, Simona; Shimizu, Masayoshi; Nicoloso, Milena S.; Cimmino, Amelia; Alder, Hansjuerg; Herlea, Vlad; Rassenti, Laura Z.; Rai, Kanti R.; Kipps, Thomas J.; Keating, Michael J.
2010-01-01
Cancer is a genetic disease in which the interplay between alterations in protein-coding genes and non-coding RNAs (ncRNAs) plays a fundamental role. In recent years, the full coding component of the human genome was sequenced in various cancers, whereas such attempts related to ncRNAs are still fragmentary. We screened genomic DNAs for sequence variations in 148 microRNAs (miRNAs) and ultraconserved regions (UCRs) loci in patients with chronic lymphocytic leukemia (CLL) or colorectal cancer (CRC) by Sanger technique and further tried to elucidate the functional consequences of some of these variations. We found sequence variations in miRNAs in both sporadic and familial CLL cases, mutations of UCRs in CLLs and CRCs and, in certain instances, detected functional effects of these variations. Furthermore, by integrating our data with previously published data on miRNA sequence variations, we have created a catalog of DNA sequence variations in miRNAs/ultraconserved genes in human cancers. These findings argue that ncRNAs are targeted by both germ line and somatic mutations as well as by single-nucleotide polymorphisms with functional significance for human tumorigenesis. Sequence variations in ncRNA loci are frequent and some have functional and biological significance. Such information can be exploited to further investigate on a genome-wide scale the frequency of genetic variations in ncRNAs and their functional meaning, as well as for the development of new diagnostic and prognostic markers for leukemias and carcinomas. PMID:19926640
Barclay, Sarah F; Rand, Casey M; Borch, Lauren A; Nguyen, Lisa; Gray, Paul A; Gibson, William T; Wilson, Richard J A; Gordon, Paul M K; Aung, Zaw; Berry-Kravis, Elizabeth M; Ize-Ludlow, Diego; Weese-Mayer, Debra E; Bech-Hansen, N Torben
2015-08-25
Rapid-onset Obesity with Hypothalamic Dysfunction, Hypoventilation, and Autonomic Dysregulation (ROHHAD) is thought to be a genetic disease caused by de novo mutations, though causative mutations have yet to be identified. We searched for de novo coding mutations among a carefully-diagnosed and clinically homogeneous cohort of 35 ROHHAD patients. We sequenced the exomes of seven ROHHAD trios, plus tumours from four of these patients and the unaffected monozygotic (MZ) twin of one (discovery cohort), to identify constitutional and somatic de novo sequence variants. We further analyzed this exome data to search for candidate genes under autosomal dominant and recessive models, and to identify structural variations. Candidate genes were tested by exome or Sanger sequencing in a replication cohort of 28 ROHHAD singletons. The analysis of the trio-based exomes found 13 de novo variants. However, no two patients had de novo variants in the same gene, and additional patient exomes and mutation analysis in the replication cohort did not provide strong genetic evidence to implicate any of these sequence variants in ROHHAD. Somatic comparisons revealed no coding differences between any blood and tumour samples, or between the two discordant MZ twins. Neither autosomal dominant nor recessive analysis yielded candidate genes for ROHHAD, and we did not identify any potentially causative structural variations. Clinical exome sequencing is highly unlikely to be a useful diagnostic test in patients with true ROHHAD. As ROHHAD has a high risk for fatality if not properly managed, it remains imperative to expand the search for non-exomic genetic risk factors, as well as to investigate other possible mechanisms of disease. In so doing, we will be able to confirm objectively the ROHHAD diagnosis and to contribute to our understanding of obesity, respiratory control, hypothalamic function, and autonomic regulation.
USDA-ARS?s Scientific Manuscript database
Genomic structural variations are an important source of genetic diversity. Copy number variations (CNVs), gains and losses of large regions of genomic sequence between individuals of a species, are known to be associated with both diseases and phenotypic traits. Deeply sequenced genomes are often u...
Streptococcus mutans clonal variation revealed by multilocus sequence typing.
Nakano, Kazuhiko; Lapirattanakul, Jinthana; Nomura, Ryota; Nemoto, Hirotoshi; Alaluusua, Satu; Grönroos, Lisa; Vaara, Martti; Hamada, Shigeyuki; Ooshima, Takashi; Nakagawa, Ichiro
2007-08-01
Streptococcus mutans is the major pathogen of dental caries, a biofilm-dependent infectious disease, and occasionally causes infective endocarditis. S. mutans strains have been classified into four serotypes (c, e, f, and k). However, little is known about the S. mutans population, including the clonal relationships among strains of S. mutans, in relation to the particular clones that cause systemic diseases. To address this issue, we have developed a multilocus sequence typing (MLST) scheme for S. mutans. Eight housekeeping gene fragments were sequenced from each of 102 S. mutans isolates collected from the four serotypes in Japan and Finland. Between 14 and 23 alleles per locus were identified, allowing us theoretically to distinguish more than 1.2 x 10(10) sequence types. We identified 92 sequence types in these 102 isolates, indicating that S. mutans contains a diverse population. Whereas serotype c strains were widely distributed in the dendrogram, serotype e, f, and k strains were differentiated into clonal complexes. Therefore, we conclude that the ancestral strain of S. mutans was serotype c. No geographic specificity was identified. However, the distribution of the collagen-binding protein gene (cnm) and direct evidence of mother-to-child transmission were clearly evident. In conclusion, the superior discriminatory capacity of this MLST scheme for S. mutans may have important practical implications.
Clonal evolution and tumor-initiating cells: New dimensions in cancer patient treatment.
Apostoli, Anthony J; Ailles, Laurie
2016-01-01
Human cancer is not a uniform disease but a plethora of disparate tumor types and subtypes. The differences that exist between individual tumors (intertumoral heterogeneity) present a significant roadblock to the eradication of cancer. It has also become increasingly clear that variations across individual tumors (intratumoral heterogeneity) have important implications to cancer progression and treatment efficacy. Therefore, in order to improve patient care and develop novel chemotherapeutics, the evolving tumor landscape needs to be further explored. Next-generation sequencing (NGS) technologies are revolutionizing the cancer research arena by providing state-of-the-art, high-speed methods of genome sequencing at single-nucleotide resolution, thus enabling an unprecedented detection of tumor-specific genetic abnormalities. These anomalies can be quantified to reveal specific frequencies of DNA alterations that correspond to distinct clonal populations within a given tumor. As such, NGS approaches have also been utilized to explore the heterogeneous landscape of patient tumors as well as to match metastatic and/or recurrent growths and patient-derived engrafts. By sequencing in this manner--through time so to speak--cancer researchers can track shifting clonal populations, make important inferences about tumor evolution and potentially identify tumor subclones that could be viably targeted. This exciting new territory has important implications for the competing clonal evolution and cancer stem cell models of tumor heterogeneity, and also offers a new dimension for cancer treatment and profound hope for patients in the coming years.
Cavusoglu, M; Ciloglu, T; Serinagaoglu, Y; Kamasak, M; Erogul, O; Akcam, T
2008-08-01
In this paper, 'snore regularity' is studied in terms of the variations of snoring sound episode durations, separations and average powers in simple snorers and in obstructive sleep apnoea (OSA) patients. The goal was to explore the possibility of distinguishing among simple snorers and OSA patients using only sleep sound recordings of individuals and to ultimately eliminate the need for spending a whole night in the clinic for polysomnographic recording. Sequences that contain snoring episode durations (SED), snoring episode separations (SES) and average snoring episode powers (SEP) were constructed from snoring sound recordings of 30 individuals (18 simple snorers and 12 OSA patients) who were also under polysomnographic recording in Gülhane Military Medical Academy Sleep Studies Laboratory (GMMA-SSL), Ankara, Turkey. Snore regularity is quantified in terms of mean, standard deviation and coefficient of variation values for the SED, SES and SEP sequences. In all three of these sequences, OSA patients' data displayed a higher variation than those of simple snorers. To exclude the effects of slow variations in the base-line of these sequences, new sequences that contain the coefficient of variation of the sample values in a 'short' signal frame, i.e., short time coefficient of variation (STCV) sequences, were defined. The mean, the standard deviation and the coefficient of variation values calculated from the STCV sequences displayed a stronger potential to distinguish among simple snorers and OSA patients than those obtained from the SED, SES and SEP sequences themselves. Spider charts were used to jointly visualize the three parameters, i.e., the mean, the standard deviation and the coefficient of variation values of the SED, SES and SEP sequences, and the corresponding STCV sequences as two-dimensional plots. Our observations showed that the statistical parameters obtained from the SED and SES sequences, and the corresponding STCV sequences, possessed a strong potential to distinguish among simple snorers and OSA patients, both marginally, i.e., when the parameters are examined individually, and jointly. The parameters obtained from the SEP sequences and the corresponding STCV sequences, on the other hand, did not have a strong discrimination capability. However, the joint behaviour of these parameters showed some potential to distinguish among simple snorers and OSA patients.
2011-01-01
Background Most information on genomic variations and their associations with phenotypes are covered exclusively in scientific publications rather than in structured databases. These texts commonly describe variations using natural language; database identifiers are seldom mentioned. This complicates the retrieval of variations, associated articles, as well as information extraction, e. g. the search for biological implications. To overcome these challenges, procedures to map textual mentions of variations to database identifiers need to be developed. Results This article describes a workflow for normalization of variation mentions, i.e. the association of them to unique database identifiers. Common pitfalls in the interpretation of single nucleotide polymorphism (SNP) mentions are highlighted and discussed. The developed normalization procedure achieves a precision of 98.1 % and a recall of 67.5% for unambiguous association of variation mentions with dbSNP identifiers on a text corpus based on 296 MEDLINE abstracts containing 527 mentions of SNPs. The annotated corpus is freely available at http://www.scai.fraunhofer.de/snp-normalization-corpus.html. Conclusions Comparable approaches usually focus on variations mentioned on the protein sequence and neglect problems for other SNP mentions. The results presented here indicate that normalizing SNPs described on DNA level is more difficult than the normalization of SNPs described on protein level. The challenges associated with normalization are exemplified with ambiguities and errors, which occur in this corpus. PMID:21992066
Ansari, Israr-ul H.; Allen, Todd; Berical, Andrew; Stock, Peter G.; Barin, Burc; Striker, Rob
2013-01-01
Hepatitis C virus (HCV) replication is limited by cyclophilin inhibitors but it remains unclear how viral genetic variations influence susceptibility to cyclosporine (cyclosporine A, CsA), a cyclophilin inhibitor. In this study HCV from liver transplant patients was sequenced before and after CsA exposure. Phenotypic analysis of NS5A sequence was performed by using HCV sub genomic replicon to determine CsA susceptibility. The data indicates an atypical proline at position 328 in NS5A causes increases CsA sensitivity both in the context of genotype 1a and 1b residues. Point mutants mimicking other naturally occurring residues at this position also increased (Ala) or decreased (Arg) replicon sensitivity to CsA relative to the typical threonine (genotype 1a) or serine (genotype 1b) at this position. This work has implications for treatment of HCV by cyclophilin inhibitors. PMID:23290631
A polygenic burden of rare disruptive mutations in schizophrenia
Purcell, Shaun M.; Moran, Jennifer L.; Fromer, Menachem; Ruderfer, Douglas; Solovieff, Nadia; Roussos, Panos; O’Dushlaine, Colm; Chambert, Kimberly; Bergen, Sarah E.; Kähler, Anna; Duncan, Laramie; Stahl, Eli; Genovese, Giulio; Fernández, Esperanza; Collins, Mark O; Komiyama, Noboru H.; Choudhary, Jyoti S.; Magnusson, Patrik K. E.; Banks, Eric; Shakir, Khalid; Garimella, Kiran; Fennell, Tim; de Pristo, Mark; Grant, Seth G.N.; Haggarty, Stephen; Gabriel, Stacey; Scolnick, Edward M.; Lander, Eric S.; Hultman, Christina; Sullivan, Patrick F.; McCarroll, Steven A.; Sklar, Pamela
2014-01-01
By analyzing the exome sequences of 2,536 schizophrenia cases and 2,543 controls, we have demonstrated a polygenic burden primarily arising from rare (<1/10,000), disruptive mutations distributed across many genes. Especially enriched genesets included the voltage-gated calcium ion channel and the signaling complex formed by the activity-regulated cytoskeleton-associated (ARC) scaffold protein of the postsynaptic density (PSD), sets previously implicated by genome-wide association studies (GWAS) and copy-number variation (CNV) studies. Similar to reports in autism, targets of the fragile × mental retardation protein (FMRP, product of FMR1) were enriched for case mutations. No individual gene-based test achieved significance after correction for multiple testing and we did not detect any alleles of moderately low frequency (~0.5-1%) and moderately large effect. Taken together, these data suggest that population-based exome sequencing can discover risk alleles and complements established gene mapping paradigms in neuropsychiatric disease. PMID:24463508
Zheng, Yanying; Liu, Li; Sun, Yi; Chen, Jie; Wang, Jianrong; Zhu, Changle; Lai, Rensheng; Xie, Ling
2016-07-30
BAT-26 is one of the representative markers for microsatellite instability evaluation and presents different polymorphisms in different ethnic populations. The current knowledge of its comparative polymorphism between healthy individuals and cancer patients in the Chinese population is insufficient. This study aims to analyze germline polymorphic variations of BAT-26 between healthy individuals and cancer patients in Chinese from Jiangsu province and the associated cancer risk implications. The various BAT-26 alleles and their percentages in cervical cells from 500 healthy women were assessed by direct sequencing. Twenty of these samples were also analyzed by fragment analysis. BAT-26 of blood DNA from 24 healthy individuals and 247 cancer patients was analyzed by fragment analysis. Compared with the sequencing results, 122.6-122.9 bp, 123.4-123.8 bp and 124.1-124.8 bp corresponded to the A25, A26 and A27 alleles, respectively. The 524 healthy individuals showed 4.58%, 92.18% and 3.24% of A25, A26 and A27, respectively. The variant alleles A18, A24, A28, A29 and A32 were only found in cancer patients, accounting for 0.81%, 0.40%, 0.40%, 0.40% and 0.40%, respectively; the A25, A26 and A27 alleles in cancer patients accounted for 6.48%, 77.33% and 13.77%. Healthy individuals had a stable BAT-26 profile within the quasimonomorphic variation range (QMVR), but cancer patients harbored variant alleles outside QMVR and showed a trend from quasimonomorph to polymonomorph, suggesting that variant alleles of BAT-26 in germline cells may be regarded as a potential marker of higher cancer risk in the Chinese population from Jiangsu province.
Synaptic, transcriptional and chromatin genes disrupted in autism.
De Rubeis, Silvia; He, Xin; Goldberg, Arthur P; Poultney, Christopher S; Samocha, Kaitlin; Cicek, A Erucment; Kou, Yan; Liu, Li; Fromer, Menachem; Walker, Susan; Singh, Tarinder; Klei, Lambertus; Kosmicki, Jack; Shih-Chen, Fu; Aleksic, Branko; Biscaldi, Monica; Bolton, Patrick F; Brownfeld, Jessica M; Cai, Jinlu; Campbell, Nicholas G; Carracedo, Angel; Chahrour, Maria H; Chiocchetti, Andreas G; Coon, Hilary; Crawford, Emily L; Curran, Sarah R; Dawson, Geraldine; Duketis, Eftichia; Fernandez, Bridget A; Gallagher, Louise; Geller, Evan; Guter, Stephen J; Hill, R Sean; Ionita-Laza, Juliana; Jimenz Gonzalez, Patricia; Kilpinen, Helena; Klauck, Sabine M; Kolevzon, Alexander; Lee, Irene; Lei, Irene; Lei, Jing; Lehtimäki, Terho; Lin, Chiao-Feng; Ma'ayan, Avi; Marshall, Christian R; McInnes, Alison L; Neale, Benjamin; Owen, Michael J; Ozaki, Noriio; Parellada, Mara; Parr, Jeremy R; Purcell, Shaun; Puura, Kaija; Rajagopalan, Deepthi; Rehnström, Karola; Reichenberg, Abraham; Sabo, Aniko; Sachse, Michael; Sanders, Stephan J; Schafer, Chad; Schulte-Rüther, Martin; Skuse, David; Stevens, Christine; Szatmari, Peter; Tammimies, Kristiina; Valladares, Otto; Voran, Annette; Li-San, Wang; Weiss, Lauren A; Willsey, A Jeremy; Yu, Timothy W; Yuen, Ryan K C; Cook, Edwin H; Freitag, Christine M; Gill, Michael; Hultman, Christina M; Lehner, Thomas; Palotie, Aaarno; Schellenberg, Gerard D; Sklar, Pamela; State, Matthew W; Sutcliffe, James S; Walsh, Christiopher A; Scherer, Stephen W; Zwick, Michael E; Barett, Jeffrey C; Cutler, David J; Roeder, Kathryn; Devlin, Bernie; Daly, Mark J; Buxbaum, Joseph D
2014-11-13
The genetic architecture of autism spectrum disorder involves the interplay of common and rare variants and their impact on hundreds of genes. Using exome sequencing, here we show that analysis of rare coding variation in 3,871 autism cases and 9,937 ancestry-matched or parental controls implicates 22 autosomal genes at a false discovery rate (FDR) < 0.05, plus a set of 107 autosomal genes strongly enriched for those likely to affect risk (FDR < 0.30). These 107 genes, which show unusual evolutionary constraint against mutations, incur de novo loss-of-function mutations in over 5% of autistic subjects. Many of the genes implicated encode proteins for synaptic formation, transcriptional regulation and chromatin-remodelling pathways. These include voltage-gated ion channels regulating the propagation of action potentials, pacemaking and excitability-transcription coupling, as well as histone-modifying enzymes and chromatin remodellers-most prominently those that mediate post-translational lysine methylation/demethylation modifications of histones.
Major histocompatibility complex variation in the endangered Przewalski's horse.
Hedrick, P W; Parker, K M; Miller, E L; Miller, P S
1999-01-01
The major histocompatibility complex (MHC) is a fundamental part of the vertebrate immune system, and the high variability in many MHC genes is thought to play an essential role in recognition of parasites. The Przewalski's horse is extinct in the wild and all the living individuals descend from 13 founders, most of whom were captured around the turn of the century. One of the primary genetic concerns in endangered species is whether they have ample adaptive variation to respond to novel selective factors. In examining 14 Przewalski's horses that are broadly representative of the living animals, we found six different class II DRB major histocompatibility sequences. The sequences showed extensive nonsynonymous variation, concentrated in the putative antigen-binding sites, and little synonymous variation. Individuals had from two to four sequences as determined by single-stranded conformation polymorphism (SSCP) analysis. On the basis of the SSCP data, phylogenetic analysis of the nucleotide sequences, and segregation in a family group, we conclude that four of these sequences are from one gene (although one sequence codes for a nonfunctional allele because it contains a stop codon) and two other sequences are from another gene. The position of the stop codon is at the same amino-acid position as in a closely related sequence from the domestic horse. Because other organisms have extensive variation at homologous loci, the Przewalski's horse may have quite low variation in this important adaptive region. PMID:10430594
Genetic Variation in the Acorn Barnacle from Allozymes to Population Genomics
Flight, Patrick A.; Rand, David M.
2012-01-01
Understanding the patterns of genetic variation within and among populations is a central problem in population and evolutionary genetics. We examine this question in the acorn barnacle, Semibalanus balanoides, in which the allozyme loci Mpi and Gpi have been implicated in balancing selection due to varying selective pressures at different spatial scales. We review the patterns of genetic variation at the Mpi locus, compare this to levels of population differentiation at mtDNA and microsatellites, and place these data in the context of genome-wide variation from high-throughput sequencing of population samples spanning the North Atlantic. Despite considerable geographic variation in the patterns of selection at the Mpi allozyme, this locus shows rather low levels of population differentiation at ecological and trans-oceanic scales (FST ∼ 5%). Pooled population sequencing was performed on samples from Rhode Island (RI), Maine (ME), and Southwold, England (UK). Analysis of more than 650 million reads identified approximately 335,000 high-quality SNPs in 19 million base pairs of the S. balanoides genome. Much variation is shared across the Atlantic, but there are significant examples of strong population differentiation among samples from RI, ME, and UK. An FST outlier screen of more than 22,000 contigs provided a genome-wide context for interpretation of earlier studies on allozymes, mtDNA, and microsatellites. FST values for allozymes, mtDNA and microsatellites are close to the genome-wide average for random SNPs, with the exception of the trans-Atlantic FST for mtDNA. The majority of FST outliers were unique between individual pairs of populations, but some genes show shared patterns of excess differentiation. These data indicate that gene flow is high, that selection is strong on a subset of genes, and that a variety of genes are experiencing diversifying selection at large spatial scales. This survey of polymorphism in S. balanoides provides a number of genomic tools that promise to make this a powerful model for ecological genomics of the rocky intertidal. PMID:22767487
Using chaos to generate variations on movement sequences
NASA Astrophysics Data System (ADS)
Bradley, Elizabeth; Stuart, Joshua
1998-12-01
We describe a method for introducing variations into predefined motion sequences using a chaotic symbol-sequence reordering technique. A progression of symbols representing the body positions in a dance piece, martial arts form, or other motion sequence is mapped onto a chaotic trajectory, establishing a symbolic dynamics that links the movement sequence and the attractor structure. A variation on the original piece is created by generating a trajectory with slightly different initial conditions, inverting the mapping, and using special corpus-based graph-theoretic interpolation schemes to smooth any abrupt transitions. Sensitive dependence guarantees that the variation is different from the original; the attractor structure and the symbolic dynamics guarantee that the two resemble one another in both aesthetic and mathematical senses.
Genetic Variation in Cardiomyopathy and Cardiovascular Disorders.
McNally, Elizabeth M; Puckelwartz, Megan J
2015-01-01
With the wider deployment of massively-parallel, next-generation sequencing, it is now possible to survey human genome data for research and clinical purposes. The reduced cost of producing short-read sequencing has now shifted the burden to data analysis. Analysis of genome sequencing remains challenged by the complexity of the human genome, including redundancy and the repetitive nature of genome elements and the large amount of variation in individual genomes. Public databases of human genome sequences greatly facilitate interpretation of common and rare genetic variation, although linking database sequence information to detailed clinical information is limited by privacy and practical issues. Genetic variation is a rich source of knowledge for cardiovascular disease because many, if not all, cardiovascular disorders are highly heritable. The role of rare genetic variation in predicting risk and complications of cardiovascular diseases has been well established for hypertrophic and dilated cardiomyopathy, where the number of genes that are linked to these disorders is growing. Bolstered by family data, where genetic variants segregate with disease, rare variation can be linked to specific genetic variation that offers profound diagnostic information. Understanding genetic variation in cardiomyopathy is likely to help stratify forms of heart failure and guide therapy. Ultimately, genetic variation may be amenable to gene correction and gene editing strategies.
[Hydrologic variability and sensitivity based on Hurst coefficient and Bartels statistic].
Lei, Xu; Xie, Ping; Wu, Zi Yi; Sang, Yan Fang; Zhao, Jiang Yan; Li, Bin Bin
2018-04-01
Due to the global climate change and frequent human activities in recent years, the pure stochastic components of hydrological sequence is mixed with one or several of the variation ingredients, including jump, trend, period and dependency. It is urgently needed to clarify which indices should be used to quantify the degree of their variability. In this study, we defined the hydrological variability based on Hurst coefficient and Bartels statistic, and used Monte Carlo statistical tests to test and analyze their sensitivity to different variants. When the hydrological sequence had jump or trend variation, both Hurst coefficient and Bartels statistic could reflect the variation, with the Hurst coefficient being more sensitive to weak jump or trend variation. When the sequence had period, only the Bartels statistic could detect the mutation of the sequence. When the sequence had a dependency, both the Hurst coefficient and the Bartels statistics could reflect the variation, with the latter could detect weaker dependent variations. For the four variations, both the Hurst variability and Bartels variability increased with the increases of variation range. Thus, they could be used to measure the variation intensity of the hydrological sequence. We analyzed the temperature series of different weather stations in the Lancang River basin. Results showed that the temperature of all stations showed the upward trend or jump, indicating that the entire basin had experienced warming in recent years and the temperature variability in the upper and lower reaches was much higher. This case study showed the practicability of the proposed method.
Phenotypic and mtDNA variation in Philippine Kappaphycus cottonii (Gigartinales, Rhodophyta).
Dumilag, Richard V; Gallardo, William George M; Garcia, Christian Philip C; You, YeaEun; Chaves, Alyssa Keren G; Agahan, Lance
2017-11-09
Members of the carrageenan-producing seaweeds of the genus Kappapphycus have a complicated taxonomic history particularly with regard to species identification. Many taxonomic challenges in this group have been currently addressed with the use of mtDNA sequences. The phylogenetic status and genetic diversity of one of the lesser known species, Kappaphycus cottonii, have repeatedly come into question. This study explored the genetic variation in Philippine K. cottonii using the mtDNA COI-5P gene and cox2-3 spacer sequences. The six phenotypic forms in K. cottonii did not correspond to the observed genetic variability; hinting at the greater involvement of environmental factors in determining changes to the morphology of this alga. Our results revealed that the Philippine K. cottonii has the richest number of haplotypes that have been detected, so far, for any Kappaphycus species. Our inferred phylogenetic trees suggested two lineages: a lineage, which exclusively includes K. cottonii and another lineage comprising the four known Kappaphycus species: K. alvarezii, K. inermis, K. malesianus, and K. striatus. The dichotomy supports the apparent synamorphy for each of these lineages (the strictly terete thalli, lack of protuberances, and the presence of a hyphal central core in the latter group, while the opposite of these morphologies in K. cottonii). These findings shed new light on understanding the evolutionary history of the genus. Assessing the breadth of the phenotypic and genetic variation in K. cottonii has implications for the conservation and management of the overall Kappaphycus genetic resources, especially in the Philippines.
van Hal, Sebastiaan J.; Steen, Jason A.; Espedido, Björn A.; Grimmond, Sean M.; Cooper, Matthew A.; Holden, Matthew T. G.; Bentley, Stephen D.; Gosbell, Iain B.; Jensen, Slade O.
2014-01-01
Objectives To obtain an expanded understanding of antibiotic resistance evolution in vivo, particularly in the context of vancomycin exposure. Methods The whole genomes of six consecutive methicillin-resistant Staphylococcus aureus blood culture isolates (ST239-MRSA-III) from a single patient exposed to various antimicrobials (over a 77 day period) were sequenced and analysed. Results Variant analysis revealed the existence of non-susceptible sub-populations derived from a common susceptible ancestor, with the predominant circulating clone(s) selected for by type and duration of antimicrobial exposure. Conclusions This study highlights the dynamic nature of bacterial evolution and that non-susceptible sub-populations can emerge from clouds of variation upon antimicrobial exposure. Diagnostically, this has direct implications for sample selection when using whole-genome sequencing as a tool to guide clinical therapy. In the context of bacteraemia, deep sequencing of bacterial DNA directly from patient blood samples would avoid culture ‘bias’ and identify mutations associated with circulating non-susceptible sub-populations, some of which may confer cross-resistance to alternate therapies. PMID:24047554
van Hal, Sebastiaan J; Steen, Jason A; Espedido, Björn A; Grimmond, Sean M; Cooper, Matthew A; Holden, Matthew T G; Bentley, Stephen D; Gosbell, Iain B; Jensen, Slade O
2014-02-01
To obtain an expanded understanding of antibiotic resistance evolution in vivo, particularly in the context of vancomycin exposure. The whole genomes of six consecutive methicillin-resistant Staphylococcus aureus blood culture isolates (ST239-MRSA-III) from a single patient exposed to various antimicrobials (over a 77 day period) were sequenced and analysed. Variant analysis revealed the existence of non-susceptible sub-populations derived from a common susceptible ancestor, with the predominant circulating clone(s) selected for by type and duration of antimicrobial exposure. This study highlights the dynamic nature of bacterial evolution and that non-susceptible sub-populations can emerge from clouds of variation upon antimicrobial exposure. Diagnostically, this has direct implications for sample selection when using whole-genome sequencing as a tool to guide clinical therapy. In the context of bacteraemia, deep sequencing of bacterial DNA directly from patient blood samples would avoid culture 'bias' and identify mutations associated with circulating non-susceptible sub-populations, some of which may confer cross-resistance to alternate therapies.
NASA Astrophysics Data System (ADS)
Wang, J.; Nathan, R.; Horne, A.
2017-12-01
Traditional approaches to characterize water-dependent ecosystem outcomes in response to flow have been based on time-averaged hydrological indicators, however there is increasing recognition for the need to characterize ecological processes that are highly dependent on the sequencing of flow conditions (i.e. floods and droughts). This study considers the representation of flow regimes when considering assessment of ecological outcomes, and in particular, the need to account for sequencing and variability of flow. We conducted two case studies - one in the largely unregulated Ovens River catchment and one in the highly regulated Murray River catchment (both located in south-eastern Australia) - to explore the importance of flow sequencing to the condition of a typical long-lived ecological asset in Australia, the River Red Gum forests. In the first, the Ovens River case study, the implications of representing climate change using different downscaling methods (annual scaling, monthly scaling, quantile mapping, and weather generator method) on the sequencing of flows and resulting ecological outcomes were considered. In the second, the Murray River catchment, sequencing within a historic drought period was considered by systematically making modest adjustments on an annual basis to the hydrological records. In both cases, the condition of River Red Gum forests was assessed using an ecological model that incorporates transitions between ecological conditions in response to sequences of required flow components. The results of both studies show the importance of considering how hydrological alterations are represented when assessing ecological outcomes. The Ovens case study showed that there is significant variation in the predicted ecological outcomes when different downscaling techniques are applied. Similarly, the analysis in the Murray case study showed that the drought as it historically occurred provided one of the best possible outcomes for River Red Gum forests when compared to other re-arrangements of flow within the same drought. These results have implications for the way we represent climate change impacts and drought risk assessments where ecological outcomes are a key management objective.
Rare variation facilitates inferences of fine-scale population structure in humans.
O'Connor, Timothy D; Fu, Wenqing; Mychaleckyj, Josyf C; Logsdon, Benjamin; Auer, Paul; Carlson, Christopher S; Leal, Suzanne M; Smith, Joshua D; Rieder, Mark J; Bamshad, Michael J; Nickerson, Deborah A; Akey, Joshua M
2015-03-01
Understanding the genetic structure of human populations has important implications for the design and interpretation of disease mapping studies and reconstructing human evolutionary history. To date, inferences of human population structure have primarily been made with common variants. However, recent large-scale resequencing studies have shown an abundance of rare variation in humans, which may be particularly useful for making inferences of fine-scale population structure. To this end, we used an information theory framework and extensive coalescent simulations to rigorously quantify the informativeness of rare and common variation to detect signatures of fine-scale population structure. We show that rare variation affords unique insights into patterns of recent population structure. Furthermore, to empirically assess our theoretical findings, we analyzed high-coverage exome sequences in 6,515 European and African American individuals. As predicted, rare variants are more informative than common polymorphisms in revealing a distinct cluster of European-American individuals, and subsequent analyses demonstrate that these individuals are likely of Ashkenazi Jewish ancestry. Our results provide new insights into the population structure using rare variation, which will be an important factor to account for in rare variant association studies. © The Author 2014. Published by Oxford University Press on behalf of the Society for Molecular Biology and Evolution.
Mutation Detection with Next-Generation Resequencing through a Mediator Genome
DOE Office of Scientific and Technical Information (OSTI.GOV)
Wurtzel, Omri; Dori-Bachash, Mally; Pietrokovski, Shmuel
2010-12-31
The affordability of next generation sequencing (NGS) is transforming the field of mutation analysis in bacteria. The genetic basis for phenotype alteration can be identified directly by sequencing the entire genome of the mutant and comparing it to the wild-type (WT) genome, thus identifying acquired mutations. A major limitation for this approach is the need for an a-priori sequenced reference genome for the WT organism, as the short reads of most current NGS approaches usually prohibit de-novo genome assembly. To overcome this limitation we propose a general framework that utilizes the genome of relative organisms as mediators for comparing WTmore » and mutant bacteria. Under this framework, both mutant and WT genomes are sequenced with NGS, and the short sequencing reads are mapped to the mediator genome. Variations between the mutant and the mediator that recur in the WT are ignored, thus pinpointing the differences between the mutant and the WT. To validate this approach we sequenced the genome of Bdellovibrio bacteriovorus 109J, an obligatory bacterial predator, and its prey-independent mutant, and compared both to the mediator species Bdellovibrio bacteriovorus HD100. Although the mutant and the mediator sequences differed in more than 28,000 nucleotide positions, our approach enabled pinpointing the single causative mutation. Experimental validation in 53 additional mutants further established the implicated gene. Our approach extends the applicability of NGS-based mutant analyses beyond the domain of available reference genomes.« less
Handel, Stephen; Todd, Sean K; Zoidis, Ann M
2009-06-01
The hierarchical organization of the male humpback whale song has been well documented. However, it is unknown how singers keep these intricate songs intact over multiple repetitions or how they learn variations that occur sequentially during each mating season. Rather than focus on the sequence of sounds within a song, results presented here demonstrate that the individual sounds are organized into rhythmic groups that make the production and perception of the lengthy songs tractable by yielding a set of simple groups that, although arranged in rigid order, can be repeated multiple times to generate the entire song.
Fokkema, Ivo F A C; den Dunnen, Johan T; Taschner, Peter E M
2005-08-01
The completion of the human genome project has initiated, as well as provided the basis for, the collection and study of all sequence variation between individuals. Direct access to up-to-date information on sequence variation is currently provided most efficiently through web-based, gene-centered, locus-specific databases (LSDBs). We have developed the Leiden Open (source) Variation Database (LOVD) software approaching the "LSDB-in-a-Box" idea for the easy creation and maintenance of a fully web-based gene sequence variation database. LOVD is platform-independent and uses PHP and MySQL open source software only. The basic gene-centered and modular design of the database follows the recommendations of the Human Genome Variation Society (HGVS) and focuses on the collection and display of DNA sequence variations. With minimal effort, the LOVD platform is extendable with clinical data. The open set-up should both facilitate and promote functional extension with scripts written by the community. The LOVD software is freely available from the Leiden Muscular Dystrophy pages (www.DMD.nl/LOVD/). To promote the use of LOVD, we currently offer curators the possibility to set up an LSDB on our Leiden server. (c) 2005 Wiley-Liss, Inc.
Ridge, Perry G; Maxwell, Taylor J; Foutz, Spencer J; Bailey, Matthew H; Corcoran, Christopher D; Tschanz, JoAnn T; Norton, Maria C; Munger, Ronald G; O'Brien, Elizabeth; Kerber, Richard A; Cawthon, Richard M; Kauwe, John S K
2014-01-01
The mitochondria are essential organelles and are the location of cellular respiration, which is responsible for the majority of ATP production. Each cell contains multiple mitochondria, and each mitochondrion contains multiple copies of its own circular genome. The ratio of mitochondrial genomes to nuclear genomes is referred to as mitochondrial copy number. Decreases in mitochondrial copy number are known to occur in many tissues as people age, and in certain diseases. The regulation of mitochondrial copy number by nuclear genes has been studied extensively. While mitochondrial variation has been associated with longevity and some of the diseases known to have reduced mitochondrial copy number, the role that the mitochondrial genome itself has in regulating mitochondrial copy number remains poorly understood. We analyzed the complete mitochondrial genomes from 1007 individuals randomly selected from the Cache County Study on Memory Health and Aging utilizing the inferred evolutionary history of the mitochondrial haplotypes present in our dataset to identify sequence variation and mitochondrial haplotypes associated with changes in mitochondrial copy number. Three variants belonging to mitochondrial haplogroups U5A1 and T2 were significantly associated with higher mitochondrial copy number in our dataset. We identified three variants associated with higher mitochondrial copy number and suggest several hypotheses for how these variants influence mitochondrial copy number by interacting with known regulators of mitochondrial copy number. Our results are the first to report sequence variation in the mitochondrial genome that causes changes in mitochondrial copy number. The identification of these variants that increase mtDNA copy number has important implications in understanding the pathological processes that underlie these phenotypes.
Neocortical malformation as consequence of nonadaptive regulation of neuronogenetic sequence
NASA Technical Reports Server (NTRS)
Caviness, V. S. Jr; Takahashi, T.; Nowakowski, R. S.
2000-01-01
Variations in the structure of the neocortex induced by single gene mutations may be extreme or subtle. They differ from variations in neocortical structure encountered across and within species in that these "normal" structural variations are adaptive (both structurally and behaviorally), whereas those associated with disorders of development are not. Here we propose that they also differ in principle in that they represent disruptions of molecular mechanisms that are not normally regulatory to variations in the histogenetic sequence. We propose an algorithm for the operation of the neuronogenetic sequence in relation to the overall neocortical histogenetic sequence and highlight the restriction point of the G1 phase of the cell cycle as the master regulatory control point for normal coordinate structural variation across species and importantly within species. From considerations based on the anatomic evidence from neocortical malformation in humans, we illustrate in principle how this overall sequence appears to be disrupted by molecular biological linkages operating principally outside the control mechanisms responsible for the normal structural variation of the neocortex. MRDD Research Reviews 6:22-33, 2000. Copyright 2000 Wiley-Liss, Inc.
Piton, Amélie; Redin, Claire; Mandel, Jean-Louis
2013-08-08
Because of the unbalanced sex ratio (1.3-1.4 to 1) observed in intellectual disability (ID) and the identification of large ID-affected families showing X-linked segregation, much attention has been focused on the genetics of X-linked ID (XLID). Mutations causing monogenic XLID have now been reported in over 100 genes, most of which are commonly included in XLID diagnostic gene panels. Nonetheless, the boundary between true mutations and rare non-disease-causing variants often remains elusive. The sequencing of a large number of control X chromosomes, required for avoiding false-positive results, was not systematically possible in the past. Such information is now available thanks to large-scale sequencing projects such as the National Heart, Lung, and Blood (NHLBI) Exome Sequencing Project, which provides variation information on 10,563 X chromosomes from the general population. We used this NHLBI cohort to systematically reassess the implication of 106 genes proposed to be involved in monogenic forms of XLID. We particularly question the implication in XLID of ten of them (AGTR2, MAGT1, ZNF674, SRPX2, ATP6AP2, ARHGEF6, NXF5, ZCCHC12, ZNF41, and ZNF81), in which truncating variants or previously published mutations are observed at a relatively high frequency within this cohort. We also highlight 15 other genes (CCDC22, CLIC2, CNKSR2, FRMPD4, HCFC1, IGBP1, KIAA2022, KLF8, MAOA, NAA10, NLGN3, RPL10, SHROOM4, ZDHHC15, and ZNF261) for which replication studies are warranted. We propose that similar reassessment of reported mutations (and genes) with the use of data from large-scale human exome sequencing would be relevant for a wide range of other genetic diseases. Copyright © 2013 The American Society of Human Genetics. Published by Elsevier Inc. All rights reserved.
NASA Astrophysics Data System (ADS)
Manikyamba, C.; Said, Nuru; Santosh, M.; Saha, Abhishek; Ganguly, Sohini; Subramanyam, K. S. V.
2018-05-01
Phanerozoic boninites record enrichments of U over Th, giving Th/U: 0.5-1.6, relative to intraoceanic island arc tholeiites (IAT) where Th/U averages 2.6. Uranium enrichment is attributed to incorporation of shallow, oxidized fluids, U-rich but Th-poor, from the slab into the melt column of boninites which form in near-trench to forearc settings of suprasubduction zone ophiolites. Well preserved Archean komatiite-tholeiite, plume-derived, oceanic volcanic sequences have primary magmatic Th/U ratios of 4.4-3.6, and Archean convergent margin IAT volcanic sequences, having REE and HFSE compositions similar to Phanerozoic IAT equivalents, preserve primary Th/U of 4-3.6. The best preserved Archean boninites of the 3.0 Ga Olondo and 2.7 Ga Gadwal greenstone belts, hosted in convergent margin ophiolite sequences, also show relative enrichments of U over Th, with low average Th/U ∼3 relative to coeval IAT, and Phanerozoic counterparts which are devoid of crustal contamination and therefore erupted in an intraoceanic setting, with minimal contemporaneous submarine hydrothermal alteration. Later enrichment of U is unlikely as Th-U-Nb-LREE patterns are coherent in these boninites whereas secondary effects induce dispersion of Th/U ratios. The variation in Th/U ratios from Archean to Phanerozoic boninites of greenstone belts to ophiolitic sequences reflect on genesis of boninitic lavas at different tectono-thermal regimes. Consequently, if the explanation for U enrichment in Phanerozoic boninites also applies to Archean examples, the implication is that U was soluble in oxygenated Archean marine water up to 600 Ma before the proposed great oxygenation event (GOE) at ∼2.4 Ga. This interpretation is consistent with large Ce anomalies in some hydrothermally altered Archean volcanic sequences aged 3.0-2.7 Ga.
Genomic Sequence Variation Markup Language (GSVML).
Nakaya, Jun; Kimura, Michio; Hiroi, Kaei; Ido, Keisuke; Yang, Woosung; Tanaka, Hiroshi
2010-02-01
With the aim of making good use of internationally accumulated genomic sequence variation data, which is increasing rapidly due to the explosive amount of genomic research at present, the development of an interoperable data exchange format and its international standardization are necessary. Genomic Sequence Variation Markup Language (GSVML) will focus on genomic sequence variation data and human health applications, such as gene based medicine or pharmacogenomics. We developed GSVML through eight steps, based on case analysis and domain investigations. By focusing on the design scope to human health applications and genomic sequence variation, we attempted to eliminate ambiguity and to ensure practicability. We intended to satisfy the requirements derived from the use case analysis of human-based clinical genomic applications. Based on database investigations, we attempted to minimize the redundancy of the data format, while maximizing the data covering range. We also attempted to ensure communication and interface ability with other Markup Languages, for exchange of omics data among various omics researchers or facilities. The interface ability with developing clinical standards, such as the Health Level Seven Genotype Information model, was analyzed. We developed the human health-oriented GSVML comprising variation data, direct annotation, and indirect annotation categories; the variation data category is required, while the direct and indirect annotation categories are optional. The annotation categories contain omics and clinical information, and have internal relationships. For designing, we examined 6 cases for three criteria as human health application and 15 data elements for three criteria as data formats for genomic sequence variation data exchange. The data format of five international SNP databases and six Markup Languages and the interface ability to the Health Level Seven Genotype Model in terms of 317 items were investigated. GSVML was developed as a potential data exchanging format for genomic sequence variation data exchange focusing on human health applications. The international standardization of GSVML is necessary, and is currently underway. GSVML can be applied to enhance the utilization of genomic sequence variation data worldwide by providing a communicable platform between clinical and research applications. Copyright 2009 Elsevier Ireland Ltd. All rights reserved.
ShatterProof: operational detection and quantification of chromothripsis.
Govind, Shaylan K; Zia, Amin; Hennings-Yeomans, Pablo H; Watson, John D; Fraser, Michael; Anghel, Catalina; Wyatt, Alexander W; van der Kwast, Theodorus; Collins, Colin C; McPherson, John D; Bristow, Robert G; Boutros, Paul C
2014-03-19
Chromothripsis, a newly discovered type of complex genomic rearrangement, has been implicated in the evolution of several types of cancers. To date, it has been described in bone cancer, SHH-medulloblastoma and acute myeloid leukemia, amongst others, however there are still no formal or automated methods for detecting or annotating it in high throughput sequencing data. As such, findings of chromothripsis are difficult to compare and many cases likely escape detection altogether. We introduce ShatterProof, a software tool for detecting and quantifying chromothriptic events. ShatterProof takes structural variation calls (translocations, copy-number variations, short insertions and loss of heterozygosity) produced by any algorithm and using an operational definition of chromothripsis performs robust statistical tests to accurately predict the presence and location of chromothriptic events. Validation of our tool was conducted using clinical data sets including matched normal, prostate cancer samples in addition to the colorectal cancer and SCLC data sets used in the original description of chromothripsis. ShatterProof is computationally efficient, having low memory requirements and near linear computation time. This allows it to become a standard component of sequencing analysis pipelines, enabling researchers to routinely and accurately assess samples for chromothripsis. Source code and documentation can be found at http://search.cpan.org/~sgovind/Shatterproof.
Edelson, Benjamin S; Best, Timothy P; Olenyuk, Bogdan; Nickols, Nicholas G; Doss, Raymond M; Foister, Shane; Heckel, Alexander; Dervan, Peter B
2004-01-01
A pivotal step forward in chemical approaches to controlling gene expression is the development of sequence-specific DNA-binding molecules that can enter live cells and traffic to nuclei unaided. DNA-binding polyamides are a class of programmable, sequence-specific small molecules that have been shown to influence a wide variety of protein-DNA interactions. We have synthesized over 100 polyamide-fluorophore conjugates and assayed their nuclear uptake profiles in 13 mammalian cell lines. The compiled dataset, comprising 1300 entries, establishes a benchmark for the nuclear localization of polyamide-dye conjugates. Compounds in this series were chosen to provide systematic variation in several structural variables, including dye composition and placement, molecular weight, charge, ordering of the aromatic and aliphatic amino-acid building blocks and overall shape. Nuclear uptake does not appear to be correlated with polyamide molecular weight or with the number of imidazole residues, although the positions of imidazole residues affect nuclear access properties significantly. Generally negative determinants for nuclear access include the presence of a beta-Ala-tail residue and the lack of a cationic alkyl amine moiety, whereas the presence of an acetylated 2,4-diaminobutyric acid-turn is a positive factor for nuclear localization. We discuss implications of these data on the design of polyamide-dye conjugates for use in biological systems.
Heritable alteration of DNA methylation induced by whole-chromosome aneuploidy in wheat.
Gao, Lihong; Diarso, Moussa; Zhang, Ai; Zhang, Huakun; Dong, Yuzhu; Liu, Lixia; Lv, Zhenling; Liu, Bao
2016-01-01
Aneuploidy causes changes in gene expression and phenotypes in all organisms studied. A previous study in the model plant Arabidopsis thaliana showed that aneuploidy-generated phenotypic changes can be inherited to euploid progenies and implicated an epigenetic underpinning of the heritable variations. Based on an analysis by amplified fragment length polymorphism and methylation-sensitive amplified fragment length polymorphism markers, we found that although genetic changes at the nucleotide sequence level were negligible, extensive changes in cytosine DNA methylation patterns occurred in all studied homeologous group 1 whole-chromosome aneuploid lines of common wheat (Triticum aestivum), with monosomic 1A showing the greatest amount of methylation changes. The changed methylation patterns were inherited by euploid progenies derived from the aneuploid parents. The aneuploidy-induced DNA methylation alterations and their heritability were verified at selected loci by bisulfite sequencing. Our data have provided empirical evidence supporting earlier suggestions that heritability of aneuploidy-generated, but aneuploidy-independent, phenotypic variations may have an epigenetic basis. That at least one type of aneuploidy - monosomic 1A - was able to cause significant epigenetic divergence of the aneuploid plants and their euploid progenies also lends support to recent suggestions that aneuploidy may have played an important and protracted role in polyploid genome evolution. © 2015 The Authors. New Phytologist © 2015 New Phytologist Trust.
Saad, Rama; Rizkallah, Mariam R; Aziz, Ramy K
2012-11-30
The influence of resident gut microbes on xenobiotic metabolism has been investigated at different levels throughout the past five decades. However, with the advance in sequencing and pyrotagging technologies, addressing the influence of microbes on xenobiotics had to evolve from assessing direct metabolic effects on toxins and botanicals by conventional culture-based techniques to elucidating the role of community composition on drugs metabolic profiles through DNA sequence-based phylogeny and metagenomics. Following the completion of the Human Genome Project, the rapid, substantial growth of the Human Microbiome Project (HMP) opens new horizons for studying how microbiome compositional and functional variations affect drug action, fate, and toxicity (pharmacomicrobiomics), notably in the human gut. The HMP continues to characterize the microbial communities associated with the human gut, determine whether there is a common gut microbiome profile shared among healthy humans, and investigate the effect of its alterations on health. Here, we offer a glimpse into the known effects of the gut microbiota on xenobiotic metabolism, with emphasis on cases where microbiome variations lead to different therapeutic outcomes. We discuss a few examples representing how the microbiome interacts with human metabolic enzymes in the liver and intestine. In addition, we attempt to envisage a roadmap for the future implications of the HMP on therapeutics and personalized medicine.
Accurate typing of short tandem repeats from genome-wide sequencing data and its applications.
Fungtammasan, Arkarachai; Ananda, Guruprasad; Hile, Suzanne E; Su, Marcia Shu-Wei; Sun, Chen; Harris, Robert; Medvedev, Paul; Eckert, Kristin; Makova, Kateryna D
2015-05-01
Short tandem repeats (STRs) are implicated in dozens of human genetic diseases and contribute significantly to genome variation and instability. Yet profiling STRs from short-read sequencing data is challenging because of their high sequencing error rates. Here, we developed STR-FM, short tandem repeat profiling using flank-based mapping, a computational pipeline that can detect the full spectrum of STR alleles from short-read data, can adapt to emerging read-mapping algorithms, and can be applied to heterogeneous genetic samples (e.g., tumors, viruses, and genomes of organelles). We used STR-FM to study STR error rates and patterns in publicly available human and in-house generated ultradeep plasmid sequencing data sets. We discovered that STRs sequenced with a PCR-free protocol have up to ninefold fewer errors than those sequenced with a PCR-containing protocol. We constructed an error correction model for genotyping STRs that can distinguish heterozygous alleles containing STRs with consecutive repeat numbers. Applying our model and pipeline to Illumina sequencing data with 100-bp reads, we could confidently genotype several disease-related long trinucleotide STRs. Utilizing this pipeline, for the first time we determined the genome-wide STR germline mutation rate from a deeply sequenced human pedigree. Additionally, we built a tool that recommends minimal sequencing depth for accurate STR genotyping, depending on repeat length and sequencing read length. The required read depth increases with STR length and is lower for a PCR-free protocol. This suite of tools addresses the pressing challenges surrounding STR genotyping, and thus is of wide interest to researchers investigating disease-related STRs and STR evolution. © 2015 Fungtammasan et al.; Published by Cold Spring Harbor Laboratory Press.
Kretschmer, Rafael; Bertocchi, Natasha Avila; Degrandi, Tiago Marafiga; de Oliveira, Edivaldo Herculano Corrêa; Cioffi, Marcelo de Bello; Garnero, Analía del Valle; Gunski, Ricardo José
2017-01-01
Birds are characterized by a low proportion of repetitive DNA in their genome when compared to other vertebrates. Among birds, species belonging to Piciformes order, such as woodpeckers, show a relatively higher amount of these sequences. The aim of this study was to analyze the distribution of different classes of repetitive DNA—including microsatellites, telomere sequences and 18S rDNA—in the karyotype of three Picidae species (Aves, Piciformes)—Colaptes melanochloros (2n = 84), Colaptes campestris (2n = 84) and Melanerpes candidus (2n = 64)–by means of fluorescence in situ hybridization. Clusters of 18S rDNA were found in one microchromosome pair in each of the three species, coinciding to a region of (CGG)10 sequence accumulation. Interstitial telomeric sequences were found in some macrochromosomes pairs, indicating possible regions of fusions, which can be related to variation of diploid number in the family. Only one, from the 11 different microsatellite sequences used, did not produce any signals. Both species of genus Colaptes showed a similar distribution of microsatellite sequences, with some difference when compared to M. candidus. Microsatellites were found preferentially in the centromeric and telomeric regions of micro and macrochromosomes. However, some sequences produced patterns of interstitial bands in the Z chromosome, which corresponds to the largest element of the karyotype in all three species. This was not observed in the W chromosome of Colaptes melanochloros, which is heterochromatic in most of its length, but was not hybridized by any of the sequences used. These results highlight the importance of microsatellite sequences in differentiation of sex chromosomes, and the accumulation of these sequences is probably responsible for the enlargement of the Z chromosome. PMID:28081238
Formulaic Sequences and the Implications for Second Language Learning
ERIC Educational Resources Information Center
Xu, Qi
2016-01-01
The present paper is a review of literature in relation to formulaic sequences and the implications for second language learning. The formulaic sequence is a significant part of our language, and plays an essential role in both first and second language learning. The paper first introduces the definition, classifications, and major features of…
Bauer, Bianca S.; Forsyth, George W.; Sandmeyer, Lynne S.; Grahn, Bruce H.
2011-01-01
Mitochondrial transcription factor A (Tfam) has been implicated in the pathogenesis of retinal dysplasia in miniature schnauzer dogs and it has been proposed that affected dogs have altered mitochondrial numbers, size, and morphology. To test these hypotheses the Tfam gene of affected and normal miniature schnauzer dogs with retinal dysplasia was sequenced and lymphocyte mitochondria were quantified, measured, and the morphology was compared in normal and affected dogs using transmission electron microscopy. For Tfam sequencing, retina, retinal pigment epithelium (RPE), and whole blood samples were collected. Total RNA was isolated from the retina and RPE and reverse transcribed to make cDNA. Genomic DNA was extracted from white blood cell pellets obtained from the whole blood samples. The Tfam coding sequence, 5′ promoter region, intron1 and the 3′ non-coding sequence of normal and affected dogs were amplified using polymerase chain reaction (PCR), cloned and sequenced. For electron microscopy, lymphocytes from affected and normal dogs were photographed and the mitochondria within each cross-section were identified, quantified, and the mitochondrial area (μm2) per lymphocyte cross-section was calculated. Lastly, using a masked technique, mitochondrial morphology was compared between the 2 groups. Sequencing of the miniature schnauzer Tfam gene revealed no functional sequence variation between affected and normal dogs. Lymphocyte and mitochondrial area, mitochondrial quantification, and morphology assessment also revealed no significant difference between the 2 groups. Further investigation into other candidate genes or factors causing retinal dysplasia in the miniature schnauzer is warranted. PMID:21731185
Bauer, Bianca S; Forsyth, George W; Sandmeyer, Lynne S; Grahn, Bruce H
2011-04-01
Mitochondrial transcription factor A (Tfam) has been implicated in the pathogenesis of retinal dysplasia in miniature schnauzer dogs and it has been proposed that affected dogs have altered mitochondrial numbers, size, and morphology. To test these hypotheses the Tfam gene of affected and normal miniature schnauzer dogs with retinal dysplasia was sequenced and lymphocyte mitochondria were quantified, measured, and the morphology was compared in normal and affected dogs using transmission electron microscopy. For Tfam sequencing, retina, retinal pigment epithelium (RPE), and whole blood samples were collected. Total RNA was isolated from the retina and RPE and reverse transcribed to make cDNA. Genomic DNA was extracted from white blood cell pellets obtained from the whole blood samples. The Tfam coding sequence, 5' promoter region, intron1 and the 3' non-coding sequence of normal and affected dogs were amplified using polymerase chain reaction (PCR), cloned and sequenced. For electron microscopy, lymphocytes from affected and normal dogs were photographed and the mitochondria within each cross-section were identified, quantified, and the mitochondrial area (μm²) per lymphocyte cross-section was calculated. Lastly, using a masked technique, mitochondrial morphology was compared between the 2 groups. Sequencing of the miniature schnauzer Tfam gene revealed no functional sequence variation between affected and normal dogs. Lymphocyte and mitochondrial area, mitochondrial quantification, and morphology assessment also revealed no significant difference between the 2 groups. Further investigation into other candidate genes or factors causing retinal dysplasia in the miniature schnauzer is warranted.
G-CNV: A GPU-Based Tool for Preparing Data to Detect CNVs with Read-Depth Methods.
Manconi, Andrea; Manca, Emanuele; Moscatelli, Marco; Gnocchi, Matteo; Orro, Alessandro; Armano, Giuliano; Milanesi, Luciano
2015-01-01
Copy number variations (CNVs) are the most prevalent types of structural variations (SVs) in the human genome and are involved in a wide range of common human diseases. Different computational methods have been devised to detect this type of SVs and to study how they are implicated in human diseases. Recently, computational methods based on high-throughput sequencing (HTS) are increasingly used. The majority of these methods focus on mapping short-read sequences generated from a donor against a reference genome to detect signatures distinctive of CNVs. In particular, read-depth based methods detect CNVs by analyzing genomic regions with significantly different read-depth from the other ones. The pipeline analysis of these methods consists of four main stages: (i) data preparation, (ii) data normalization, (iii) CNV regions identification, and (iv) copy number estimation. However, available tools do not support most of the operations required at the first two stages of this pipeline. Typically, they start the analysis by building the read-depth signal from pre-processed alignments. Therefore, third-party tools must be used to perform most of the preliminary operations required to build the read-depth signal. These data-intensive operations can be efficiently parallelized on graphics processing units (GPUs). In this article, we present G-CNV, a GPU-based tool devised to perform the common operations required at the first two stages of the analysis pipeline. G-CNV is able to filter low-quality read sequences, to mask low-quality nucleotides, to remove adapter sequences, to remove duplicated read sequences, to map the short-reads, to resolve multiple mapping ambiguities, to build the read-depth signal, and to normalize it. G-CNV can be efficiently used as a third-party tool able to prepare data for the subsequent read-depth signal generation and analysis. Moreover, it can also be integrated in CNV detection tools to generate read-depth signals.
Extraordinary Sequence Divergence at Tsga8, an X-linked Gene Involved in Mouse Spermiogenesis
Good, Jeffrey M.; Vanderpool, Dan; Smith, Kimberly L.; Nachman, Michael W.
2011-01-01
The X chromosome plays an important role in both adaptive evolution and speciation. We used a molecular evolutionary screen of X-linked genes potentially involved in reproductive isolation in mice to identify putative targets of recurrent positive selection. We then sequenced five very rapidly evolving genes within and between several closely related species of mice in the genus Mus. All five genes were involved in male reproduction and four of the genes showed evidence of recurrent positive selection. The most remarkable evolutionary patterns were found at Testis-specific gene a8 (Tsga8), a spermatogenesis-specific gene expressed during postmeiotic chromatin condensation and nuclear transformation. Tsga8 was characterized by extremely high levels of insertion–deletion variation of an alanine-rich repetitive motif in natural populations of Mus domesticus and M. musculus, differing in length from the reference mouse genome by up to 89 amino acids (27% of the total protein length). This population-level variation was coupled with striking divergence in protein sequence and length between closely related mouse species. Although no clear orthologs had previously been described for Tsga8 in other mammalian species, we have identified a highly divergent hypothetical gene on the rat X chromosome that shares clear orthology with the 5′ and 3′ ends of Tsga8. Further inspection of this ortholog verified that it is expressed in rat testis and shares remarkable similarity with mouse Tsga8 across several general features of the protein sequence despite no conservation of nucleotide sequence across over 60% of the rat-coding domain. Overall, Tsga8 appears to be one of the most rapidly evolving genes to have been described in rodents. We discuss the potential evolutionary causes and functional implications of this extraordinary divergence and the possible contribution of Tsga8 and the other four genes we examined to reproductive isolation in mice. PMID:21186189
Sabree, Zakee L; Hansen, Allison K; Moran, Nancy A
2012-01-01
Starting in 2003, numerous studies using culture-independent methodologies to characterize the gut microbiota of honey bees have retrieved a consistent and distinctive set of eight bacterial species, based on near identity of the 16S rRNA gene sequences. A recent study [Mattila HR, Rios D, Walker-Sperling VE, Roeselers G, Newton ILG (2012) Characterization of the active microbiotas associated with honey bees reveals healthier and broader communities when colonies are genetically diverse. PLoS ONE 7(3): e32962], using pyrosequencing of the V1-V2 hypervariable region of the 16S rRNA gene, reported finding entirely novel bacterial species in honey bee guts, and used taxonomic assignments from these reads to predict metabolic activities based on known metabolisms of cultivable species. To better understand this discrepancy, we analyzed the Mattila et al. pyrotag dataset. In contrast to the conclusions of Mattila et al., we found that the large majority of pyrotag sequences belonged to clusters for which representative sequences were identical to sequences from previously identified core species of the bee microbiota. On average, they represent 95% of the bacteria in each worker bee in the Mattila et al. dataset, a slightly lower value than that found in other studies. Some colonies contain small proportions of other bacteria, mostly species of Enterobacteriaceae. Reanalysis of the Mattila et al. dataset also did not support a relationship between abundances of Bifidobacterium and of putative pathogens or a significant difference in gut communities between colonies from queens that were singly or multiply mated. Additionally, consistent with previous studies, the dataset supports the occurrence of considerable strain variation within core species, even within single colonies. The roles of these bacteria within bees, or the implications of the strain variation, are not yet clear.
Molecular Population Genetics of the Alcohol Dehydrogenase Gene Region of DROSOPHILA MELANOGASTER
Aquadro, Charles F.; Desse, Susan F.; Bland, Molly M.; Langley, Charles H.; Laurie-Ahlberg, Cathy C.
1986-01-01
Variation in the DNA restriction map of a 13-kb region of chromosome II including the alcohol dehydrogenase structural gene (Adh) was examined in Drosophila melanogaster from natural populations. Detailed analysis of 48 D. melanogaster lines representing four eastern United States populations revealed extensive DNA sequence variation due to base substitutions, insertions and deletions. Cloning of this region from several lines allowed characterization of length variation as due to unique sequence insertions or deletions [nine sizes; 21–200 base pairs (bp)] or transposable element insertions (several sizes, 340 bp to 10.2 kb, representing four different elements). Despite this extensive variation in sequences flanking the Adh gene, only one length polymorphism is clearly associated with altered Adh expression (a copia element approximately 250 bp 5' to the distal transcript start site). Nonetheless, the frequency spectra of transposable elements within and between Drosophila species suggests they are slightly deleterious. Strong nonrandom associations are observed among Adh region sequence variants, ADH allozyme (Fast vs. Slow), ADH enzyme activity and the chromosome inversion ln(2L) t. Phylogenetic analysis of restriction map haplotypes suggest that the major twofold component of ADH activity variation (high vs. low, typical of Fast and Slow allozymes, respectively) is due to sequence variation tightly linked to and possibly distinct from that underlying the allozyme difference. The patterns of nucleotide and haplotype variation for Fast and Slow allozyme lines are consistent with the recent increase in frequency and spread of the Fast haplotype associated with high ADH activity. These data emphasize the important role of evolutionary history and strong nonrandom associations among tightly linked sequence variation as determinants of the patterns of variation observed in natural populations. PMID:3026893
Variation, Repetition, And Choice
Abreu-Rodrigues, Josele; Lattal, Kennon A; dos Santos, Cristiano V; Matos, Ricardo A
2005-01-01
Experiment 1 investigated the controlling properties of variability contingencies on choice between repeated and variable responding. Pigeons were exposed to concurrent-chains schedules with two alternatives. In the REPEAT alternative, reinforcers in the terminal link depended on a single sequence of four responses. In the VARY alternative, a response sequence in the terminal link was reinforced only if it differed from the n previous sequences (lag criterion). The REPEAT contingency generated low, constant levels of sequence variation whereas the VARY contingency produced levels of sequence variation that increased with the lag criterion. Preference for the REPEAT alternative tended to increase directly with the degree of variation required for reinforcement. Experiment 2 examined the potential confounding effects in Experiment 1 of immediacy of reinforcement by yoking the interreinforcer intervals in the REPEAT alternative to those in the VARY alternative. Again, preference for REPEAT was a function of the lag criterion. Choice between varying and repeating behavior is discussed with respect to obtained behavioral variability, probability of reinforcement, delay of reinforcement, and switching within a sequence. PMID:15828592
The genetic breakdown of sporophytic self-incompatibility in Tolpis coronopifolia (Asteraceae).
Koseva, Boryana; Crawford, Daniel J; Brown, Keely E; Mort, Mark E; Kelly, John K
2017-12-01
Angiosperm diversity has been shaped by mating system evolution, with the most common transition from outcrossing to self-fertilizing. To investigate the genetic basis of this transition, we performed crosses between two species endemic to the Canary Islands, the self-compatible (SC) species Tolpis coronopifolia and its self-incompatible (SI) relative Tolpis santosii. We scored self-compatibility as self-seed set of recombinant plants within two F 2 populations. To map and genetically characterize the breakdown of SI, we built a draft genome sequence of T. coronopifolia, genotyped F 2 plants using multiplexed shotgun genotyping (MSG), and located MSG markers to the genome sequence. We identified a single quantitative trait locus (QTL) that explains nearly all variation in self-seed set in both F 2 populations. To identify putative causal genetic variants within the QTL, we performed transcriptome sequencing on mature floral tissue from both SI and SC species, constructed a transcriptome for each species, and then located each predicted transcript to the T. coronopifolia genome sequence. We annotated each predicted gene within the QTL and found two strong candidates for SI breakdown. Each gene has a coding sequence insertion/deletion mutation within the SC species that produces a truncated protein. Homologs of each gene have been implicated in pollen development, pollen germination, and pollen tube growth in other species. © 2017 The Authors. New Phytologist © 2017 New Phytologist Trust.
Mejia-Velasquez, Paula J; Dilcher, David L; Jaramillo, Carlos A; Fortini, Lucas B; Manchester, Steven R
2012-11-01
Reconstruction of floristic patterns during the early diversification of angiosperms is impeded by the scarce fossil record, especially in tropical latitudes. Here we collected quantitative palynological data from a stratigraphic sequence in tropical South America to provide floristic and climatic insights into such tropical environments during the Early Cretaceous. We reconstructed the floristic composition of an Aptian-Albian tropical sequence from central Colombia using quantitative palynology (rarefied species richness and abundance) and used it to infer its predominant climatic conditions. Additionally, we compared our results with available quantitative data from three other sequences encompassing 70 floristic assemblages to determine latitudinal diversity patterns. Abundance of humidity indicators was higher than that of aridity indicators (61% vs. 10%). Additionally, we found an angiosperm latitudinal diversity gradient (LDG) for the Aptian, but not for the Albian, and an inverted LDG of the overall diversity for the Albian. Angiosperm species turnover during the Albian, however, was higher in humid tropics. There were humid climates in northwestern South America during the Aptian-Albian interval contrary to the widespread aridity expected for the tropical belt. The Albian inverted overall LDG is produced by a faster increase in per-sample angiosperm and pteridophyte diversity in temperate latitudes. However, humid tropical sequences had higher rates of floristic turnover suggesting a higher degree of morphological variation than in temperate regions.
Mejia-Velasquez, Paula J.; Dilcher, David L.; Jaramillo, Carlos A.; Fortini, Lucas B.; Manchester, Steven R.
2012-01-01
Premise of the study: Reconstruction of floristic patterns during the early diversification of angiosperms is impeded by the scarce fossil record, especially in tropical latitudes. Here we collected quantitative palynological data from a stratigraphic sequence in tropical South America to provide floristic and climatic insights into such tropical environments during the Early Cretaceous. Methods: We reconstructed the floristic composition of an Aptian-Albian tropical sequence from central Colombia using quantitative palynology (rarefied species richness and abundance) and used it to infer its predominant climatic conditions. Additionally, we compared our results with available quantitative data from three other sequences encompassing 70 floristic assemblages to determine latitudinal diversity patterns. Key results: Abundance of humidity indicators was higher than that of aridity indicators (61% vs. 10%). Additionally, we found an angiosperm latitudinal diversity gradient (LDG) for the Aptian, but not for the Albian, and an inverted LDG of the overall diversity for the Albian. Angiosperm species turnover during the Albian, however, was higher in humid tropics. Conclusions: There were humid climates in northwestern South America during the Aptian-Albian interval contrary to the widespread aridity expected for the tropical belt. The Albian inverted overall LDG is produced by a faster increase in per-sample angiosperm and pteridophyte diversity in temperate latitudes. However, humid tropical sequences had higher rates of floristic turnover suggesting a higher degree of morphological variation than in temperate regions.
Spiske, M.; Jaffe, B.E.
2009-01-01
Storms and associated surges are major coast-shaping processes. Nevertheless, no typical sequences for storm surge deposits in different coastal settings have been established. This study interprets a coarse-grained hurricane ridge deposit on the island of Bonaire, Netherlands Antilles. The sequence was deposited during Hurricane Lenny in November 1999. Insight is gained into the hydrodynamics of surge flow by interpreting textural trends, particle imbrication, and deposit geometry. Vertical textural variations, caused by time-dependent hydrodynamic changes, were used to subdivide the deposit into depositional units that correspond to different stages of the surge, such as setup, peak, and return flow. Particle size and imbrication trends and geometry of the units reflect landward bed-load transport of components during the setup, a nondirectional flow with sediment falling out of suspension during the peak, and a seaward bedload transport during the return flow. Formation of a ridge during setup affected the texture of the return flow unit. Changing angles of imbrication reflect alternating flow velocities during each phase. Normal grading during setup and inverse grading during return flow are caused by decelerating and accelerating flow, respectively. Hence, the interpreted deposit seems to represent the first described complete hurricane surge sequence from a carbonate environment. ?? 2009 Geological Society of America.
Kumar, Pankaj; Chaitanya, Pasumarthy S; Nagarajaram, Hampapathalu A
2011-01-01
PSSRdb (Polymorphic Simple Sequence Repeats database) (http://www.cdfd.org.in/PSSRdb/) is a relational database of polymorphic simple sequence repeats (PSSRs) extracted from 85 different species of prokaryotes. Simple sequence repeats (SSRs) are the tandem repeats of nucleotide motifs of the sizes 1-6 bp and are highly polymorphic. SSR mutations in and around coding regions affect transcription and translation of genes. Such changes underpin phase variations and antigenic variations seen in some bacteria. Although SSR-mediated phase variation and antigenic variations have been well-studied in some bacteria there seems a lot of other species of prokaryotes yet to be investigated for SSR mediated adaptive and other evolutionary advantages. As a part of our on-going studies on SSR polymorphism in prokaryotes we compared the genome sequences of various strains and isolates available for 85 different species of prokaryotes and extracted a number of SSRs showing length variations and created a relational database called PSSRdb. This database gives useful information such as location of PSSRs in genomes, length variation across genomes, the regions harboring PSSRs, etc. The information provided in this database is very useful for further research and analysis of SSRs in prokaryotes.
Attwood, Stephen W.; Fatih, Farrah A.; Upatham, E. Suchart
2008-01-01
Background Schistosomiasis in humans along the lower Mekong River has proven a persistent public health problem in the region. The causative agent is the parasite Schistosoma mekongi (Trematoda: Digenea). A new transmission focus is reported, as well as the first study of genetic variation among S. mekongi populations. The aim is to confirm the identity of the species involved at each known focus of Mekong schistosomiasis transmission, to examine historical relationships among the populations and related taxa, and to provide data for use (a priori) in further studies of the origins, radiation, and future dispersal capabilities of S. mekongi. Methodology/Principal Findings DNA sequence data are presented for four populations of S. mekongi from Cambodia and southern Laos, three of which were distinguishable at the COI (cox1) and 12S (rrnS) mitochondrial loci sampled. A phylogeny was estimated for these populations and the other members of the Schistosoma sinensium group. The study provides new DNA sequence data for three new populations and one new locus/population combination. A Bayesian approach is used to estimate divergence dates for events within the S. sinensium group and among the S. mekongi populations. Conclusions/Significance The date estimates are consistent with phylogeographical hypotheses describing a Pliocene radiation of the S. sinensium group and a mid-Pleistocene invasion of Southeast Asia by S. mekongi. The date estimates also provide Bayesian priors for future work on the evolution of S. mekongi. The public health implications of S. mekongi transmission outside the lower Mekong River are also discussed. PMID:18350111
Kim, Sang Hu; Clark, Shawn T.; Surendra, Anuradha; Copeland, Julia K.; Wang, Pauline W.; Ammar, Ron; Collins, Cathy; Tullis, D. Elizabeth; Nislow, Corey; Hwang, David M.; Guttman, David S.; Cowen, Leah E.
2015-01-01
The microbiome shapes diverse facets of human biology and disease, with the importance of fungi only beginning to be appreciated. Microbial communities infiltrate diverse anatomical sites as with the respiratory tract of healthy humans and those with diseases such as cystic fibrosis, where chronic colonization and infection lead to clinical decline. Although fungi are frequently recovered from cystic fibrosis patient sputum samples and have been associated with deterioration of lung function, understanding of species and population dynamics remains in its infancy. Here, we coupled high-throughput sequencing of the ribosomal RNA internal transcribed spacer 1 (ITS1) with phenotypic and genotypic analyses of fungi from 89 sputum samples from 28 cystic fibrosis patients. Fungal communities defined by sequencing were concordant with those defined by culture-based analyses of 1,603 isolates from the same samples. Different patients harbored distinct fungal communities. There were detectable trends, however, including colonization with Candida and Aspergillus species, which was not perturbed by clinical exacerbation or treatment. We identified considerable inter- and intra-species phenotypic variation in traits important for host adaptation, including antifungal drug resistance and morphogenesis. While variation in drug resistance was largely between species, striking variation in morphogenesis emerged within Candida species. Filamentation was uncoupled from inducing cues in 28 Candida isolates recovered from six patients. The filamentous isolates were resistant to the filamentation-repressive effects of Pseudomonas aeruginosa, implicating inter-kingdom interactions as the selective force. Genome sequencing revealed that all but one of the filamentous isolates harbored mutations in the transcriptional repressor NRG1; such mutations were necessary and sufficient for the filamentous phenotype. Six independent nrg1 mutations arose in Candida isolates from different patients, providing a poignant example of parallel evolution. Together, this combined clinical-genomic approach provides a high-resolution portrait of the fungal microbiome of cystic fibrosis patient lungs and identifies a genetic basis of pathogen adaptation. PMID:26588216
Kim, Sang Hu; Clark, Shawn T; Surendra, Anuradha; Copeland, Julia K; Wang, Pauline W; Ammar, Ron; Collins, Cathy; Tullis, D Elizabeth; Nislow, Corey; Hwang, David M; Guttman, David S; Cowen, Leah E
2015-11-01
The microbiome shapes diverse facets of human biology and disease, with the importance of fungi only beginning to be appreciated. Microbial communities infiltrate diverse anatomical sites as with the respiratory tract of healthy humans and those with diseases such as cystic fibrosis, where chronic colonization and infection lead to clinical decline. Although fungi are frequently recovered from cystic fibrosis patient sputum samples and have been associated with deterioration of lung function, understanding of species and population dynamics remains in its infancy. Here, we coupled high-throughput sequencing of the ribosomal RNA internal transcribed spacer 1 (ITS1) with phenotypic and genotypic analyses of fungi from 89 sputum samples from 28 cystic fibrosis patients. Fungal communities defined by sequencing were concordant with those defined by culture-based analyses of 1,603 isolates from the same samples. Different patients harbored distinct fungal communities. There were detectable trends, however, including colonization with Candida and Aspergillus species, which was not perturbed by clinical exacerbation or treatment. We identified considerable inter- and intra-species phenotypic variation in traits important for host adaptation, including antifungal drug resistance and morphogenesis. While variation in drug resistance was largely between species, striking variation in morphogenesis emerged within Candida species. Filamentation was uncoupled from inducing cues in 28 Candida isolates recovered from six patients. The filamentous isolates were resistant to the filamentation-repressive effects of Pseudomonas aeruginosa, implicating inter-kingdom interactions as the selective force. Genome sequencing revealed that all but one of the filamentous isolates harbored mutations in the transcriptional repressor NRG1; such mutations were necessary and sufficient for the filamentous phenotype. Six independent nrg1 mutations arose in Candida isolates from different patients, providing a poignant example of parallel evolution. Together, this combined clinical-genomic approach provides a high-resolution portrait of the fungal microbiome of cystic fibrosis patient lungs and identifies a genetic basis of pathogen adaptation.
Reicher, S; Seroussi, E; Weller, J I; Rosov, A; Gootwine, E
2012-07-01
Polymorphisms in mitochondrial DNA (mtDNA) protein- and tRNA-coding genes were shown to be associated with various diseases in humans as well as with production and reproduction traits in livestock. Alignment of full length mitochondria sequences from the 5 known ovine haplogroups: HA (n = 3), HB (n = 5), HC (n = 3), HD (n = 2), and HE (n = 2; GenBank accession nos. HE577847-50 and 11 published complete ovine mitochondria sequences) revealed sequence variation in 10 out of the 13 protein coding mtDNA sequences. Twenty-six of the 245 variable sites found in the protein coding sequences represent non-synonymous mutations. Sequence variation was observed also in 8 out of the 22 tRNA mtDNA sequences. On the basis of the mtDNA control region and cytochrome b partial sequences along with information on maternal lineages within an Afec-Assaf flock, 1,126 Afec-Assaf ewes were assigned to mitochondrial haplogroups HA, HB, and HC, with frequencies of 0.43, 0.43, and 0.14, respectively. Analysis of birth weight and growth rate records of lamb (n = 1286) and productivity from 4,993 lambing records revealed no association between mitochondrial haplogroup affiliation and female longevity, lambs perinatal survival rate, birth weight, and daily growth rate of lambs up to 150 d that averaged 1,664 d, 88.3%, 4.5 kg, and 320 g/d, respectively. However, significant (P < 0.0001) differences among the haplogroups were found for prolificacy of ewes, with prolificacies (mean ± SE) of 2.14 ± 0.04, 2.25 ± 0.04, and 2.30 ± 0.06 lamb born/ewe lambing for the HA, HB, and the HC haplogroups, respectively. Our results highlight the ovine mitogenome genetic variation in protein- and tRNA coding genes and suggest that sequence variation in ovine mtDNA is associated with variation in ewe prolificacy.
HIV-1 sequence variation between isolates from mother-infant transmission pairs
DOE Office of Scientific and Technical Information (OSTI.GOV)
Wike, C.M.; Daniels, M.R.; Furtado, M.
1991-12-31
To examine the sequence diversity of human immunodeficiency virus type 1 (HIV-1) between known transmission sets, sequences from the V3 and V4-V5 region of the env gene from 4 mother-infant pairs were analyzed. The mean interpatient sequence variation between isolates from linked mother-infant pairs was comparable to the sequence diversity found between isolates from other close contacts. The mean intrapatient variation was significantly less in the infants` isolates then the isolates from both their mothers and other characterized intrapatient sequence sets. In addition, a distinct and characteristic difference in the glycosylation pattern preceding the V3 loop was found between eachmore » linked transmission pair. These findings indicate that selection of specific genotypic variants, which may play a role in some direct transmission sets, and the duration of infection are important factors in the degree of diversity seen between the sequence sets.« less
Extent, Causes, and Consequences of Small RNA Expression Variation in Human Adipose Tissue
Knights, Andrew J.; Abreu-Goodger, Cei; van de Bunt, Martijn; Guerra-Assunção, José Afonso; Bartonicek, Nenad; van Dongen, Stijn; Mägi, Reedik; Nisbet, James; Barrett, Amy; Rantalainen, Mattias; Nica, Alexandra C.; Quail, Michael A.; Small, Kerrin S.; Glass, Daniel; Enright, Anton J.; Winn, John; Deloukas, Panos; Dermitzakis, Emmanouil T.; McCarthy, Mark I.; Spector, Timothy D.; Durbin, Richard; Lindgren, Cecilia M.
2012-01-01
Small RNAs are functional molecules that modulate mRNA transcripts and have been implicated in the aetiology of several common diseases. However, little is known about the extent of their variability within the human population. Here, we characterise the extent, causes, and effects of naturally occurring variation in expression and sequence of small RNAs from adipose tissue in relation to genotype, gene expression, and metabolic traits in the MuTHER reference cohort. We profiled the expression of 15 to 30 base pair RNA molecules in subcutaneous adipose tissue from 131 individuals using high-throughput sequencing, and quantified levels of 591 microRNAs and small nucleolar RNAs. We identified three genetic variants and three RNA editing events. Highly expressed small RNAs are more conserved within mammals than average, as are those with highly variable expression. We identified 14 genetic loci significantly associated with nearby small RNA expression levels, seven of which also regulate an mRNA transcript level in the same region. In addition, these loci are enriched for variants significant in genome-wide association studies for body mass index. Contrary to expectation, we found no evidence for negative correlation between expression level of a microRNA and its target mRNAs. Trunk fat mass, body mass index, and fasting insulin were associated with more than twenty small RNA expression levels each, while fasting glucose had no significant associations. This study highlights the similar genetic complexity and shared genetic control of small RNA and mRNA transcripts, and gives a quantitative picture of small RNA expression variation in the human population. PMID:22589741
Doucet, S.M.; McDonald, D.B.; Foster, M.S.; Clay, R.P.
2007-01-01
Lek-mating Long-tailed Manakins (Chiroxiphia linearis) exhibit an unusual pattern of delayed plumage maturation. Each year, males progress through a series of predefinitive plumages before attaining definitive plumage in their fifth calendar year. Females also exhibit variation in plumage coloration, with some females displaying male-like plumage characteristics. Using data from mist-net captures in northwest Costa Rica (n = 1,315) and museum specimens from throughout the range of Long-tailed Manakins (n = 585), we documented the plumage sequence progression of males, explored variation in female plumage, and described the timing of molt in this species. Males progressed through a series of age-specific predefinitive plumages, which enabled the accurate aging of predefinitive-plumaged males in the field; this predefinitive plumage sequence is the basis for age-related status-signaling in these males. Females tended to acquire red coloration in the crown as they aged. However, colorful plumage in females may be a byproduct of selection on bright male plumage. Females exhibited an early peak of molt activity from February to April, little molt from May through July, and a second, more pronounced peak of molt activity in October. By contrast, males in older predefinitive-plumage stages and males in definitive plumage exhibited comparable unimodal distributions in molt activity beginning in June and peaking between July and October. Our data are consistent with selective pressure to avoid the costs of molt-breeding overlap in females and older males. Our findings have important implications for social organization and signaling in Long- tailed Manakins, and for the evolution of delayed plumage maturation in birds.
Maryam, J; Babar, M E; Bao, Zhang; Nadeem, A
2016-10-01
Modern molecular interventions are dynamic gears for breeding animals with superior genetic make-up. These scientific efforts lead us toward sustainable dairy herds with improved milk production in terms of yield and quality. Many of candidate genes have been dissected at molecular level, and suitable genetic markers have been identified in cattle, but this work has not been validated in buffaloes so far. Stearoyl-coenzyme A desaturase (SCD) has been a potential candidate gene for fat content of milk. Genomic analysis of SCD revealed a total of six variations that were identified through DNA sequencing of animals with lower and higher butter fat %age. After statistical analysis, genotype AB of p.K158I could be associated (P value <0.0001) with higher milk fat %age (10.5 ± 0.5464). This SNP was validated on larger data set by cleaved amplified polymorphic sequences (CAPS) by using DdeI. To scrutinize the functional consequences of p.K158I, 3D protein structure of SCD was predicted by homology modeling and this variation was found located in the vicinity of functional domain and a part of transmembrane helix of this membrane integrated protein. This is a first report toward genetic screening of SCD gene at molecular level in buffalo. This report illustrates the implication of SCD gene and in particular p.K158I variation, in imparting its effect on milk fat %age, which can be targeted in selection of superior dairy buffaloes.
Talkowski, Michael E; Ernst, Carl; Heilbut, Adrian; Chiang, Colby; Hanscom, Carrie; Lindgren, Amelia; Kirby, Andrew; Liu, Shangtao; Muddukrishna, Bhavana; Ohsumi, Toshiro K; Shen, Yiping; Borowsky, Mark; Daly, Mark J; Morton, Cynthia C; Gusella, James F
2011-04-08
The contribution of balanced chromosomal rearrangements to complex disorders remains unclear because they are not detected routinely by genome-wide microarrays and clinical localization is imprecise. Failure to consider these events bypasses a potentially powerful complement to single nucleotide polymorphism and copy-number association approaches to complex disorders, where much of the heritability remains unexplained. To capitalize on this genetic resource, we have applied optimized sequencing and analysis strategies to test whether these potentially high-impact variants can be mapped at reasonable cost and throughput. By using a whole-genome multiplexing strategy, rearrangement breakpoints could be delineated at a fraction of the cost of standard sequencing. For rearrangements already mapped regionally by karyotyping and fluorescence in situ hybridization, a targeted approach enabled capture and sequencing of multiple breakpoints simultaneously. Importantly, this strategy permitted capture and unique alignment of up to 97% of repeat-masked sequences in the targeted regions. Genome-wide analyses estimate that only 3.7% of bases should be routinely omitted from genomic DNA capture experiments. Illustrating the power of these approaches, the rearrangement breakpoints were rapidly defined to base pair resolution and revealed unexpected sequence complexity, such as co-occurrence of inversion and translocation as an underlying feature of karyotypically balanced alterations. These findings have implications ranging from genome annotation to de novo assemblies and could enable sequencing screens for structural variations at a cost comparable to that of microarrays in standard clinical practice. Copyright © 2011 The American Society of Human Genetics. Published by Elsevier Inc. All rights reserved.
The Genetic Architecture of Natural Variation in Recombination Rate in Drosophila melanogaster.
Hunter, Chad M; Huang, Wen; Mackay, Trudy F C; Singh, Nadia D
2016-04-01
Meiotic recombination ensures proper chromosome segregation in many sexually reproducing organisms. Despite this crucial function, rates of recombination are highly variable within and between taxa, and the genetic basis of this variation remains poorly understood. Here, we exploit natural variation in the inbred, sequenced lines of the Drosophila melanogaster Genetic Reference Panel (DGRP) to map genetic variants affecting recombination rate. We used a two-step crossing scheme and visible markers to measure rates of recombination in a 33 cM interval on the X chromosome and in a 20.4 cM interval on chromosome 3R for 205 DGRP lines. Though we cannot exclude that some biases exist due to viability effects associated with the visible markers used in this study, we find ~2-fold variation in recombination rate among lines. Interestingly, we further find that recombination rates are uncorrelated between the two chromosomal intervals. We performed a genome-wide association study to identify genetic variants associated with recombination rate in each of the two intervals surveyed. We refined our list of candidate variants and genes associated with recombination rate variation and selected twenty genes for functional assessment. We present strong evidence that five genes are likely to contribute to natural variation in recombination rate in D. melanogaster; these genes lie outside the canonical meiotic recombination pathway. We also find a weak effect of Wolbachia infection on recombination rate and we confirm the interchromosomal effect. Our results highlight the magnitude of population variation in recombination rate present in D. melanogaster and implicate new genetic factors mediating natural variation in this quantitative trait.
The Genetic Architecture of Natural Variation in Recombination Rate in Drosophila melanogaster
Hunter, Chad M.; Huang, Wen; Mackay, Trudy F. C.; Singh, Nadia D.
2016-01-01
Meiotic recombination ensures proper chromosome segregation in many sexually reproducing organisms. Despite this crucial function, rates of recombination are highly variable within and between taxa, and the genetic basis of this variation remains poorly understood. Here, we exploit natural variation in the inbred, sequenced lines of the Drosophila melanogaster Genetic Reference Panel (DGRP) to map genetic variants affecting recombination rate. We used a two-step crossing scheme and visible markers to measure rates of recombination in a 33 cM interval on the X chromosome and in a 20.4 cM interval on chromosome 3R for 205 DGRP lines. Though we cannot exclude that some biases exist due to viability effects associated with the visible markers used in this study, we find ~2-fold variation in recombination rate among lines. Interestingly, we further find that recombination rates are uncorrelated between the two chromosomal intervals. We performed a genome-wide association study to identify genetic variants associated with recombination rate in each of the two intervals surveyed. We refined our list of candidate variants and genes associated with recombination rate variation and selected twenty genes for functional assessment. We present strong evidence that five genes are likely to contribute to natural variation in recombination rate in D. melanogaster; these genes lie outside the canonical meiotic recombination pathway. We also find a weak effect of Wolbachia infection on recombination rate and we confirm the interchromosomal effect. Our results highlight the magnitude of population variation in recombination rate present in D. melanogaster and implicate new genetic factors mediating natural variation in this quantitative trait. PMID:27035832
Methylation of avpr1a in the cortex of wild prairie voles: effects of CpG position and polymorphism
Maguire, S. M.; Phelps, S. M.
2017-01-01
DNA methylation can cause stable changes in neuronal gene expression, but we know little about its role in individual differences in the wild. In this study, we focus on the vasopressin 1a receptor (avpr1a), a gene extensively implicated in vertebrate social behaviour, and explore natural variation in DNA methylation, genetic polymorphism and neuronal gene expression among 30 wild prairie voles (Microtus ochrogaster). Examination of CpG density across 8 kb of the locus revealed two distinct CpG islands overlapping promoter and first exon, characterized by few CpG polymorphisms. We used a targeted bisulfite sequencing approach to measure DNA methylation across approximately 3 kb of avpr1a in the retrosplenial cortex, a brain region implicated in male space use and sexual fidelity. We find dramatic variation in methylation across the avrp1a locus, with pronounced diversity near the exon–intron boundary and in a genetically variable putative enhancer within the intron. Among our wild voles, differences in cortical avpr1a expression correlate with DNA methylation in this putative enhancer, but not with the methylation status of the promoter. We also find an unusually high number of polymorphic CpG sites (polyCpGs) in this focal enhancer. One polyCpG within this enhancer (polyCpG 2170) may drive variation in expression either by disrupting transcription factor binding motifs or by changing local DNA methylation and chromatin silencing. Our results contradict some assumptions made within behavioural epigenetics, but are remarkably concordant with genome-wide studies of gene regulation. PMID:28280564
Equivalent Indels – Ambiguous Functional Classes and Redundancy in Databases
Assmus, Jens; Kleffe, Jürgen; Schmitt, Armin O.; Brockmann, Gudrun A.
2013-01-01
There is considerable interest in studying sequenced variations. However, while the positions of substitutions are uniquely identifiable by sequence alignment, the location of insertions and deletions still poses problems. Each insertion and deletion causes a change of sequence. Yet, due to low complexity or repetitive sequence structures, the same indel can sometimes be annotated in different ways. Two indels which differ in allele sequence and position can be one and the same, i.e. the alternative sequence of the whole chromosome is identical in both cases and, therefore, the two deletions are biologically equivalent. In such a case, it is impossible to identify the exact position of an indel merely based on sequence alignment. Thus, variation entries in a mutation database are not necessarily uniquely defined. We prove the existence of a contiguous region around an indel in which all deletions of the same length are biologically identical. Databases often show only one of several possible locations for a given variation. Furthermore, different data base entries can represent equivalent variation events. We identified 1,045,590 such problematic entries of insertions and deletions out of 5,860,408 indel entries in the current human database of Ensembl. Equivalent indels are found in sequence regions of different functions like exons, introns or 5' and 3' UTRs. One and the same variation can be assigned to several different functional classifications of which only one is correct. We implemented an algorithm that determines for each indel database entry its complete set of equivalent indels which is uniquely characterized by the indel itself and a given interval of the reference sequence. PMID:23658777
McCutchen-Maloney, Sandra L.
2002-01-01
DNA mutation binding proteins alone and as chimeric proteins with nucleases are used with solid supports to detect DNA sequence variations, DNA mutations and single nucleotide polymorphisms. The solid supports may be flow cytometry beads, DNA chips, glass slides or DNA dips sticks. DNA molecules are coupled to solid supports to form DNA-support complexes. Labeled DNA is used with unlabeled DNA mutation binding proteins such at TthMutS to detect DNA sequence variations, DNA mutations and single nucleotide length polymorphisms by binding which gives an increase in signal. Unlabeled DNA is utilized with labeled chimeras to detect DNA sequence variations, DNA mutations and single nucleotide length polymorphisms by nuclease activity of the chimera which gives a decrease in signal.
Setoh, Yin Xiang; Amarilla, Alberto A; Peng, Nias Y; Slonchak, Andrii; Periasamy, Parthiban; Figueiredo, Luiz T M; Aquino, Victor H; Khromykh, Alexander A
2018-01-01
Rocio virus (ROCV) is an arbovirus belonging to the genus Flavivirus, family Flaviviridae. We present an updated sequence of ROCV strain SPH 34675 (GenBank: AY632542.4), the only available full genome sequence prior to this study. Using next-generation sequencing of the entire genome, we reveal substantial sequence variation from the prototype sequence, with 30 nucleotide differences amounting to 14 amino acid changes, as well as significant changes to predicted 3'UTR RNA structures. Our results present an updated and corrected sequence of a potential emerging human-virulent flavivirus uniquely indigenous to Brazil (GenBank: MF461639).
Numerical studies of asymmetric adiabatic accretion flow - The effect of velocity gradients
NASA Technical Reports Server (NTRS)
Taam, Ronald E.; Fryxell, B. A.
1989-01-01
A numerical study of the time variation of the angular momentum and mass capture rates for a central object accreting from a uniform medium with a velocity gradient transverse to the direction of the mean flow is presented, covering a range of velocity asymmetries and Mach numbers in the incident flow. It is found that the mass accretion rate in a given evolutionary sequence varies in an irregular manner, with the matter accreting onto the central object from either a continuously moving accretion wake or from an accretion disk. The implications of the results from the study of short-term fluctuations observed in the pulse period and luminosity of X-ray pulsars are discussed.
Holeski, Liza M; Monnahan, Patrick; Koseva, Boryana; McCool, Nick; Lindroth, Richard L; Kelly, John K
2014-03-13
Genotyping-by-sequencing methods have vastly improved the resolution and accuracy of genetic linkage maps by increasing both the number of marker loci as well as the number of individuals genotyped at these loci. Using restriction-associated DNA sequencing, we construct a dense linkage map for a panel of recombinant inbred lines derived from a cross between divergent ecotypes of Mimulus guttatus. We used this map to estimate recombination rate across the genome and to identify quantitative trait loci for the production of several secondary compounds (PPGs) of the phenylpropanoid pathway implicated in defense against herbivores. Levels of different PPGs are correlated across recombinant inbred lines suggesting joint regulation of the phenylpropanoid pathway. However, the three quantitative trait loci identified in this study each act on a distinct PPG. Finally, we map three putative genomic inversions differentiating the two parental populations, including a previously characterized inversion that contributes to life-history differences between the annual/perennial ecotypes. Copyright © 2014 Holeski et al.
Fine-scale population structure and the era of next-generation sequencing.
Henn, Brenna M; Gravel, Simon; Moreno-Estrada, Andres; Acevedo-Acevedo, Suehelay; Bustamante, Carlos D
2010-10-15
Fine-scale population structure characterizes most continents and is especially pronounced in non-cosmopolitan populations. Roughly half of the world's population remains non-cosmopolitan and even populations within cities often assort along ethnic and linguistic categories. Barriers to random mating can be ecologically extreme, such as the Sahara Desert, or cultural, such as the Indian caste system. In either case, subpopulations accumulate genetic differences if the barrier is maintained over multiple generations. Genome-wide polymorphism data, initially with only a few hundred autosomal microsatellites, have clearly established differences in allele frequency not only among continental regions, but also within continents and within countries. We review recent evidence from the analysis of genome-wide polymorphism data for genetic boundaries delineating human population structure and the main demographic and genomic processes shaping variation, and discuss the implications of population structure for the distribution and discovery of disease-causing genetic variants, in the light of the imminent availability of sequencing data for a multitude of diverse human genomes.
Towers, Rebecca J.; Fagan, Peter K.; Talay, Susanne R.; Currie, Bart J.; Sriprakash, Kadaba S.; Walker, Mark J.; Chhatwal, Gursharan S.
2003-01-01
Streptococcal fibronectin-binding protein is an important virulence factor involved in colonization and invasion of epithelial cells and tissues by Streptococcus pyogenes. In order to investigate the mechanisms involved in the evolution of sfbI, the sfbI genes from 54 strains were sequenced. Thirty-four distinct alleles were identified. Three principal mechanisms appear to have been involved in the evolution of sfbI. The amino-terminal aromatic amino acid-rich domain is the most variable region and is apparently generated by intergenic recombination of horizontally acquired DNA cassettes, resulting in a genetic mosaic in this region. Two distinct and divergent sequence types that shared only 61 to 70% identity were identified in the central proline-rich region, while variation at the 3′ end of the gene is due to deletion or duplication of defined repeat units. Potential antigenic and functional variabilities in SfbI imply significant selective pressure in vivo with direct implications for the microbial pathogenesis of S. pyogenes. PMID:14662917
Bowhead whale (Balaena mysticetus) songs in the Chukchi Sea between October 2007 and May 2008.
Delarue, Julien; Laurinolli, Marjo; Martin, Bruce
2009-12-01
This paper reports on the acoustic detection of bowhead whale (Balaena mysticetus) songs from the Bering-Chukchi-Beaufort stock, including the first recordings of songs in the fall and early winter. Bowhead whale songs were detected almost continuously in the Chukchi Sea between October 30, 2007 and January 1, 2008 and twice from April 16 to May 5, 2008 during a long-term deployment of five acoustic recorders moored off Point Lay and Wainwright, AK, between October 21, 2007 and August 3, 2008. Two complex and four simple songs were detected. The complex songs consisted of highly stereotyped sequences of four units. The simple songs were primarily made of sequences of two to three moan types whose repetition patterns were constant over short periods but more variable over time. Multiple song types were recorded simultaneously and there is evidence of synchronized song variation over time. The implications of the spatiotemporal distribution of song detection with respect to the migratory and mating behavior of western Arctic bowheads are discussed.
Fine-scale patterns of population stratification confound rare variant association tests.
O'Connor, Timothy D; Kiezun, Adam; Bamshad, Michael; Rich, Stephen S; Smith, Joshua D; Turner, Emily; Leal, Suzanne M; Akey, Joshua M
2013-01-01
Advances in next-generation sequencing technology have enabled systematic exploration of the contribution of rare variation to Mendelian and complex diseases. Although it is well known that population stratification can generate spurious associations with common alleles, its impact on rare variant association methods remains poorly understood. Here, we performed exhaustive coalescent simulations with demographic parameters calibrated from exome sequence data to evaluate the performance of nine rare variant association methods in the presence of fine-scale population structure. We find that all methods have an inflated spurious association rate for parameter values that are consistent with levels of differentiation typical of European populations. For example, at a nominal significance level of 5%, some test statistics have a spurious association rate as high as 40%. Finally, we empirically assess the impact of population stratification in a large data set of 4,298 European American exomes. Our results have important implications for the design, analysis, and interpretation of rare variant genome-wide association studies.
USDA-ARS?s Scientific Manuscript database
Deep sequencing of viruses isolated from infected hosts is an efficient way to measure population-genetic variation and can reveal patterns of dispersal and natural selection. In this study, we mined existing Illumina sequence reads to investigate single-nucleotide polymorphisms (SNPs) within two RN...
Human Genome Sequencing in Health and Disease
Gonzaga-Jauregui, Claudia; Lupski, James R.; Gibbs, Richard A.
2013-01-01
Following the “finished,” euchromatic, haploid human reference genome sequence, the rapid development of novel, faster, and cheaper sequencing technologies is making possible the era of personalized human genomics. Personal diploid human genome sequences have been generated, and each has contributed to our better understanding of variation in the human genome. We have consequently begun to appreciate the vastness of individual genetic variation from single nucleotide to structural variants. Translation of genome-scale variation into medically useful information is, however, in its infancy. This review summarizes the initial steps undertaken in clinical implementation of personal genome information, and describes the application of whole-genome and exome sequencing to identify the cause of genetic diseases and to suggest adjuvant therapies. Better analysis tools and a deeper understanding of the biology of our genome are necessary in order to decipher, interpret, and optimize clinical utility of what the variation in the human genome can teach us. Personal genome sequencing may eventually become an instrument of common medical practice, providing information that assists in the formulation of a differential diagnosis. We outline herein some of the remaining challenges. PMID:22248320
Association of Amine-Receptor DNA Sequence Variants with Associative Learning in the Honeybee.
Lagisz, Malgorzata; Mercer, Alison R; de Mouzon, Charlotte; Santos, Luana L S; Nakagawa, Shinichi
2016-03-01
Octopamine- and dopamine-based neuromodulatory systems play a critical role in learning and learning-related behaviour in insects. To further our understanding of these systems and resulting phenotypes, we quantified DNA sequence variations at six loci coding octopamine-and dopamine-receptors and their association with aversive and appetitive learning traits in a population of honeybees. We identified 79 polymorphic sequence markers (mostly SNPs and a few insertions/deletions) located within or close to six candidate genes. Intriguingly, we found that levels of sequence variation in the protein-coding regions studied were low, indicating that sequence variation in the coding regions of receptor genes critical to learning and memory is strongly selected against. Non-coding and upstream regions of the same genes, however, were less conserved and sequence variations in these regions were weakly associated with between-individual differences in learning-related traits. While these associations do not directly imply a specific molecular mechanism, they suggest that the cross-talk between dopamine and octopamine signalling pathways may influence olfactory learning and memory in the honeybee.
NASA Astrophysics Data System (ADS)
Sheynkman, Gloria M.; Shortreed, Michael R.; Cesnik, Anthony J.; Smith, Lloyd M.
2016-06-01
Mass spectrometry-based proteomics has emerged as the leading method for detection, quantification, and characterization of proteins. Nearly all proteomic workflows rely on proteomic databases to identify peptides and proteins, but these databases typically contain a generic set of proteins that lack variations unique to a given sample, precluding their detection. Fortunately, proteogenomics enables the detection of such proteomic variations and can be defined, broadly, as the use of nucleotide sequences to generate candidate protein sequences for mass spectrometry database searching. Proteogenomics is experiencing heightened significance due to two developments: (a) advances in DNA sequencing technologies that have made complete sequencing of human genomes and transcriptomes routine, and (b) the unveiling of the tremendous complexity of the human proteome as expressed at the levels of genes, cells, tissues, individuals, and populations. We review here the field of human proteogenomics, with an emphasis on its history, current implementations, the types of proteomic variations it reveals, and several important applications.
Sheynkman, Gloria M.; Shortreed, Michael R.; Cesnik, Anthony J.; Smith, Lloyd M.
2016-01-01
Mass spectrometry–based proteomics has emerged as the leading method for detection, quantification, and characterization of proteins. Nearly all proteomic workflows rely on proteomic databases to identify peptides and proteins, but these databases typically contain a generic set of proteins that lack variations unique to a given sample, precluding their detection. Fortunately, proteogenomics enables the detection of such proteomic variations and can be defined, broadly, as the use of nucleotide sequences to generate candidate protein sequences for mass spectrometry database searching. Proteogenomics is experiencing heightened significance due to two developments: (a) advances in DNA sequencing technologies that have made complete sequencing of human genomes and transcriptomes routine, and (b) the unveiling of the tremendous complexity of the human proteome as expressed at the levels of genes, cells, tissues, individuals, and populations. We review here the field of human proteogenomics, with an emphasis on its history, current implementations, the types of proteomic variations it reveals, and several important applications. PMID:27049631
Temporal analysis of mtDNA variation reveals decreased genetic diversity in least terns
Draheim, Hope M.; Baird, Patricia; Haig, Susan M.
2012-01-01
The Least Tern (Sternula antillarum) has undergone large population declines over the last century as a result of direct and indirect anthropogenic factors. The genetic implications of these declines are unknown. We used historical museum specimens (pre-1960) and contemporary (2001–2005) samples to examine range-wide phylogeographic patterns and investigate potential loss in the species' genetic variation. We obtained sequences (522 bp) of the mitochondrial gene for NADH dehydrogenase subunit 6 (ND6) from 268 individuals from across the species' range. Phylogeographic analysis revealed no association with geography or traditional subspecies designations. However, we detected potential reductions in genetic diversity in contemporary samples from California and the Atlantic coast Least Tern from that in historical samples, suggesting that current genetic diversity in Least Tern populations is lower than in their pre-1960 counterparts. Our results offer unique insights into changes in the Least Tern's genetic diversity over the past century and highlight the importance and utility of museum specimens in studies of conservation genetics.
Global variation in CYP2C8–CYP2C9 functional haplotypes
Speed, William C; Kang, Soonmo Peter; Tuck, David P; Harris, Lyndsay N; Kidd, Kenneth K
2009-01-01
We have studied the global frequency distributions of 10 single nucleotide polymorphisms (SNPs) across 132 kb of CYP2C8 and CYP2C9 in ∼2500 individuals representing 45 populations. Five of the SNPs were in noncoding sequences; the other five involved the more common missense variants (four in CYP2C8, one in CYP2C9) that change amino acids in the gene products. One haplotype containing two CYP2C8 coding variants and one CYP2C9 coding variant reaches an average frequency of 10% in Europe; a set of haplotypes with a different CYP2C8 coding variant reaches 17% in Africa. In both cases these haplotypes are found in other regions of the world at <1%. This considerable geographic variation in haplotype frequencies impacts the interpretation of CYP2C8/CYP2C9 association studies, and has pharmacogenomic implications for drug interactions. PMID:19381162
Characterization of genetic sequence variation of 58 STR loci in four major population groups.
Novroski, Nicole M M; King, Jonathan L; Churchill, Jennifer D; Seah, Lay Hong; Budowle, Bruce
2016-11-01
Massively parallel sequencing (MPS) can identify sequence variation within short tandem repeat (STR) alleles as well as their nominal allele lengths that traditionally have been obtained by capillary electrophoresis. Using the MiSeq FGx Forensic Genomics System (Illumina), STRait Razor, and in-house excel workbooks, genetic variation was characterized within STR repeat and flanking regions of 27 autosomal, 7 X-chromosome and 24 Y-chromosome STR markers in 777 unrelated individuals from four population groups. Seven hundred and forty six autosomal, 227 X-chromosome, and 324 Y-chromosome STR alleles were identified by sequence compared with 357 autosomal, 107 X-chromosome, and 189 Y-chromosome STR alleles that were identified by length. Within the observed sequence variation, 227 autosomal, 156 X-chromosome, and 112 Y-chromosome novel alleles were identified and described. One hundred and seventy six autosomal, 123 X-chromosome, and 93 Y-chromosome sequence variants resided within STR repeat regions, and 86 autosomal, 39 X-chromosome, and 20 Y-chromosome variants were located in STR flanking regions. Three markers, D18S51, DXS10135, and DYS385a-b had 1, 4, and 1 alleles, respectively, which contained both a novel repeat region variant and a flanking sequence variant in the same nucleotide sequence. There were 50 markers that demonstrated a relative increase in diversity with the variant sequence alleles compared with those of traditional nominal length alleles. These population data illustrate the genetic variation that exists in the commonly used STR markers in the selected population samples and provide allele frequencies for statistical calculations related to STR profiling with MPS data. Copyright © 2016 Elsevier Ireland Ltd. All rights reserved.
Variation block-based genomics method for crop plants.
Kim, Yul Ho; Park, Hyang Mi; Hwang, Tae-Young; Lee, Seuk Ki; Choi, Man Soo; Jho, Sungwoong; Hwang, Seungwoo; Kim, Hak-Min; Lee, Dongwoo; Kim, Byoung-Chul; Hong, Chang Pyo; Cho, Yun Sung; Kim, Hyunmin; Jeong, Kwang Ho; Seo, Min Jung; Yun, Hong Tai; Kim, Sun Lim; Kwon, Young-Up; Kim, Wook Han; Chun, Hye Kyung; Lim, Sang Jong; Shin, Young-Ah; Choi, Ik-Young; Kim, Young Sun; Yoon, Ho-Sung; Lee, Suk-Ha; Lee, Sunghoon
2014-06-15
In contrast with wild species, cultivated crop genomes consist of reshuffled recombination blocks, which occurred by crossing and selection processes. Accordingly, recombination block-based genomics analysis can be an effective approach for the screening of target loci for agricultural traits. We propose the variation block method, which is a three-step process for recombination block detection and comparison. The first step is to detect variations by comparing the short-read DNA sequences of the cultivar to the reference genome of the target crop. Next, sequence blocks with variation patterns are examined and defined. The boundaries between the variation-containing sequence blocks are regarded as recombination sites. All the assumed recombination sites in the cultivar set are used to split the genomes, and the resulting sequence regions are termed variation blocks. Finally, the genomes are compared using the variation blocks. The variation block method identified recurring recombination blocks accurately and successfully represented block-level diversities in the publicly available genomes of 31 soybean and 23 rice accessions. The practicality of this approach was demonstrated by the identification of a putative locus determining soybean hilum color. We suggest that the variation block method is an efficient genomics method for the recombination block-level comparison of crop genomes. We expect that this method will facilitate the development of crop genomics by bringing genomics technologies to the field of crop breeding.
2012-01-01
The influence of resident gut microbes on xenobiotic metabolism has been investigated at different levels throughout the past five decades. However, with the advance in sequencing and pyrotagging technologies, addressing the influence of microbes on xenobiotics had to evolve from assessing direct metabolic effects on toxins and botanicals by conventional culture-based techniques to elucidating the role of community composition on drugs metabolic profiles through DNA sequence-based phylogeny and metagenomics. Following the completion of the Human Genome Project, the rapid, substantial growth of the Human Microbiome Project (HMP) opens new horizons for studying how microbiome compositional and functional variations affect drug action, fate, and toxicity (pharmacomicrobiomics), notably in the human gut. The HMP continues to characterize the microbial communities associated with the human gut, determine whether there is a common gut microbiome profile shared among healthy humans, and investigate the effect of its alterations on health. Here, we offer a glimpse into the known effects of the gut microbiota on xenobiotic metabolism, with emphasis on cases where microbiome variations lead to different therapeutic outcomes. We discuss a few examples representing how the microbiome interacts with human metabolic enzymes in the liver and intestine. In addition, we attempt to envisage a roadmap for the future implications of the HMP on therapeutics and personalized medicine. PMID:23194438
Cazaux, Benoîte; Catalan, Josette; Claude, Julien; Britton-Davidian, Janice
2014-01-01
The house mouse, Mus musculus domesticus, shows extraordinary chromosomal diversity driven by fixation of Robertsonian (Rb) translocations. The high frequency of this rearrangement, which involves the centromeric regions, has been ascribed to the architecture of the satellite sequence (high quantity and homogeneity). This promotes centromere-related translocations through unequal recombination and gene conversion. A characteristic feature of Rb variation in this subspecies is the non-random contribution of different chromosomes to the translocation frequency, which, in turn, depends on the chromosome size. Here, the association between satellite quantity and Rb frequency was tested by PRINS of the minor satellite which is the sequence involved in the translocation breakpoints. Five chromosomes with different translocation frequencies were selected and analyzed among wild house mice from 8 European localities. Using a relative quantitative measurement per chromosome, the analysis detected a large variability in signal size most of which was observed between individuals and/or localities. The chromosomes differed significantly in the quantity of the minor satellite, but these differences were not correlated with their translocation frequency. However, the data uncovered a marginally significant correlation between the quantity of the minor satellite and chromosome size. The implications of these results on the evolution of the chromosomal architecture in the house mouse are discussed. © 2014 S. Karger AG, Basel.
Polymorphism in the Eruption Sequence of Primary Dentition: A Cross-sectional Study
Bhojraj, Nandlal; Narayanappa
2017-01-01
Introduction Primary teeth have shown wide variations in their eruption time among different population. Population specific eruption ages are provided as mean with standard deviations or median ages with its percentile range. This alone will be insufficient for prediction of tooth eruption sequence because they provide no information on the frequency of sequence variation within the pairs of teeth. Norms of polymorphic variation in the eruption sequence can be more useful. Aim This study aims at providing norms for the sequence polymorphism in primary teeth among the children of Mysore population. Materials and Methods A cross-sectional study was designed with 1392 children, recruited from December 2015 to June 2016 by simple random sampling method. Tooth was recorded as present or absent. Across the entire possible intra quadrant tooth pair, cases of present-present, absent-absent, present-absent and absent-present and were counted and computed as percentages. Results Sequence polymorphisms were more common in 82-84 pairs of teeth. Significant polymorphic reverse sequence was observed in 52-54 (9%), 82-84 (35%) in males and 82-84 (18%) in females. There was no polymorphism in maxillary arch in females. Conclusion The present study provides the baseline data values for sequence variation in primary teeth eruption. To the best of investigators knowledge, there are no previous studies describing the sequence polymorphism in primary teeth in Indian population. The results of this study helps in assessment of eruption sequence problems in paediatric dentistry and in evaluation and prediction of tooth eruption sequence in individual child. PMID:28658912
Ashfaq, Muhammad; Hebert, Paul D N; Mirza, M Sajjad; Khan, Arif M; Mansoor, Shahid; Shah, Ghulam S; Zafar, Yusuf
2014-01-01
Although whiteflies (Bemisia tabaci complex) are an important pest of cotton in Pakistan, its taxonomic diversity is poorly understood. As DNA barcoding is an effective tool for resolving species complexes and analyzing species distributions, we used this approach to analyze genetic diversity in the B. tabaci complex and map the distribution of B. tabaci lineages in cotton growing areas of Pakistan. Sequence diversity in the DNA barcode region (mtCOI-5') was examined in 593 whiteflies from Pakistan to determine the number of whitefly species and their distributions in the cotton-growing areas of Punjab and Sindh provinces. These new records were integrated with another 173 barcode sequences for B. tabaci, most from India, to better understand regional whitefly diversity. The Barcode Index Number (BIN) System assigned the 766 sequences to 15 BINs, including nine from Pakistan. Representative specimens of each Pakistan BIN were analyzed for mtCOI-3' to allow their assignment to one of the putative species in the B. tabaci complex recognized on the basis of sequence variation in this gene region. This analysis revealed the presence of Asia II 1, Middle East-Asia Minor 1, Asia 1, Asia II 5, Asia II 7, and a new lineage "Pakistan". The first two taxa were found in both Punjab and Sindh, but Asia 1 was only detected in Sindh, while Asia II 5, Asia II 7 and "Pakistan" were only present in Punjab. The haplotype networks showed that most haplotypes of Asia II 1, a species implicated in transmission of the cotton leaf curl virus, occurred in both India and Pakistan. DNA barcodes successfully discriminated cryptic species in B. tabaci complex. The dominant haplotypes in the B. tabaci complex were shared by India and Pakistan. Asia II 1 was previously restricted to Punjab, but is now the dominant lineage in southern Sindh; its southward spread may have serious implications for cotton plantations in this region.
Mosaic PPM1D mutations are associated with predisposition to breast and ovarian cancer.
Ruark, Elise; Snape, Katie; Humburg, Peter; Loveday, Chey; Bajrami, Ilirjana; Brough, Rachel; Rodrigues, Daniel Nava; Renwick, Anthony; Seal, Sheila; Ramsay, Emma; Duarte, Silvana Del Vecchio; Rivas, Manuel A; Warren-Perry, Margaret; Zachariou, Anna; Campion-Flora, Adriana; Hanks, Sandra; Murray, Anne; Ansari Pour, Naser; Douglas, Jenny; Gregory, Lorna; Rimmer, Andrew; Walker, Neil M; Yang, Tsun-Po; Adlard, Julian W; Barwell, Julian; Berg, Jonathan; Brady, Angela F; Brewer, Carole; Brice, Glen; Chapman, Cyril; Cook, Jackie; Davidson, Rosemarie; Donaldson, Alan; Douglas, Fiona; Eccles, Diana; Evans, D Gareth; Greenhalgh, Lynn; Henderson, Alex; Izatt, Louise; Kumar, Ajith; Lalloo, Fiona; Miedzybrodzka, Zosia; Morrison, Patrick J; Paterson, Joan; Porteous, Mary; Rogers, Mark T; Shanley, Susan; Walker, Lisa; Gore, Martin; Houlston, Richard; Brown, Matthew A; Caufield, Mark J; Deloukas, Panagiotis; McCarthy, Mark I; Todd, John A; Turnbull, Clare; Reis-Filho, Jorge S; Ashworth, Alan; Antoniou, Antonis C; Lord, Christopher J; Donnelly, Peter; Rahman, Nazneen
2013-01-17
Improved sequencing technologies offer unprecedented opportunities for investigating the role of rare genetic variation in common disease. However, there are considerable challenges with respect to study design, data analysis and replication. Using pooled next-generation sequencing of 507 genes implicated in the repair of DNA in 1,150 samples, an analytical strategy focused on protein-truncating variants (PTVs) and a large-scale sequencing case-control replication experiment in 13,642 individuals, here we show that rare PTVs in the p53-inducible protein phosphatase PPM1D are associated with predisposition to breast cancer and ovarian cancer. PPM1D PTV mutations were present in 25 out of 7,781 cases versus 1 out of 5,861 controls (P = 1.12 × 10(-5)), including 18 mutations in 6,912 individuals with breast cancer (P = 2.42 × 10(-4)) and 12 mutations in 1,121 individuals with ovarian cancer (P = 3.10 × 10(-9)). Notably, all of the identified PPM1D PTVs were mosaic in lymphocyte DNA and clustered within a 370-base-pair region in the final exon of the gene, carboxy-terminal to the phosphatase catalytic domain. Functional studies demonstrate that the mutations result in enhanced suppression of p53 in response to ionizing radiation exposure, suggesting that the mutant alleles encode hyperactive PPM1D isoforms. Thus, although the mutations cause premature protein truncation, they do not result in the simple loss-of-function effect typically associated with this class of variant, but instead probably have a gain-of-function effect. Our results have implications for the detection and management of breast and ovarian cancer risk. More generally, these data provide new insights into the role of rare and of mosaic genetic variants in common conditions, and the use of sequencing in their identification.
Mosaic PPM1D mutations are associated with predisposition to breast and ovarian cancer
Ruark, Elise; Snape, Katie; Humburg, Peter; Loveday, Chey; Bajrami, Ilirjana; Brough, Rachel; Rodrigues, Daniel Nava; Renwick, Anthony; Seal, Sheila; Ramsay, Emma; Duarte, Silvana Del Vecchio; Rivas, Manuel A.; Warren-Perry, Margaret; Zachariou, Anna; Campion-Flora, Adriana; Hanks, Sandra; Murray, Anne; Pour, Naser Ansari; Douglas, Jenny; Gregory, Lorna; Rimmer, Andrew; Walker, Neil M.; Yang, Tsun-Po; Adlard, Julian W.; Barwell, Julian; Berg, Jonathan; Brady, Angela F.; Brewer, Carole; Brice, Glen; Chapman, Cyril; Cook, Jackie; Davidson, Rosemarie; Donaldson, Alan; Douglas, Fiona; Eccles, Diana; Evans, D. Gareth; Greenhalgh, Lynn; Henderson, Alex; Izatt, Louise; Kumar, Ajith; Lalloo, Fiona; Miedzybrodzka, Zosia; Morrison, Patrick J.; Paterson, Joan; Porteous, Mary; Rogers, Mark T.; Shanley, Susan; Walker, Lisa; Gore, Martin; Houlston, Richard; Brown, Matthew A.; Caufield, Mark J.; Deloukas, Panagiotis; McCarthy, Mark I.; Todd, John A.; Turnbull, Clare; Reis-Filho, Jorge S.; Ashworth, Alan; Antoniou, Antonis C.; Lord, Christopher J.; Donnelly, Peter; Rahman, Nazneen
2013-01-01
Improved sequencing technologies offer unprecedented opportunities for investigating the role of rare genetic variation in common disease. However, there are considerable challenges with respect to study design, data analysis and replication1. Here, using pooled next-generation sequencing of 507 genes implicated in the repair of DNA in 1,150 samples, an analytical strategy focussed on protein truncating variants (PTVs) and a large-scale sequencing case-control replication experiment in 13,642 individuals, we show that rare PTVs in the p53 inducible protein phosphatase PPM1D are associated with predisposition to breast cancer and to ovarian cancer. PPM1D PTV mutations were present in 25/7781 cases vs 1/5861 controls; P=1.12×10−5, which included 18 mutations in 6,912 individuals with breast cancer; P = 2.42×10−4 and 12 mutations in 1,121 individuals with ovarian cancer; P = 3.10×10−9. Notably, all the identified PPM1D PTVs were mosaic in lymphocyte DNA and clustered within a 370 bp region in the final exon of the gene, C-terminal to the phosphatase catalytic domain. Functional studies demonstrated that the mutations result in enhanced suppression of p53 in response to ionising radiation exposure, suggesting the mutant alleles encode hyperactive PPM1D isoforms. Thus, although the mutations cause premature protein truncation, they do not result in the simple loss-of-function typically associated with this class of variant, but instead likely have a gain-of-function effect. Our results have implications for the detection and management of breast and ovarian cancer risk. More generally, these data provide new insights into the role of rare and of mosaic genetic variants in common conditions, and the utility of sequencing in their identification. PMID:23242139
Spuesens, Emiel B M; Oduber, Minoushka; Hoogenboezem, Theo; Sluijter, Marcel; Hartwig, Nico G; van Rossum, Annemarie M C; Vink, Cornelis
2009-07-01
The gene encoding major adhesin protein P1 of Mycoplasma pneumoniae, MPN141, contains two DNA sequence stretches, designated RepMP2/3 and RepMP4, which display variation among strains. This variation allows strains to be differentiated into two major P1 genotypes (1 and 2) and several variants. Interestingly, multiple versions of the RepMP2/3 and RepMP4 elements exist at other sites within the bacterial genome. Because these versions are closely related in sequence, but not identical, it has been hypothesized that they have the capacity to recombine with their counterparts within MPN141, and thereby serve as a source of sequence variation of the P1 protein. In order to determine the variation within the RepMP2/3 and RepMP4 elements, both within the bacterial genome and among strains, we analysed the DNA sequences of all RepMP2/3 and RepMP4 elements within the genomes of 23 M. pneumoniae strains. Our data demonstrate that: (i) recombination is likely to have occurred between two RepMP2/3 elements in four of the strains, and (ii) all previously described P1 genotypes can be explained by inter-RepMP recombination events. Moreover, the difference between the two major P1 genotypes was reflected in all RepMP elements, such that subtype 1 and 2 strains can be differentiated on the basis of sequence variation in each RepMP element. This implies that subtype 1 and subtype 2 strains represent evolutionarily diverged strain lineages. Finally, a classification scheme is proposed in which the P1 genotype of M. pneumoniae isolates can be described in a sequence-based, universal fashion.
Cohen, Paul A; Flowers, Nicola; Tong, Stephen; Hannan, Natalie; Pertile, Mark D; Hui, Lisa
2016-08-24
Non-invasive prenatal testing (NIPT) identifies fetal aneuploidy by sequencing cell-free DNA in the maternal plasma. Pre-symptomatic maternal malignancies have been incidentally detected during NIPT based on abnormal genomic profiles. This low coverage sequencing approach could have potential for ovarian cancer screening in the non-pregnant population. Our objective was to investigate whether plasma DNA sequencing with a clinical whole genome NIPT platform can detect early- and late-stage high-grade serous ovarian carcinomas (HGSOC). This is a case control study of prospectively-collected biobank samples comprising preoperative plasma from 32 women with HGSOC (16 'early cancer' (FIGO I-II) and 16 'advanced cancer' (FIGO III-IV)) and 32 benign controls. Plasma DNA from cases and controls were sequenced using a commercial NIPT platform and chromosome dosage measured. Sequencing data were blindly analyzed with two methods: (1) Subchromosomal changes were called using an open source algorithm WISECONDOR (WIthin-SamplE COpy Number aberration DetectOR). Genomic gains or losses ≥ 15 Mb were prespecified as "screen positive" calls, and mapped to recurrent copy number variations reported in an ovarian cancer genome atlas. (2) Selected whole chromosome gains or losses were reported using the routine NIPT pipeline for fetal aneuploidy. We detected 13/32 cancer cases using the subchromosomal analysis (sensitivity 40.6 %, 95 % CI, 23.7-59.4 %), including 6/16 early and 7/16 advanced HGSOC cases. Two of 32 benign controls had subchromosomal gains ≥ 15 Mb (specificity 93.8 %, 95 % CI, 79.2-99.2 %). Twelve of the 13 true positive cancer cases exhibited specific recurrent changes reported in HGSOC tumors. The NIPT pipeline resulted in one "monosomy 18" call from the cancer group, and two "monosomy X" calls in the controls. Low coverage plasma DNA sequencing used for prenatal testing detected 40.6 % of all HGSOC, including 38 % of early stage cases. Our findings demonstrate the potential of a high throughput sequencing platform to screen for early HGSOC in plasma based on characteristic multiple segmental chromosome gains and losses. The performance of this approach may be further improved by refining bioinformatics algorithms and targeting selected cancer copy number variations.
Child Development and Structural Variation in the Human Genome
ERIC Educational Resources Information Center
Zhang, Ying; Haraksingh, Rajini; Grubert, Fabian; Abyzov, Alexej; Gerstein, Mark; Weissman, Sherman; Urban, Alexander E.
2013-01-01
Structural variation of the human genome sequence is the insertion, deletion, or rearrangement of stretches of DNA sequence sized from around 1,000 to millions of base pairs. Over the past few years, structural variation has been shown to be far more common in human genomes than previously thought. Very little is currently known about the effects…
Mtambo, Jupiter; Madder, Maxime; Van Bortel, Wim; Chaka, George; Berkvens, Dirk; Backeljau, Thierry
2007-01-01
Studies in the biology, ecology and behaviour of R. appendiculatus in Zambia have shown considerable variation within and between populations often associated with their geographical origin. We studied variation in the mitochondrial COI (mtCOI) gene of adult R. appendiculatus ticks originating from the Eastern and Southern provinces of Zambia. Rhipicephalus appendiculatus ticks from the two provinces were placed into two groups on the mtCOI sequence data tree. One group comprised all haplotypes of specimens from the Eastern province plateau districts of Chipata and Petauke. The second group consisted of a single haplotype of specimens from the Southern province districts and Nyimba, an Eastern province district on the fringes of the valley. This variation provides additional evidence to the earlier observations in the 12S rDNA and ITS2 data for the geographic subdivision of R. appendiculatus from Southern province and Eastern province plateau. The geographic subdivision further corresponds with differences in body size and diapause between R. appendiculatus from these geographic areas. The possible implications of these findings on the epidemiology of East Coast fever (ECF) the disease for which R. appendiculatus is one of the vectors are discussed.
Radiobiological Implications of Fukushima Nuclear Accident for Personalized Medical Approach.
Fukunaga, Hisanori; Yokoya, Akinari; Taki, Yasuyuki; Prise, Kevin M
2017-05-01
On March 11, 2011, a devastating earthquake and subsequent tsunami caused serious damage to areas of the Pacific coast in Fukushima prefecture and prompted fears among the residents about a possible meltdown of the Fukushima Daiichi Nuclear Power Plant reactors. As of 2017, over six years have passed since the Fukushima nuclear crisis and yet the full ramifications of the biological exposures to this accidental release of radioactive substances remain unclear. Furthermore, although several genetic studies have determined that the variation in radiation sensitivity among different individuals is wider than expected, personalized medical approaches for Fukushima victims have seemed to be insufficient. In this commentary, we discuss radiobiological issues arising from low-dose radiation exposure, from the cell-based to the population level. We also introduce the scientific utility of the Integrative Japanese Genome Variation Database (iJGVD), an online database released by the Tohoku Medical Megabank Organization, Tohoku University that covered the whole genome sequences of 2,049 healthy individuals in the northeastern part of Japan in 2016. Here we propose a personalized radiation risk assessment and medical approach, which considers the genetic variation of radiation sensitivity among individuals, for next-step developments in radiological protection.
Wang, Yiqin; Picard, Martin; Gu, Zhenglong
2016-10-01
Increasing clinical and biochemical evidence implicate mitochondrial dysfunction in the pathophysiology of Autism Spectrum Disorder (ASD), but little is known about the biological basis for this connection. A possible cause of ASD is the genetic variation in the mitochondrial DNA (mtDNA) sequence, which has yet to be thoroughly investigated in large genomic studies of ASD. Here we evaluated mtDNA variation, including the mixture of different mtDNA molecules in the same individual (i.e., heteroplasmy), using whole-exome sequencing data from mother-proband-sibling trios from simplex families (n = 903) where only one child is affected by ASD. We found that heteroplasmic mutations in autistic probands were enriched at non-polymorphic mtDNA sites (P = 0.0015), which were more likely to confer deleterious effects than heteroplasmies at polymorphic mtDNA sites. Accordingly, we observed a ~1.5-fold enrichment of nonsynonymous mutations (P = 0.0028) as well as a ~2.2-fold enrichment of predicted pathogenic mutations (P = 0.0016) in autistic probands compared to their non-autistic siblings. Both nonsynonymous and predicted pathogenic mutations private to probands conferred increased risk of ASD (Odds Ratio, OR[95% CI] = 1.87[1.14-3.11] and 2.55[1.26-5.51], respectively), and their influence on ASD was most pronounced in families with probands showing diminished IQ and/or impaired social behavior compared to their non-autistic siblings. We also showed that the genetic transmission pattern of mtDNA heteroplasmies with high pathogenic potential differed between mother-autistic proband pairs and mother-sibling pairs, implicating developmental and possibly in utero contributions. Taken together, our genetic findings substantiate pathogenic mtDNA mutations as a potential cause for ASD and synergize with recent work calling attention to their unique metabolic phenotypes for diagnosis and treatment of children with ASD.
Sampson, Juliana K.; Sheth, Nihar U.; Koparde, Vishal N.; Scalora, Allison F.; Serrano, Myrna G.; Lee, Vladimir; Roberts, Catherine H.; Jameson-Lee, Max; Ferreira-Gonzalez, Andrea; Manjili, Masoud H.; Buck, Gregory A.; Neale, Michael C.; Toor, Amir A.
2016-01-01
Summary Whole exome sequencing (WES) was performed on stem cell transplant donor-recipient (D-R) pairs to determine the extent of potential antigenic variation at a molecular level. In a small cohort of D-R pairs, a high frequency of sequence variation was observed between the donor and recipient exomes independent of human leucocyte antigen (HLA) matching. Nonsynonymous, nonconservative single nucleotide polymorphisms were approximately twice as frequent in HLA-matched unrelated, compared with related D-R pairs. When mapped to individual chromosomes, these polymorphic nucleotides were uniformly distributed across the entire exome. In conclusion, WES reveals extensive nucleotide sequence variation in the exomes of HLA-matched donors and recipients. PMID:24749631
An, Z; Tang, Z; Ma, B; Mason, A S; Guo, Y; Yin, J; Gao, C; Wei, L; Li, J; Fu, D
2014-07-01
Although many studies have shown that transposable element (TE) activation is induced by hybridisation and polyploidisation in plants, much less is known on how different types of TE respond to hybridisation, and the impact of TE-associated sequences on gene function. We investigated the frequency and regularity of putative transposon activation for different types of TE, and determined the impact of TE-associated sequence variation on the genome during allopolyploidisation. We designed different types of TE primers and adopted the Inter-Retrotransposon Amplified Polymorphism (IRAP) method to detect variation in TE-associated sequences during the process of allopolyploidisation between Brassica rapa (AA) and Brassica oleracea (CC), and in successive generations of self-pollinated progeny. In addition, fragments with TE insertions were used to perform Blast2GO analysis to characterise the putative functions of the fragments with TE insertions. Ninety-two primers amplifying 548 loci were used to detect variation in sequences associated with four different orders of TE sequences. TEs could be classed in ascending frequency into LTR-REs, TIRs, LINEs, SINEs and unknown TEs. The frequency of novel variation (putative activation) detected for the four orders of TEs was highest from the F1 to F2 generations, and lowest from the F2 to F3 generations. Functional annotation of sequences with TE insertions showed that genes with TE insertions were mainly involved in metabolic processes and binding, and preferentially functioned in organelles. TE variation in our study severely disturbed the genetic compositions of the different generations, resulting in inconsistencies in genetic clustering. Different types of TE showed different patterns of variation during the process of allopolyploidisation. © 2013 German Botanical Society and The Royal Botanical Society of the Netherlands.
Mining sequence variations in representative polyploid sugarcane germplasm accessions
DOE Office of Scientific and Technical Information (OSTI.GOV)
Yang, Xiping; Song, Jian; You, Qian
Sugarcane (Saccharum spp.) is one of the most important economic crops because of its high sugar production and biofuel potential. Due to the high polyploid level and complex genome of sugarcane, it has been a huge challenge to investigate genomic sequence variations, which are critical for identifying alleles contributing to important agronomic traits. In order to mine the genetic variations in sugarcane, genotyping by sequencing (GBS), was used to genotype 14 representative Saccharum complex accessions. GBS is a method to generate a large number of markers, enabled by next generation sequencing (NGS) and the genome complexity reduction using restriction enzymes.more » To use GBS for high throughput genotyping highly polyploid sugarcane, the GBS analysis pipelines in 14 Saccharum complex accessions were established by evaluating different alignment methods, sequence variants callers, and sequence depth for single nucleotide polymorphism (SNP) filtering. By using the established pipeline, a total of 76,251 non-redundant SNPs, 5642 InDels, 6380 presence/absence variants (PAVs), and 826 copy number variations (CNVs) were detected among the 14 accessions. In addition, non-reference based universal network enabled analysis kit and Stacks de novo called 34,353 and 109,043 SNPs, respectively. In the 14 accessions, the percentages of single dose SNPs ranged from 38.3% to 62.3% with an average of 49.6%, much more than the portions of multiple dosage SNPs. Concordantly called SNPs were used to evaluate the phylogenetic relationship among the 14 accessions. The results showed that the divergence time between the Erianthus genus and the Saccharum genus was more than 10 million years ago (MYA). The Saccharum species separated from their common ancestors ranging from 0.19 to 1.65 MYA. The GBS pipelines including the reference sequences, alignment methods, sequence variant callers, and sequence depth were recommended and discussed for the Saccharum complex and other related species. A large number of sequence variations were discovered in the Saccharum complex, including SNPs, InDels, PAVs, and CNVs. Genome-wide SNPs were further used to illustrate sequence features of polyploid species and demonstrated the divergence of different species in the Saccharum complex. The results of this study showed that GBS was an effective NGS-based method to discover genomic sequence variations in highly polyploid and heterozygous species.« less
Mining sequence variations in representative polyploid sugarcane germplasm accessions
Yang, Xiping; Song, Jian; You, Qian; ...
2017-08-09
Sugarcane (Saccharum spp.) is one of the most important economic crops because of its high sugar production and biofuel potential. Due to the high polyploid level and complex genome of sugarcane, it has been a huge challenge to investigate genomic sequence variations, which are critical for identifying alleles contributing to important agronomic traits. In order to mine the genetic variations in sugarcane, genotyping by sequencing (GBS), was used to genotype 14 representative Saccharum complex accessions. GBS is a method to generate a large number of markers, enabled by next generation sequencing (NGS) and the genome complexity reduction using restriction enzymes.more » To use GBS for high throughput genotyping highly polyploid sugarcane, the GBS analysis pipelines in 14 Saccharum complex accessions were established by evaluating different alignment methods, sequence variants callers, and sequence depth for single nucleotide polymorphism (SNP) filtering. By using the established pipeline, a total of 76,251 non-redundant SNPs, 5642 InDels, 6380 presence/absence variants (PAVs), and 826 copy number variations (CNVs) were detected among the 14 accessions. In addition, non-reference based universal network enabled analysis kit and Stacks de novo called 34,353 and 109,043 SNPs, respectively. In the 14 accessions, the percentages of single dose SNPs ranged from 38.3% to 62.3% with an average of 49.6%, much more than the portions of multiple dosage SNPs. Concordantly called SNPs were used to evaluate the phylogenetic relationship among the 14 accessions. The results showed that the divergence time between the Erianthus genus and the Saccharum genus was more than 10 million years ago (MYA). The Saccharum species separated from their common ancestors ranging from 0.19 to 1.65 MYA. The GBS pipelines including the reference sequences, alignment methods, sequence variant callers, and sequence depth were recommended and discussed for the Saccharum complex and other related species. A large number of sequence variations were discovered in the Saccharum complex, including SNPs, InDels, PAVs, and CNVs. Genome-wide SNPs were further used to illustrate sequence features of polyploid species and demonstrated the divergence of different species in the Saccharum complex. The results of this study showed that GBS was an effective NGS-based method to discover genomic sequence variations in highly polyploid and heterozygous species.« less
Queen, Rachel A.; Steyn, Jannetta S.; Lord, Phillip
2017-01-01
Mitochondrial DNA (mtDNA) mutations are well recognized as an important cause of inherited disease. Diseases caused by mtDNA mutations exhibit a high degree of clinical heterogeneity with a complex genotype-phenotype relationship, with many such mutations exhibiting incomplete penetrance. There is evidence that the spectrum of mutations causing mitochondrial disease might differ between different mitochondrial lineages (haplogroups) seen in different global populations. This would point to the importance of sequence context in the expression of mutations. To explore this possibility, we looked for mutations which are known to cause disease in humans, in animals of other species unaffected by mtDNA disease. The mt-tRNA genes are the location of many pathogenic mutations, with the m.3243A>G mutation on the mt-tRNA-Leu(UUR) being the most frequently seen mutation in humans. This study looked for the presence of m.3243A>G in 2784 sequences from 33 species, as well as any of the other mutations reported in association with disease located on mt-tRNA-Leu(UUR). We report a number of disease associated variations found on mt-tRNA-Leu(UUR) in other chordates, as the major population variant, with m.3243A>G being seen in 6 species. In these, we also found a number of mutations which appear compensatory and which could prevent the pathogenicity associated with this change in humans. This work has important implications for the discovery and diagnosis of mtDNA mutations in non-European populations. In addition, it might provide a partial explanation for the conflicting results in the literature that examines the role of mtDNA variants in complex traits. PMID:29161289
Deep sequencing reveals cell-type-specific patterns of single-cell transcriptome variation.
Dueck, Hannah; Khaladkar, Mugdha; Kim, Tae Kyung; Spaethling, Jennifer M; Francis, Chantal; Suresh, Sangita; Fisher, Stephen A; Seale, Patrick; Beck, Sheryl G; Bartfai, Tamas; Kuhn, Bernhard; Eberwine, James; Kim, Junhyong
2015-06-09
Differentiation of metazoan cells requires execution of different gene expression programs but recent single-cell transcriptome profiling has revealed considerable variation within cells of seeming identical phenotype. This brings into question the relationship between transcriptome states and cell phenotypes. Additionally, single-cell transcriptomics presents unique analysis challenges that need to be addressed to answer this question. We present high quality deep read-depth single-cell RNA sequencing for 91 cells from five mouse tissues and 18 cells from two rat tissues, along with 30 control samples of bulk RNA diluted to single-cell levels. We find that transcriptomes differ globally across tissues with regard to the number of genes expressed, the average expression patterns, and within-cell-type variation patterns. We develop methods to filter genes for reliable quantification and to calibrate biological variation. All cell types include genes with high variability in expression, in a tissue-specific manner. We also find evidence that single-cell variability of neuronal genes in mice is correlated with that in rats consistent with the hypothesis that levels of variation may be conserved. Single-cell RNA-sequencing data provide a unique view of transcriptome function; however, careful analysis is required in order to use single-cell RNA-sequencing measurements for this purpose. Technical variation must be considered in single-cell RNA-sequencing studies of expression variation. For a subset of genes, biological variability within each cell type appears to be regulated in order to perform dynamic functions, rather than solely molecular noise.
Population genetic implications from sequence variation in four Y chromosome genes.
Shen, P; Wang, F; Underhill, P A; Franco, C; Yang, W H; Roxas, A; Sung, R; Lin, A A; Hyman, R W; Vollrath, D; Davis, R W; Cavalli-Sforza, L L; Oefner, P J
2000-06-20
Some insight into human evolution has been gained from the sequencing of four Y chromosome genes. Primary genomic sequencing determined gene SMCY to be composed of 27 exons that comprise 4,620 bp of coding sequence. The unfinished sequencing of the 5' portion of gene UTY1 was completed by primer walking, and a total of 20 exons were found. By using denaturing HPLC, these two genes, as well as DBY and DFFRY, were screened for polymorphic sites in 53-72 representatives of the five continents. A total of 98 variants were found, yielding nucleotide diversity estimates of 2.45 x 10(-5), 5. 07 x 10(-5), and 8.54 x 10(-5) for the coding regions of SMCY, DFFRY, and UTY1, respectively, with no variant having been observed in DBY. In agreement with most autosomal genes, diversity estimates for the noncoding regions were about 2- to 3-fold higher and ranged from 9. 16 x 10(-5) to 14.2 x 10(-5) for the four genes. Analysis of the frequencies of derived alleles for all four genes showed that they more closely fit the expectation of a Luria-Delbrück distribution than a distribution expected under a constant population size model, providing evidence for exponential population growth. Pairwise nucleotide mismatch distributions date the occurrence of population expansion to approximately 28,000 years ago. This estimate is in accord with the spread of Aurignacian technology and the disappearance of the Neanderthals.
Lacerra, Giuseppina; Fiorito, Mirella; Musollino, Gennaro; Di Noce, Francesca; Esposito, Maria; Nigro, Vincenzo; Gaudiano, Carlo; Carestia, Clementina
2004-10-01
The alpha-globin chains are encoded by two duplicated genes (HBA2 and HBA1, 5'-3') showing overall sequence homology >96% and average CG content >60%. alpha-Thalassemia, the most prevalent worldwide autosomal recessive disorder, is a hereditary anemia caused by sequence variations of these genes in about 25% of carriers. We evaluated the overall sensitivity and suitability of DHPLC and DG-DGGE in scanning both the alpha-globin genes by carrying out a retrospective analysis of 19 variant alleles in 29 genotypes. The HBA2 alleles c.1A>G, c.79G>A, and c.281T>G, and the HBA1 allele c.475C>A were new. Three pathogenic sequence variations were associated in cis with nonpathogenic variations in all families studied; they were the HBA2 variation c.2T>C associated with c.-24C>G, and the HBA2 variations c.391G>C and c.427T>C, both associated with c.565G>A. We set up original experimental conditions for DHPLC and DG-DGGE and analyzed 10 normal subjects, 46 heterozygotes, seven homozygotes, seven compound heterozygotes, and six compound heterozygotes for a hybrid gene. Both the methodologies gave reproducible results and no false-positive was detected. DHPLC showed 100% sensitivity and DG-DGGE nearly 90%. About 100% of the sequence from the cap site to the polyA addition site could be scanned by DHPLC, about 87% by DG-DGGE. It is noteworthy that the three most common pathogenic sequence variations (HBA2 alleles c.2T>C, c.95+2_95+6del, and c.523A>G) were unambiguously detected by both the methodologies. Genotype diagnosis must be confirmed with PCR sequencing of single amplicons or with an allele-specific method. This study can be helpful for scanning genes with high CG content and offers a model suitable for duplicated genes with high homology. Copyright 2004 Wiley-Liss, Inc.
Read clouds uncover variation in complex regions of the human genome
Bishara, Alex; Liu, Yuling; Weng, Ziming; Kashef-Haghighi, Dorna; Newburger, Daniel E.; West, Robert; Sidow, Arend; Batzoglou, Serafim
2015-01-01
Although an increasing amount of human genetic variation is being identified and recorded, determining variants within repeated sequences of the human genome remains a challenge. Most population and genome-wide association studies have therefore been unable to consider variation in these regions. Core to the problem is the lack of a sequencing technology that produces reads with sufficient length and accuracy to enable unique mapping. Here, we present a novel methodology of using read clouds, obtained by accurate short-read sequencing of DNA derived from long fragment libraries, to confidently align short reads within repeat regions and enable accurate variant discovery. Our novel algorithm, Random Field Aligner (RFA), captures the relationships among the short reads governed by the long read process via a Markov Random Field. We utilized a modified version of the Illumina TruSeq synthetic long-read protocol, which yielded shallow-sequenced read clouds. We test RFA through extensive simulations and apply it to discover variants on the NA12878 human sample, for which shallow TruSeq read cloud sequencing data are available, and on an invasive breast carcinoma genome that we sequenced using the same method. We demonstrate that RFA facilitates accurate recovery of variation in 155 Mb of the human genome, including 94% of 67 Mb of segmental duplication sequence and 96% of 11 Mb of transcribed sequence, that are currently hidden from short-read technologies. PMID:26286554
ACTG: novel peptide mapping onto gene models.
Choi, Seunghyuk; Kim, Hyunwoo; Paek, Eunok
2017-04-15
In many proteogenomic applications, mapping peptide sequences onto genome sequences can be very useful, because it allows us to understand origins of the gene products. Existing software tools either take the genomic position of a peptide start site as an input or assume that the peptide sequence exactly matches the coding sequence of a given gene model. In case of novel peptides resulting from genomic variations, especially structural variations such as alternative splicing, these existing tools cannot be directly applied unless users supply information about the variant, either its genomic position or its transcription model. Mapping potentially novel peptides to genome sequences, while allowing certain genomic variations, requires introducing novel gene models when aligning peptide sequences to gene structures. We have developed a new tool called ACTG (Amino aCids To Genome), which maps peptides to genome, assuming all possible single exon skipping, junction variation allowing three edit distances from the original splice sites, exon extension and frame shift. In addition, it can also consider SNVs (single nucleotide variations) during mapping phase if a user provides the VCF (variant call format) file as an input. Available at http://prix.hanyang.ac.kr/ACTG/search.jsp . eunokpaek@hanyang.ac.kr. Supplementary data are available at Bioinformatics online. © The Author 2016. Published by Oxford University Press. All rights reserved. For Permissions, please e-mail: journals.permissions@oup.com
Schuenzel, Erin L.; Scally, Mark; Bromley, Robin E.; Stouthamer, Richard
2014-01-01
Homologous recombination plays an important role in the structuring of genetic variation of many bacteria; however, its importance in adaptive evolution is not well established. We investigated the association of intersubspecific homologous recombination (IHR) with the shift to a novel host (mulberry) by the plant-pathogenic bacterium Xylella fastidiosa. Mulberry leaf scorch was identified about 25 years ago in native red mulberry in the eastern United States and has spread to introduced white mulberry in California. Comparing a sequence of 8 genes (4,706 bp) from 21 mulberry-type isolates to published data (352 isolates representing all subspecies), we confirmed previous indications that the mulberry isolates define a group distinct from the 4 subspecies, and we propose naming the taxon X. fastidiosa subsp. morus. The ancestry of its gene sequences was mixed, with 4 derived from X. fastidiosa subsp. fastidiosa (introduced from Central America), 3 from X. fastidiosa subsp. multiplex (considered native to the United States), and 1 chimeric, demonstrating that this group originated by large-scale IHR. The very low within-type genetic variation (0.08% site polymorphism), plus the apparent inability of native X. fastidiosa subsp. multiplex to infect mulberry, suggests that this host shift was achieved after strong selection acted on genetic variants created by IHR. Sequence data indicate that a single ancestral IHR event gave rise not only to X. fastidiosa subsp. morus but also to the X. fastidiosa subsp. multiplex recombinant group which infects several hosts but is the only type naturally infecting blueberry, thus implicating this IHR in the invasion of at least two novel native hosts, mulberry and blueberry. PMID:24610840
Nunney, Leonard; Schuenzel, Erin L; Scally, Mark; Bromley, Robin E; Stouthamer, Richard
2014-05-01
Homologous recombination plays an important role in the structuring of genetic variation of many bacteria; however, its importance in adaptive evolution is not well established. We investigated the association of intersubspecific homologous recombination (IHR) with the shift to a novel host (mulberry) by the plant-pathogenic bacterium Xylella fastidiosa. Mulberry leaf scorch was identified about 25 years ago in native red mulberry in the eastern United States and has spread to introduced white mulberry in California. Comparing a sequence of 8 genes (4,706 bp) from 21 mulberry-type isolates to published data (352 isolates representing all subspecies), we confirmed previous indications that the mulberry isolates define a group distinct from the 4 subspecies, and we propose naming the taxon X. fastidiosa subsp. morus. The ancestry of its gene sequences was mixed, with 4 derived from X. fastidiosa subsp. fastidiosa (introduced from Central America), 3 from X. fastidiosa subsp. multiplex (considered native to the United States), and 1 chimeric, demonstrating that this group originated by large-scale IHR. The very low within-type genetic variation (0.08% site polymorphism), plus the apparent inability of native X. fastidiosa subsp. multiplex to infect mulberry, suggests that this host shift was achieved after strong selection acted on genetic variants created by IHR. Sequence data indicate that a single ancestral IHR event gave rise not only to X. fastidiosa subsp. morus but also to the X. fastidiosa subsp. multiplex recombinant group which infects several hosts but is the only type naturally infecting blueberry, thus implicating this IHR in the invasion of at least two novel native hosts, mulberry and blueberry.
Fishman, G A; Stone, E M; Grover, S; Derlacki, D J; Haines, H L; Hockey, R R
1999-04-01
To report the spectrum of ophthalmic findings in patients with Stargardt dystrophy or fundus flavimaculatus who have a specific sequence variation in the ABCR gene. Twenty-nine patients with Stargardt dystrophy or fundus flavimaculatus from different pedigrees were identified with possible disease-causing sequence variations in the ABCR gene from a group of 66 patients who were screened for sequence variations in this gene. Patients underwent a routine ocular examination, including slitlamp biomicroscopy and a dilated fundus examination. Fluorescein angiography was performed on 22 patients, and electroretinographic measurements were obtained on 24 of 29 patients. Kinetic visual fields were measured with a Goldmann perimeter in 26 patients. Single-strand conformation polymorphism analysis and DNA sequencing were used to identify variations in coding sequences of the ABCR gene. Three clinical phenotypes were observed among these 29 patients. In phenotype I, 9 of 12 patients had a sequence change in exon 42 of the ABCR gene in which the amino acid glutamic acid was substituted for glycine (Gly1961Glu). In only 4 of these 9 patients was a second possible disease-causing mutation found on the other ABCR allele. In addition to an atrophic-appearing macular lesion, phenotype I was characterized by localized perifoveal yellowish white flecks, the absence of a dark choroid, and normal electroretinographic amplitudes. Phenotype II consisted of 10 patients who showed a dark choroid and more diffuse yellowish white flecks in the fundus. None exhibited the Gly1961Glu change. Phenotype III consisted of 7 patients who showed extensive atrophic-appearing changes of the retinal pigment epithelium. Electroretinographic cone and rod amplitudes were reduced. One patient showed the Gly1961Glu change. A wide variation in clinical phenotype can occur in patients with sequence changes in the ABCR gene. In individual patients, a certain phenotype seems to be associated with the presence of a Gly1961Glu change in exon 42 of the ABCR gene. The identification of correlations between specific mutations in the ABCR gene and clinical phenotypes will better facilitate the counseling of patients on their visual prognosis. This information will also likely be important for future therapeutic trials in patients with Stargardt dystrophy.
Ben Lazhar-Ajroud, Wafa; Caruso, Aurore; Mezghani, Maha; Bouallegue, Maryem; Tastard, Emmanuelle; Denis, Françoise; Rouault, Jacques-Deric; Makni, Hanem; Capy, Pierre; Chénais, Benoît; Makni, Mohamed; Casse, Nathalie
2016-08-01
Genomic variation among species is commonly driven by transposable element (TE) invasion; thus, the pattern of TEs in a genome allows drawing an evolutionary history of the studied species. This paper reports in vitro and in silico detection and characterization of irritans mariner-like elements (MLEs) in the genome and transcriptome of Bactrocera oleae (Rossi) (Diptera: Tephritidae). Eleven irritans MLE sequences have been isolated in vitro using terminal inverted repeats (TIRs) as primers, and 215 have been extracted in silico from the sequenced genome of B. oleae. Additionally, the sequenced genomes of Bactrocera tryoni (Froggatt) and Bactrocera cucurbitae (Diptera: Tephritidae) have been explored to identify irritans MLEs. A total of 129 sequences from B. tryoni have been extracted, while the genome of B. cucurbitae appears probably devoid of irritans MLEs. All detected irritans MLEs are defective due to several mutations and are clustered together in a monophyletic group suggesting a common ancestor. The evolutionary history and dynamics of these TEs are discussed in relation with the phylogenetic distribution of their hosts. The knowledge on the structure, distribution, dynamic, and evolution of irritans MLEs in Bactrocera species contributes to the understanding of both their evolutionary history and the invasion history of their hosts. This could also be the basis for genetic control strategies using transposable elements.
DNA Metabarcoding of Amazonian Ichthyoplankton Swarms.
Maggia, M E; Vigouroux, Y; Renno, J F; Duponchelle, F; Desmarais, E; Nunez, J; García-Dávila, C; Carvajal-Vallejos, F M; Paradis, E; Martin, J F; Mariac, C
2017-01-01
Tropical rainforests harbor extraordinary biodiversity. The Amazon basin is thought to hold 30% of all river fish species in the world. Information about the ecology, reproduction, and recruitment of most species is still lacking, thus hampering fisheries management and successful conservation strategies. One of the key understudied issues in the study of population dynamics is recruitment. Fish larval ecology in tropical biomes is still in its infancy owing to identification difficulties. Molecular techniques are very promising tools for the identification of larvae at the species level. However, one of their limits is obtaining individual sequences with large samples of larvae. To facilitate this task, we developed a new method based on the massive parallel sequencing capability of next generation sequencing (NGS) coupled with hybridization capture. We focused on the mitochondrial marker cytochrome oxidase I (COI). The results obtained using the new method were compared with individual larval sequencing. We validated the ability of the method to identify Amazonian catfish larvae at the species level and to estimate the relative abundance of species in batches of larvae. Finally, we applied the method and provided evidence for strong temporal variation in reproductive activity of catfish species in the Ucayalí River in the Peruvian Amazon. This new time and cost effective method enables the acquisition of large datasets, paving the way for a finer understanding of reproductive dynamics and recruitment patterns of tropical fish species, with major implications for fisheries management and conservation.
Chen, Xin; Wu, Qiong; Sun, Ruimin; Zhang, Louxin
2012-01-01
The discovery of single-nucleotide polymorphisms (SNPs) has important implications in a variety of genetic studies on human diseases and biological functions. One valuable approach proposed for SNP discovery is based on base-specific cleavage and mass spectrometry. However, it is still very challenging to achieve the full potential of this SNP discovery approach. In this study, we formulate two new combinatorial optimization problems. While both problems are aimed at reconstructing the sample sequence that would attain the minimum number of SNPs, they search over different candidate sequence spaces. The first problem, denoted as SNP - MSP, limits its search to sequences whose in silico predicted mass spectra have all their signals contained in the measured mass spectra. In contrast, the second problem, denoted as SNP - MSQ, limits its search to sequences whose in silico predicted mass spectra instead contain all the signals of the measured mass spectra. We present an exact dynamic programming algorithm for solving the SNP - MSP problem and also show that the SNP - MSQ problem is NP-hard by a reduction from a restricted variation of the 3-partition problem. We believe that an efficient solution to either problem above could offer a seamless integration of information in four complementary base-specific cleavage reactions, thereby improving the capability of the underlying biotechnology for sensitive and accurate SNP discovery.
VaDiR: an integrated approach to Variant Detection in RNA.
Neums, Lisa; Suenaga, Seiji; Beyerlein, Peter; Anders, Sara; Koestler, Devin; Mariani, Andrea; Chien, Jeremy
2018-02-01
Advances in next-generation DNA sequencing technologies are now enabling detailed characterization of sequence variations in cancer genomes. With whole-genome sequencing, variations in coding and non-coding sequences can be discovered. But the cost associated with it is currently limiting its general use in research. Whole-exome sequencing is used to characterize sequence variations in coding regions, but the cost associated with capture reagents and biases in capture rate limit its full use in research. Additional limitations include uncertainty in assigning the functional significance of the mutations when these mutations are observed in the non-coding region or in genes that are not expressed in cancer tissue. We investigated the feasibility of uncovering mutations from expressed genes using RNA sequencing datasets with a method called Variant Detection in RNA(VaDiR) that integrates 3 variant callers, namely: SNPiR, RVBoost, and MuTect2. The combination of all 3 methods, which we called Tier 1 variants, produced the highest precision with true positive mutations from RNA-seq that could be validated at the DNA level. We also found that the integration of Tier 1 variants with those called by MuTect2 and SNPiR produced the highest recall with acceptable precision. Finally, we observed a higher rate of mutation discovery in genes that are expressed at higher levels. Our method, VaDiR, provides a possibility of uncovering mutations from RNA sequencing datasets that could be useful in further functional analysis. In addition, our approach allows orthogonal validation of DNA-based mutation discovery by providing complementary sequence variation analysis from paired RNA/DNA sequencing datasets.
Intra-isolate genome variation in arbuscular mycorrhizal fungi persists in the transcriptome.
Boon, E; Zimmerman, E; Lang, B F; Hijri, M
2010-07-01
Arbuscular mycorrhizal fungi (AMF) are heterokaryotes with an unusual genetic makeup. Substantial genetic variation occurs among nuclei within a single mycelium or isolate. AMF reproduce through spores that contain varying fractions of this heterogeneous population of nuclei. It is not clear whether this genetic variation on the genome level actually contributes to the AMF phenotype. To investigate the extent to which polymorphisms in nuclear genes are transcribed, we analysed the intra-isolate genomic and cDNA sequence variation of two genes, the large subunit ribosomal RNA (LSU rDNA) of Glomus sp. DAOM-197198 (previously known as G. intraradices) and the POL1-like sequence (PLS) of Glomus etunicatum. For both genes, we find high sequence variation at the genome and transcriptome level. Reconstruction of LSU rDNA secondary structure shows that all variants are functional. Patterns of PLS sequence polymorphism indicate that there is one functional gene copy, PLS2, which is preferentially transcribed, and one gene copy, PLS1, which is a pseudogene. This is the first study that investigates AMF intra-isolate variation at the transcriptome level. In conclusion, it is possible that, in AMF, multiple nuclear genomes contribute to a single phenotype.
Whole-Genome Sequence Variation among Multiple Isolates of Pseudomonas aeruginosa
Spencer, David H.; Kas, Arnold; Smith, Eric E.; Raymond, Christopher K.; Sims, Elizabeth H.; Hastings, Michele; Burns, Jane L.; Kaul, Rajinder; Olson, Maynard V.
2003-01-01
Whole-genome shotgun sequencing was used to study the sequence variation of three Pseudomonas aeruginosa isolates, two from clonal infections of cystic fibrosis patients and one from an aquatic environment, relative to the genomic sequence of reference strain PAO1. The majority of the PAO1 genome is represented in these strains; however, at least three prominent islands of PAO1-specific sequence are apparent. Conversely, ∼10% of the sequencing reads derived from each isolate fail to align with the PAO1 backbone. While average sequence variation among all strains is roughly 0.5%, regions of pronounced differences were evident in whole-genome scans of nucleotide diversity. We analyzed two such divergent loci, the pyoverdine and O-antigen biosynthesis regions, by complete resequencing. A thorough analysis of isolates collected over time from one of the cystic fibrosis patients revealed independent mutations resulting in the loss of O-antigen synthesis alternating with a mucoid phenotype. Overall, we conclude that most of the PAO1 genome represents a core P. aeruginosa backbone sequence while the strains addressed in this study possess additional genetic material that accounts for at least 10% of their genomes. Approximately half of these additional sequences are novel. PMID:12562802
Chen, Fen; Li, Juan; Sugiyama, Hiromu; Zhou, Dong-Hui; Song, Hui-Qun; Zhao, Guang-Hui; Zhu, Xing-Quan
2015-02-01
The present study examined sequence variability in the mitochondrial (mt) protein-coding genes cytochrome b (cytb), NADH dehydrogenase subunits 2 and 6 (nad2 and nad6) among 24 isolates of Schistosoma japonicum from different endemic regions in the Philippines, Japan and China. The complete cytb, nad2 and nad6 genes were amplified and sequenced separately from individual schistosome. Sequence variations for isolates from the Philippines were 0-0.5% for cytb, 0-0.6% for nad2, and 0-0.9% for nad6. Variation was 0-0.5%, 0.1-0.8%, 0-0.7% for corresponding genes for schistosome samples from mainland China. For worms in Japan, genetic variations were 0-0.2%, 0.1-0.2% and 0 for the three genes, respectively. Sequence variations were 0-1.0%, 0-1.8% and 0-1.1% for cytb, nad2 and nad6, respectively, among schistosome isolates from different geographical strains in the Philippines, Japan and China. Of the three countries, lowest sequence variations were found between isolates from mainland China and the Philippines and highest were detected between Japan and the Philippines in three mtDNA genes. Phylogenetic analyses based on the combined sequences of cytb, nad2 and nad6 revealed that all isolates in the Philippines clustered together sistered to samples from Yunnan and Zhejiang provinces in China, while isolates from Yamanashi in Japan were in a solitary clade. These results demonstrated the usefulness of the combined three mtDNA sequences for studying genetic diversity and population structure among S. japonicum isolates from the Philippines, China and Japan.
NASA Astrophysics Data System (ADS)
Abdullatif, Osman; Abdlmutalib, Ammar; Ahmed, Jarrah; Abdelgadir, Mohamed; Adam, Ammar
2017-04-01
The Permian-Triassic Khuff Formation carbonate reservoirs (and equivalents) in the Middle East are estimated to contain about 15-20 % of the world's gas reserves. Excellently exposed outcropping Khuff strata in central Saudi Arabia provide good outcrop equivalents to the Khuff Formation in the subsurface. The Khuff Formation is composed of five members and from bottom to top are Ash Shiqqah, Huqayl, Duhaysan, Midnab and Khartam members. The Carbonates lithofacies dominate with minor terrestrial clastics, and the paleoenvironments vary from terrestrial, sabkha, tidal-intertidal and open marine environments. This study investigates the relationship between lithostratigraphy, sequence stratigraphy and chemostratigraphy by integration of both field and laboratory sedimentological and chemical elements data. The vertical chemical elements profiles along the Khuff members show variations in their chemical elements content with the variation in lithofacies types, staking pattern, depositional and stratigraphic pattern. The chemostratigraphic distribution of the chemical elements also showed variation within and between the Khuff members. There is a general agreement between chemostratigraphic analyses based on vertical profiles and binary cross plots. The Khuff members and their stratigraphic boundaries can be differentiated based on their chemostratigraphic signatures. Moreover, the lithofacies and depositional paleoenvironmental of different Khuff members can be identified based on their chemical element contents. Chemostratigraphic zones or clusters are markedly established indicating different lithofacies and depositional paleoenvironments. Terrestrial, channel, lacustrine, shoreline to open marine carbonate lithofacies, as building blocks of sequence stratigraphy, all may be distinguished based on their chemical signatures. These outcrop analog results might be of significance to lithofacies, paleoenvironmental, stratigraphic identification, classification and correlation of Khuff Formation in the subsurface. The results might also provide guides and application to reservoir Khuff Formation identification, layering and zonation in the subsurface.
Ullah, Muhammad Ikram; Ahmad, Arsalan; Raza, Syed Irfan; Amar, Ali; Ali, Amjad; Bhatti, Attya; John, Peter; Mohyuddin, Aisha; Ahmad, Wasim; Hassan, Muhammad Jawad
2015-10-01
Amyotrophic lateral sclerosis (ALS) is a neurodegenerative disorder affecting upper motor neurons in the brain and lower motor neurons in the brain stem and spinal cord, resulting in fatal paralysis. It has been found to be associated with frontotemporal lobar degeneration (FTLD). In the present study, we have described homozygosity mapping and gene sequencing in a consanguineous autosomal recessive Pakistani family showing non-juvenile ALS without signs of FTLD. Gene mapping was carried out in all recruited family members using microsatellite markers, and linkage was established with sigma non-opioid intracellular receptor 1 (SIGMAR1) gene at chromosome 9p13.2. Gene sequencing of SIGMAR1 revealed a novel 3'-UTR nucleotide variation c.672*31A>G (rs4879809) segregating with disease in this family. The C9ORF72 repeat region in intron 1, previously implicated in a related phenotype, was excluded through linkage, and further confirmation of exclusion was obtained by amplifying intron 1 of C9ORF72 with multiple primers in affected individuals and controls. In silico analysis was carried out to explore the possible role of 3'-UTR variant of SIGMAR1 in ALS. The Regulatory RNA motif and Element Finder program revealed disturbance in miRNA (hsa-miR-1205) binding site due to this variation. ESEFinder analysis showed new SRSF1 and SRSF1-IgM-BRCA1 binding sites with significant scores due to this variation. Our results indicate that the 3'-UTR SIGMAR1 variant c.672*31A>G may have a role in the pathogenesis of ALS in this family.
Canter, Jeffrey A.; Norris, Patrick R.; Moore, Jason H.; Jenkins, Judith M.; Morris, John A.
2007-01-01
Objective: To determine whether specific genetic variations in the mtDNA that impact energy production and free-radical generation are potential new risk factors for in-hospital mortality after severe trauma. Summary Background Data: Each of the 3 mitochondrial DNA polymorphisms selected for this study (at positions 4216, 10398, 4917) alter the amino acid sequence of different key subunits of Complex I in the electron transport chain. They have been previously implicated in phenotypes involving tissues with high-energy demand, such as the brain and retina. Methods: Seven hundred forty-five consecutive patients admitted to the trauma intensive care unit at Vanderbilt University Medical Center between April 11, 2005, and February 27, 2006, were potentially eligible for this study. Under an Institutional Review Board-approved protocol (which excluded patients <18 years of age and prisoners), 666 patients had DNA extracted from a blood sample. Detailed demographic and clinical covariates were also obtained (including age, gender, ethnicity, lactate measurements, and injury severity score). A flurogenic 5′ nuclease allelic discrimination Taqman assay and the ABI 7900HT Sequence Detection System (v2.1) was used to genotype the T4216C, A10398G, and A4917G polymorphisms. The primary outcome was in-hospital mortality. Results: Multivariate logistic regression analysis revealed that the 4216T allele was a significant independent predictor of in-hospital mortality (OR = 2.63, 95% CI 1.14–6.07, P = 0.02) after adjustment for age, gender, injury severity score, highest lactate level, mechanism of injury, and the 10398 polymorphism. Conclusions: Variation in the mtDNA, specifically the 4216T allele, appears to increase the risk of in-hospital mortality after severe injury. PMID:17717444
ErbB polymorphisms: insights and implications for response to targeted cancer therapeutics.
Alaoui-Jamali, Moulay A; Morand, Grégoire B; da Silva, Sabrina Daniela
2015-01-01
Advances in high-throughput genomic-scanning have expanded the repertory of genetic variations in DNA sequences encoding ErbB tyrosine kinase receptors in humans, including single nucleotide polymorphisms (SNPs), polymorphic repetitive elements, microsatellite variations, small-scale insertions and deletions. The ErbB family members: EGFR, ErbB2, ErbB3, and ErbB4 receptors are established as drivers of many aspects of tumor initiation and progression to metastasis. This knowledge has provided rationales for the development of an arsenal of anti-ErbB therapeutics, ranging from small molecule kinase inhibitors to monoclonal antibodies. Anti-ErbB agents are becoming the cornerstone therapeutics for the management of cancers that overexpress hyperactive variants of ErbB receptors, in particular ErbB2-positive breast cancer and non-small cell lung carcinomas. However, their clinical benefit has been limited to a subset of patients due to a wide heterogeneity in drug response despite the expression of the ErbB targets, attributed to intrinsic (primary) and to acquired (secondary) resistance. Somatic mutations in ErbB tyrosine kinase domains have been extensively investigated in preclinical and clinical setting as determinants for either high sensitivity or resistance to anti-ErbB therapeutics. In contrast, only scant information is available on the impact of SNPs, which are widespread in genes encoding ErbB receptors, on receptor structure and activity, and their predictive values for drug susceptibility. This review aims to briefly update polymorphic variations in genes encoding ErbB receptors based on recent advances in deep sequencing technologies, and to address challenging issues for a better understanding of the functional impact of single versus combined SNPs in ErbB genes to receptor topology, receptor-drug interaction, and drug susceptibility. The potential of exploiting SNPs in the era of stratified targeted therapeutics is discussed.
Prudent, James R.; Hall, Jeff G.; Lyamichev, Victor L.; Brow, Mary Ann D.; Dahlberg, James E.
2007-12-11
The present invention relates to means for the detection and characterization of nucleic acid sequences, as well as variations in nucleic acid sequences. The present invention also relates to methods for forming a nucleic acid cleavage structure on a target sequence and cleaving the nucleic acid cleavage structure in a site-specific manner. The structure-specific nuclease activity of a variety of enzymes is used to cleave the target-dependent cleavage structure, thereby indicating the presence of specific nucleic acid sequences or specific variations thereof.
Invasive cleavage of nucleic acids
Prudent, James R.; Hall, Jeff G.; Lyamichev, Victor I.; Brow, Mary Ann D.; Dahlberg, James E.
1999-01-01
The present invention relates to means for the detection and characterization of nucleic acid sequences, as well as variations in nucleic acid sequences. The present invention also relates to methods for forming a nucleic acid cleavage structure on a target sequence and cleaving the nucleic acid cleavage structure in a site-specific manner. The structure-specific nuclease activity of a variety of enzymes is used to cleave the target-dependent cleavage structure, thereby indicating the presence of specific nucleic acid sequences or specific variations thereof.
Invasive cleavage of nucleic acids
Prudent, James R.; Hall, Jeff G.; Lyamichev, Victor I.; Brow, Mary Ann D.; Dahlberg, James E.
2002-01-01
The present invention relates to means for the detection and characterization of nucleic acid sequences, as well as variations in nucleic acid sequences. The present invention also relates to methods for forming a nucleic acid cleavage structure on a target sequence and cleaving the nucleic acid cleavage structure in a site-specific manner. The structure-specific nuclease activity of a variety of enzymes is used to cleave the target-dependent cleavage structure, thereby indicating the presence of specific nucleic acid sequences or specific variations thereof.
Prudent, James R.; Hall, Jeff G.; Lyamichev, Victor I.; Brow; Mary Ann D.; Dahlberg, James E.
2010-11-09
The present invention relates to means for the detection and characterization of nucleic acid sequences, as well as variations in nucleic acid sequences. The present invention also relates to methods for forming a nucleic acid cleavage structure on a target sequence and cleaving the nucleic acid cleavage structure in a site-specific manner. The structure-specific nuclease activity of a variety of enzymes is used to cleave the target-dependent cleavage structure, thereby indicating the presence of specific nucleic acid sequences or specific variations thereof.
Prudent, James R.; Hall, Jeff G.; Lyamichev, Victor I.; Brow, Mary Ann D.; Dahlberg, James E.
2000-01-01
The present invention relates to means for the detection and characterization of nucleic acid sequences, as well as variations in nucleic acid sequences. The present invention also relates to methods for forming a nucleic acid cleavage structure on a target sequence and cleaving the nucleic acid cleavage structure in a site-specific manner. The structure-specific nuclease activity of a variety of enzymes is used to cleave the target-dependent cleavage structure, thereby indicating the presence of specific nucleic acid sequences or specific variations thereof.
Prudent, James R.; Hall, Jeff G.; Lyamichev, Victor I.; Brow, Mary Ann; Dahlberg, James E.
2005-04-05
The present invention relates to means for the detection and characterization of nucleic acid sequences, as well as variations in nucleic acid sequences. The present invention also relates to methods for forming a nucleic acid cleavage structure on a target sequence and cleaving the nucleic acid cleavage structure in a site-specific manner. The structure-specific nuclease activity of a variety of enzymes is used to cleave the target-dependent cleavage structure, thereby indicating the presence of specific nucleic acid sequences or specific variations thereof.
Kumar, Girish; Kocour, Martin; Kunal, Swaraj Priyaranjan
2016-05-01
In order to assess the DNA sequence variation and phylogenetic relationship among five tuna species (Auxis thazard, Euthynnus affinis, Katsuwonus pelamis, Thunnus tonggol, and T. albacares) out of all four tuna genera, partial sequences of the mitochondrial DNA (mtDNA) D-loop region were analyzed. The estimate of intra-specific sequence variation in studied species was low, ranging from 0.027 to 0.080 [Kimura's two parameter distance (K2P)], whereas values of inter-specific variation ranged from 0.049 to 0.491. The longtail tuna (T. tonggol) and yellowfin tuna (T. albacares) were found to share a close relationship (K2P = 0.049) while skipjack tuna (K. pelamis) was most divergent studied species. Phylogenetic analysis using Maximum-Likelihood (ML) and Neighbor-Joining (NJ) methods supported the monophyletic origin of Thunnus species. Similarly, phylogeny of Auxis and Euthynnus species substantiate the monophyly. However, results showed a distinct origin of K. pelamis from genus Thunnus as well as Auxis and Euthynnus. Thus, the mtDNA D-loop region sequence data supports the polyphyletic origin of tuna species.
The Mouse Genomes Project: a repository of inbred laboratory mouse strain genomes.
Adams, David J; Doran, Anthony G; Lilue, Jingtao; Keane, Thomas M
2015-10-01
The Mouse Genomes Project was initiated in 2009 with the goal of using next-generation sequencing technologies to catalogue molecular variation in the common laboratory mouse strains, and a selected set of wild-derived inbred strains. The initial sequencing and survey of sequence variation in 17 inbred strains was completed in 2011 and included comprehensive catalogue of single nucleotide polymorphisms, short insertion/deletions, larger structural variants including their fine scale architecture and landscape of transposable element variation, and genomic sites subject to post-transcriptional alteration of RNA. From this beginning, the resource has expanded significantly to include 36 fully sequenced inbred laboratory mouse strains, a refined and updated data processing pipeline, and new variation querying and data visualisation tools which are available on the project's website ( http://www.sanger.ac.uk/resources/mouse/genomes/ ). The focus of the project is now the completion of de novo assembled chromosome sequences and strain-specific gene structures for the core strains. We discuss how the assembled chromosomes will power comparative analysis, data access tools and future directions of mouse genetics.
Sampson, Juliana K; Sheth, Nihar U; Koparde, Vishal N; Scalora, Allison F; Serrano, Myrna G; Lee, Vladimir; Roberts, Catherine H; Jameson-Lee, Max; Ferreira-Gonzalez, Andrea; Manjili, Masoud H; Buck, Gregory A; Neale, Michael C; Toor, Amir A
2014-08-01
Whole exome sequencing (WES) was performed on stem cell transplant donor-recipient (D-R) pairs to determine the extent of potential antigenic variation at a molecular level. In a small cohort of D-R pairs, a high frequency of sequence variation was observed between the donor and recipient exomes independent of human leucocyte antigen (HLA) matching. Nonsynonymous, nonconservative single nucleotide polymorphisms were approximately twice as frequent in HLA-matched unrelated, compared with related D-R pairs. When mapped to individual chromosomes, these polymorphic nucleotides were uniformly distributed across the entire exome. In conclusion, WES reveals extensive nucleotide sequence variation in the exomes of HLA-matched donors and recipients. © 2014 John Wiley & Sons Ltd.
2011-01-01
Background Integration of genomic variation with phenotypic information is an effective approach for uncovering genotype-phenotype associations. This requires an accurate identification of the different types of variation in individual genomes. Results We report the integration of the whole genome sequence of a single Holstein Friesian bull with data from single nucleotide polymorphism (SNP) and comparative genomic hybridization (CGH) array technologies to determine a comprehensive spectrum of genomic variation. The performance of resequencing SNP detection was assessed by combining SNPs that were identified to be either in identity by descent (IBD) or in copy number variation (CNV) with results from SNP array genotyping. Coding insertions and deletions (indels) were found to be enriched for size in multiples of 3 and were located near the N- and C-termini of proteins. For larger indels, a combination of split-read and read-pair approaches proved to be complementary in finding different signatures. CNVs were identified on the basis of the depth of sequenced reads, and by using SNP and CGH arrays. Conclusions Our results provide high resolution mapping of diverse classes of genomic variation in an individual bovine genome and demonstrate that structural variation surpasses sequence variation as the main component of genomic variability. Better accuracy of SNP detection was achieved with little loss of sensitivity when algorithms that implemented mapping quality were used. IBD regions were found to be instrumental for calculating resequencing SNP accuracy, while SNP detection within CNVs tended to be less reliable. CNV discovery was affected dramatically by platform resolution and coverage biases. The combined data for this study showed that at a moderate level of sequencing coverage, an ensemble of platforms and tools can be applied together to maximize the accurate detection of sequence and structural variants. PMID:22082336
Hundreds of variants clustered in genomic loci and biological pathways affect human height
Lango Allen, Hana; Estrada, Karol; Lettre, Guillaume; Berndt, Sonja I.; Weedon, Michael N.; Rivadeneira, Fernando; Willer, Cristen J.; Jackson, Anne U.; Vedantam, Sailaja; Raychaudhuri, Soumya; Ferreira, Teresa; Wood, Andrew R.; Weyant, Robert J.; Segrè, Ayellet V.; Speliotes, Elizabeth K.; Wheeler, Eleanor; Soranzo, Nicole; Park, Ju-Hyun; Yang, Jian; Gudbjartsson, Daniel; Heard-Costa, Nancy L.; Randall, Joshua C.; Qi, Lu; Smith, Albert Vernon; Mägi, Reedik; Pastinen, Tomi; Liang, Liming; Heid, Iris M.; Luan, Jian'an; Thorleifsson, Gudmar; Winkler, Thomas W.; Goddard, Michael E.; Lo, Ken Sin; Palmer, Cameron; Workalemahu, Tsegaselassie; Aulchenko, Yurii S.; Johansson, Åsa; Zillikens, M.Carola; Feitosa, Mary F.; Esko, Tõnu; Johnson, Toby; Ketkar, Shamika; Kraft, Peter; Mangino, Massimo; Prokopenko, Inga; Absher, Devin; Albrecht, Eva; Ernst, Florian; Glazer, Nicole L.; Hayward, Caroline; Hottenga, Jouke-Jan; Jacobs, Kevin B.; Knowles, Joshua W.; Kutalik, Zoltán; Monda, Keri L.; Polasek, Ozren; Preuss, Michael; Rayner, Nigel W.; Robertson, Neil R.; Steinthorsdottir, Valgerdur; Tyrer, Jonathan P.; Voight, Benjamin F.; Wiklund, Fredrik; Xu, Jianfeng; Zhao, Jing Hua; Nyholt, Dale R.; Pellikka, Niina; Perola, Markus; Perry, John R.B.; Surakka, Ida; Tammesoo, Mari-Liis; Altmaier, Elizabeth L.; Amin, Najaf; Aspelund, Thor; Bhangale, Tushar; Boucher, Gabrielle; Chasman, Daniel I.; Chen, Constance; Coin, Lachlan; Cooper, Matthew N.; Dixon, Anna L.; Gibson, Quince; Grundberg, Elin; Hao, Ke; Junttila, M. Juhani; Kaplan, Lee M.; Kettunen, Johannes; König, Inke R.; Kwan, Tony; Lawrence, Robert W.; Levinson, Douglas F.; Lorentzon, Mattias; McKnight, Barbara; Morris, Andrew P.; Müller, Martina; Ngwa, Julius Suh; Purcell, Shaun; Rafelt, Suzanne; Salem, Rany M.; Salvi, Erika; Sanna, Serena; Shi, Jianxin; Sovio, Ulla; Thompson, John R.; Turchin, Michael C.; Vandenput, Liesbeth; Verlaan, Dominique J.; Vitart, Veronique; White, Charles C.; Ziegler, Andreas; Almgren, Peter; Balmforth, Anthony J.; Campbell, Harry; Citterio, Lorena; De Grandi, Alessandro; Dominiczak, Anna; Duan, Jubao; Elliott, Paul; Elosua, Roberto; Eriksson, Johan G.; Freimer, Nelson B.; Geus, Eco J.C.; Glorioso, Nicola; Haiqing, Shen; Hartikainen, Anna-Liisa; Havulinna, Aki S.; Hicks, Andrew A.; Hui, Jennie; Igl, Wilmar; Illig, Thomas; Jula, Antti; Kajantie, Eero; Kilpeläinen, Tuomas O.; Koiranen, Markku; Kolcic, Ivana; Koskinen, Seppo; Kovacs, Peter; Laitinen, Jaana; Liu, Jianjun; Lokki, Marja-Liisa; Marusic, Ana; Maschio, Andrea; Meitinger, Thomas; Mulas, Antonella; Paré, Guillaume; Parker, Alex N.; Peden, John F.; Petersmann, Astrid; Pichler, Irene; Pietiläinen, Kirsi H.; Pouta, Anneli; Ridderstråle, Martin; Rotter, Jerome I.; Sambrook, Jennifer G.; Sanders, Alan R.; Schmidt, Carsten Oliver; Sinisalo, Juha; Smit, Jan H.; Stringham, Heather M.; Walters, G.Bragi; Widen, Elisabeth; Wild, Sarah H.; Willemsen, Gonneke; Zagato, Laura; Zgaga, Lina; Zitting, Paavo; Alavere, Helene; Farrall, Martin; McArdle, Wendy L.; Nelis, Mari; Peters, Marjolein J.; Ripatti, Samuli; van Meurs, Joyce B.J.; Aben, Katja K.; Ardlie, Kristin G; Beckmann, Jacques S.; Beilby, John P.; Bergman, Richard N.; Bergmann, Sven; Collins, Francis S.; Cusi, Daniele; den Heijer, Martin; Eiriksdottir, Gudny; Gejman, Pablo V.; Hall, Alistair S.; Hamsten, Anders; Huikuri, Heikki V.; Iribarren, Carlos; Kähönen, Mika; Kaprio, Jaakko; Kathiresan, Sekar; Kiemeney, Lambertus; Kocher, Thomas; Launer, Lenore J.; Lehtimäki, Terho; Melander, Olle; Mosley, Tom H.; Musk, Arthur W.; Nieminen, Markku S.; O'Donnell, Christopher J.; Ohlsson, Claes; Oostra, Ben; Palmer, Lyle J.; Raitakari, Olli; Ridker, Paul M.; Rioux, John D.; Rissanen, Aila; Rivolta, Carlo; Schunkert, Heribert; Shuldiner, Alan R.; Siscovick, David S.; Stumvoll, Michael; Tönjes, Anke; Tuomilehto, Jaakko; van Ommen, Gert-Jan; Viikari, Jorma; Heath, Andrew C.; Martin, Nicholas G.; Montgomery, Grant W.; Province, Michael A.; Kayser, Manfred; Arnold, Alice M.; Atwood, Larry D.; Boerwinkle, Eric; Chanock, Stephen J.; Deloukas, Panos; Gieger, Christian; Grönberg, Henrik; Hall, Per; Hattersley, Andrew T.; Hengstenberg, Christian; Hoffman, Wolfgang; Lathrop, G.Mark; Salomaa, Veikko; Schreiber, Stefan; Uda, Manuela; Waterworth, Dawn; Wright, Alan F.; Assimes, Themistocles L.; Barroso, Inês; Hofman, Albert; Mohlke, Karen L.; Boomsma, Dorret I.; Caulfield, Mark J.; Cupples, L.Adrienne; Erdmann, Jeanette; Fox, Caroline S.; Gudnason, Vilmundur; Gyllensten, Ulf; Harris, Tamara B.; Hayes, Richard B.; Jarvelin, Marjo-Riitta; Mooser, Vincent; Munroe, Patricia B.; Ouwehand, Willem H.; Penninx, Brenda W.; Pramstaller, Peter P.; Quertermous, Thomas; Rudan, Igor; Samani, Nilesh J.; Spector, Timothy D.; Völzke, Henry; Watkins, Hugh; Wilson, James F.; Groop, Leif C.; Haritunians, Talin; Hu, Frank B.; Kaplan, Robert C.; Metspalu, Andres; North, Kari E.; Schlessinger, David; Wareham, Nicholas J.; Hunter, David J.; O'Connell, Jeffrey R.; Strachan, David P.; Wichmann, H.-Erich; Borecki, Ingrid B.; van Duijn, Cornelia M.; Schadt, Eric E.; Thorsteinsdottir, Unnur; Peltonen, Leena; Uitterlinden, André; Visscher, Peter M.; Chatterjee, Nilanjan; Loos, Ruth J.F.; Boehnke, Michael; McCarthy, Mark I.; Ingelsson, Erik; Lindgren, Cecilia M.; Abecasis, Gonçalo R.; Stefansson, Kari; Frayling, Timothy M.; Hirschhorn, Joel N
2010-01-01
Most common human traits and diseases have a polygenic pattern of inheritance: DNA sequence variants at many genetic loci influence phenotype. Genome-wide association (GWA) studies have identified >600 variants associated with human traits1, but these typically explain small fractions of phenotypic variation, raising questions about the utility of further studies. Here, using 183,727 individuals, we show that hundreds of genetic variants, in at least 180 loci, influence adult height, a highly heritable and classic polygenic trait2,3. The large number of loci reveals patterns with important implications for genetic studies of common human diseases and traits. First, the 180 loci are not random, but instead are enriched for genes that are connected in biological pathways (P=0.016), and that underlie skeletal growth defects (P<0.001). Second, the likely causal gene is often located near the most strongly associated variant: in 13 of 21 loci containing a known skeletal growth gene, that gene was closest to the associated variant. Third, at least 19 loci have multiple independently associated variants, suggesting that allelic heterogeneity is a frequent feature of polygenic traits, that comprehensive explorations of already-discovered loci should discover additional variants, and that an appreciable fraction of associated loci may have been identified. Fourth, associated variants are enriched for likely functional effects on genes, being over-represented amongst variants that alter amino acid structure of proteins and expression levels of nearby genes. Our data explain ∼10% of the phenotypic variation in height, and we estimate that unidentified common variants of similar effect sizes would increase this figure to ∼16% of phenotypic variation (∼20% of heritable variation). Although additional approaches are needed to fully dissect the genetic architecture of polygenic human traits, our findings indicate that GWA studies can identify large numbers of loci that implicate biologically relevant genes and pathways. PMID:20881960
Beyond DNA: integrating inclusive inheritance into an extended theory of evolution.
Danchin, Étienne; Charmantier, Anne; Champagne, Frances A; Mesoudi, Alex; Pujol, Benoit; Blanchet, Simon
2011-06-17
Many biologists are calling for an 'extended evolutionary synthesis' that would 'modernize the modern synthesis' of evolution. Biological information is typically considered as being transmitted across generations by the DNA sequence alone, but accumulating evidence indicates that both genetic and non-genetic inheritance, and the interactions between them, have important effects on evolutionary outcomes. We review the evidence for such effects of epigenetic, ecological and cultural inheritance and parental effects, and outline methods that quantify the relative contributions of genetic and non-genetic heritability to the transmission of phenotypic variation across generations. These issues have implications for diverse areas, from the question of missing heritability in human complex-trait genetics to the basis of major evolutionary transitions.
Whole-Genome Sequences of Thirteen Isolates of Borrelia burgdorferi
DOE Office of Scientific and Technical Information (OSTI.GOV)
Schutzer S. E.; Dunn J.; Fraser-Liggett, C. M.
2011-02-01
Borrelia burgdorferi is a causative agent of Lyme disease in North America and Eurasia. The first complete genome sequence of B. burgdorferi strain 31, available for more than a decade, has assisted research on the pathogenesis of Lyme disease. Because a single genome sequence is not sufficient to understand the relationship between genotypic and geographic variation and disease phenotype, we determined the whole-genome sequences of 13 additional B. burgdorferi isolates that span the range of natural variation. These sequences should allow improved understanding of pathogenesis and provide a foundation for novel detection, diagnosis, and prevention strategies.
Karas, Vlad O; Sinnott-Armstrong, Nicholas A; Varghese, Vici; Shafer, Robert W; Greenleaf, William J; Sherlock, Gavin
2018-01-01
Abstract Much of the within species genetic variation is in the form of single nucleotide polymorphisms (SNPs), typically detected by whole genome sequencing (WGS) or microarray-based technologies. However, WGS produces mostly uninformative reads that perfectly match the reference, while microarrays require genome-specific reagents. We have developed Diff-seq, a sequencing-based mismatch detection assay for SNP discovery without the requirement for specialized nucleic-acid reagents. Diff-seq leverages the Surveyor endonuclease to cleave mismatched DNA molecules that are generated after cross-annealing of a complex pool of DNA fragments. Sequencing libraries enriched for Surveyor-cleaved molecules result in increased coverage at the variant sites. Diff-seq detected all mismatches present in an initial test substrate, with specific enrichment dependent on the identity and context of the variation. Application to viral sequences resulted in increased observation of variant alleles in a biologically relevant context. Diff-Seq has the potential to increase the sensitivity and efficiency of high-throughput sequencing in the detection of variation. PMID:29361139
RSAT 2015: Regulatory Sequence Analysis Tools
Medina-Rivera, Alejandra; Defrance, Matthieu; Sand, Olivier; Herrmann, Carl; Castro-Mondragon, Jaime A.; Delerce, Jeremy; Jaeger, Sébastien; Blanchet, Christophe; Vincens, Pierre; Caron, Christophe; Staines, Daniel M.; Contreras-Moreira, Bruno; Artufel, Marie; Charbonnier-Khamvongsa, Lucie; Hernandez, Céline; Thieffry, Denis; Thomas-Chollier, Morgane; van Helden, Jacques
2015-01-01
RSAT (Regulatory Sequence Analysis Tools) is a modular software suite for the analysis of cis-regulatory elements in genome sequences. Its main applications are (i) motif discovery, appropriate to genome-wide data sets like ChIP-seq, (ii) transcription factor binding motif analysis (quality assessment, comparisons and clustering), (iii) comparative genomics and (iv) analysis of regulatory variations. Nine new programs have been added to the 43 described in the 2011 NAR Web Software Issue, including a tool to extract sequences from a list of coordinates (fetch-sequences from UCSC), novel programs dedicated to the analysis of regulatory variants from GWAS or population genomics (retrieve-variation-seq and variation-scan), a program to cluster motifs and visualize the similarities as trees (matrix-clustering). To deal with the drastic increase of sequenced genomes, RSAT public sites have been reorganized into taxon-specific servers. The suite is well-documented with tutorials and published protocols. The software suite is available through Web sites, SOAP/WSDL Web services, virtual machines and stand-alone programs at http://www.rsat.eu/. PMID:25904632
Potenza, L; Cafiero, M A; Camarda, A; La Salandra, G; Cucchiarini, L; Dachà, M
2009-10-01
In the present work mites previously identified as Dermanyssus gallinae De Geer (Acari, Mesostigmata) using morphological keys were investigated by molecular tools. The complete internal transcribed spacer 1 (ITS1), 5.8S ribosomal DNA, and ITS2 region of the ribosomal DNA from mites were amplified and sequenced to examine the level of sequence variations and to explore the feasibility of using this region in the identification of this mite. Conserved primers located at the 3'end of 18S and at the 5'start of 28S rRNA genes were used first, and amplified fragments were sequenced. Sequence analyses showed no variation in 5.8S and ITS2 region while slight intraspecific variations involving substitutions as well as deletions concentrated in the ITS1 region. Based on the sequence analyses a nested PCR of the ITS2 region followed by RFLP analyses has been set up in the attempt to provide a rapid molecular diagnostic tool of D. gallinae.
Mapping and phasing of structural variation in patient genomes using nanopore sequencing.
Cretu Stancu, Mircea; van Roosmalen, Markus J; Renkens, Ivo; Nieboer, Marleen M; Middelkamp, Sjors; de Ligt, Joep; Pregno, Giulia; Giachino, Daniela; Mandrile, Giorgia; Espejo Valle-Inclan, Jose; Korzelius, Jerome; de Bruijn, Ewart; Cuppen, Edwin; Talkowski, Michael E; Marschall, Tobias; de Ridder, Jeroen; Kloosterman, Wigard P
2017-11-06
Despite improvements in genomics technology, the detection of structural variants (SVs) from short-read sequencing still poses challenges, particularly for complex variation. Here we analyse the genomes of two patients with congenital abnormalities using the MinION nanopore sequencer and a novel computational pipeline-NanoSV. We demonstrate that nanopore long reads are superior to short reads with regard to detection of de novo chromothripsis rearrangements. The long reads also enable efficient phasing of genetic variations, which we leveraged to determine the parental origin of all de novo chromothripsis breakpoints and to resolve the structure of these complex rearrangements. Additionally, genome-wide surveillance of inherited SVs reveals novel variants, missed in short-read data sets, a large proportion of which are retrotransposon insertions. We provide a first exploration of patient genome sequencing with a nanopore sequencer and demonstrate the value of long-read sequencing in mapping and phasing of SVs for both clinical and research applications.
Variations analysis of NLGN3 and NLGN4X gene in Chinese autism patients.
Xu, Xiaojuan; Xiong, Zhimin; Zhang, Lusi; Liu, Yalan; Lu, Lina; Peng, Yu; Guo, Hui; Zhao, Jingping; Xia, Kun; Hu, Zhengmao
2014-06-01
Autism is a neurodevelopmental disorder clinically characterized by impairment of social interaction, deficits in verbal communication, as well as stereotypic and repetitive behaviors. Several studies have implicated that abnormal synaptogenesis was involved in the incidence of autism. Neuroligins are postsynaptic cell adhesion molecules and interacted with neurexins to regulate the fine balance between excitation and inhibition of synapses. Recently, mutation analysis, cellular and mice models hinted neuroligin mutations probably affected synapse maturation and function. In this study, four missense variations [p.G426S (NLGN3), p.G84R (NLGN4X), p.Q162 K (NLGN4X) and p.A283T (NLGN4X)] in four different unrelated patients have been identified by PCR and direct sequencing. These four missense variations were absent in the 453 controls and have not been reported in 1000 Genomes Project. Bioinformatic analysis of the four missense variations revealed that p.G84R and p.A283T were "Probably Damaging". The variations may cause abnormal synaptic homeostasis and therefore trigger the patients more predisposed to autism. By case-control analysis, we identified the common SNPs (rs3747333 and rs3747334) in the NLGN4X gene significantly associated with risk for autism [p = 5.09E-005; OR 4.685 (95% CI 2.073-10.592)]. Our data provided a further evidence for the involvement of NLGN3 and NLGN4X gene in the pathogenesis of autism in Chinese population.
Human structural variation: mechanisms of chromosome rearrangements
Weckselblatt, Brooke; Rudd, M. Katharine
2015-01-01
Chromosome structural variation (SV) is a normal part of variation in the human genome, but some classes of SV can cause neurodevelopmental disorders. Analysis of the DNA sequence at SV breakpoints can reveal mutational mechanisms and risk factors for chromosome rearrangement. Large-scale SV breakpoint studies have become possible recently owing to advances in next-generation sequencing (NGS) including whole-genome sequencing (WGS). These findings have shed light on complex forms of SV such as triplications, inverted duplications, insertional translocations, and chromothripsis. Sequence-level breakpoint data resolve SV structure and determine how genes are disrupted, fused, and/or misregulated by breakpoints. Recent improvements in breakpoint sequencing have also revealed non-allelic homologous recombination (NAHR) between paralogous long interspersed nuclear element (LINE) or human endogenous retrovirus (HERV) repeats as a cause of deletions, duplications, and translocations. This review covers the genomic organization of simple and complex constitutional SVs, as well as the molecular mechanisms of their formation. PMID:26209074
Comparative analysis of vaginal microbiota sampling using 16S rRNA gene analysis.
Virtanen, Seppo; Kalliala, Ilkka; Nieminen, Pekka; Salonen, Anne
2017-01-01
Molecular methods such as next-generation sequencing are actively being employed to characterize the vaginal microbiota in health and disease. Previous studies have focused on characterizing the biological variation in the microbiota, and less is known about how factors related to sampling contribute to the results. Our aim was to investigate the impact of a sampling device and anatomical sampling site on the quantitative and qualitative outcomes relevant for vaginal microbiota research. We sampled 10 Finnish women representing diverse clinical characteristics with flocked swabs, the Evalyn® self-sampling device, sterile plastic spatulas and a cervical brush that were used to collect samples from fornix, vaginal wall and cervix. Samples were compared on DNA and protein yield, bacterial load, and microbiota diversity and species composition based on Illumina MiSeq sequencing of the 16S rRNA gene. We quantified the relative contributions of sampling variables versus intrinsic variables in the overall microbiota variation, and evaluated the microbiota profiles using several commonly employed metrics such as alpha and beta diversity as well as abundance of major bacterial genera and species. The total DNA yield was strongly dependent on the sampling device and to a lesser extent on the anatomical site of sampling. The sampling strategy did not affect the protein yield or the bacterial load. All tested sampling methods produced highly comparable microbiota profiles based on MiSeq sequencing. The sampling method explained only 2% (p-value = 0.89) of the overall microbiota variation, markedly surpassed by intrinsic factors such as clinical status (microscopy for bacterial vaginosis 53%, p = 0.0001), bleeding (19%, p = 0.0001), and the variation between subjects (11%, p-value 0.0001). The results indicate that different sampling strategies yield comparable vaginal microbiota composition and diversity. Hence, past and future vaginal microbiota studies employing different sampling strategies should be comparable in the absence of other technical confounders. The Evalyn® self-sampling device performed equally well compared to samples taken by a clinician, and hence offers a good-quality microbiota sample without the need for a gynecological examination. The amount of collected sample as well as the DNA and protein yield varied across the sampling techniques, which may have practical implications for study design.
Comparative analysis of vaginal microbiota sampling using 16S rRNA gene analysis
Kalliala, Ilkka; Nieminen, Pekka; Salonen, Anne
2017-01-01
Background Molecular methods such as next-generation sequencing are actively being employed to characterize the vaginal microbiota in health and disease. Previous studies have focused on characterizing the biological variation in the microbiota, and less is known about how factors related to sampling contribute to the results. Our aim was to investigate the impact of a sampling device and anatomical sampling site on the quantitative and qualitative outcomes relevant for vaginal microbiota research. We sampled 10 Finnish women representing diverse clinical characteristics with flocked swabs, the Evalyn® self-sampling device, sterile plastic spatulas and a cervical brush that were used to collect samples from fornix, vaginal wall and cervix. Samples were compared on DNA and protein yield, bacterial load, and microbiota diversity and species composition based on Illumina MiSeq sequencing of the 16S rRNA gene. We quantified the relative contributions of sampling variables versus intrinsic variables in the overall microbiota variation, and evaluated the microbiota profiles using several commonly employed metrics such as alpha and beta diversity as well as abundance of major bacterial genera and species. Results The total DNA yield was strongly dependent on the sampling device and to a lesser extent on the anatomical site of sampling. The sampling strategy did not affect the protein yield or the bacterial load. All tested sampling methods produced highly comparable microbiota profiles based on MiSeq sequencing. The sampling method explained only 2% (p-value = 0.89) of the overall microbiota variation, markedly surpassed by intrinsic factors such as clinical status (microscopy for bacterial vaginosis 53%, p = 0.0001), bleeding (19%, p = 0.0001), and the variation between subjects (11%, p-value 0.0001). Conclusions The results indicate that different sampling strategies yield comparable vaginal microbiota composition and diversity. Hence, past and future vaginal microbiota studies employing different sampling strategies should be comparable in the absence of other technical confounders. The Evalyn® self-sampling device performed equally well compared to samples taken by a clinician, and hence offers a good-quality microbiota sample without the need for a gynecological examination. The amount of collected sample as well as the DNA and protein yield varied across the sampling techniques, which may have practical implications for study design. PMID:28723942
Read clouds uncover variation in complex regions of the human genome.
Bishara, Alex; Liu, Yuling; Weng, Ziming; Kashef-Haghighi, Dorna; Newburger, Daniel E; West, Robert; Sidow, Arend; Batzoglou, Serafim
2015-10-01
Although an increasing amount of human genetic variation is being identified and recorded, determining variants within repeated sequences of the human genome remains a challenge. Most population and genome-wide association studies have therefore been unable to consider variation in these regions. Core to the problem is the lack of a sequencing technology that produces reads with sufficient length and accuracy to enable unique mapping. Here, we present a novel methodology of using read clouds, obtained by accurate short-read sequencing of DNA derived from long fragment libraries, to confidently align short reads within repeat regions and enable accurate variant discovery. Our novel algorithm, Random Field Aligner (RFA), captures the relationships among the short reads governed by the long read process via a Markov Random Field. We utilized a modified version of the Illumina TruSeq synthetic long-read protocol, which yielded shallow-sequenced read clouds. We test RFA through extensive simulations and apply it to discover variants on the NA12878 human sample, for which shallow TruSeq read cloud sequencing data are available, and on an invasive breast carcinoma genome that we sequenced using the same method. We demonstrate that RFA facilitates accurate recovery of variation in 155 Mb of the human genome, including 94% of 67 Mb of segmental duplication sequence and 96% of 11 Mb of transcribed sequence, that are currently hidden from short-read technologies. © 2015 Bishara et al.; Published by Cold Spring Harbor Laboratory Press.
Grievink, Liat Shavit; Penny, David; Hendy, Mike D; Holland, Barbara R
2009-01-01
Correction to Shavit Grievink L, Penny D, Hendy MD, Holland BR: LineageSpecificSeqgen: generating sequence data with lineage-specific variation in the proportion of variable sites. BMC Evol Biol 2008, 8(1):317.
BayesPI-BAR: a new biophysical model for characterization of regulatory sequence variations
Wang, Junbai; Batmanov, Kirill
2015-01-01
Sequence variations in regulatory DNA regions are known to cause functionally important consequences for gene expression. DNA sequence variations may have an essential role in determining phenotypes and may be linked to disease; however, their identification through analysis of massive genome-wide sequencing data is a great challenge. In this work, a new computational pipeline, a Bayesian method for protein–DNA interaction with binding affinity ranking (BayesPI-BAR), is proposed for quantifying the effect of sequence variations on protein binding. BayesPI-BAR uses biophysical modeling of protein–DNA interactions to predict single nucleotide polymorphisms (SNPs) that cause significant changes in the binding affinity of a regulatory region for transcription factors (TFs). The method includes two new parameters (TF chemical potentials or protein concentrations and direct TF binding targets) that are neglected by previous methods. The new method is verified on 67 known human regulatory SNPs, of which 47 (70%) have predicted true TFs ranked in the top 10. Importantly, the performance of BayesPI-BAR, which uses principal component analysis to integrate multiple predictions from various TF chemical potentials, is found to be better than that of existing programs, such as sTRAP and is-rSNP, when evaluated on the same SNPs. BayesPI-BAR is a publicly available tool and is able to carry out parallelized computation, which helps to investigate a large number of TFs or SNPs and to detect disease-associated regulatory sequence variations in the sea of genome-wide noncoding regions. PMID:26202972
Melendrez, Melanie C.; Lange, Rachel K.; Cohan, Frederick M.; Ward, David M.
2011-01-01
Previous research has shown that sequences of 16S rRNA genes and 16S-23S rRNA internal transcribed spacer regions may not have enough genetic resolution to define all ecologically distinct Synechococcus populations (ecotypes) inhabiting alkaline, siliceous hot spring microbial mats. To achieve higher molecular resolution, we studied sequence variation in three protein-encoding loci sampled by PCR from 60°C and 65°C sites in the Mushroom Spring mat (Yellowstone National Park, WY). Sequences were analyzed using the ecotype simulation (ES) and AdaptML algorithms to identify putative ecotypes. Between 4 and 14 times more putative ecotypes were predicted from variation in protein-encoding locus sequences than from variation in 16S rRNA and 16S-23S rRNA internal transcribed spacer sequences. The number of putative ecotypes predicted depended on the number of sequences sampled and the molecular resolution of the locus. Chao estimates of diversity indicated that few rare ecotypes were missed. Many ecotypes hypothesized by sequence analyses were different in their habitat specificities, suggesting different adaptations to temperature or other parameters that vary along the flow channel. PMID:21169433
Fietz, Katharina; Rye Hintze, Christian Olaf; Skovrind, Mikkel; Kjærgaard Nielsen, Tue; Limborg, Morten T; Krag, Marcus A; Palsbøll, Per J; Hestbjerg Hansen, Lars; Rask Møller, Peter; Gilbert, M Thomas P
2018-05-02
Deciphering the mechanisms governing population genetic divergence and local adaptation across heterogeneous environments is a central theme in marine ecology and conservation. While population divergence and ecological adaptive potential are classically viewed at the genetic level, it has recently been argued that their microbiomes may also contribute to population genetic divergence. We explored whether this might be plausible along the well-described environmental gradient of the Baltic Sea in two species of sand lance (Ammodytes tobianus and Hyperoplus lanceolatus). Specifically, we assessed both their population genetic and gut microbial composition variation and investigated not only which environmental parameters correlate with the observed variation, but whether host genome also correlates with microbiome variation. We found a clear genetic structure separating the high-salinity North Sea from the low-salinity Baltic Sea sand lances. The observed genetic divergence was not simply a function of isolation by distance, but correlated with environmental parameters, such as salinity, sea surface temperature, and, in the case of A. tobianus, possibly water microbiota. Furthermore, we detected two distinct genetic groups in Baltic A. tobianus that might represent sympatric spawning types. Investigation of possible drivers of gut microbiome composition variation revealed that host species identity was significantly correlated with the microbial community composition of the gut. A potential influence of host genetic factors on gut microbiome composition was further confirmed by the results of a constrained analysis of principal coordinates. The host genetic component was among the parameters that best explain observed variation in gut microbiome composition. Our findings have relevance for the population structure of two commercial species but also provide insights into potentially relevant genomic and microbial factors with regards to sand lance adaptation across the North Sea-Baltic Sea environmental gradient. Furthermore, our findings support the hypothesis that host genetics may play a role in regulating the gut microbiome at both the interspecific and intraspecific levels. As sequencing costs continue to drop, we anticipate that future studies that include full genome and microbiome sequencing will be able to explore the full relationship and its potential adaptive implications for these species.
In Silico Detection of Sequence Variations Modifying Transcriptional Regulation
Andersen, Malin C; Engström, Pär G; Lithwick, Stuart; Arenillas, David; Eriksson, Per; Lenhard, Boris; Wasserman, Wyeth W; Odeberg, Jacob
2008-01-01
Identification of functional genetic variation associated with increased susceptibility to complex diseases can elucidate genes and underlying biochemical mechanisms linked to disease onset and progression. For genes linked to genetic diseases, most identified causal mutations alter an encoded protein sequence. Technological advances for measuring RNA abundance suggest that a significant number of undiscovered causal mutations may alter the regulation of gene transcription. However, it remains a challenge to separate causal genetic variations from linked neutral variations. Here we present an in silico driven approach to identify possible genetic variation in regulatory sequences. The approach combines phylogenetic footprinting and transcription factor binding site prediction to identify variation in candidate cis-regulatory elements. The bioinformatics approach has been tested on a set of SNPs that are reported to have a regulatory function, as well as background SNPs. In the absence of additional information about an analyzed gene, the poor specificity of binding site prediction is prohibitive to its application. However, when additional data is available that can give guidance on which transcription factor is involved in the regulation of the gene, the in silico binding site prediction improves the selection of candidate regulatory polymorphisms for further analyses. The bioinformatics software generated for the analysis has been implemented as a Web-based application system entitled RAVEN (regulatory analysis of variation in enhancers). The RAVEN system is available at http://www.cisreg.ca for all researchers interested in the detection and characterization of regulatory sequence variation. PMID:18208319
Brown, J. R.; Beckenbach, K.; Beckenbach, A. T.; Smith, M. J.
1996-01-01
The extent of mtDNA length variation and heteroplasmy as well as DNA sequences of the control region and two tRNA genes were determined for four North American sturgeon species: Acipenser transmontanus, A. medirostris, A. fulvescens and A. oxyrhnychus. Across the Continental Divide, a division in the occurrence of length variation and heteroplasmy was observed that was concordant with species biogeography as well as with phylogenies inferred from restriction fragment length polymorphisms (RFLP) of whole mtDNA and pairwise comparisons of unique sequences of the control region. In all species, mtDNA length variation was due to repeated arrays of 78-82-bp sequences each containing a D-loop strand synthesis termination associated sequence (TAS). Individual repeats showed greater sequence conservation within individuals and species rather than between species, which is suggestive of concerted evolution. Differences in the frequencies of multiple copy genomes and heteroplasmy among the four species may be ascribed to differences in the rates of recurrent mutation. A mechanism that may offset the high rate of mutation for increased copy number is suggested on the basis that an increase in the number of functional TAS motifs might reduce the frequency of successfully initiated H-strand replications. PMID:8852850
Mitochondrial DNA Sequence Variation in North Atlantic Long-Finned Pilot Whales, Globicephala melas
1994-06-01
Delphinapterus leucas : mitochondrial DNA sequence variation within and among North American populations. M.Sc. thesis. McMaster University. Brown, G.G...Delphinapteras leucas ) (Brennin 1992), minke whales {Balaenoptera acutorostratd) (Wada et al. 1991), bottlenose dolphins {Tursiops truncatus) (Dowling & Brown
Widespread Transient Hoogsteen Base-Pairs in Canonical Duplex DNA with Variable Energetics
Alvey, Heidi S.; Gottardo, Federico L.; Nikolova, Evgenia N.; Al-Hashimi, Hashim M.
2015-01-01
Hoogsteen base-pairing involves a 180 degree rotation of the purine base relative to Watson-Crick base-pairing within DNA duplexes, creating alternative DNA conformations that can play roles in recognition, damage induction, and replication. Here, using Nuclear Magnetic Resonance R1ρ relaxation dispersion, we show that transient Hoogsteen base-pairs occur across more diverse sequence and positional contexts than previously anticipated. We observe sequence-specific variations in Hoogsteen base-pair energetic stabilities that are comparable to variations in Watson-Crick base-pair stability, with Hoogsteen base-pairs being more abundant for energetically less favorable Watson-Crick base-pairs. Our results suggest that the variations in Hoogsteen stabilities and rates of formation are dominated by variations in Watson-Crick base pair stability, suggesting a late transition state for the Watson-Crick to Hoogsteen conformational switch. The occurrence of sequence and position-dependent Hoogsteen base-pairs provide a new potential mechanism for achieving sequence-dependent DNA transactions. PMID:25185517
CNV-seq, a new method to detect copy number variation using high-throughput sequencing.
Xie, Chao; Tammi, Martti T
2009-03-06
DNA copy number variation (CNV) has been recognized as an important source of genetic variation. Array comparative genomic hybridization (aCGH) is commonly used for CNV detection, but the microarray platform has a number of inherent limitations. Here, we describe a method to detect copy number variation using shotgun sequencing, CNV-seq. The method is based on a robust statistical model that describes the complete analysis procedure and allows the computation of essential confidence values for detection of CNV. Our results show that the number of reads, not the length of the reads is the key factor determining the resolution of detection. This favors the next-generation sequencing methods that rapidly produce large amount of short reads. Simulation of various sequencing methods with coverage between 0.1x to 8x show overall specificity between 91.7 - 99.9%, and sensitivity between 72.2 - 96.5%. We also show the results for assessment of CNV between two individual human genomes.
Koole, Cassandra; Savage, Emilia E.; Christopoulos, Arthur; Miller, Laurence J.
2013-01-01
The glucagon-like peptide-1 receptor (GLP-1R) controls the physiological responses to the incretin hormone glucagon-like peptide-1 and is a major therapeutic target for the treatment of type 2 diabetes, owing to the broad range of effects that are mediated upon its activation. These include the promotion of glucose-dependent insulin secretion, increased insulin biosynthesis, preservation of β-cell mass, improved peripheral insulin action, and promotion of weight loss. Regulation of GLP-1R function is complex, with multiple endogenous and exogenous peptides that interact with the receptor that result in the activation of numerous downstream signaling cascades. The current understanding of GLP-1R signaling and regulation is limited, with the desired spectrum of signaling required for the ideal therapeutic outcome still to be determined. In addition, there are several single-nucleotide polymorphisms (used in this review as defining a natural change of single nucleotide in the receptor sequence; clinically, this is viewed as a single-nucleotide polymorphism only if the frequency of the mutation occurs in 1% or more of the population) distributed within the coding sequence of the receptor protein that have the potential to produce differential responses for distinct ligands. In this review, we discuss the current understanding of GLP-1R function, in particular highlighting recent advances in the field on ligand-directed signal bias, allosteric modulation, and probe dependence and the implications of these behaviors for drug discovery and development. PMID:23864649
Wang, Jin; Dong, Hongping; Chionh, Yok Hian; McBee, Megan E.; Sirirungruang, Sasilada; Cunningham, Richard P.; Shi, Pei-Yong; Dedon, Peter C.
2016-01-01
The misincorporation of 2′-deoxyribonucleotides (dNs) into RNA has important implications for the function of non-coding RNAs, the translational fidelity of coding RNAs and the mutagenic evolution of viral RNA genomes. However, quantitative appreciation for the degree to which dN misincorporation occurs is limited by the lack of analytical tools. Here, we report a method to hydrolyze RNA to release 2′-deoxyribonucleotide-ribonucleotide pairs (dNrN) that are then quantified by chromatography-coupled mass spectrometry (LC-MS). Using this platform, we found misincorporated dNs occurring at 1 per 103 to 105 ribonucleotide (nt) in mRNA, rRNAs and tRNA in human cells, Escherichia coli, Saccharomyces cerevisiae and, most abundantly, in the RNA genome of dengue virus. The frequency of dNs varied widely among organisms and sequence contexts, and partly reflected the in vitro discrimination efficiencies of different RNA polymerases against 2′-deoxyribonucleoside 5′-triphosphates (dNTPs). Further, we demonstrate a strong link between dN frequencies in RNA and the balance of dNTPs and ribonucleoside 5′-triphosphates (rNTPs) in the cellular pool, with significant stress-induced variation of dN incorporation. Potential implications of dNs in RNA are discussed, including the possibilities of dN incorporation in RNA as a contributing factor in viral evolution and human disease, and as a host immune defense mechanism against viral infections. PMID:27365049
Fernández, Cecilia S; Bruque, Carlos D; Taboas, Melisa; Buzzalino, Noemí D; Espeche, Lucia D; Pasqualini, Titania; Charreau, Eduardo H; Alba, Liliana G; Ghiringhelli, Pablo D; Dain, Liliana
2015-09-01
The aim of the current study was to search for the presence of genetic variants in the CYP21A2 Z promoter regulatory region in patients with congenital adrenal hyperplasia due to 21-hydroxylase deficiency. Screening of the 10 most frequent pseudogene-derived mutations was followed by direct sequencing of the entire coding sequence, the proximal promoter, and a distal regulatory region in DNA samples from patients with at least one non-determined allele. We report three non-classical patients that presented a novel genetic variant-g.15626A>G-within the Z promoter regulatory region. In all the patients, the novel variant was found in cis with the mild, less frequent, p.P482S mutation located in the exon 10 of the CYP21A2 gene. The putative pathogenic implication of the novel variant was assessed by in silico analyses and in vitro assays. Topological analyses showed differences in the curvature and bendability of the DNA region bearing the novel variant. By performing functional studies, a significantly decreased activity of a reporter gene placed downstream from the regulatory region was found by the G transition. Our results may suggest that the activity of an allele bearing the p.P482S mutation may be influenced by the misregulated CYP21A2 transcriptional activity exerted by the Z promoter A>G variation.
Sikora, Martin; Carpenter, Meredith L.; Moreno-Estrada, Andres; Henn, Brenna M.; Underhill, Peter A.; Sánchez-Quinto, Federico; Zara, Ilenia; Pitzalis, Maristella; Sidore, Carlo; Busonero, Fabio; Maschio, Andrea; Angius, Andrea; Jones, Chris; Mendoza-Revilla, Javier; Nekhrizov, Georgi; Dimitrova, Diana; Theodossiev, Nikola; Harkins, Timothy T.; Keller, Andreas; Maixner, Frank; Zink, Albert; Abecasis, Goncalo; Sanna, Serena; Cucca, Francesco; Bustamante, Carlos D.
2014-01-01
Genome sequencing of the 5,300-year-old mummy of the Tyrolean Iceman, found in 1991 on a glacier near the border of Italy and Austria, has yielded new insights into his origin and relationship to modern European populations. A key finding of that study was an apparent recent common ancestry with individuals from Sardinia, based largely on the Y chromosome haplogroup and common autosomal SNP variation. Here, we compiled and analyzed genomic datasets from both modern and ancient Europeans, including genome sequence data from over 400 Sardinians and two ancient Thracians from Bulgaria, to investigate this result in greater detail and determine its implications for the genetic structure of Neolithic Europe. Using whole-genome sequencing data, we confirm that the Iceman is, indeed, most closely related to Sardinians. Furthermore, we show that this relationship extends to other individuals from cultural contexts associated with the spread of agriculture during the Neolithic transition, in contrast to individuals from a hunter-gatherer context. We hypothesize that this genetic affinity of ancient samples from different parts of Europe with Sardinians represents a common genetic component that was geographically widespread across Europe during the Neolithic, likely related to migrations and population expansions associated with the spread of agriculture. PMID:24809476
Genomic and chromatin features shaping meiotic double-strand break formation and repair in mice
Jasin, Maria; Lange, Julian
2017-01-01
ABSTRACT The SPO11-generated DNA double-strand breaks (DSBs) that initiate meiotic recombination occur non-randomly across genomes, but mechanisms shaping their distribution and repair remain incompletely understood. Here, we expand on recent studies of nucleotide-resolution DSB maps in mouse spermatocytes. We find that trimethylation of histone H3 lysine 36 around DSB hotspots is highly correlated, both spatially and quantitatively, with trimethylation of H3 lysine 4, consistent with coordinated formation and action of both PRDM9-dependent histone modifications. In contrast, the DSB-responsive kinase ATM contributes independently of PRDM9 to controlling hotspot activity, and combined action of ATM and PRDM9 can explain nearly two-thirds of the variation in DSB frequency between hotspots. DSBs were modestly underrepresented in most repetitive sequences such as segmental duplications and transposons. Nonetheless, numerous DSBs form within repetitive sequences in each meiosis and some classes of repeats are preferentially targeted. Implications of these findings are discussed for evolution of PRDM9 and its role in hybrid strain sterility in mice. Finally, we document the relationship between mouse strain-specific DNA sequence variants within PRDM9 recognition motifs and attendant differences in recombination outcomes. Our results provide further insights into the complex web of factors that influence meiotic recombination patterns. PMID:28820351
The roles of WRN and BLM RecQ helicases in the Alternative Lengthening of Telomeres
Mendez-Bermudez, Aaron; Hidalgo-Bravo, Alberto; Cotton, Victoria E.; Gravani, Athanasia; Jeyapalan, Jennie N.; Royle, Nicola J.
2012-01-01
Approximately 10% of all cancers, but a higher proportion of sarcomas, use the recombination-based alternative lengthening of telomeres (ALT) to maintain telomeres. Two RecQ helicase genes, BLM and WRN, play important roles in homologous recombination repair and they have been implicated in telomeric recombination activity, but their precise roles in ALT are unclear. Using analysis of sequence variation present in human telomeres, we found that a WRN– ALT+ cell line lacks the class of complex telomere mutations attributed to inter-telomeric recombination in other ALT+ cell lines. This suggests that WRN facilitates inter-telomeric recombination when there are sequence differences between the donor and recipient molecules or that sister-telomere interactions are suppressed in the presence of WRN and this promotes inter-telomeric recombination. Depleting BLM in the WRN– ALT+ cell line increased the mutation frequency at telomeres and at the MS32 minisatellite, which is a marker of ALT. The absence of complex telomere mutations persisted in BLM-depleted clones, and there was a clear increase in sequence homogenization across the telomere and MS32 repeat arrays. These data indicate that BLM suppresses unequal sister chromatid interactions that result in excessive homogenization at MS32 and at telomeres in ALT+ cells. PMID:22989712
The roles of WRN and BLM RecQ helicases in the Alternative Lengthening of Telomeres.
Mendez-Bermudez, Aaron; Hidalgo-Bravo, Alberto; Cotton, Victoria E; Gravani, Athanasia; Jeyapalan, Jennie N; Royle, Nicola J
2012-11-01
Approximately 10% of all cancers, but a higher proportion of sarcomas, use the recombination-based alternative lengthening of telomeres (ALT) to maintain telomeres. Two RecQ helicase genes, BLM and WRN, play important roles in homologous recombination repair and they have been implicated in telomeric recombination activity, but their precise roles in ALT are unclear. Using analysis of sequence variation present in human telomeres, we found that a WRN- ALT+ cell line lacks the class of complex telomere mutations attributed to inter-telomeric recombination in other ALT+ cell lines. This suggests that WRN facilitates inter-telomeric recombination when there are sequence differences between the donor and recipient molecules or that sister-telomere interactions are suppressed in the presence of WRN and this promotes inter-telomeric recombination. Depleting BLM in the WRN- ALT+ cell line increased the mutation frequency at telomeres and at the MS32 minisatellite, which is a marker of ALT. The absence of complex telomere mutations persisted in BLM-depleted clones, and there was a clear increase in sequence homogenization across the telomere and MS32 repeat arrays. These data indicate that BLM suppresses unequal sister chromatid interactions that result in excessive homogenization at MS32 and at telomeres in ALT+ cells.
Molecular mechanisms of epigenetic variation in plants.
Fujimoto, Ryo; Sasaki, Taku; Ishikawa, Ryo; Osabe, Kenji; Kawanabe, Takahiro; Dennis, Elizabeth S
2012-01-01
Natural variation is defined as the phenotypic variation caused by spontaneous mutations. In general, mutations are associated with changes of nucleotide sequence, and many mutations in genes that can cause changes in plant development have been identified. Epigenetic change, which does not involve alteration to the nucleotide sequence, can also cause changes in gene activity by changing the structure of chromatin through DNA methylation or histone modifications. Now there is evidence based on induced or spontaneous mutants that epigenetic changes can cause altering plant phenotypes. Epigenetic changes have occurred frequently in plants, and some are heritable or metastable causing variation in epigenetic status within or between species. Therefore, heritable epigenetic variation as well as genetic variation has the potential to drive natural variation.
The study of human Y chromosome variation through ancient DNA.
Kivisild, Toomas
2017-05-01
High throughput sequencing methods have completely transformed the study of human Y chromosome variation by offering a genome-scale view on genetic variation retrieved from ancient human remains in context of a growing number of high coverage whole Y chromosome sequence data from living populations from across the world. The ancient Y chromosome sequences are providing us the first exciting glimpses into the past variation of male-specific compartment of the genome and the opportunity to evaluate models based on previously made inferences from patterns of genetic variation in living populations. Analyses of the ancient Y chromosome sequences are challenging not only because of issues generally related to ancient DNA work, such as DNA damage-induced mutations and low content of endogenous DNA in most human remains, but also because of specific properties of the Y chromosome, such as its highly repetitive nature and high homology with the X chromosome. Shotgun sequencing of uniquely mapping regions of the Y chromosomes to sufficiently high coverage is still challenging and costly in poorly preserved samples. To increase the coverage of specific target SNPs capture-based methods have been developed and used in recent years to generate Y chromosome sequence data from hundreds of prehistoric skeletal remains. Besides the prospects of testing directly as how much genetic change in a given time period has accompanied changes in material culture the sequencing of ancient Y chromosomes allows us also to better understand the rate at which mutations accumulate and get fixed over time. This review considers genome-scale evidence on ancient Y chromosome diversity that has recently started to accumulate in geographic areas favourable to DNA preservation. More specifically the review focuses on examples of regional continuity and change of the Y chromosome haplogroups in North Eurasia and in the New World.
Detection of nucleic acid sequences by invader-directed cleavage
Brow, Mary Ann D.; Hall, Jeff Steven Grotelueschen; Lyamichev, Victor; Olive, David Michael; Prudent, James Robert
1999-01-01
The present invention relates to means for the detection and characterization of nucleic acid sequences, as well as variations in nucleic acid sequences. The present invention also relates to methods for forming a nucleic acid cleavage structure on a target sequence and cleaving the nucleic acid cleavage structure in a site-specific manner. The 5' nuclease activity of a variety of enzymes is used to cleave the target-dependent cleavage structure, thereby indicating the presence of specific nucleic acid sequences or specific variations thereof. The present invention further relates to methods and devices for the separation of nucleic acid molecules based by charge.
Rocha, Leonardo de Souza; Falqueto, Aloisio; Dos Santos, Claudiney Biral; Grimaldi, Gabriel Júnior; Cupolillo, Elisa
2011-09-01
Lutzomyia longipalpis (Diptera: Psychodidae) is the principal vector of American visceral leishmaniasis. Several studies have indicated that the Lu. longipalpis population structure is complex. It has been suggested that genetic divergence caused by genetic drift, selection, or both may affect the vectorial capacity of Lu. longipalpis. However, it remains unclear whether genetic differences among Lu. longipalpis populations are directly implicated in the transmission features of visceral leishmaniasis. We evaluated the genetic composition and the patterns of genetic differentiation among Lu. longipalpis populations collected from regions with different patterns of transmission of visceral leishmaniasis by analyzing the sequence variation in the mitochondrial cytochrome b gene. Furthermore, we investigated the temporal distribution of haplotypes and compared our results with those obtained in a previous study. Our data indicate that there are differences in the haplotype composition and that there has been significant differentiation between the analyzed populations. Our results reveal that measures used to control visceral leishmaniasis might have influenced the genetic composition of the vector population. This finding raises important questions concerning the epidemiology of visceral leishmaniasis, because these differences in the genetic structures among populations of Lu. longipalpis may have implications with respect to their efficiency as vectors for visceral leishmaniasis.
Litim, Nadhir; Labrie, Yvan; Desjardins, Sylvie; Ouellette, Geneviève; Plourde, Karine; Belleau, Pascal; Durocher, Francine
2013-02-01
The majority of genes associated with breast cancer susceptibility, including BRCA1 and BRCA2 genes, are involved in DNA repair mechanisms. Moreover, among the genes recently associated with an increased susceptibility to breast cancer, four are Fanconi Anemia (FA) genes: FANCD1/BRCA2, FANCJ/BACH1/BRIP1, FANCN/PALB2 and FANCO/RAD51C. FANCA is implicated in DNA repair and has been shown to interact directly with BRCA1. It has been proposed that the formation of FANCA/G (dependent upon the phosphorylation of FANCA) and FANCB/L sub-complexes altogether with FANCM, represent the initial step for DNA repair activation and subsequent formation of other sub-complexes leading to ubiquitination of FANCD2 and FANCI. As only approximately 25% of inherited breast cancers are attributable to BRCA1/2 mutations, FANCA therefore becomes an attractive candidate for breast cancer susceptibility. We thus analyzed FANCA gene in 97 high-risk French Canadian non-BRCA1/2 breast cancer individuals by direct sequencing as well as in 95 healthy control individuals from the same population. Among a total of 85 sequence variants found in either or both series, 28 are coding variants and 19 of them are missense variations leading to amino acid change. Three of the amino acid changes, namely Thr561Met, Cys625Ser and particularly Ser1088Phe, which has been previously reported to be associated with FA, are predicted to be damaging by the SIFT and PolyPhen softwares. cDNA amplification revealed significant expression of 4 alternative splicing events (insertion of an intronic portion of intron 10, and the skipping of exons 11, 30 and 31). In silico analyzes of relevant genomic variants have been performed in order to identify potential variations involved in the expression of these spliced transcripts. Sequence variants in FANCA could therefore be potential spoilers of the Fanconi-BRCA pathway and as a result, they could in turn have an impact in non-BRCA1/2 breast cancer families. Copyright © 2012 Federation of European Biochemical Societies. Published by Elsevier B.V. All rights reserved.
Genome sequence of Phytophthora ramorum: implications for management
Brett Tyler; Sucheta Tripathy; Nik Grunwald; Kurt Lamour; Kelly Ivors; Matteo Garbelotto; Daniel Rokhsar; Nik Putnam; Igor Grigoriev; Jeffrey Boore
2006-01-01
A draft genome sequence has been determined for Phytophthora ramorum, together with a draft sequence of the soybean pathogen Phytophthora sojae. The P. ramorum genome was sequenced to a depth of 7-fold coverage, while the P. sojae genome was sequenced to a depth of 9-fold coverage. The genome...
Positive Selection Underlies Faster-Z Evolution of Gene Expression in Birds
Dean, Rebecca; Harrison, Peter W.; Wright, Alison E.; Zimmer, Fabian; Mank, Judith E.
2015-01-01
The elevated rate of evolution for genes on sex chromosomes compared with autosomes (Fast-X or Fast-Z evolution) can result either from positive selection in the heterogametic sex or from nonadaptive consequences of reduced relative effective population size. Recent work in birds suggests that Fast-Z of coding sequence is primarily due to relaxed purifying selection resulting from reduced relative effective population size. However, gene sequence and gene expression are often subject to distinct evolutionary pressures; therefore, we tested for Fast-Z in gene expression using next-generation RNA-sequencing data from multiple avian species. Similar to studies of Fast-Z in coding sequence, we recover clear signatures of Fast-Z in gene expression; however, in contrast to coding sequence, our data indicate that Fast-Z in expression is due to positive selection acting primarily in females. In the soma, where gene expression is highly correlated between the sexes, we detected Fast-Z in both sexes, although at a higher rate in females, suggesting that many positively selected expression changes in females are also expressed in males. In the gonad, where intersexual correlations in expression are much lower, we detected Fast-Z for female gene expression, but crucially, not males. This suggests that a large amount of expression variation is sex-specific in its effects within the gonad. Taken together, our results indicate that Fast-Z evolution of gene expression is the product of positive selection acting on recessive beneficial alleles in the heterogametic sex. More broadly, our analysis suggests that the adaptive potential of Z chromosome gene expression may be much greater than that of gene sequence, results which have important implications for the role of sex chromosomes in speciation and sexual selection. PMID:26067773
Variation in Symbiodinium ITS2 sequence assemblages among coral colonies.
Stat, Michael; Bird, Christopher E; Pochon, Xavier; Chasqui, Luis; Chauka, Leonard J; Concepcion, Gregory T; Logan, Dan; Takabayashi, Misaki; Toonen, Robert J; Gates, Ruth D
2011-01-05
Endosymbiotic dinoflagellates in the genus Symbiodinium are fundamentally important to the biology of scleractinian corals, as well as to a variety of other marine organisms. The genus Symbiodinium is genetically and functionally diverse and the taxonomic nature of the union between Symbiodinium and corals is implicated as a key trait determining the environmental tolerance of the symbiosis. Surprisingly, the question of how Symbiodinium diversity partitions within a species across spatial scales of meters to kilometers has received little attention, but is important to understanding the intrinsic biological scope of a given coral population and adaptations to the local environment. Here we address this gap by describing the Symbiodinium ITS2 sequence assemblages recovered from colonies of the reef building coral Montipora capitata sampled across Kāne'ohe Bay, Hawai'i. A total of 52 corals were sampled in a nested design of Coral Colony(Site(Region)) reflecting spatial scales of meters to kilometers. A diversity of Symbiodinium ITS2 sequences was recovered with the majority of variance partitioning at the level of the Coral Colony. To confirm this result, the Symbiodinium ITS2 sequence diversity in six M. capitata colonies were analyzed in much greater depth with 35 to 55 clones per colony. The ITS2 sequences and quantitative composition recovered from these colonies varied significantly, indicating that each coral hosted a different assemblage of Symbiodinium. The diversity of Symbiodinium ITS2 sequence assemblages retrieved from individual colonies of M. capitata here highlights the problems inherent in interpreting multi-copy and intra-genomically variable molecular markers, and serves as a context for discussing the utility and biological relevance of assigning species names based on Symbiodinium ITS2 genotyping.
Variation in Symbiodinium ITS2 Sequence Assemblages among Coral Colonies
Stat, Michael; Bird, Christopher E.; Pochon, Xavier; Chasqui, Luis; Chauka, Leonard J.; Concepcion, Gregory T.; Logan, Dan; Takabayashi, Misaki; Toonen, Robert J.; Gates, Ruth D.
2011-01-01
Endosymbiotic dinoflagellates in the genus Symbiodinium are fundamentally important to the biology of scleractinian corals, as well as to a variety of other marine organisms. The genus Symbiodinium is genetically and functionally diverse and the taxonomic nature of the union between Symbiodinium and corals is implicated as a key trait determining the environmental tolerance of the symbiosis. Surprisingly, the question of how Symbiodinium diversity partitions within a species across spatial scales of meters to kilometers has received little attention, but is important to understanding the intrinsic biological scope of a given coral population and adaptations to the local environment. Here we address this gap by describing the Symbiodinium ITS2 sequence assemblages recovered from colonies of the reef building coral Montipora capitata sampled across Kāne'ohe Bay, Hawai'i. A total of 52 corals were sampled in a nested design of Coral Colony(Site(Region)) reflecting spatial scales of meters to kilometers. A diversity of Symbiodinium ITS2 sequences was recovered with the majority of variance partitioning at the level of the Coral Colony. To confirm this result, the Symbiodinium ITS2 sequence diversity in six M. capitata colonies were analyzed in much greater depth with 35 to 55 clones per colony. The ITS2 sequences and quantitative composition recovered from these colonies varied significantly, indicating that each coral hosted a different assemblage of Symbiodinium. The diversity of Symbiodinium ITS2 sequence assemblages retrieved from individual colonies of M. capitata here highlights the problems inherent in interpreting multi-copy and intra-genomically variable molecular markers, and serves as a context for discussing the utility and biological relevance of assigning species names based on Symbiodinium ITS2 genotyping. PMID:21246044
How life changes itself: the Read-Write (RW) genome.
Shapiro, James A
2013-09-01
The genome has traditionally been treated as a Read-Only Memory (ROM) subject to change by copying errors and accidents. In this review, I propose that we need to change that perspective and understand the genome as an intricately formatted Read-Write (RW) data storage system constantly subject to cellular modifications and inscriptions. Cells operate under changing conditions and are continually modifying themselves by genome inscriptions. These inscriptions occur over three distinct time-scales (cell reproduction, multicellular development and evolutionary change) and involve a variety of different processes at each time scale (forming nucleoprotein complexes, epigenetic formatting and changes in DNA sequence structure). Research dating back to the 1930s has shown that genetic change is the result of cell-mediated processes, not simply accidents or damage to the DNA. This cell-active view of genome change applies to all scales of DNA sequence variation, from point mutations to large-scale genome rearrangements and whole genome duplications (WGDs). This conceptual change to active cell inscriptions controlling RW genome functions has profound implications for all areas of the life sciences. © 2013 Elsevier B.V. All rights reserved.
Negotiating towards a next turn: phonetic resources for 'doing the same'.
Sikveland, Rein Ove
2012-03-01
This paper investigates hearers' use of response tokens (back-channels), in maintaining and differentiating their actions. Initial observations suggest that hearers produce a sequence of phonetically similar responses to disengage from the current topic, and dissimilar responses to engage with the current topic. This is studied systematically by combining detailed interactional and phonetic analysis in a collection of naturally-occurring talk in Norwegian. The interactional analysis forms the basis for labeling actions as maintained ('doing the same') and differentiated ('NOT doing the same'), which is then used as a basis for phonetic analysis. The phonetic analysis shows that certain phonetic characteristics, including pitch, loudness, voice quality and articulatory characteristics, are associated with 'doing the same', as different from 'NOT doing the same'. Interactional analysis gives further evidence of how this differentiation is of systematic relevance in the negotiations of a next turn. This paper addresses phonetic variation and variability by focusing on the relationship between sequence and phonetics in the turn-by-turn development of meaning. This has important implications for linguistic/phonetic research, and for the study of back-channels.
Al-Bustan, Suzanne A; Al-Serri, Ahmad; Annice, Babitha G; Alnaqeeb, Majed A; Al-Kandari, Wafa Y; Dashti, Mohammed
2018-01-01
The role interethnic genetic differences play in plasma lipid level variation across populations is a global health concern. Several genes involved in lipid metabolism and transport are strong candidates for the genetic association with lipid level variation especially lipoprotein lipase (LPL). The objective of this study was to re-sequence the full LPL gene in Kuwaiti Arabs, analyse the sequence variation and identify variants that could attribute to variation in plasma lipid levels for further genetic association. Samples (n = 100) of an Arab ethnic group from Kuwait were analysed for sequence variation by Sanger sequencing across the 30 Kb LPL gene and its flanking sequences. A total of 293 variants including 252 single nucleotide polymorphisms (SNPs) and 39 insertions/deletions (InDels) were identified among which 47 variants (32 SNPs and 15 InDels) were novel to Kuwaiti Arabs. This study is the first to report sequence data and analysis of frequencies of variants at the LPL gene locus in an Arab ethnic group with a novel "rare" variant (LPL:g.18704C>A) significantly associated to HDL (B = -0.181; 95% CI (-0.357, -0.006); p = 0.043), TG (B = 0.134; 95% CI (0.004-0.263); p = 0.044) and VLDL (B = 0.131; 95% CI (-0.001-0.263); p = 0.043) levels. Sequence variation in Kuwaiti Arabs was compared to other populations and was found to be similar with regards to the number of SNPs, InDels and distribution of the number of variants across the LPL gene locus and minor allele frequency (MAF). Moreover, comparison of the identified variants and their MAF with other reports provided a list of 46 potential variants across the LPL gene to be considered for future genetic association studies. The findings warrant further investigation into the association of g.18704C>A with lipid levels in other ethnic groups and with clinical manifestations of dyslipidemia.
Al-Serri, Ahmad; Annice, Babitha G.; Alnaqeeb, Majed A.; Al-Kandari, Wafa Y.; Dashti, Mohammed
2018-01-01
The role interethnic genetic differences play in plasma lipid level variation across populations is a global health concern. Several genes involved in lipid metabolism and transport are strong candidates for the genetic association with lipid level variation especially lipoprotein lipase (LPL). The objective of this study was to re-sequence the full LPL gene in Kuwaiti Arabs, analyse the sequence variation and identify variants that could attribute to variation in plasma lipid levels for further genetic association. Samples (n = 100) of an Arab ethnic group from Kuwait were analysed for sequence variation by Sanger sequencing across the 30 Kb LPL gene and its flanking sequences. A total of 293 variants including 252 single nucleotide polymorphisms (SNPs) and 39 insertions/deletions (InDels) were identified among which 47 variants (32 SNPs and 15 InDels) were novel to Kuwaiti Arabs. This study is the first to report sequence data and analysis of frequencies of variants at the LPL gene locus in an Arab ethnic group with a novel “rare” variant (LPL:g.18704C>A) significantly associated to HDL (B = -0.181; 95% CI (-0.357, -0.006); p = 0.043), TG (B = 0.134; 95% CI (0.004–0.263); p = 0.044) and VLDL (B = 0.131; 95% CI (-0.001–0.263); p = 0.043) levels. Sequence variation in Kuwaiti Arabs was compared to other populations and was found to be similar with regards to the number of SNPs, InDels and distribution of the number of variants across the LPL gene locus and minor allele frequency (MAF). Moreover, comparison of the identified variants and their MAF with other reports provided a list of 46 potential variants across the LPL gene to be considered for future genetic association studies. The findings warrant further investigation into the association of g.18704C>A with lipid levels in other ethnic groups and with clinical manifestations of dyslipidemia. PMID:29438437
Demidov, German; Simakova, Tamara; Vnuchkova, Julia; Bragin, Anton
2016-10-22
Multiplex polymerase chain reaction (PCR) is a common enrichment technique for targeted massive parallel sequencing (MPS) protocols. MPS is widely used in biomedical research and clinical diagnostics as the fast and accurate tool for the detection of short genetic variations. However, identification of larger variations such as structure variants and copy number variations (CNV) is still being a challenge for targeted MPS. Some approaches and tools for structural variants detection were proposed, but they have limitations and often require datasets of certain type, size and expected number of amplicons affected by CNVs. In the paper, we describe novel algorithm for high-resolution germinal CNV detection in the PCR-enriched targeted sequencing data and present accompanying tool. We have developed a machine learning algorithm for the detection of large duplications and deletions in the targeted sequencing data generated with PCR-based enrichment step. We have performed verification studies and established the algorithm's sensitivity and specificity. We have compared developed tool with other available methods applicable for the described data and revealed its higher performance. We showed that our method has high specificity and sensitivity for high-resolution copy number detection in targeted sequencing data using large cohort of samples.
Ismail, Nurul-Ain; Adilah-Amrannudin, Nurul; Hamsidi, Mayamin; Ismail, Rodziah; Dom, Nazri Che; Ahmad, Abu Hassan; Mastuki, Mohd Fahmi; Camalxaman, Siti Nazrina
2017-11-07
The global expansion of Ae. albopictus from its native range in Southeast Asia has been implicated in the recent emergence of dengue endemicity in Malaysia. Genetic variability studies of Ae. albopictus are currently lacking in the Malaysian setting, yet are crucial to enhancing the existing vector control strategies. The study was conducted to establish the genetic variability of maternally inherited mitochondrial DNA encoding for cytochrome oxidase subunit 1 (CO1) gene in Ae. albopictus. Twelve localities were selected in the Subang Jaya district based on temporal indices utilizing 120 mosquito samples. Genetic polymorphism and phylogenetic analysis were conducted to unveil the genetic variability and geographic origins of Ae. albopictus. The haplotype network was mapped to determine the genealogical relationship of sequences among groups of population in the Asian region. Comparison of Malaysian CO1 sequences with sequences derived from five Asian countries revealed genetically distinct Ae. albopictus populations. Phylogenetic analysis revealed that all sequences from other Asian countries descended from the same genetic lineage as the Malaysian sequences. Noteworthy, our study highlights the discovery of 20 novel haplotypes within the Malaysian population which to date had not been reported. These findings could help determine the genetic variation of this invasive species, which in turn could possibly improve the current dengue vector surveillance strategies, locally and regionally. © The Authors 2017. Published by Oxford University Press on behalf of Entomological Society of America. All rights reserved. For Permissions, please email: journals.permissions@oup.com.
Characterization of microRNAs Expressed during Secondary Wall Biosynthesis in Acacia mangium
Ong, Seong Siang; Wickneswari, Ratnam
2012-01-01
MicroRNAs (miRNAs) play critical regulatory roles by acting as sequence specific guide during secondary wall formation in woody and non-woody species. Although thousands of plant miRNAs have been sequenced, there is no comprehensive view of miRNA mediated gene regulatory network to provide profound biological insights into the regulation of xylem development. Herein, we report the involvement of six highly conserved amg-miRNA families (amg-miR166, amg-miR172, amg-miR168, amg-miR159, amg-miR394, and amg-miR156) as the potential regulatory sequences of secondary cell wall biosynthesis. Within this highly conserved amg-miRNA family, only amg-miR166 exhibited strong differences in expression between phloem and xylem tissue. The functional characterization of amg-miR166 targets in various tissues revealed three groups of HD-ZIP III: ATHB8, ATHB15, and REVOLUTA which play pivotal roles in xylem development. Although these three groups vary in their functions, -psRNA target analysis indicated that miRNA target sequences of the nine different members of HD-ZIP III are always conserved. We found that precursor structures of amg-miR166 undergo exhaustive sequence variation even within members of the same family. Gene expression analysis showed three key lignin pathway genes: C4H, CAD, and CCoAOMT were upregulated in compression wood where a cascade of miRNAs was downregulated. This study offers a comprehensive analysis on the involvement of highly conserved miRNAs implicated in the secondary wall formation of woody plants. PMID:23251324
Liu, G H; Zhou, W; Nisbet, A J; Xu, M J; Zhou, D H; Zhao, G H; Wang, S K; Song, H Q; Lin, R Q; Zhu, X Q
2014-03-01
Trichuris trichiura and Trichuris suis parasitize (at the adult stage) the caeca of humans and pigs, respectively, causing trichuriasis. Despite these parasites being of human and animal health significance, causing considerable socio-economic losses globally, little is known of the molecular characteristics of T. trichiura and T. suis from China. In the present study, the entire first and second internal transcribed spacer (ITS-1 and ITS-2) regions of nuclear ribosomal DNA (rDNA) of T. trichiura and T. suis from China were amplified by polymerase chain reaction (PCR), the representative amplicons were cloned and sequenced, and sequence variation in the ITS rDNA was examined. The ITS rDNA sequences for the T. trichiura and T. suis samples were 1222-1267 bp and 1339-1353 bp in length, respectively. Sequence analysis revealed that the ITS-1, 5.8S and ITS-2 rDNAs of both whipworms were 600-627 bp and 655-661 bp, 154 bp, and 468-486 bp and 530-538 bp in size, respectively. Sequence variation in ITS rDNA within and among T. trichiura and T. suis was examined. Excluding nucleotide variations in the simple sequence repeats, the intra-species sequence variation in the ITS-1 was 0.2-1.7% within T. trichiura, and 0-1.5% within T. suis. For ITS-2 rDNA, the intra-species sequence variation was 0-1.3% within T. trichiura and 0.2-1.7% within T. suis. The inter-species sequence differences between the two whipworms were 60.7-65.3% for ITS-1 and 59.3-61.5% for ITS-2. These results demonstrated that the ITS rDNA sequences provide additional genetic markers for the characterization and differentiation of the two whipworms. These data should be useful for studying the epidemiology and population genetics of T. trichiura and T. suis, as well as for the diagnosis of trichuriasis in humans and pigs.
Mikaeili, F; Mirhendi, H; Mohebali, M; Hosseini, M; Sharbatkhori, M; Zarei, Z; Kia, E B
2015-07-01
The study was conducted to determine the sequence variation in two mitochondrial genes, namely cytochrome c oxidase 1 (pcox1) and NADH dehydrogenase 1 (pnad1) within and among isolates of Toxocara cati, Toxocara canis and Toxascaris leonina. Genomic DNA was extracted from 32 isolates of T. cati, 9 isolates of T. canis and 19 isolates of T. leonina collected from cats and dogs in different geographical areas of Iran. Mitochondrial genes were amplified by polymerase chain reaction (PCR) and sequenced. Sequence data were aligned using the BioEdit software and compared with published sequences in GenBank. Phylogenetic analysis was performed using Bayesian inference and maximum likelihood methods. Based on pairwise comparison, intra-species genetic diversity within Iranian isolates of T. cati, T. canis and T. leonina amounted to 0-2.3%, 0-1.3% and 0-1.0% for pcox1 and 0-2.0%, 0-1.7% and 0-2.6% for pnad1, respectively. Inter-species sequence variation among the three ascaridoid nematodes was significantly higher, being 9.5-16.6% for pcox1 and 11.9-26.7% for pnad1. Sequence and phylogenetic analysis of the pcox1 and pnad1 genes indicated that there is significant genetic diversity within and among isolates of T. cati, T. canis and T. leonina from different areas of Iran, and these genes can be used for studying genetic variation of ascaridoid nematodes.
Copy number variation of individual cattle genomes using next-generation sequencing
USDA-ARS?s Scientific Manuscript database
Copy number variations (CNVs) affect a wide range of phenotypic traits; however, CNVs in or near segmental duplication regions are often intractable. Using a read depth approach based on next-generation sequencing, we examined genome-wide copy number differences among five taurine (three Angus, one ...
Copy number variation of individual cattle genomes using next-generation sequencing
USDA-ARS?s Scientific Manuscript database
Copy Number Variations (CNVs) affect a wide range of phenotypic traits; however, CNVs in or near segmental duplication regions are often difficult to track. Using a read depth approach based on next generation sequencing, we examined genome-wide copy number differences among five taurine (three Angu...
A high-resolution cattle CNV map by population-scale genome sequencing
USDA-ARS?s Scientific Manuscript database
Copy Number Variations (CNVs) are common genomic structural variations that have been linked to human diseases and phenotypic traits. Prior studies in cattle have produced low-resolution CNV maps. We constructed a draft, high-resolution map of cattle CNVs based on whole genome sequencing data from 7...
Maize HapMap2 identifies extant variation from a genome in flux
USDA-ARS?s Scientific Manuscript database
The maize genome is the largest, most diverse and complex plant genome sequenced to date. Using high-throughput sequencing to access genetic variation and a population genetics model to score the polymorphisms, we characterize and unite the diversity of the world’s key breeding germplasm, wild rela...
RSAT 2015: Regulatory Sequence Analysis Tools.
Medina-Rivera, Alejandra; Defrance, Matthieu; Sand, Olivier; Herrmann, Carl; Castro-Mondragon, Jaime A; Delerce, Jeremy; Jaeger, Sébastien; Blanchet, Christophe; Vincens, Pierre; Caron, Christophe; Staines, Daniel M; Contreras-Moreira, Bruno; Artufel, Marie; Charbonnier-Khamvongsa, Lucie; Hernandez, Céline; Thieffry, Denis; Thomas-Chollier, Morgane; van Helden, Jacques
2015-07-01
RSAT (Regulatory Sequence Analysis Tools) is a modular software suite for the analysis of cis-regulatory elements in genome sequences. Its main applications are (i) motif discovery, appropriate to genome-wide data sets like ChIP-seq, (ii) transcription factor binding motif analysis (quality assessment, comparisons and clustering), (iii) comparative genomics and (iv) analysis of regulatory variations. Nine new programs have been added to the 43 described in the 2011 NAR Web Software Issue, including a tool to extract sequences from a list of coordinates (fetch-sequences from UCSC), novel programs dedicated to the analysis of regulatory variants from GWAS or population genomics (retrieve-variation-seq and variation-scan), a program to cluster motifs and visualize the similarities as trees (matrix-clustering). To deal with the drastic increase of sequenced genomes, RSAT public sites have been reorganized into taxon-specific servers. The suite is well-documented with tutorials and published protocols. The software suite is available through Web sites, SOAP/WSDL Web services, virtual machines and stand-alone programs at http://www.rsat.eu/. © The Author(s) 2015. Published by Oxford University Press on behalf of Nucleic Acids Research.
Benz, Matthias R; Bongartz, Georg; Froehlich, Johannes M; Winkel, David; Boll, Daniel T; Heye, Tobias
2018-07-01
The aim was to investigate the variation of the arterial input function (AIF) within and between various DCE MRI sequences. A dynamic flow-phantom and steady signal reference were scanned on a 3T MRI using fast low angle shot (FLASH) 2d, FLASH3d (parallel imaging factor (P) = P0, P2, P4), volumetric interpolated breath-hold examination (VIBE) (P = P0, P3, P2 × 2, P2 × 3, P3 × 2), golden-angle radial sparse parallel imaging (GRASP), and time-resolved imaging with stochastic trajectories (TWIST). Signal over time curves were normalized and quantitatively analyzed by full width half maximum (FWHM) measurements to assess variation within and between sequences. The coefficient of variation (CV) for the steady signal reference ranged from 0.07-0.8%. The non-accelerated gradient echo FLASH2d, FLASH3d, and VIBE sequences showed low within sequence variation with 2.1%, 1.0%, and 1.6%. The maximum FWHM CV was 3.2% for parallel imaging acceleration (VIBE P2 × 3), 2.7% for GRASP and 9.1% for TWIST. The FWHM CV between sequences ranged from 8.5-14.4% for most non-accelerated/accelerated gradient echo sequences except 6.2% for FLASH3d P0 and 0.3% for FLASH3d P2; GRASP FWHM CV was 9.9% versus 28% for TWIST. MRI acceleration techniques vary in reproducibility and quantification of the AIF. Incomplete coverage of the k-space with TWIST as a representative of view-sharing techniques showed the highest variation within sequences and might be less suited for reproducible quantification of the AIF. Copyright © 2018 Elsevier B.V. All rights reserved.
Schoeman, Elizna M; Lopez, Genghis H; McGowan, Eunike C; Millard, Glenda M; O'Brien, Helen; Roulis, Eileen V; Liew, Yew-Wah; Martin, Jacqueline R; McGrath, Kelli A; Powley, Tanya; Flower, Robert L; Hyland, Catherine A
2017-04-01
Blood group single nucleotide polymorphism genotyping probes for a limited range of polymorphisms. This study investigated whether massively parallel sequencing (also known as next-generation sequencing), with a targeted exome strategy, provides an extended blood group genotype and the extent to which massively parallel sequencing correctly genotypes in homologous gene systems, such as RH and MNS. Donor samples (n = 28) that were extensively phenotyped and genotyped using single nucleotide polymorphism typing, were analyzed using the TruSight One Sequencing Panel and MiSeq platform. Genes for 28 protein-based blood group systems, GATA1, and KLF1 were analyzed. Copy number variation analysis was used to characterize complex structural variants in the GYPC and RH systems. The average sequencing depth per target region was 66.2 ± 39.8. Each sample harbored on average 43 ± 9 variants, of which 10 ± 3 were used for genotyping. For the 28 samples, massively parallel sequencing variant sequences correctly matched expected sequences based on single nucleotide polymorphism genotyping data. Copy number variation analysis defined the Rh C/c alleles and complex RHD hybrids. Hybrid RHD*D-CE-D variants were correctly identified, but copy number variation analysis did not confidently distinguish between D and CE exon deletion versus rearrangement. The targeted exome sequencing strategy employed extended the range of blood group genotypes detected compared with single nucleotide polymorphism typing. This single-test format included detection of complex MNS hybrid cases and, with copy number variation analysis, defined RH hybrid genes along with the RHCE*C allele hitherto difficult to resolve by variant detection. The approach is economical compared with whole-genome sequencing and is suitable for a red blood cell reference laboratory setting. © 2017 AABB.
NASA Astrophysics Data System (ADS)
Liu, Xingxing; Sun, Youbin; Vandenberghe, Jef; Li, Ying; An, Zhisheng
2018-06-01
Sedimentary sequences that developed on river terraces have been widely investigated to reconstruct high-resolution palaeoclimatic changes since the last deglaciation. However, frequent changes in sedimentary facies make palaeoenvironmental interpretation of grain-size variations relatively complicated. In this paper, we employed multiple grain-size parameters to discriminate the sedimentary characteristics of aeolian and fluvial facies in the Dadiwan (DDW) section on the western Chinese Loess Plateau. We found that wind and fluvial dynamics have quite different impacts on the grain-size compositions, with distinctive imprints on the distribution pattern. By using a lognormal distribution fitting approach, two major grain-size components sensitive to aeolian and fluvial processes, respectively, were distinguished from the grain-size compositions of the DDW terrace deposits. The fine grain-size component (GSC2) represents mixing of long-distance aeolian and short-distance fluvial inputs, whilst the coarse grain-size component (GSC3) is mainly transported by wind from short-distance sources. Thus GSC3 can be used to infer the wind intensity. Grain-size variations reveal that the wind intensity experienced a stepwise shift from large-amplitude variations during the last deglaciation to small-amplitude oscillations in the Holocene, corresponding well to climate changes from regional to global context.
Rašić, Gordana; Schama, Renata; Powell, Rosanna; Maciel-de Freitas, Rafael; Endersby-Harshman, Nancy M; Filipović, Igor; Sylvestre, Gabriel; Máspero, Renato C; Hoffmann, Ary A
2015-01-01
Dengue is the most prevalent global arboviral disease that affects over 300 million people every year. Brazil has the highest number of dengue cases in the world, with the most severe epidemics in the city of Rio de Janeiro (Rio). The effective control of dengue is critically dependent on the knowledge of population genetic structuring in the primary dengue vector, the mosquito Aedes aegypti. We analyzed mitochondrial and nuclear genomewide single nucleotide polymorphism markers generated via Restriction-site Associated DNA sequencing, as well as traditional microsatellite markers in Ae. aegypti from Rio. We found four divergent mitochondrial lineages and a strong spatial structuring of mitochondrial variation, in contrast to the overall nuclear homogeneity across Rio. Despite a low overall differentiation in the nuclear genome, we detected strong spatial structure for variation in over 20 genes that have a significantly altered expression in response to insecticides, xenobiotics, and pathogens, including the novel biocontrol agent Wolbachia. Our results indicate that high genetic diversity, spatially unconstrained admixing likely mediated by male dispersal, along with locally heterogeneous genetic variation that could affect insecticide resistance and mosquito vectorial capacity, set limits to the effectiveness of measures to control dengue fever in Rio. PMID:26495042
Identification of rare genetic variation of NLRP1 gene in familial multiple sclerosis.
Maver, Ales; Lavtar, Polona; Ristić, Smiljana; Stopinšek, Sanja; Simčič, Saša; Hočevar, Keli; Sepčić, Juraj; Drulović, Jelena; Pekmezović, Tatjana; Novaković, Ivana; Alenka, Hodžić; Rudolf, Gorazd; Šega, Saša; Starčević-Čizmarević, Nada; Palandačić, Anja; Zamolo, Gordana; Kapović, Miljenko; Likar, Tina; Peterlin, Borut
2017-06-16
The genetic etiology and the contribution of rare genetic variation in multiple sclerosis (MS) has not yet been elucidated. Although familial forms of MS have been described, no convincing rare and penetrant variants have been reported to date. We aimed to characterize the contribution of rare genetic variation in familial and sporadic MS and have identified a family with two sibs affected by concomitant MS and malignant melanoma (MM). We performed whole exome sequencing in this primary family and 38 multiplex MS families and 44 sporadic MS cases and performed transcriptional and immunologic assessment of the identified variants. We identified a potentially causative homozygous missense variant in NLRP1 gene (Gly587Ser) in the primary family. Further possibly pathogenic NLRP1 variants were identified in the expanded cohort of patients. Stimulation of peripheral blood mononuclear cells from MS patients with putatively pathogenic NLRP1 variants showed an increase in IL-1B gene expression and active cytokine IL-1β production, as well as global activation of NLRP1-driven immunologic pathways. We report a novel familial association of MS and MM, and propose a possible underlying genetic basis in NLRP1 gene. Furthermore, we provide initial evidence of the broader implications of NLRP1-related pathway dysfunction in MS.
Bavarva, Jasmin H.; Tae, Hongseok; McIver, Lauren; Garner, Harold R.
2014-01-01
Although the connection between cancer and cigarette smoke is well established, nicotine is not characterized as a carcinogen. Here, we used exome sequencing to identify nicotine and oxidative stress-induced somatic mutations in normal human epithelial cells and its correlation with cancer. We identified over 6,400 SNVs, indels and microsatellites in each of the stress exposed cells relative to the control, of which, 2,159 were consistently observed at all nicotine doses. These included 429 nsSNVs including 158 novel and 79 cancer-associated. Over 80% of consistently nicotine induced variants overlap with variations detected in oxidative stressed cells, indicating that nicotine induced genomic alterations could be mediated through oxidative stress. Nicotine induced mutations were distributed across 1,585 genes, of which 49% were associated with cancer. MUC family genes were among the top mutated genes. Analysis of 591 lung carcinoma tumor exomes from The Cancer Genome Atlas (TCGA) revealed that 20% of non-small-cell lung cancer tumors in smokers have mutations in at least one of the MUC4, MUC6 or MUC12 genes in contrast to only 6% in non-smokers. These results indicate that nicotine induces genomic variations, promotes instability potentially mediated by oxidative stress, implicating nicotine in carcinogenesis, and establishes MUC genes as potential targets. PMID:24947164
Vectors as Epidemiological Sentinels: Patterns of Within-Tick Borrelia burgdorferi Diversity
Walter, Katharine S.; Carpi, Giovanna; Evans, Benjamin R.; Caccone, Adalgisa; Diuk-Wasser, Maria A.
2016-01-01
Hosts including humans, other vertebrates, and arthropods, are frequently infected with heterogeneous populations of pathogens. Within-host pathogen diversity has major implications for human health, epidemiology, and pathogen evolution. However, pathogen diversity within-hosts is difficult to characterize and little is known about the levels and sources of within-host diversity maintained in natural populations of disease vectors. Here, we examine genomic variation of the Lyme disease bacteria, Borrelia burgdorferi (Bb), in 98 individual field-collected tick vectors as a model for study of within-host processes. Deep population sequencing reveals extensive and previously undocumented levels of Bb variation: the majority (~70%) of ticks harbor mixed strain infections, which we define as levels Bb diversity pre-existing in a diverse inoculum. Within-tick diversity is thus a sample of the variation present within vertebrate hosts. Within individual ticks, we detect signatures of positive selection. Genes most commonly under positive selection across ticks include those involved in dissemination in vertebrate hosts and evasion of the vertebrate immune complement. By focusing on tick-borne Bb, we show that vectors can serve as epidemiological and evolutionary sentinels: within-vector pathogen diversity can be a useful and unbiased way to survey circulating pathogen diversity and identify evolutionary processes occurring in natural transmission cycles. PMID:27414806
Lemieux, Jacob E; Kyes, Sue A; Otto, Thomas D; Feller, Avi I; Eastman, Richard T; Pinches, Robert A; Berriman, Matthew; Su, Xin-zhuan; Newbold, Chris I
2013-01-01
Spatial relationships within the eukaryotic nucleus are essential for proper nuclear function. In Plasmodium falciparum, the repositioning of chromosomes has been implicated in the regulation of the expression of genes responsible for antigenic variation, and the formation of a single, peri-nuclear nucleolus results in the clustering of rDNA. Nevertheless, the precise spatial relationships between chromosomes remain poorly understood, because, until recently, techniques with sufficient resolution have been lacking. Here we have used chromosome conformation capture and second-generation sequencing to study changes in chromosome folding and spatial positioning that occur during switches in var gene expression. We have generated maps of chromosomal spatial affinities within the P. falciparum nucleus at 25 Kb resolution, revealing a structured nucleolus, an absence of chromosome territories, and confirming previously identified clustering of heterochromatin foci. We show that switches in var gene expression do not appear to involve interaction with a distant enhancer, but do result in local changes at the active locus. These maps reveal the folding properties of malaria chromosomes, validate known physical associations, and characterize the global landscape of spatial interactions. Collectively, our data provide critical information for a better understanding of gene expression regulation and antigenic variation in malaria parasites. PMID:23980881
DOE Office of Scientific and Technical Information (OSTI.GOV)
Stevens, F. J.; Pokkuluri, P. R.; Schiffer, M.
2000-12-19
The antibody light chain variable domain (V{sub L}){sup 1} and myelin protein zero (MPZ) are representatives of the functionally diverse immunoglobulin superfamily. The V{sub L} is a subunit of the antigen-binding component of antibodies, while MPZ is the major membrane-linked constituent of the myelin sheaths that coat peripheral nerves. Despite limited amino acid sequence homology, the conformations of the core structures of the two proteins are largely superimposable. Amino acid variations in V{sub L} account for various conformational disease outcomes, including amyloidosis. However, the specific amino acid changes in V{sub L} that are responsible for disease have been obscured bymore » multiple concurrent primary structure alterations. Recently, certain demyelination disorders have been linked to point mutations and single amino acid polymorphisms in MPZ. We demonstrate here that some pathogenic variations in MPZ correspond to changes suspected of determining amyloidosis in V{sub L}. This unanticipated observation suggests that studies of the biophysical origin of conformational disease in one member of a superfamily of homologous proteins may have implications throughout the superfamily. In some cases, findings may account for overt disease; in other cases, due to the natural repertoire of inherited polymorphisms, variations in a representative protein may predict subclinical impairment of homologous proteins.« less
Spiekman, Stephan N. F.; Werneburg, Ingmar
2017-01-01
Development in marsupials is specialized towards an extremely short gestation and highly altricial newborns. As a result, marsupial neonates display morphological adaptations at birth related to functional constraints. However, little is known about the variability of marsupial skull development and its relation to morphological diversity. We studied bony skull development in five marsupial species. The relative timing of the onset of ossification was compared to literature data and the ossification sequence of the marsupial ancestor was reconstructed using squared-change parsimony. The high range of variation in the onset of ossification meant that no patterns could be observed that differentiate species. This finding challenges traditional studies concentrating on the onset of ossification as a marker for phylogeny or as a functional proxy. Our study presents observations on the developmental timing of cranial bone-to-bone contacts and their evolutionary implications. Although certain bone contacts display high levels of variation, connections of early and late development are quite conserved and informative. Bones that surround the oral cavity are generally the first to connect and the bones of the occipital region are among the last. We conclude that bone contact is preferable over onset of ossification for studying cranial bone development. PMID:28233826
Africa: continent of genome contrasts with implications for biomedical research and health.
Ramsay, Michèle
2012-08-31
The genomic architecture of African populations is poorly understood and there is considerable variation between ethno-linguistic groups. Genome-wide approaches have been extensively applied to search for genetic associations to complex traits in Europeans, but rarely in Africans. This is largely attributed to lower levels of funding, poor infrastructure and public health systems, and to the small pool of trained scientists. High levels of genetic variation and underlying population structure in Africans present significant challenges, but lower levels of linkage disequilibrium provide an opportunity for more effective localisation of causal variants. High throughput technologies, including dense genotyping arrays, genome sequencing and epigenome studies, together with plummeting costs, are making research more affordable, even for African scientists. Understanding the interactions between genome structure and environmental influences is essential to interpreting their contributions to the increase in infectious diseases and non-communicable diseases, exacerbated by adverse environments and lifestyle choices. The unique genome dynamics in African populations have an important role to play in understanding human health and susceptibility to disease. Copyright © 2012. Published by Elsevier B.V.
Oliveros, R; Cutillas, C; De Rojas, M; Arias, P
2000-12-01
Adult worms of Trichuris ovis and T. globulosa were collected from Ovis aries (sheep) and Capra hircus (goats). T. suis was isolated from Sus scrofa domestica (swine) and T. leporis was isolated from Lepus europaeus (rabbits) in Spain. Genomic DNA was isolated and a ribosomal internal transcribed spacer (ITS2) was amplified and sequenced using polymerase-chain-reaction (PCR) techniques. The ITS2 of T. ovis and T. globulosa was 407 nucleotides in length and had a GC content of about 62%. Furthermore, the ITS2 of T. suis and T. leporis was 534 and 418 nucleotides in length and had a GC content of about 64.8% and 62.4%, respectively. There was evidence of slight variation in the sequence within individuals of all species analyzed, indicating intraindividual variation in the sequence of different copies of the ribosomal DNA. Furthermore, low-level intraspecific variation was detected. Sequence analyses of ITS2 products of T. ovis and T. globulosa demonstrated no sequence difference between them. Nevertheless, differences were detected between the ITS2 sequences of T. suis, T. leporis, and T. ovis, indicating that Trichuris species can reliably be differentiated by their ITS2 sequences and PCR-linked restriction-fragment-length polymorphism (RFLP).
High levels of variation in Salix lignocellulose genes revealed using poplar genomic resources
2013-01-01
Background Little is known about the levels of variation in lignin or other wood related genes in Salix, a genus that is being increasingly used for biomass and biofuel production. The lignin biosynthesis pathway is well characterized in a number of species, including the model tree Populus. We aimed to transfer the genomic resources already available in Populus to its sister genus Salix to assess levels of variation within genes involved in wood formation. Results Amplification trials for 27 gene regions were undertaken in 40 Salix taxa. Twelve of these regions were sequenced. Alignment searches of the resulting sequences against reference databases, combined with phylogenetic analyses, showed the close similarity of these Salix sequences to Populus, confirming homology of the primer regions and indicating a high level of conservation within the wood formation genes. However, all sequences were found to vary considerably among Salix species, mainly as SNPs with a smaller number of insertions-deletions. Between 25 and 176 SNPs per kbp per gene region (in predicted exons) were discovered within Salix. Conclusions The variation found is sizeable but not unexpected as it is based on interspecific and not intraspecific comparison; it is comparable to interspecific variation in Populus. The characterisation of genetic variation is a key process in pre-breeding and for the conservation and exploitation of genetic resources in Salix. This study characterises the variation in several lignocellulose gene markers for such purposes. PMID:23924375
Nolan, Danielle; Carlson, Martha
2016-06-01
Genetic heterogeneity in neurologic disorders has been an obstacle to phenotype-based diagnostic testing. The authors hypothesized that information compiled via whole exome sequencing will improve clinical diagnosis and management of pediatric neurology patients. The authors performed a retrospective chart review of patients evaluated in the University of Michigan Pediatric Neurology clinic between 6/2011 and 6/2015. The authors recorded previous diagnostic testing, indications for whole exome sequencing, and whole exome sequencing results. Whole exome sequencing was recommended for 135 patients and obtained in 53 patients. Insurance barriers often precluded whole exome sequencing. The most common indication for whole exome sequencing was neurodevelopmental disorders. Whole exome sequencing improved the presumptive diagnostic rate in the patient cohort from 25% to 48%. Clinical implications included family planning, medication selection, and systemic investigation. Compared to current second tier testing, whole exome sequencing can result in lower long-term charges and more timely diagnosis. Overcoming barriers related to whole exome sequencing insurance authorization could allow for more efficient and fruitful diagnostic neurological evaluations. © The Author(s) 2016.
DNA Metabarcoding of Amazonian Ichthyoplankton Swarms
Maggia, M. E.; Vigouroux, Y.; Renno, J. F.; Duponchelle, F.; Desmarais, E.; Nunez, J.; García-Dávila, C.; Carvajal-Vallejos, F. M.; Paradis, E.; Martin, J. F.; Mariac, C.
2017-01-01
Tropical rainforests harbor extraordinary biodiversity. The Amazon basin is thought to hold 30% of all river fish species in the world. Information about the ecology, reproduction, and recruitment of most species is still lacking, thus hampering fisheries management and successful conservation strategies. One of the key understudied issues in the study of population dynamics is recruitment. Fish larval ecology in tropical biomes is still in its infancy owing to identification difficulties. Molecular techniques are very promising tools for the identification of larvae at the species level. However, one of their limits is obtaining individual sequences with large samples of larvae. To facilitate this task, we developed a new method based on the massive parallel sequencing capability of next generation sequencing (NGS) coupled with hybridization capture. We focused on the mitochondrial marker cytochrome oxidase I (COI). The results obtained using the new method were compared with individual larval sequencing. We validated the ability of the method to identify Amazonian catfish larvae at the species level and to estimate the relative abundance of species in batches of larvae. Finally, we applied the method and provided evidence for strong temporal variation in reproductive activity of catfish species in the Ucayalí River in the Peruvian Amazon. This new time and cost effective method enables the acquisition of large datasets, paving the way for a finer understanding of reproductive dynamics and recruitment patterns of tropical fish species, with major implications for fisheries management and conservation. PMID:28095487
Qin, Tian; Zhou, Haijian; Ren, Hongyu; Guan, Hong; Li, Machao; Zhu, Bingqing; Shao, Zhujun
2014-04-01
Legionella pneumophila serogroup 1 causes Legionnaires' disease. Water systems contaminated with Legionella are the implicated sources of Legionnaires' disease. This study analyzed L. pneumophila serogroup 1 strains in China using sequence-based typing. Strains were isolated from cooling towers (n = 96), hot springs (n = 42), and potable water systems (n = 26). Isolates from cooling towers, hot springs, and potable water systems were divided into 25 sequence types (STs; index of discrimination [IOD], 0.711), 19 STs (IOD, 0.934), and 3 STs (IOD, 0.151), respectively. The genetic variation among the potable water isolates was lower than that among cooling tower and hot spring isolates. ST1 was the predominant type, accounting for 49.4% of analyzed strains (n = 81), followed by ST154. With the exception of two strains, all potable water isolates (92.3%) belonged to ST1. In contrast, 53.1% (51/96) and only 14.3% (6/42) of cooling tower and hot spring, respectively, isolates belonged to ST1. There were differences in the distributions of clone groups among the water sources. The comparisons among L. pneumophila strains isolated in China, Japan, and South Korea revealed that similar clones (ST1 complex and ST154 complex) exist in these countries. In conclusion, in China, STs had several unique allelic profiles, and ST1 was the most prevalent sequence type of environmental L. pneumophila serogroup 1 isolates, similar to its prevalence in Japan and South Korea.
Investigation of the role of TCF4 rare sequence variants in schizophrenia.
Basmanav, F Buket; Forstner, Andreas J; Fier, Heide; Herms, Stefan; Meier, Sandra; Degenhardt, Franziska; Hoffmann, Per; Barth, Sandra; Fricker, Nadine; Strohmaier, Jana; Witt, Stephanie H; Ludwig, Michael; Schmael, Christine; Moebus, Susanne; Maier, Wolfgang; Mössner, Rainald; Rujescu, Dan; Rietschel, Marcella; Lange, Christoph; Nöthen, Markus M; Cichon, Sven
2015-07-01
Transcription factor 4 (TCF4) is one of the most robust of all reported schizophrenia risk loci and is supported by several genetic and functional lines of evidence. While numerous studies have implicated common genetic variation at TCF4 in schizophrenia risk, the role of rare, small-sized variants at this locus-such as single nucleotide variants and short indels which are below the resolution of chip-based arrays requires further exploration. The aim of the present study was to investigate the association between rare TCF4 sequence variants and schizophrenia. Exon-targeted resequencing was performed in 190 German schizophrenia patients. Six rare variants at the coding exons and flanking sequences of the TCF4 gene were identified, including two missense variants and one splice site variant. These six variants were then pooled with nine additional rare variants identified in 379 European participants of the 1000 Genomes Project, and all 15 variants were genotyped in an independent German sample (n = 1,808 patients; n = 2,261 controls). These data were then analyzed using six statistical methods developed for the association analysis of rare variants. No significant association (P < 0.05) was found. However, the results from our association and power analyses suggest that further research into the possible involvement of rare TCF4 sequence variants in schizophrenia risk is warranted by the assessment of larger cohorts with higher statistical power to identify rare variant associations. © 2015 Wiley Periodicals, Inc.
Shafiei, Reza; Sarkari, Bahador; Sadjjadi, Seyed Mahmuod; Mowlavi, Gholam Reza; Moshfe, Abdolali
2014-01-01
The current study aimed to find out the morphometric and genotypic divergences of the flukes isolated from different hosts in a newly emerging focus of human fascioliasis in Iran. Adult Fasciola spp. were collected from 34 cattle, 13 sheep, and 11 goats from Kohgiluyeh and Boyer-Ahmad province, southwest of Iran. Genomic DNA was extracted from the flukes and PCR-RFLP was used to characterize the isolates. The ITS1, ITS2, and mitochondrial genes (mtDNA) of NDI and COI from individual liver flukes were amplified and the amplicons were sequenced. Genetic variation within and between the species was evaluated by comparing the sequences. Moreover, morphometric characteristics of flukes were measured through a computer image analysis system. Based on RFLP profile, from the total of 58 isolates, 41 isolates (from cattle, sheep, and goat) were identified as Fasciola hepatica, while 17 isolates from cattle were identified as Fasciola gigantica. Comparison of the ITS1 and ITS2 sequences showed six and seven single-base substitutions, resulting in segregation of the specimens into two different genotypes. The sequences of COI markers showed seven DNA polymorphic sites for F. hepatica and 35 DNA polymorphic sites for F. gigantica. Morphological diversity of the two species was observed in linear, ratios, and areas measurements. The findings have implications for studying the population genetics, epidemiology, and control of the disease. PMID:25018891
Castro-Prieto, Aines; Wachter, Bettina; Melzheimer, Joerg; Thalwitzer, Susanne; Sommer, Simone
2011-01-01
The genes of the major histocompatibility complex (MHC) are a key component of the mammalian immune system and have become important molecular markers for fitness-related genetic variation in wildlife populations. Currently, no information about the MHC sequence variation and constitution in African leopards exists. In this study, we isolated and characterized genetic variation at the adaptively most important region of MHC class I and MHC class II-DRB genes in 25 free-ranging African leopards from Namibia and investigated the mechanisms that generate and maintain MHC polymorphism in the species. Using single-stranded conformation polymorphism analysis and direct sequencing, we detected 6 MHC class I and 6 MHC class II-DRB sequences, which likely correspond to at least 3 MHC class I and 3 MHC class II-DRB loci. Amino acid sequence variation in both MHC classes was higher or similar in comparison to other reported felids. We found signatures of positive selection shaping the diversity of MHC class I and MHC class II-DRB loci during the evolutionary history of the species. A comparison of MHC class I and MHC class II-DRB sequences of the leopard to those of other felids revealed a trans-species mode of evolution. In addition, the evolutionary relationships of MHC class II-DRB sequences between African and Asian leopard subspecies are discussed.
NASA Astrophysics Data System (ADS)
Damnati, B.
1993-05-01
Sedimentological and geochemical analyses have been carried out on lacustrine deposits of East Africa, at Lake Magadi (2°S, 36°E, Kenya) and at Green Crater Lake (0°S, 36°E, Kenya), to determine the parameters controlling climatic and environmental dynamics during late Pleistocene and Holocene. These sedimentary sequences were collected with a stationary piston corer. At Lake Magadi (Fig. 1), sedimentary and geochemical control show three phases of lake level variation which corresponds to climatic change occurring during the last 40 thousand years. These phases were defined by three lithostratigraphic units. Laminated deposits of Lake Magadi were formed during a wet period. Analysis of these laminae define two microfacies: a dark lamina, characterised by lacustrine organic matter and a light lamina enriched in detritus, carbonates (CaCO 3) and magadiite (NaSi 7O 13(OH) 3, 3H 2O). The formation and preservation of each couplet was favoured by climatic contrast, lake stratification and various origin of the sediments (autochthon and allochthon) in the drainage basin. Therefore a relative chronology can be derived from laminae counting and the duration of deposition of each couplet. Spectral analysis applied on variation of the laminae thickness, shows the existence of three main periods, 4-7 years, 8-14 years and 18-30 years, respectively (Fig. 2). These cyclicites of the lacustrine environment precise former determinations established on more recent lacustrine sequences from East Africa. They are related to the global climatic cycle (quasi-biannual oscillations, El Nino Southern Oscillations and the sun spot cycles). At Green Crater Lake, the study of the sedimentary sequence was completed by physico-chemical analysis of the waters and interface sediments which demonstrate the carbonate, sodium, bicarbonate composition and the thermal and chemical stratification of the modern lake. The sedimentary sequence is characterized by volcanic deposits overlain by physico-chemical analysis of the lake waters and interface sediments which demonstrate the carbonate, sodium, bicarbonate composition and the thermal and chemical stratification of the modern lake. The sedimentary sequence is characterized by volcanic deposits overlain by silt and clays deposited before 7400 years B.P., followed by loweing of the lake level at 3000 years B. P. Results from lake Magadi document the occurrence of a wet period starting at about 12,000 years B. P. The methodology applied on modern Green Crater lake provides base of interpretative models for other Holocene sequence lacustrine systems of intertropical zones.
Specific material recognition by small peptides mediated by the interfacial solvent structure.
Schneider, Julian; Ciacchi, Lucio Colombi
2012-02-01
We present evidence that specific material recognition by small peptides is governed by local solvent density variations at solid/liquid interfaces, sensed by the side-chain residues with atomic-scale precision. In particular, we unveil the origin of the selectivity of the binding motif RKLPDA for Ti over Si using a combination of metadynamics and steered molecular dynamics simulations, obtaining adsorption free energies and adhesion forces in quantitative agreement with corresponding experiments. For an accurate description, we employ realistic models of the natively oxidized surfaces which go beyond the commonly used perfect crystal surfaces. These results have profound implications for nanotechnology and materials science applications, offering a previously missing structure-function relationship for the rational design of materials-selective peptide sequences. © 2011 American Chemical Society
Conservation genomics of threatened animal species.
Steiner, Cynthia C; Putnam, Andrea S; Hoeck, Paquita E A; Ryder, Oliver A
2013-01-01
The genomics era has opened up exciting possibilities in the field of conservation biology by enabling genomic analyses of threatened species that previously were limited to model organisms. Next-generation sequencing (NGS) and the collection of genome-wide data allow for more robust studies of the demographic history of populations and adaptive variation associated with fitness and local adaptation. Genomic analyses can also advance management efforts for threatened wild and captive populations by identifying loci contributing to inbreeding depression and disease susceptibility, and predicting fitness consequences of introgression. However, the development of genomic tools in wild species still carries multiple challenges, particularly those associated with computational and sampling constraints. This review provides an overview of the most significant applications of NGS and the implications and limitations of genomic studies in conservation.
The evolution of transcriptional regulation in eukaryotes
NASA Technical Reports Server (NTRS)
Wray, Gregory A.; Hahn, Matthew W.; Abouheif, Ehab; Balhoff, James P.; Pizer, Margaret; Rockman, Matthew V.; Romano, Laura A.
2003-01-01
Gene expression is central to the genotype-phenotype relationship in all organisms, and it is an important component of the genetic basis for evolutionary change in diverse aspects of phenotype. However, the evolution of transcriptional regulation remains understudied and poorly understood. Here we review the evolutionary dynamics of promoter, or cis-regulatory, sequences and the evolutionary mechanisms that shape them. Existing evidence indicates that populations harbor extensive genetic variation in promoter sequences, that a substantial fraction of this variation has consequences for both biochemical and organismal phenotype, and that some of this functional variation is sorted by selection. As with protein-coding sequences, rates and patterns of promoter sequence evolution differ considerably among loci and among clades for reasons that are not well understood. Studying the evolution of transcriptional regulation poses empirical and conceptual challenges beyond those typically encountered in analyses of coding sequence evolution: promoter organization is much less regular than that of coding sequences, and sequences required for the transcription of each locus reside at multiple other loci in the genome. Because of the strong context-dependence of transcriptional regulation, sequence inspection alone provides limited information about promoter function. Understanding the functional consequences of sequence differences among promoters generally requires biochemical and in vivo functional assays. Despite these challenges, important insights have already been gained into the evolution of transcriptional regulation, and the pace of discovery is accelerating.
Parallel gene analysis with allele-specific padlock probes and tag microarrays
Banér, Johan; Isaksson, Anders; Waldenström, Erik; Jarvius, Jonas; Landegren, Ulf; Nilsson, Mats
2003-01-01
Parallel, highly specific analysis methods are required to take advantage of the extensive information about DNA sequence variation and of expressed sequences. We present a scalable laboratory technique suitable to analyze numerous target sequences in multiplexed assays. Sets of padlock probes were applied to analyze single nucleotide variation directly in total genomic DNA or cDNA for parallel genotyping or gene expression analysis. All reacted probes were then co-amplified and identified by hybridization to a standard tag oligonucleotide array. The technique was illustrated by analyzing normal and pathogenic variation within the Wilson disease-related ATP7B gene, both at the level of DNA and RNA, using allele-specific padlock probes. PMID:12930977
2014-01-01
Background Neisseria meningitidis expresses type four pili (Tfp) which are important for colonisation and virulence. Tfp have been considered as one of the most variable structures on the bacterial surface due to high frequency gene conversion, resulting in amino acid sequence variation of the major pilin subunit (PilE). Meningococci express either a class I or a class II pilE gene and recent work has indicated that class II pilins do not undergo antigenic variation, as class II pilE genes encode conserved pilin subunits. The purpose of this work was to use whole genome sequences to further investigate the frequency and variability of the class II pilE genes in meningococcal isolate collections. Results We analysed over 600 publically available whole genome sequences of N. meningitidis isolates to determine the sequence and genomic organization of pilE. We confirmed that meningococcal strains belonging to a limited number of clonal complexes (ccs, namely cc1, cc5, cc8, cc11 and cc174) harbour a class II pilE gene which is conserved in terms of sequence and chromosomal context. We also identified pilS cassettes in all isolates with class II pilE, however, our analysis indicates that these do not serve as donor sequences for pilE/pilS recombination. Furthermore, our work reveals that the class II pilE locus lacks the DNA sequence motifs that enable (G4) or enhance (Sma/Cla repeat) pilin antigenic variation. Finally, through analysis of pilin genes in commensal Neisseria species we found that meningococcal class II pilE genes are closely related to pilE from Neisseria lactamica and Neisseria polysaccharea, suggesting horizontal transfer among these species. Conclusions Class II pilins can be defined by their amino acid sequence and genomic context and are present in meningococcal isolates which have persisted and spread globally. The absence of G4 and Sma/Cla sequences adjacent to the class II pilE genes is consistent with the lack of pilin subunit variation in these isolates, although horizontal transfer may generate class II pilin diversity. This study supports the suggestion that high frequency antigenic variation of pilin is not universal in pathogenic Neisseria. PMID:24690385
Somatic Genetic Variation in Solid Pseudopapillary Tumor of the Pancreas by Whole Exome Sequencing
Guo, Meng; Luo, Guopei; Jin, Kaizhou; Long, Jiang; Cheng, He; Lu, Yu; Wang, Zhengshi; Yang, Chao; Xu, Jin; Ni, Quanxing; Yu, Xianjun; Liu, Chen
2017-01-01
Solid pseudopapillary tumor of the pancreas (SPT) is a rare pancreatic disease with a unique clinical manifestation. Although CTNNB1 gene mutations had been universally reported, genetic variation profiles of SPT are largely unidentified. We conducted whole exome sequencing in nine SPT patients to probe the SPT-specific insertions and deletions (indels) and single nucleotide polymorphisms (SNPs). In total, 54 SNPs and 41 indels of prominent variations were demonstrated through parallel exome sequencing. We detected that CTNNB1 mutations presented throughout all patients studied (100%), and a higher count of SNPs was particularly detected in patients with older age, larger tumor, and metastatic disease. By aggregating 95 detected variation events and viewing the interconnections among each of the genes with variations, CTNNB1 was identified as the core portion in the network, which might collaborate with other events such as variations of USP9X, EP400, HTT, MED12, and PKD1 to regulate tumorigenesis. Pathway analysis showed that the events involved in other cancers had the potential to influence the progression of the SNPs count. Our study revealed an insight into the variation of the gene encoding region underlying solid-pseudopapillary neoplasm tumorigenesis. The detection of these variations might partly reflect the potential molecular mechanism. PMID:28054945
Thermal and acid tolerant beta-xylosidases, genes encoding, related organisms, and methods
Thompson, David N [Idaho Falls, ID; Thompson, Vicki S [Idaho Falls, ID; Schaller, Kastli D [Ammon, ID; Apel, William A [Jackson, WY; Lacey, Jeffrey A [Idaho Falls, ID; Reed, David W [Idaho Falls, ID
2011-04-12
Isolated and/or purified polypeptides and nucleic acid sequences encoding polypeptides from Alicyclobacillus acidocaldarius and variations thereof are provided. Further provided are methods of at least partially degrading xylotriose and/or xylobiose using isolated and/or purified polypeptides and nucleic acid sequences encoding polypeptides from Alicyclobacillus acidocaldarius and variations thereof.
USDA-ARS?s Scientific Manuscript database
Little is known about genetic variation of Lymantria dispar multiple nucleopolyhedrovirus (LdMNPV; Baculoviridae: Alphabaculovirus) at the nucleotide sequence level. To obtain a more comprehensive view of genetic diversity among isolates of LdMNPV, partial sequences of the lef-8 gene were generated...
DOE Office of Scientific and Technical Information (OSTI.GOV)
Gordon, Sean
2013-03-01
Sean Gordon of the USDA on Natural variation in Brachypodium disctachyon: Deep Sequencing of Highly Diverse Natural Accessions at the 8th Annual Genomics of Energy Environment Meeting on March 27, 2013 in Walnut Creek, CA.
Sequence variation of the feline immunodeficiency virus genome and its clinical relevance.
Stickney, A L; Dunowska, M; Cave, N J
2013-06-08
The ongoing evolution of feline immunodeficiency virus (FIV) has resulted in the existence of a diverse continuum of viruses. FIV isolates differ with regards to their mutation and replication rates, plasma viral loads, cell tropism and the ability to induce apoptosis. Clinical disease in FIV-infected cats is also inconsistent. Genomic sequence variation of FIV is likely to be responsible for some of the variation in viral behaviour. The specific genetic sequences that influence these key viral properties remain to be determined. With knowledge of the specific key determinants of pathogenicity, there is the potential for veterinarians in the future to apply this information for prognostic purposes. Genomic sequence variation of FIV also presents an obstacle to effective vaccine development. Most challenge studies demonstrate acceptable efficacy of a dual-subtype FIV vaccine (Fel-O-Vax FIV) against FIV infection under experimental settings; however, vaccine efficacy in the field still remains to be proven. It is important that we discover the key determinants of immunity induced by this vaccine; such data would compliment vaccine field efficacy studies and provide the basis to make informed recommendations on its use.
Nouvel, Laurent X; Vultos, Tiago Dos; Kassa-Kelembho, Eric; Rauzier, Jean; Gicquel, Brigitte
2007-01-01
Background Previous studies have suggested that variations in DNA repair genes of W-Beijing strains may have led to transient mutator phenotypes which in turn may have contributed to host adaptation of this strain family. Single nucleotide polymorphism (SNP) in the DNA repair gene mutT1 was identified in MDR-prone strains from the Central African Republic. A Mycobacteriumtuberculosis H37Rv mutant inactivated in two DNA repair genes, namely ada/alkA and ogt, was shown to display a hypermutator phenotype. We then looked for polymorphisms in these genes in Central African Republic strains (CAR). Results In this study, 55 MDR and 194 non-MDR strains were analyzed. Variations in DNA repair genes ada/alkA and ogt were identified. Among them, by comparison to M. tuberculosis published sequences, we found a non-sense variation in ada/alkA gene which was also observed in M. bovis AF2122 strain. SNPs that are present in the adjacent regions to the amber variation are different in M. bovis and in M. tuberculosis strain. Conclusion An Amber codon was found in the ada/alkA locus of clustered M. tuberculosis isolates and in M. bovis strain AF2122. This is likely due to convergent evolution because SNP differences between strains are incompatible with horizontal transfer of an entire gene. This suggests that such a variation may confer a selective advantage and be implicated in hypermutator phenotype expression, which in turn contributes to adaptation to environmental changes. PMID:17506895
Selection of a DNA barcode for Nectriaceae from fungal whole-genomes.
Zeng, Zhaoqing; Zhao, Peng; Luo, Jing; Zhuang, Wenying; Yu, Zhihe
2012-01-01
A DNA barcode is a short segment of sequence that is able to distinguish species. A barcode must ideally contain enough variation to distinguish every individual species and be easily obtained. Fungi of Nectriaceae are economically important and show high species diversity. To establish a standard DNA barcode for this group of fungi, the genomes of Neurospora crassa and 30 other filamentous fungi were compared. The expect value was treated as a criterion to recognize homologous sequences. Four candidate markers, Hsp90, AAC, CDC48, and EF3, were tested for their feasibility as barcodes in the identification of 34 well-established species belonging to 13 genera of Nectriaceae. Two hundred and fifteen sequences were analyzed. Intra- and inter-specific variations and the success rate of PCR amplification and sequencing were considered as important criteria for estimation of the candidate markers. Ultimately, the partial EF3 gene met the requirements for a good DNA barcode: No overlap was found between the intra- and inter-specific pairwise distances. The smallest inter-specific distance of EF3 gene was 3.19%, while the largest intra-specific distance was 1.79%. In addition, there was a high success rate in PCR and sequencing for this gene (96.3%). CDC48 showed sufficiently high sequence variation among species, but the PCR and sequencing success rate was 84% using a single pair of primers. Although the Hsp90 and AAC genes had higher PCR and sequencing success rates (96.3% and 97.5%, respectively), overlapping occurred between the intra- and inter-specific variations, which could lead to misidentification. Therefore, we propose the EF3 gene as a possible DNA barcode for the nectriaceous fungi.
Boussaha, Mekki; Michot, Pauline; Letaief, Rabia; Hozé, Chris; Fritz, Sébastien; Grohs, Cécile; Esquerré, Diane; Duchesne, Amandine; Philippe, Romain; Blanquet, Véronique; Phocas, Florence; Floriot, Sandrine; Rocha, Dominique; Klopp, Christophe; Capitan, Aurélien; Boichard, Didier
2016-11-15
In recent years, several bovine genome sequencing projects were carried out with the aim of developing genomic tools to improve dairy and beef production efficiency and sustainability. In this study, we describe the first French cattle genome variation dataset obtained by sequencing 274 whole genomes representing several major dairy and beef breeds. This dataset contains over 28 million single nucleotide polymorphisms (SNPs) and small insertions and deletions. Comparisons between sequencing results and SNP array genotypes revealed a very high genotype concordance rate, which indicates the good quality of our data. To our knowledge, this is the first large-scale catalog of small genomic variations in French dairy and beef cattle. This resource will contribute to the study of gene functions and population structure and also help to improve traits through genotype-guided selection.
Richardson, David S; Westerdahl, Helena
2003-12-01
The Great reed warbler (GRW) and the Seychelles warbler (SW) are congeners with markedly different demographic histories. The GRW is a normal outbred bird species while the SW population remains isolated and inbred after undergoing a severe population bottleneck. We examined variation at Major Histocompatibility Complex (MHC) class I exon 3 using restriction fragment length polymorphism, denaturing gradient gel electrophoresis and DNA sequencing. Although genetic variation was higher in the GRW, considerable variation has been maintained in the SW. The ten exon 3 sequences found in the SW were as diverged from each other as were a random sub-sample of the 67 sequences from the GRW. There was evidence for balancing selection in both species, and the phylogenetic analysis showing that the exon 3 sequences did not separate according to species, was consistent with transspecies evolution of the MHC.
Zhao, Min; Wang, Qingguo; Wang, Quan; Jia, Peilin; Zhao, Zhongming
2013-01-01
Copy number variation (CNV) is a prevalent form of critical genetic variation that leads to an abnormal number of copies of large genomic regions in a cell. Microarray-based comparative genome hybridization (arrayCGH) or genotyping arrays have been standard technologies to detect large regions subject to copy number changes in genomes until most recently high-resolution sequence data can be analyzed by next-generation sequencing (NGS). During the last several years, NGS-based analysis has been widely applied to identify CNVs in both healthy and diseased individuals. Correspondingly, the strong demand for NGS-based CNV analyses has fuelled development of numerous computational methods and tools for CNV detection. In this article, we review the recent advances in computational methods pertaining to CNV detection using whole genome and whole exome sequencing data. Additionally, we discuss their strengths and weaknesses and suggest directions for future development.
2013-01-01
Copy number variation (CNV) is a prevalent form of critical genetic variation that leads to an abnormal number of copies of large genomic regions in a cell. Microarray-based comparative genome hybridization (arrayCGH) or genotyping arrays have been standard technologies to detect large regions subject to copy number changes in genomes until most recently high-resolution sequence data can be analyzed by next-generation sequencing (NGS). During the last several years, NGS-based analysis has been widely applied to identify CNVs in both healthy and diseased individuals. Correspondingly, the strong demand for NGS-based CNV analyses has fuelled development of numerous computational methods and tools for CNV detection. In this article, we review the recent advances in computational methods pertaining to CNV detection using whole genome and whole exome sequencing data. Additionally, we discuss their strengths and weaknesses and suggest directions for future development. PMID:24564169
Identification of the sequence variations of 15 autosomal STR loci in a Chinese population.
Chen, Wenjing; Cheng, Jianding; Ou, Xueling; Chen, Yong; Tong, Dayue; Sun, Hongyu
2014-01-01
DNA sequence variation including base(s) changes and insertion or deletion in the primer binding region may cause a null allele and, if this changes the length of the amplified fragment out of the allelic ladder, off-ladder (OL) alleles may be detected. In order to provide accurate and reliable DNA evidence for forensic DNA analysis, it is essential to clarify sequence variations in prevalently used STR loci. Suspected null alleles and OL alleles of PlowerPlex16® System from 21,934 unrelated Chinese individuals were verified by alternative systems and sequenced. A total of 17 cases with null alleles were identified, including 12 kinds of point mutations in 16 cases and a 19-base deletion in one case. The total frequency of null alleles was 7.751 × 10(-4). Eight hundred and forty-four OL alleles classified as being of 97 different kinds were observed at 15 STR loci of the PowerPlex®16 system except vWA. All the frequencies of OL alleles were under 0.01. Null alleles should be confirmed by alternative primers and OL alleles should be named appropriately. Particular attention should be paid to sequence variation, since incorrect designation could lead to false conclusions.
The diploid genome sequence of an Asian individual
Wang, Jun; Wang, Wei; Li, Ruiqiang; Li, Yingrui; Tian, Geng; Goodman, Laurie; Fan, Wei; Zhang, Junqing; Li, Jun; Zhang, Juanbin; Guo, Yiran; Feng, Binxiao; Li, Heng; Lu, Yao; Fang, Xiaodong; Liang, Huiqing; Du, Zhenglin; Li, Dong; Zhao, Yiqing; Hu, Yujie; Yang, Zhenzhen; Zheng, Hancheng; Hellmann, Ines; Inouye, Michael; Pool, John; Yi, Xin; Zhao, Jing; Duan, Jinjie; Zhou, Yan; Qin, Junjie; Ma, Lijia; Li, Guoqing; Yang, Zhentao; Zhang, Guojie; Yang, Bin; Yu, Chang; Liang, Fang; Li, Wenjie; Li, Shaochuan; Li, Dawei; Ni, Peixiang; Ruan, Jue; Li, Qibin; Zhu, Hongmei; Liu, Dongyuan; Lu, Zhike; Li, Ning; Guo, Guangwu; Zhang, Jianguo; Ye, Jia; Fang, Lin; Hao, Qin; Chen, Quan; Liang, Yu; Su, Yeyang; san, A.; Ping, Cuo; Yang, Shuang; Chen, Fang; Li, Li; Zhou, Ke; Zheng, Hongkun; Ren, Yuanyuan; Yang, Ling; Gao, Yang; Yang, Guohua; Li, Zhuo; Feng, Xiaoli; Kristiansen, Karsten; Wong, Gane Ka-Shu; Nielsen, Rasmus; Durbin, Richard; Bolund, Lars; Zhang, Xiuqing; Li, Songgang; Yang, Huanming; Wang, Jian
2009-01-01
Here we present the first diploid genome sequence of an Asian individual. The genome was sequenced to 36-fold average coverage using massively parallel sequencing technology. We aligned the short reads onto the NCBI human reference genome to 99.97% coverage, and guided by the reference genome, we used uniquely mapped reads to assemble a high-quality consensus sequence for 92% of the Asian individual's genome. We identified approximately 3 million single-nucleotide polymorphisms (SNPs) inside this region, of which 13.6% were not in the dbSNP database. Genotyping analysis showed that SNP identification had high accuracy and consistency, indicating the high sequence quality of this assembly. We also carried out heterozygote phasing and haplotype prediction against HapMap CHB and JPT haplotypes (Chinese and Japanese, respectively), sequence comparison with the two available individual genomes (J. D. Watson and J. C. Venter), and structural variation identification. These variations were considered for their potential biological impact. Our sequence data and analyses demonstrate the potential usefulness of next-generation sequencing technologies for personal genomics. PMID:18987735
Ashfaq, Muhammad; Hebert, Paul D. N.; Mirza, M. Sajjad; Khan, Arif M.; Mansoor, Shahid; Shah, Ghulam S.; Zafar, Yusuf
2014-01-01
Background Although whiteflies (Bemisia tabaci complex) are an important pest of cotton in Pakistan, its taxonomic diversity is poorly understood. As DNA barcoding is an effective tool for resolving species complexes and analyzing species distributions, we used this approach to analyze genetic diversity in the B. tabaci complex and map the distribution of B. tabaci lineages in cotton growing areas of Pakistan. Methods/Principal Findings Sequence diversity in the DNA barcode region (mtCOI-5′) was examined in 593 whiteflies from Pakistan to determine the number of whitefly species and their distributions in the cotton-growing areas of Punjab and Sindh provinces. These new records were integrated with another 173 barcode sequences for B. tabaci, most from India, to better understand regional whitefly diversity. The Barcode Index Number (BIN) System assigned the 766 sequences to 15 BINs, including nine from Pakistan. Representative specimens of each Pakistan BIN were analyzed for mtCOI-3′ to allow their assignment to one of the putative species in the B. tabaci complex recognized on the basis of sequence variation in this gene region. This analysis revealed the presence of Asia II 1, Middle East-Asia Minor 1, Asia 1, Asia II 5, Asia II 7, and a new lineage “Pakistan”. The first two taxa were found in both Punjab and Sindh, but Asia 1 was only detected in Sindh, while Asia II 5, Asia II 7 and “Pakistan” were only present in Punjab. The haplotype networks showed that most haplotypes of Asia II 1, a species implicated in transmission of the cotton leaf curl virus, occurred in both India and Pakistan. Conclusions DNA barcodes successfully discriminated cryptic species in B. tabaci complex. The dominant haplotypes in the B. tabaci complex were shared by India and Pakistan. Asia II 1 was previously restricted to Punjab, but is now the dominant lineage in southern Sindh; its southward spread may have serious implications for cotton plantations in this region. PMID:25099936
Zhang, J R; Norris, S J
1998-08-01
The Lyme disease spirochete Borrelia burgdorferi possesses 15 silent vls cassettes and a vls expression site (vlsE) encoding a surface-exposed lipoprotein. Segments of the silent vls cassettes have been shown to recombine with the vlsE cassette region in the mammalian host, resulting in combinatorial antigenic variation. Despite promiscuous recombination within the vlsE cassette region, the 5' and 3' coding sequences of vlsE that flank the cassette region are not subject to sequence variation during these recombination events. The segments of the silent vls cassettes recombine in the vlsE cassette region through a unidirectional process such that the sequence and organization of the silent vls loci are not affected. As a result of recombination, the previously expressed segments are replaced by incoming segments and apparently degraded. These results provide evidence for a gene conversion mechanism in VlsE antigenic variation.
NASA Technical Reports Server (NTRS)
Rai, Man Mohan (Inventor); Madavan, Nateri K. (Inventor)
2007-01-01
A method and system for data modeling that incorporates the advantages of both traditional response surface methodology (RSM) and neural networks is disclosed. The invention partitions the parameters into a first set of s simple parameters, where observable data are expressible as low order polynomials, and c complex parameters that reflect more complicated variation of the observed data. Variation of the data with the simple parameters is modeled using polynomials; and variation of the data with the complex parameters at each vertex is analyzed using a neural network. Variations with the simple parameters and with the complex parameters are expressed using a first sequence of shape functions and a second sequence of neural network functions. The first and second sequences are multiplicatively combined to form a composite response surface, dependent upon the parameter values, that can be used to identify an accurate mode
Spuesens, Emiel B M; van de Kreeke, Nick; Estevão, Silvia; Hoogenboezem, Theo; Sluijter, Marcel; Hartwig, Nico G; van Rossum, Annemarie M C; Vink, Cornelis
2011-02-01
Mycoplasma pneumoniae is a human pathogen that causes a range of respiratory tract infections. The first step in infection is adherence of the bacteria to the respiratory epithelium. This step is mediated by a specialized organelle, which contains several proteins (cytadhesins) that have an important function in adherence. Two of these cytadhesins, P40 and P90, represent the proteolytic products from a single 130 kDa protein precursor, which is encoded by the MPN142 gene. Interestingly, MPN142 contains a repetitive DNA element, termed RepMP5, of which homologues are found at seven other loci within the M. pneumoniae genome. It has been hypothesized that these RepMP5 elements, which are similar but not identical in sequence, recombine with their counterpart within MPN142 and thereby provide a source of sequence variation for this gene. As this variation may give rise to amino acid changes within P40 and P90, the recombination between RepMP5 elements may constitute the basis of antigenic variation and, possibly, immune evasion by M. pneumoniae. To investigate the sequence variation of MPN142 in relation to inter-RepMP5 recombination, we determined the sequences of all RepMP5 elements in a collection of 25 strains. The results indicate that: (i) inter-RepMP5 recombination events have occurred in seven of the strains, and (ii) putative RepMP5 recombination events involving MPN142 have induced amino acid changes in a surface-exposed part of the P40 protein in two of the strains. We conclude that recombination between RepMP5 elements is a common phenomenon that may lead to sequence variation of MPN142-encoded proteins.
Dynamics of actin evolution in dinoflagellates.
Kim, Sunju; Bachvaroff, Tsvetan R; Handy, Sara M; Delwiche, Charles F
2011-04-01
Dinoflagellates have unique nuclei and intriguing genome characteristics with very high DNA content making complete genome sequencing difficult. In dinoflagellates, many genes are found in multicopy gene families, but the processes involved in the establishment and maintenance of these gene families are poorly understood. Understanding the dynamics of gene family evolution in dinoflagellates requires comparisons at different evolutionary scales. Studies of closely related species provide fine-scale information relative to species divergence, whereas comparisons of more distantly related species provides broad context. We selected the actin gene family as a highly expressed conserved gene previously studied in dinoflagellates. Of the 142 sequences determined in this study, 103 were from the two closely related species, Dinophysis acuminata and D. caudata, including full length and partial cDNA sequences as well as partial genomic amplicons. For these two Dinophysis species, at least three types of sequences could be identified. Most copies (79%) were relatively similar and in nucleotide trees, the sequences formed two bushy clades corresponding to the two species. In comparisons within species, only eight to ten nucleotide differences were found between these copies. The two remaining types formed clades containing sequences from both species. One type included the most similar sequences in between-species comparisons with as few as 12 nucleotide differences between species. The second type included the most divergent sequences in comparisons between and within species with up to 93 nucleotide differences between sequences. In all the sequences, most variation occurred in synonymous sites or the 5' UnTranslated Region (UTR), although there was still limited amino acid variation between most sequences. Several potential pseudogenes were found (approximately 10% of all sequences depending on species) with incomplete open reading frames due to frameshifts or early stop codons. Overall, variation in the actin gene family fits best with the "birth and death" model of evolution based on recent duplications, pseudogenes, and incomplete lineage sorting. Divergence between species was similar to variation within species, so that actin may be too conserved to be useful for phylogenetic estimation of closely related species.
Natural Allelic Variations in Highly Polyploidy Saccharum Complex
DOE Office of Scientific and Technical Information (OSTI.GOV)
Song, Jian; Yang, Xiping; Resende, Jr., Marcio F. R.
Sugarcane ( Saccharum spp.) is an important sugar and biofuel crop with high polyploid and complex genomes. The Saccharum complex, comprised of Saccharum genus and a few related genera, are important genetic resources for sugarcane breeding. A large amount of natural variation exists within the Saccharum complex. Though understanding their allelic variation has been challenging, it is critical to dissect allelic structure and to identify the alleles controlling important traits in sugarcane. To characterize natural variations in Saccharum complex, a target enrichment sequencing approach was used to assay 12 representative germplasm accessions. In total, 55,946 highly efficient probes were designedmore » based on the sorghum genome and sugarcane unigene set targeting a total of 6 Mb of the sugarcane genome. A pipeline specifically tailored for polyploid sequence variants and genotype calling was established. BWAmem and sorghum genome approved to be an acceptable aligner and reference for sugarcane target enrichment sequence analysis, respectively. Genetic variations including 1,166,066 non-redundant SNPs, 150,421 InDels, 919 gene copy number variations, and 1,257 gene presence/absence variations were detected. SNPs from three different callers (Samtools, Freebayes, and GATK) were compared and the validation rates were nearly 90%. Based on the SNP loci of each accession and their ploidy levels, 999,258 single dosage SNPs were identified and most loci were estimated as largely homozygotes. An average of 34,397 haplotype blocks for each accession was inferred. The highest divergence time among the Saccharum spp. was estimated as 1.2 million years ago (MYA). Saccharum spp. diverged from Erianthus and Sorghum approximately 5 and 6 MYA, respectively. Furthermore, the target enrichment sequencing approach provided an effective way to discover and catalog natural allelic variation in highly polyploid or heterozygous genomes.« less
Natural Allelic Variations in Highly Polyploidy Saccharum Complex
Song, Jian; Yang, Xiping; Resende, Jr., Marcio F. R.; ...
2016-06-08
Sugarcane ( Saccharum spp.) is an important sugar and biofuel crop with high polyploid and complex genomes. The Saccharum complex, comprised of Saccharum genus and a few related genera, are important genetic resources for sugarcane breeding. A large amount of natural variation exists within the Saccharum complex. Though understanding their allelic variation has been challenging, it is critical to dissect allelic structure and to identify the alleles controlling important traits in sugarcane. To characterize natural variations in Saccharum complex, a target enrichment sequencing approach was used to assay 12 representative germplasm accessions. In total, 55,946 highly efficient probes were designedmore » based on the sorghum genome and sugarcane unigene set targeting a total of 6 Mb of the sugarcane genome. A pipeline specifically tailored for polyploid sequence variants and genotype calling was established. BWAmem and sorghum genome approved to be an acceptable aligner and reference for sugarcane target enrichment sequence analysis, respectively. Genetic variations including 1,166,066 non-redundant SNPs, 150,421 InDels, 919 gene copy number variations, and 1,257 gene presence/absence variations were detected. SNPs from three different callers (Samtools, Freebayes, and GATK) were compared and the validation rates were nearly 90%. Based on the SNP loci of each accession and their ploidy levels, 999,258 single dosage SNPs were identified and most loci were estimated as largely homozygotes. An average of 34,397 haplotype blocks for each accession was inferred. The highest divergence time among the Saccharum spp. was estimated as 1.2 million years ago (MYA). Saccharum spp. diverged from Erianthus and Sorghum approximately 5 and 6 MYA, respectively. Furthermore, the target enrichment sequencing approach provided an effective way to discover and catalog natural allelic variation in highly polyploid or heterozygous genomes.« less
NASA Astrophysics Data System (ADS)
Hooke, J. M.
2015-12-01
In spite of major physical impacts from large floods, present river management rarely takes into account the possible dynamics and variation in magnitude-impact relations over time in flood risk mapping and assessment nor incorporates feedback effects of changes into modelling. Using examples from the literature and from field measurements over several decades in two contrasting environments, a semi-arid region and a humid-temperate region, temporal variations in channel response to flood events are evaluated. The evidence demonstrates how flood physical impacts can vary at a location over time. The factors influencing that variation on differing timescales are examined. The analysis indicates the importance of morphological changes and trajectory of adjustment in relation to thresholds, and that trends in force or resistance can take place over various timescales, altering those thresholds. Sediment supply can also change with altered connectivity upstream and changes in state of hillslope-channel coupling. It demonstrates that seasonal timing and sequence of events can affect response, particularly deposition through sediment supply. Duration can also have a significant effect and modify the magnitude relation. Lack of response or deposits in some events can mean that flood frequency using such evidence is underestimated. A framework for assessment of both past and possible future changes is provided which emphasises the uncertainty and the inconstancy of the magnitude-impact relation and highlights the dynamic factors and nature of variability that should be considered in sustainable management of river channels.
Mader, Malte; Le Paslier, Marie-Christine; Bounon, Rémi; Berard, Aurélie; Vettori, Cristina; Schroeder, Hilke; Leplé, Jean-Charles; Fladung, Matthias
2016-01-01
Complete Populus genome sequences are available for the nucleus (P. trichocarpa; section Tacamahaca) and for chloroplasts (seven species), but not for mitochondria. Here, we provide the complete genome sequences of the chloroplast and the mitochondrion for the clones P. tremula W52 and P. tremula x P. alba 717-1B4 (section Populus). The organization of the chloroplast genomes of both Populus clones is described. A phylogenetic tree constructed from all available complete chloroplast DNA sequences of Populus was not congruent with the assignment of the related species to different Populus sections. In total, 3,024 variable nucleotide positions were identified among all compared Populus chloroplast DNA sequences. The 5-prime part of the LSC from trnH to atpA showed the highest frequency of variations. The variable positions included 163 positions with SNPs allowing for differentiating the two clones with P. tremula chloroplast genomes (W52, 717-1B4) from the other seven Populus individuals. These potential P. tremula-specific SNPs were displayed as a whole-plastome barcode on the P. tremula W52 chloroplast DNA sequence. Three of these SNPs and one InDel in the trnH-psbA linker were successfully validated by Sanger sequencing in an extended set of Populus individuals. The complete mitochondrial genome sequence of P. tremula is the first in the family of Salicaceae. The mitochondrial genomes of the two clones are 783,442 bp (W52) and 783,513 bp (717-1B4) in size, structurally very similar and organized as single circles. DNA sequence regions with high similarity to the W52 chloroplast sequence account for about 2% of the W52 mitochondrial genome. The mean SNP frequency was found to be nearly six fold higher in the chloroplast than in the mitochondrial genome when comparing 717-1B4 with W52. The availability of the genomic information of all three DNA-containing cell organelles will allow a holistic approach in poplar molecular breeding in the future. PMID:26800039
Variant Inferior Alveolar Nerves and Implications for Local Anesthesia
Wolf, Kevin T.; Brokaw, Everett J.; Bell, Andrea; Joy, Anita
2016-01-01
A sound knowledge of anatomical variations that could be encountered during surgical procedures is helpful in avoiding surgical complications. The current article details anomalous morphology of inferior alveolar nerves encountered during routine dissection of the craniofacial region in the Gross Anatomy laboratory. We also report variations of the lingual nerves, associated with the inferior alveolar nerves. The variations were documented and a thorough review of literature was carried out. We focus on the variations themselves, and the clinical implications that these variations present. Thorough understanding of variant anatomy of the lingual and inferior alveolar nerves may determine the success of procedural anesthesia, the etiology of pathologic processes, and the avoidance of surgical misadventure. PMID:27269666
Williams, Tony D.; Ames, Caroline E.; Kiparissis, Yiannis; Wynne-Edwards, Katherine E.
2005-01-01
We investigated the relationship between plasma and yolk oestrogens in laying female zebra finches (Taeniopygia guttata) by manipulating plasma oestradiol (E2) levels, via injection of oestradiol-17β, in a sequence-specific manner to maintain chronically high plasma levels for later-developing eggs (contrasting with the endogenous pattern of decreasing plasma E2 concentrations during laying). We report systematic variation in yolk oestrogen concentrations, in relation to laying sequence, similar to that widely reported for androgenic steroids. In sham-manipulated females, yolk E2 concentrations decreased with laying sequence. However, in E2-treated females plasma E2 levels were higher during the period of rapid yolk development of later-laid eggs, compared with control females. As a consequence, we reversed the laying-sequence-specific pattern of yolk E2: in E2-treated females, yolk E2 concentrations increased with laying-sequence. In general therefore, yolk E2 levels were a direct reflection of plasma E2 levels. However, in control females there was some inter-individual variability in the endogenous pattern of plasma E2 levels through the laying cycle which could generate variation in sequence-specific patterns of yolk hormone levels even if these primarily reflect circulating steroid levels. PMID:15695208
Barik, Suvakanta; SarkarDas, Shabari; Singh, Archita; Gautam, Vibhav; Kumar, Pramod; Majee, Manoj; Sarkar, Ananda K
2014-01-01
Similar to the majority of the microRNAs, mature miR166s are derived from multiple members of MIR166 genes (precursors) and regulate various aspects of plant development by negatively regulating their target genes (Class III HD-ZIP). The evolutionary conservation or functional diversification of miRNA166 family members remains elusive. Here, we show the phylogenetic relationships among MIR166 precursor and mature sequences from three diverse model plant species. Despite strong conservation, some mature miR166 sequences, such as ppt-miR166m, have undergone sequence variation. Critical sequence variation in ppt-miR166m has led to functional diversification, as it targets non-HD-ZIPIII gene transcript (s). MIR166 precursor sequences have diverged in a lineage specific manner, and both precursors and mature osa-miR166i/j are highly conserved. Interestingly, polycistronic MIR166s were present in Physcomitrella and Oryza but not in Arabidopsis. The nature of cis-regulatory motifs on the upstream promoter sequences of MIR166 genes indicates their possible contribution to the functional variation observed among miR166 species. Copyright © 2013 Elsevier Inc. All rights reserved.
Liu, Siyang; Huang, Shujia; Rao, Junhua; Ye, Weijian; Krogh, Anders; Wang, Jun
2015-01-01
Comprehensive recognition of genomic variation in one individual is important for understanding disease and developing personalized medication and treatment. Many tools based on DNA re-sequencing exist for identification of single nucleotide polymorphisms, small insertions and deletions (indels) as well as large deletions. However, these approaches consistently display a substantial bias against the recovery of complex structural variants and novel sequence in individual genomes and do not provide interpretation information such as the annotation of ancestral state and formation mechanism. We present a novel approach implemented in a single software package, AsmVar, to discover, genotype and characterize different forms of structural variation and novel sequence from population-scale de novo genome assemblies up to nucleotide resolution. Application of AsmVar to several human de novo genome assemblies captures a wide spectrum of structural variants and novel sequences present in the human population in high sensitivity and specificity. Our method provides a direct solution for investigating structural variants and novel sequences from de novo genome assemblies, facilitating the construction of population-scale pan-genomes. Our study also highlights the usefulness of the de novo assembly strategy for definition of genome structure.
Zheng, Yan Ying; Xie, Ling; Liu, Li; Zhang, Shu Peng; Wu, Xiao Bin; Zhu, Chang Le; Lai, Ren Sheng
2012-10-08
Colorectal cancer is one of the most common tumors with high mortality in China. Microsatellite instability (MSI) analysis is important for the diagnosis of hereditary non-polyposis colorectal cancer (HNPCC) and for the prediction of 5-FU chemotherapy efficiency of colorectal tumors, especially in terms of therapeutic response and overall survival rates. Among the MSI markers recommended by the NIH/NCI, BAT-25 has been extensively studied for its major role in MSI. BAT-25 presents different polymorphisms in different ethnic populations and studies of its polymorphisms in the Chinese population are still very limited. To analyze the frequency of constitutive polymorphic variation at the BAT-25 locus in Chinese from Jiangsu Province and its implication for locus MSI screening. The frequency of allelic variation at the BAT-25 locus of cervical cells from 500 healthy women and blood from 16 healthy males was assessed by direct sequencing. Twenty samples were also analyzed by fragment analysis. DNA extracted from blood of 94 patients with gastrointestinal cancer or endometrial cancer was analyzed by fragment analysis. After comparison with the sequencing results, the more frequent allele lengths were 126-127 bp, 128-129 bp, 129-130 bp, respectively consistent with the 24 poly(T) (T24), T25 and T26 alleles. At the BAT-25 locus, 516 healthy individuals had respectively 1.36%, 97.28% and 1.36% of the T24, T25 and T26. Whereas for the 94 cancer patients allelic frequencies were 0.53%, 1.06%, 96.8%, 1.6% for T15, T24, T25 and T26 alleles respectively. Sixteen healthy males had only the T25 allele and heterozygous T15 was only found in 1 male patient with colon cancer. We established the relation between fragment length and thymine repeats in BAT-25. The results showed that the BAT-25 locus is quasimonomorphic in Chinese from Jiangsu province. Moreover we showed that variant alleles of BAT-25 were found more likely in blood from cancer patients than in healthy individuals, suggesting the need to perform comparative studies between tumor and blood, or normal tissue, as to obtain a correct MSI identification.
Genetic variation patterns of American chestnut populations at EST-SSRs
Oliver Gailing; C. Dana Nelson
2017-01-01
The objective of this study is to analyze patterns of genetic variation at genic expressed sequence tag - simple sequence repeats (EST-SSRs) and at chloroplast DNA markers in populations of American chestnut (Castanea dentata Borkh.) to assist in conservation and breeding efforts. Allelic diversity at EST-SSRs decreased significantly from southwest to northeast along...
Thompson, David N; Thompson, Vicki S; Schaller, Kastli D; Apel, William A; Reed, David W; Lacey, Jeffrey A
2013-04-30
Isolated and/or purified polypeptides and nucleic acid sequences encoding polypeptides from Alicyclobacillus acidocaldarius and variations thereof are provided. Further provided are methods of at least partially degrading xylotriose, xylobiose, and/or arabinofuranose-substituted xylan using isolated and/or purified polypeptides and nucleic acid sequences encoding polypeptides from Alicyclobacillus acidocaldarius and variations thereof.
USDA-ARS?s Scientific Manuscript database
Copy number variations (CNVs) are large insertions, deletions or duplications in the genome that vary between members of a species and are known to affect a wide variety of phenotypic traits. In this study, we identified CNVs in a population of bulls using low coverage next-generation sequence data....
Blake, Jonathon; Riddell, Andrew; Theiss, Susanne; Gonzalez, Alexis Perez; Haase, Bettina; Jauch, Anna; Janssen, Johannes W. G.; Ibberson, David; Pavlinic, Dinko; Moog, Ute; Benes, Vladimir; Runz, Heiko
2014-01-01
Balanced chromosome abnormalities (BCAs) occur at a high frequency in healthy and diseased individuals, but cost-efficient strategies to identify BCAs and evaluate whether they contribute to a phenotype have not yet become widespread. Here we apply genome-wide mate-pair library sequencing to characterize structural variation in a patient with unclear neurodevelopmental disease (NDD) and complex de novo BCAs at the karyotype level. Nucleotide-level characterization of the clinically described BCA breakpoints revealed disruption of at least three NDD candidate genes (LINC00299, NUP205, PSMD14) that gave rise to abnormal mRNAs and could be assumed as disease-causing. However, unbiased genome-wide analysis of the sequencing data for cryptic structural variation was key to reveal an additional submicroscopic inversion that truncates the schizophrenia- and bipolar disorder-associated brain transcription factor ZNF804A as an equally likely NDD-driving gene. Deep sequencing of fluorescent-sorted wild-type and derivative chromosomes confirmed the clinically undetected BCA. Moreover, deep sequencing further validated a high accuracy of mate-pair library sequencing to detect structural variants larger than 10 kB, proposing that this approach is powerful for clinical-grade genome-wide structural variant detection. Our study supports previous evidence for a role of ZNF804A in NDD and highlights the need for a more comprehensive assessment of structural variation in karyotypically abnormal individuals and patients with neurocognitive disease to avoid diagnostic deception. PMID:24625750
Goettel, Wolfgang; Xia, Eric; Upchurch, Robert; Wang, Ming-Li; Chen, Pengyin; An, Yong-Qiang Charles
2014-04-23
Variation in seed oil composition and content among soybean varieties is largely attributed to differences in transcript sequences and/or transcript accumulation of oil production related genes in seeds. Discovery and analysis of sequence and expression variations in these genes will accelerate soybean oil quality improvement. In an effort to identify these variations, we sequenced the transcriptomes of soybean seeds from nine lines varying in oil composition and/or total oil content. Our results showed that 69,338 distinct transcripts from 32,885 annotated genes were expressed in seeds. A total of 8,037 transcript expression polymorphisms and 50,485 transcript sequence polymorphisms (48,792 SNPs and 1,693 small Indels) were identified among the lines. Effects of the transcript polymorphisms on their encoded protein sequences and functions were predicted. The studies also provided independent evidence that the lack of FAD2-1A gene activity and a non-synonymous SNP in the coding sequence of FAB2C caused elevated oleic acid and stearic acid levels in soybean lines M23 and FAM94-41, respectively. As a proof-of-concept, we developed an integrated RNA-seq and bioinformatics approach to identify and functionally annotate transcript polymorphisms, and demonstrated its high effectiveness for discovery of genetic and transcript variations that result in altered oil quality traits. The collection of transcript polymorphisms coupled with their predicted functional effects will be a valuable asset for further discovery of genes, gene variants, and functional markers to improve soybean oil quality.
Granados-Cifuentes, Camila; Bellantuono, Anthony J; Ridgway, Tyrone; Hoegh-Guldberg, Ove; Rodriguez-Lanetty, Mauricio
2013-04-08
Ecosystems worldwide are suffering the consequences of anthropogenic impact. The diverse ecosystem of coral reefs, for example, are globally threatened by increases in sea surface temperatures due to global warming. Studies to date have focused on determining genetic diversity, the sequence variability of genes in a species, as a proxy to estimate and predict the potential adaptive response of coral populations to environmental changes linked to climate changes. However, the examination of natural gene expression variation has received less attention. This variation has been implicated as an important factor in evolutionary processes, upon which natural selection can act. We acclimatized coral nubbins from six colonies of the reef-building coral Acropora millepora to a common garden in Heron Island (Great Barrier Reef, GBR) for a period of four weeks to remove any site-specific environmental effects on the physiology of the coral nubbins. By using a cDNA microarray platform, we detected a high level of gene expression variation, with 17% (488) of the unigenes differentially expressed across coral nubbins of the six colonies (jsFDR-corrected, p < 0.01). Among the main categories of biological processes found differentially expressed were transport, translation, response to stimulus, oxidation-reduction processes, and apoptosis. We found that the transcriptional profiles did not correspond to the genotype of the colony characterized using either an intron of the carbonic anhydrase gene or microsatellite loci markers. Our results provide evidence of the high inter-colony variation in A. millepora at the transcriptomic level grown under a common garden and without a correspondence with genotypic identity. This finding brings to our attention the importance of taking into account natural variation between reef corals when assessing experimental gene expression differences. The high transcriptional variation detected in this study is interpreted and discussed within the context of adaptive potential and phenotypic plasticity of reef corals. Whether this variation will allow coral reefs to survive to current challenges remains unknown.
de la Fuente, José; Díez-Delgado, Iratxe; Contreras, Marinela; Vicente, Joaquín; Cabezas-Cruz, Alejandro; Tobes, Raquel; Manrique, Marina; López, Vladimir; Romero, Beatriz; Bezos, Javier; Dominguez, Lucas; Sevilla, Iker A; Garrido, Joseba M; Juste, Ramón; Madico, Guillermo; Jones-López, Edward; Gortazar, Christian
2015-11-01
Mycobacteria of the Mycobacterium tuberculosis complex (MTBC) greatly affect humans and animals worldwide. The life cycle of mycobacteria is complex and the mechanisms resulting in pathogen infection and survival in host cells are not fully understood. Recently, comparative genomics analyses have provided new insights into the evolution and adaptation of the MTBC to survive inside the host. However, most of this information has been obtained using M. tuberculosis but not other members of the MTBC such as M. bovis and M. caprae. In this study, the genome of three M. bovis (MB1, MB3, MB4) and one M. caprae (MB2) field isolates with different lesion score, prevalence and host distribution phenotypes were sequenced. Genome sequence information was used for whole-genome and protein-targeted comparative genomics analysis with the aim of finding correlates with phenotypic variation with potential implications for tuberculosis (TB) disease risk assessment and control. At the whole-genome level the results of the first comparative genomics study of field isolates of M. bovis including M. caprae showed that as previously reported for M. tuberculosis, sequential chromosomal nucleotide substitutions were the main driver of the M. bovis genome evolution. The phylogenetic analysis provided a strong support for the M. bovis/M. caprae clade, but supported M. caprae as a separate species. The comparison of the MB1 and MB4 isolates revealed differences in genome sequence, including gene families that are important for bacterial infection and transmission, thus highlighting differences with functional implications between isolates otherwise classified with the same spoligotype. Strategic protein-targeted analysis using the ESX or type VII secretion system, proteins linking stress response with lipid metabolism, host T cell epitopes of mycobacteria, antigens and peptidoglycan assembly protein identified new genetic markers and candidate vaccine antigens that warrant further study to develop tools to evaluate risks for TB disease caused by M. bovis/M.caprae and for TB control in humans and animals.
Elbaz, Benayahu; Shoshani-Knaani, Noa; David-Assael, Ora; Mizrachy-Dagri, Talya; Mizrahi, Keren; Saul, Helen; Brook, Emil; Berezin, Irina; Shaul, Orit
2006-06-01
Zn hyperaccumulator plants sequester Zn into their shoot vacuoles. To date, the only transporters implicated in Zn sequestration into the vacuoles of hyperaccumulator plants are cation diffusion facilitators (CDFs). We investigated the expression in Arabidopsis halleri of a homolog of AtMHX, an A. thaliana tonoplast transporter that exchanges protons with Mg, Zn and Fe ions. A. halleri has a single copy of a homologous gene, encoding a protein that shares 98% sequence identity with AtMHX. Western blot analysis with vacuolar-enriched membrane fractions suggests localization of AhMHX in the tonoplast. The levels of MHX proteins are much higher in leaves of A. halleri than in leaves of the non-accumulator plant A. thaliana. At the same time, the levels of MHX transcripts are similar in leaves of the two species. This suggests that the difference in MHX levels is regulated at the post-transcriptional level. In vitro translation studies indicated that the difference between AhMHX and AtMHX expression is not likely to result from the variations in the sequence of their 5' untranslated regions (5'UTRs). The high expression of AhMHX in A. halleri leaves is constitutive and not significantly affected by the metal status of the plants. In both species, MHX transcript levels are higher in leaves than in roots, but the difference is higher in A. halleri. Metal sequestration into root vacuoles was suggested to inhibit hyperaccumulation in the shoot. Our data implicate AhMHX as a candidate gene in metal accumulation or tolerance in A. halleri.
Macular xanthophylls, lipoprotein-related genes, and age-related macular degeneration1234
Koo, Euna; Neuringer, Martha; SanGiovanni, John Paul
2014-01-01
Plant-based macular xanthophylls (MXs; lutein and zeaxanthin) and the lutein metabolite meso-zeaxanthin are the major constituents of macular pigment, a compound concentrated in retinal areas that are responsible for fine-feature visual sensation. There is an unmet need to examine the genetics of factors influencing regulatory mechanisms and metabolic fates of these 3 MXs because they are linked to processes implicated in the pathogenesis of age-related macular degeneration (AMD). In this work we provide an overview of evidence supporting a molecular basis for AMD-MX associations as they may relate to DNA sequence variation in AMD- and lipoprotein-related genes. We recognize a number of emerging research opportunities, barriers, knowledge gaps, and tools offering promise for meaningful investigation and inference in the field. Overviews on AMD- and high-density lipoprotein (HDL)–related genes encoding receptors, transporters, and enzymes affecting or affected by MXs are followed with information on localization of products from these genes to retinal cell types manifesting AMD-related pathophysiology. Evidence on the relation of each gene or gene product with retinal MX response to nutrient intake is discussed. This information is followed by a review of results from mechanistic studies testing gene-disease relations. We then present findings on relations of AMD with DNA sequence variants in MX-associated genes. Our conclusion is that AMD-associated DNA variants that influence the actions and metabolic fates of HDL system constituents should be examined further for concomitant influence on MX absorption, retinal tissue responses to MX intake, and the capacity to modify MX-associated factors and processes implicated in AMD pathogenesis. PMID:24829491
Kanduma, Esther G; Mwacharo, Joram M; Githaka, Naftaly W; Kinyanjui, Peter W; Njuguna, Joyce N; Kamau, Lucy M; Kariuki, Edward; Mwaura, Stephen; Skilton, Robert A; Bishop, Richard P
2016-06-22
The ixodid tick Rhipicephalus appendiculatus transmits the apicomplexan protozoan parasite Theileria parva, which causes East coast fever (ECF), the most economically important cattle disease in eastern and southern Africa. Recent analysis of micro- and minisatellite markers showed an absence of geographical and host-associated genetic sub-structuring amongst field populations of R. appendiculatus in Kenya. To assess further the phylogenetic relationships between field and laboratory R. appendiculatus tick isolates, this study examined sequence variations at two mitochondrial genes, cytochrome c oxidase subunit I (COI) and 12S ribosomal RNA (rRNA), and the nuclear encoded ribosomal internal transcribed spacer 2 (ITS2) of the rRNA gene, respectively. The analysis of 332 COI sequences revealed 30 polymorphic sites, which defined 28 haplotypes that were separated into two distinct haplogroups (A and B). Inclusion of previously published haplotypes in our analysis revealed a high degree of phylogenetic complexity never reported before in haplogroup A. Neither haplogroup however, showed any clustering pattern related to either the geographical sampling location, the type of tick sampled (laboratory stocks vs field populations) or the mammalian host species. This finding was supported by the results obtained from the analysis of 12S rDNA sequences. Analysis of molecular variance (AMOVA) indicated that 90.8 % of the total genetic variation was explained by the two haplogroups, providing further support for their genetic divergence. These results were, however, not replicated by the nuclear transcribed ITS2 sequences likely because of recombination between the nuclear genomes maintaining a high level of genetic sequence conservation. COI and 12S rDNA are better markers than ITS2 for studying intraspecific diversity. Based on these genes, two major genetic groups of R. appendiculatus that have gone through a demographic expansion exist in Kenya. The two groups show no phylogeographic structure or correlation with the type of host species from which the ticks were collected, nor to the evolutionary and breeding history of the species. The two lineages may have a wide geographic distribution range in eastern and southern Africa. The findings of this study may have implications for the spread and control of R. appendiculatus, and indirectly, on the transmission dynamics of ECF.
Genome sequence of Stachybotrys chartarum Strain 51-11
Stachybotrys chartarum strain 51-11 genome was sequenced by shotgun sequencing utilizing Illumina Hiseq 2000 and PacBio long read technology. Since Stachybotrys chartarum has been implicated in health impacts within water-damaged buildings, any information extracted from the geno...
Clark, Shaunna L; McClay, Joseph L; Adkins, Daniel E; Aberg, Karolina A; Kumar, Gaurav; Nerella, Sri; Xie, Linying; Collins, Ann L; Crowley, James J; Quakenbush, Corey R; Hillard, Christopher E; Gao, Guimin; Shabalin, Andrey A; Peterson, Roseann E; Copeland, William E; Silberg, Judy L; Maes, Hermine; Sullivan, Patrick F; Costello, Elizabeth J; van den Oord, Edwin J
2016-05-01
Genome-wide association study meta-analyses have robustly implicated three loci that affect susceptibility for smoking: CHRNA5\\CHRNA3\\CHRNB4, CHRNB3\\CHRNA6 and EGLN2\\CYP2A6. Functional follow-up studies of these loci are needed to provide insight into biological mechanisms. However, these efforts have been hampered by a lack of knowledge about the specific causal variant(s) involved. In this study, we prioritized variants in terms of the likelihood they account for the reported associations. We employed targeted capture of the CHRNA5\\CHRNA3\\CHRNB4, CHRNB3\\CHRNA6, and EGLN2\\CYP2A6 loci and flanking regions followed by next-generation deep sequencing (mean coverage 78×) to capture genomic variation in 363 individuals. We performed single locus tests to determine if any single variant accounts for the association, and examined if sets of (rare) variants that overlapped with biologically meaningful annotations account for the associations. In total, we investigated 963 variants, of which 71.1% were rare (minor allele frequency < 0.01), 6.02% were insertion/deletions, and 51.7% were catalogued in dbSNP141. The single variant results showed that no variant fully accounts for the association in any region. In the variant set results, CHRNB4 accounts for most of the signal with significant sets consisting of directly damaging variants. CHRNA6 explains most of the signal in the CHRNB3\\CHRNA6 locus with significant sets indicating a regulatory role for CHRNA6. Significant sets in CYP2A6 involved directly damaging variants while the significant variant sets suggested a regulatory role for EGLN2. We found that multiple variants implicating multiple processes explain the signal. Some variants can be prioritized for functional follow-up. © The Author 2015. Published by Oxford University Press on behalf of the Society for Research on Nicotine and Tobacco. All rights reserved. For permissions, please e-mail: journals.permissions@oup.com.
McClay, Joseph L.; Adkins, Daniel E.; Aberg, Karolina A.; Kumar, Gaurav; Nerella, Sri; Xie, Linying; Collins, Ann L.; Crowley, James J.; Quakenbush, Corey R.; Hillard, Christopher E.; Gao, Guimin; Shabalin, Andrey A.; Peterson, Roseann E.; Copeland, William E.; Silberg, Judy L.; Maes, Hermine; Sullivan, Patrick F.; Costello, Elizabeth J.; van den Oord, Edwin J.
2016-01-01
Abstract Introduction: Genome-wide association study meta-analyses have robustly implicated three loci that affect susceptibility for smoking: CHRNA5\\CHRNA3\\CHRNB4 , CHRNB3\\CHRNA6 and EGLN2\\CYP2A6 . Functional follow-up studies of these loci are needed to provide insight into biological mechanisms. However, these efforts have been hampered by a lack of knowledge about the specific causal variant(s) involved. In this study, we prioritized variants in terms of the likelihood they account for the reported associations. Methods: We employed targeted capture of the CHRNA5\\CHRNA3\\CHRNB4 , CHRNB3\\CHRNA6 , and EGLN2\\CYP2A6 loci and flanking regions followed by next-generation deep sequencing (mean coverage 78×) to capture genomic variation in 363 individuals. We performed single locus tests to determine if any single variant accounts for the association, and examined if sets of (rare) variants that overlapped with biologically meaningful annotations account for the associations. Results: In total, we investigated 963 variants, of which 71.1% were rare (minor allele frequency < 0.01), 6.02% were insertion/deletions, and 51.7% were catalogued in dbSNP141. The single variant results showed that no variant fully accounts for the association in any region. In the variant set results, CHRNB4 accounts for most of the signal with significant sets consisting of directly damaging variants. CHRNA6 explains most of the signal in the CHRNB3\\CHRNA6 locus with significant sets indicating a regulatory role for CHRNA6 . Significant sets in CYP2A6 involved directly damaging variants while the significant variant sets suggested a regulatory role for EGLN2 . Conclusions: We found that multiple variants implicating multiple processes explain the signal. Some variants can be prioritized for functional follow-up. PMID:26283763
Kos, Mark Z; Carless, Melanie A; Peralta, Juan; Curran, Joanne E; Quillen, Ellen E; Almeida, Marcio; Blackburn, August; Blondell, Lucy; Roalf, David R; Pogue-Geile, Michael F; Gur, Ruben C; Göring, Harald H H; Nimgaonkar, Vishwajit L; Gur, Raquel E; Almasy, Laura
2017-12-01
Schizophrenia is a serious mental illness, involving disruptions in thought and behavior, with a worldwide prevalence of about one percent. Although highly heritable, much of the genetic liability of schizophrenia is yet to be explained. We searched for susceptibility loci in multiplex, multigenerational families affected by schizophrenia, targeting protein-altering variation with in silico predicted functional effects. Exome sequencing was performed on 136 samples from eight European-American families, including 23 individuals diagnosed with schizophrenia or schizoaffective disorder. In total, 11,878 non-synonymous variants from 6,396 genes were tested for their association with schizophrenia spectrum disorders. Pathway enrichment analyses were conducted on gene-based test results, protein-protein interaction (PPI) networks, and epistatic effects. Using a significance threshold of FDR < 0.1, association was detected for rs10941112 (p = 2.1 × 10 -5 ; q-value = 0.073) in AMACR, a gene involved in fatty acid metabolism and previously implicated in schizophrenia, with significant cis effects on gene expression (p = 5.5 × 10 -4 ), including brain tissue data from the Genotype-Tissue Expression project (minimum p = 6.0 × 10 -5 ). A second SNP, rs10378 located in TMEM176A, also shows risk effects in the exome data (p = 2.8 × 10 -5 ; q-value = 0.073). PPIs among our top gene-based association results (p < 0.05; n = 359 genes) reveal significant enrichment of genes involved in NCAM-mediated neurite outgrowth (p = 3.0 × 10 -5 ), while exome-wide SNP-SNP interaction effects for rs10941112 and rs10378 indicate a potential role for kinase-mediated signaling involved in memory and learning. In conclusion, these association results implicate AMACR and TMEM176A in schizophrenia risk, whose effects may be modulated by genes involved in synaptic plasticity and neurocognitive performance. © 2017 Wiley Periodicals, Inc.
NASA Astrophysics Data System (ADS)
Cleary, Alison C.; Durbin, Edward G.; Rynearson, Tatiana A.; Bailey, Jennifer
2016-12-01
Pseudocalanus copepods are small, abundant zooplankton in the Bering Sea ecosystem that play an important role in transferring primary production to fish and other higher trophic-level predators. Four morphologically cryptic species, the primarily arctic Pseudocalanus minutus and Pseudocalanus acuspes, and the more temperate Pseudocalanus newmani and Pseudocalanus mimus, are found within the Bering Sea. Pseudocalanus are generally considered phytoplanktivores. However, their feeding is poorly known, despite their importance to the ecosystem. In situ feeding by the three most abundant Pseudocalanus congeners, P. minutus, P. newmani, and P. acuspes, was investigated by sequencing partial 18S rDNA (ribosomal Deoxyribonucleic Acid) of gut contents from 225 individuals sampled from 8 stations across the Bering Sea in May and June of 2010. The 28,456 prey 18S rDNA sequences obtained clustered into 138 distinct prey items with a 97% similarity cut-off, and included diatoms, dinoflagellates, microzooplankton, mesozooplankton, and vascular plants. Pseudocalanus diets reflected variations in the environment, with phytoplankton sequences relatively more abundant in copepods from stations with higher water-column chlorophyll a concentrations. Feeding differences were observed between species. P. acuspes diet contained relatively more heterotrophic dinoflagellate sequences, and was significantly different from that of P. minutus and P. newmani, both of which contained relatively more diatom sequences, and between which no significant difference was observed. Feeding differences between the two primarily arctic species may be a mechanism of niche partitioning between these spatially co-located congeners and may have implications for the effects of climate change on the success of these abundant zooplankters and their many predators in this ecosystem.
Pattaradilokrat, Sittiporn; Trakoolsoontorn, Chawinya; Simpalipan, Phumin; Warrit, Natapot; Kaewthamasorn, Morakot; Harnyuttanakorn, Pongchai
2018-01-22
The glutamate-rich protein (GLURP) of the malaria parasite Plasmodium falciparum is a key surface antigen that serves as a component of a clinical vaccine. Moreover, the GLURP gene is also employed routinely as a genetic marker for malarial genotyping in epidemiological studies. While extensive size polymorphisms in GLURP are well recorded, the extent of the sequence diversity of this gene is rarely investigated. The present study aimed to explore the genetic diversity of GLURP in natural populations of P. falciparum. The polymorphic C-terminal repetitive R2 region of GLURP sequences from 65 P. falciparum isolates in Thailand were generated and combined with the data from 103 worldwide isolates to generate a GLURP database. The collection was comprised of 168 alleles, encoding 105 unique GLURP subtypes, characterized by 18 types of amino acid repeat units (AAU). Of these, 28 GLURP subtypes, formed by 10 AAU types, were detected in P. falciparum in Thailand. Among them, 19 GLURP subtypes and 2 AAU types are described for the first time in the Thai parasite population. The AAU sequences were highly conserved, which is likely due to negative selection. Standard Fst analysis revealed the shared distributions of GLURP types among the P. falciparum populations, providing evidence of gene flow among the different demographic populations. Sequence diversity causing size variations in GLURP in Thai P. falciparum populations were detected, and caused by non-synonymous substitutions in repeat units and some insertion/deletion of aspartic acid or glutamic acid codons between repeat units. The P. falciparum population structure based on GLURP showed promising implications for the development of GLURP-based vaccines and for monitoring vaccine efficacy.
Krieger, Jeannette; Hett, Anne Kathrin; Fuerst, Paul A; Birstein, Vadim J; Ludwig, Arne
2006-01-01
Significant intraindividual variation in the sequence of the 18S rRNA gene is unusual in animal genomes. In a previous study, multiple 18S rRNA gene sequences were observed within individuals of eight species of sturgeon from North America but not in the North American paddlefish, Polyodon spathula, in two species of Polypterus (Polypterus delhezi and Polypterus senegalus), in other primitive fishes (Erpetoichthys calabaricus, Lepisosteus osseus, Amia calva) or in a lungfish (Protopterus sp.). These observations led to the hypothesis that this unusual genetic characteristic arose within the Acipenseriformes after the presumed divergence of the sturgeon and paddlefish families. In the present study, a survey of nearly all Eurasian acipenseriform species was conducted to examine 18S rDNA variation. Intraindividual variation was not found in the polyodontid species, the Chinese paddlefish, Psephurus gladius, but variation was detected in all Eurasian acipenserid species. The comparison of sequences from two major segments of the 18S rRNA gene and identification of sites where insertion/deletion events have occurred are placed in the context of evolutionary relationships within the Acipenseriformes and the evolution of rDNA variation in this group.
Brandstätter, Anita; Peterson, Christine T; Irwin, Jodi A; Mpoke, Solomon; Koech, Davy K; Parson, Walther; Parsons, Thomas J
2004-10-01
Large forensic mtDNA databases which adhere to strict guidelines for generation and maintenance, are not available for many populations outside of the United States and western Europe. We have established a high quality mtDNA control region sequence database for urban Nairobi as both a reference database for forensic investigations, and as a tool to examine the genetic variation of Kenyan sequences in the context of known African variation. The Nairobi sequences exhibited high variation and a low random match probability, indicating utility for forensic testing. Haplogroup identification and frequencies were compared with those reported from other published studies on African, or African-origin populations from Mozambique, Sierra Leone, and the United States, and suggest significant differences in the mtDNA compositions of the various populations. The quality of the sequence data in our study was investigated and supported using phylogenetic measures. Our data demonstrate the diversity and distinctiveness of African populations, and underline the importance of establishing additional forensic mtDNA databases of indigenous African populations.
Genomic Sequencing: Assessing The Health Care System, Policy, And Big-Data Implications
Phillips, Kathryn A.; Trosman, Julia; Kelley, Robin K.; Pletcher, Mark J.; Douglas, Michael P.; Weldon, Christine B.
2014-01-01
New genomic sequencing technologies enable the high-speed analysis of multiple genes simultaneously, including all of those in a person's genome. Sequencing is a prominent example of a “big data” technology because of the massive amount of information it produces and its complexity, diversity, and timeliness. Our objective in this article is to provide a policy primer on sequencing and illustrate how it can affect health care system and policy issues. Toward this end, we developed an easily applied classification of sequencing based on inputs, methods, and outputs. We used it to examine the implications of sequencing for three health care system and policy issues: making care more patient-centered, developing coverage and reimbursement policies, and assessing economic value. We conclude that sequencing has great promise but that policy challenges include how to optimize patient engagement as well as privacy, develop coverage policies that distinguish research from clinical uses and account for bioinformatics costs, and determine the economic value of sequencing through complex economic models that take into account multiple findings and downstream costs. PMID:25006153
Genomic sequencing: assessing the health care system, policy, and big-data implications.
Phillips, Kathryn A; Trosman, Julia R; Kelley, Robin K; Pletcher, Mark J; Douglas, Michael P; Weldon, Christine B
2014-07-01
New genomic sequencing technologies enable the high-speed analysis of multiple genes simultaneously, including all of those in a person's genome. Sequencing is a prominent example of a "big data" technology because of the massive amount of information it produces and its complexity, diversity, and timeliness. Our objective in this article is to provide a policy primer on sequencing and illustrate how it can affect health care system and policy issues. Toward this end, we developed an easily applied classification of sequencing based on inputs, methods, and outputs. We used it to examine the implications of sequencing for three health care system and policy issues: making care more patient-centered, developing coverage and reimbursement policies, and assessing economic value. We conclude that sequencing has great promise but that policy challenges include how to optimize patient engagement as well as privacy, develop coverage policies that distinguish research from clinical uses and account for bioinformatics costs, and determine the economic value of sequencing through complex economic models that take into account multiple findings and downstream costs. Project HOPE—The People-to-People Health Foundation, Inc.
Consensus generation and variant detection by Celera Assembler.
Denisov, Gennady; Walenz, Brian; Halpern, Aaron L; Miller, Jason; Axelrod, Nelson; Levy, Samuel; Sutton, Granger
2008-04-15
We present an algorithm to identify allelic variation given a Whole Genome Shotgun (WGS) assembly of haploid sequences, and to produce a set of haploid consensus sequences rather than a single consensus sequence. Existing WGS assemblers take a column-by-column approach to consensus generation, and produce a single consensus sequence which can be inconsistent with the underlying haploid alleles, and inconsistent with any of the aligned sequence reads. Our new algorithm uses a dynamic windowing approach. It detects alleles by simultaneously processing the portions of aligned reads spanning a region of sequence variation, assigns reads to their respective alleles, phases adjacent variant alleles and generates a consensus sequence corresponding to each confirmed allele. This algorithm was used to produce the first diploid genome sequence of an individual human. It can also be applied to assemblies of multiple diploid individuals and hybrid assemblies of multiple haploid organisms. Being applied to the individual human genome assembly, the new algorithm detects exactly two confirmed alleles and reports two consensus sequences in 98.98% of the total number 2,033311 detected regions of sequence variation. In 33,269 out of 460,373 detected regions of size >1 bp, it fixes the constructed errors of a mosaic haploid representation of a diploid locus as produced by the original Celera Assembler consensus algorithm. Using an optimized procedure calibrated against 1 506 344 known SNPs, it detects 438 814 new heterozygous SNPs with false positive rate 12%. The open source code is available at: http://wgs-assembler.cvs.sourceforge.net/wgs-assembler/
Takeda, Ryuta; Petrov, Anton I.; Leontis, Neocles B.; Ding, Biao
2011-01-01
Cell-to-cell trafficking of RNA is an emerging biological principle that integrates systemic gene regulation, viral infection, antiviral response, and cell-to-cell communication. A key mechanistic question is how an RNA is specifically selected for trafficking from one type of cell into another type. Here, we report the identification of an RNA motif in Potato spindle tuber viroid (PSTVd) required for trafficking from palisade mesophyll to spongy mesophyll in Nicotiana benthamiana leaves. This motif, called loop 6, has the sequence 5′-CGA-3′...5′-GAC-3′ flanked on both sides by cis Watson-Crick G/C and G/U wobble base pairs. We present a three-dimensional (3D) structural model of loop 6 that specifies all non-Watson-Crick base pair interactions, derived by isostericity-based sequence comparisons with 3D RNA motifs from the RNA x-ray crystal structure database. The model is supported by available chemical modification patterns, natural sequence conservation/variations in PSTVd isolates and related species, and functional characterization of all possible mutants for each of the loop 6 base pairs. Our findings and approaches have broad implications for studying the 3D RNA structural motifs mediating trafficking of diverse RNA species across specific cellular boundaries and for studying the structure-function relationships of RNA motifs in other biological processes. PMID:21258006
Takeda, Ryuta; Petrov, Anton I; Leontis, Neocles B; Ding, Biao
2011-01-01
Cell-to-cell trafficking of RNA is an emerging biological principle that integrates systemic gene regulation, viral infection, antiviral response, and cell-to-cell communication. A key mechanistic question is how an RNA is specifically selected for trafficking from one type of cell into another type. Here, we report the identification of an RNA motif in Potato spindle tuber viroid (PSTVd) required for trafficking from palisade mesophyll to spongy mesophyll in Nicotiana benthamiana leaves. This motif, called loop 6, has the sequence 5'-CGA-3'...5'-GAC-3' flanked on both sides by cis Watson-Crick G/C and G/U wobble base pairs. We present a three-dimensional (3D) structural model of loop 6 that specifies all non-Watson-Crick base pair interactions, derived by isostericity-based sequence comparisons with 3D RNA motifs from the RNA x-ray crystal structure database. The model is supported by available chemical modification patterns, natural sequence conservation/variations in PSTVd isolates and related species, and functional characterization of all possible mutants for each of the loop 6 base pairs. Our findings and approaches have broad implications for studying the 3D RNA structural motifs mediating trafficking of diverse RNA species across specific cellular boundaries and for studying the structure-function relationships of RNA motifs in other biological processes.
Ackerman, Joshua T.; Eagles-Smith, Collin A.; Herzog, Mark P.; Yee, Julie L.; Hartman, C. Alex
2016-01-01
Bird eggs are commonly used in contaminant monitoring programs and toxicological risk assessments, but intra-clutch variation and sampling methodology could influence interpretability. We examined the influence of egg laying sequence on egg mercury concentrations and burdens in American avocets, black-necked stilts, and Forster's terns. The average decline in mercury concentrations between the first and last egg laid was 33% for stilts, 22% for terns, and 11% for avocets, and most of this decline occurred between the first and second eggs laid (24% for stilts, 18% for terns, and 9% for avocets). Trends in egg size with egg laying order were inconsistent among species and overall differences in egg volume, mass, length, and width were <3%. We summarized the literature and, among 17 species studied, mercury concentrations generally declined by 16% between the first and second eggs laid. Despite the strong effect of egg laying sequence, most of the variance in egg mercury concentrations still occurred among clutches (75%-91%) rather than within clutches (9%-25%). Using simulations, we determined that to accurately estimate a population's mean egg mercury concentration using only a single random egg from a subset of nests, it would require sampling >60 nests to represent a large population (10% accuracy) or ≥14 nests to represent a small colony that contained <100 nests (20% accuracy).
Recombination-Mediated Host Adaptation by Avian Staphylococcus aureus
Murray, Susan; Pascoe, Ben; Méric, Guillaume; Mageiros, Leonardos; Yahara, Koji; Hitchings, Matthew D.; Friedmann, Yasmin; Wilkinson, Thomas S.; Gormley, Fraser J.; Mack, Dietrich; Bray, James E.; Lamble, Sarah; Bowden, Rory; Jolley, Keith A.; Maiden, Martin C.J.; Wendlandt, Sarah; Schwarz, Stefan; Corander, Jukka; Fitzgerald, J. Ross
2017-01-01
Staphylococcus aureus are globally disseminated among farmed chickens causing skeletal muscle infections, dermatitis, and septicaemia. The emergence of poultry-associated lineages has involved zoonotic transmission from humans to chickens but questions remain about the specific adaptations that promote proliferation of chicken pathogens. We characterized genetic variation in a population of genome-sequenced S. aureus isolates of poultry and human origin. Genealogical analysis identified a dominant poultry-associated sequence cluster within the CC5 clonal complex. Poultry and human CC5 isolates were significantly distinct from each other and more recombination events were detected in the poultry isolates. We identified 44 recombination events in 33 genes along the branch extending to the poultry-specific CC5 cluster, and 47 genes were found more often in CC5 poultry isolates compared with those from humans. Many of these gene sequences were common in chicken isolates from other clonal complexes suggesting horizontal gene transfer among poultry associated lineages. Consistent with functional predictions for putative poultry-associated genes, poultry isolates showed enhanced growth at 42 °C and greater erythrocyte lysis on chicken blood agar in comparison with human isolates. By combining phenotype information with evolutionary analyses of staphylococcal genomes, we provide evidence of adaptation, following a human-to-poultry host transition. This has important implications for the emergence and dissemination of new pathogenic clones associated with modern agriculture. PMID:28338786
Dostie, Josée; Lemire, Edmond; Bouchard, Philippe; Field, Michael; Jones, Kristie; Lorenz, Birgit; Menten, Björn; Buysse, Karen; Pattyn, Filip; Friedli, Marc; Ucla, Catherine; Rossier, Colette; Wyss, Carine; Speleman, Frank; De Paepe, Anne; Dekker, Job; Antonarakis, Stylianos E.; De Baere, Elfride
2009-01-01
To date, the contribution of disrupted potentially cis-regulatory conserved non-coding sequences (CNCs) to human disease is most likely underestimated, as no systematic screens for putative deleterious variations in CNCs have been conducted. As a model for monogenic disease we studied the involvement of genetic changes of CNCs in the cis-regulatory domain of FOXL2 in blepharophimosis syndrome (BPES). Fifty-seven molecularly unsolved BPES patients underwent high-resolution copy number screening and targeted sequencing of CNCs. Apart from three larger distant deletions, a de novo deletion as small as 7.4 kb was found at 283 kb 5′ to FOXL2. The deletion appeared to be triggered by an H-DNA-induced double-stranded break (DSB). In addition, it disrupts a novel long non-coding RNA (ncRNA) PISRT1 and 8 CNCs. The regulatory potential of the deleted CNCs was substantiated by in vitro luciferase assays. Interestingly, Chromosome Conformation Capture (3C) of a 625 kb region surrounding FOXL2 in expressing cellular systems revealed physical interactions of three upstream fragments and the FOXL2 core promoter. Importantly, one of these contains the 7.4 kb deleted fragment. Overall, this study revealed the smallest distant deletion causing monogenic disease and impacts upon the concept of mutation screening in human disease and developmental disorders in particular. PMID:19543368
Sperm Bindin Divergence under Sexual Selection and Concerted Evolution in Sea Stars.
Patiño, Susana; Keever, Carson C; Sunday, Jennifer M; Popovic, Iva; Byrne, Maria; Hart, Michael W
2016-08-01
Selection associated with competition among males or sexual conflict between mates can create positive selection for high rates of molecular evolution of gamete recognition genes and lead to reproductive isolation between species. We analyzed coding sequence and repetitive domain variation in the gene encoding the sperm acrosomal protein bindin in 13 diverse sea star species. We found that bindin has a conserved coding sequence domain structure in all 13 species, with several repeated motifs in a large central region that is similar among all sea stars in organization but highly divergent among genera in nucleotide and predicted amino acid sequence. More bindin codons and lineages showed positive selection for high relative rates of amino acid substitution in genera with gonochoric outcrossing adults (and greater expected strength of sexual selection) than in selfing hermaphrodites. That difference is consistent with the expectation that selfing (a highly derived mating system) may moderate the strength of sexual selection and limit the accumulation of bindin amino acid differences. The results implicate both positive selection on single codons and concerted evolution within the repetitive region in bindin divergence, and suggest that both single amino acid differences and repeat differences may affect sperm-egg binding and reproductive compatibility. © The Author 2016. Published by Oxford University Press on behalf of the Society for Molecular Biology and Evolution. All rights reserved. For permissions, please e-mail: journals.permissions@oup.com.
Minimal Absent Words in Four Human Genome Assemblies
Garcia, Sara P.; Pinho, Armando J.
2011-01-01
Minimal absent words have been computed in genomes of organisms from all domains of life. Here, we aim to contribute to the catalogue of human genomic variation by investigating the variation in number and content of minimal absent words within a species, using four human genome assemblies. We compare the reference human genome GRCh37 assembly, the HuRef assembly of the genome of Craig Venter, the NA12878 assembly from cell line GM12878, and the YH assembly of the genome of a Han Chinese individual. We find the variation in number and content of minimal absent words between assemblies more significant for large and very large minimal absent words, where the biases of sequencing and assembly methodologies become more pronounced. Moreover, we find generally greater similarity between the human genome assemblies sequenced with capillary-based technologies (GRCh37 and HuRef) than between the human genome assemblies sequenced with massively parallel technologies (NA12878 and YH). Finally, as expected, we find the overall variation in number and content of minimal absent words within a species to be generally smaller than the variation between species. PMID:22220210
NASA Astrophysics Data System (ADS)
Sirota, Ido; enzel, Yehouda; Lensky, Nadav G.
2017-04-01
Layered halite sequences deposited in deep basins throughout the geological record. However, analogues of such sequences are commonly studied in sallow environments. Here we study active precipitation of halite layers from the only modern analog for deep, halite-precipitating basin, the hypersaline Dead Sea. In situ observations in the Dead Sea link seasonal thermohaline stratification, halite saturation, and the characteristics of the actively forming halite layers. The spatiotemporal evolution of halite precipitation in the Dead Sea was characterized by means of monthly observations of the i) lake thermohaline stratification (temperature, salinity, and density), ii) degree of halite saturation, and iii) textural evolution of the active halite deposits. We present the observed relationships between textural characteristics of layered halite deposits (i.e. grain size, consolidation, and roughness) and the degree of saturation, which in turn reflected the limnology and hydro-climatology. The lakefloor is divided into two principle environments: A deep, hypolimnetic and a shallow, epilimnetic lakefloor. In the deeper hypolimnetic lakefloor halite continuously precipitates with seasonal variations: (a) during summer, consolidated coarse halite crystals form rough surfaces under slight super-saturation. (b) During winter, unconsolidated, fine halite crystals form smooth seafloor deposits under high supersaturation. The observations also emphasize the thought regarding seasonal alternation of halite crystallization mechanism. The shallow epilimnetic lake floor is highly influenced by the seasonal temperature variations, and by intensive summer dissolution of part of the previous year's halite deposit which results in thin sequences with annual unconformities. This emphasizes the control of temperature seasonality on the precipitated halite layers characteristics. In addition, precipitation of halite in the hypolimnetic floor, on the expense of the dissolution of the epilimnetic floor, results in lateral focusing and thickening of halite deposit in the deeper part of the basin and thinning of the deposits in shallow marginal basins.
Schadt, Eric E.; Banerjee, Onureena; Fang, Gang; Feng, Zhixing; Wong, Wing H.; Zhang, Xuegong; Kislyuk, Andrey; Clark, Tyson A.; Luong, Khai; Keren-Paz, Alona; Chess, Andrew; Kumar, Vipin; Chen-Plotkin, Alice; Sondheimer, Neal; Korlach, Jonas; Kasarskis, Andrew
2013-01-01
Current generation DNA sequencing instruments are moving closer to seamlessly sequencing genomes of entire populations as a routine part of scientific investigation. However, while significant inroads have been made identifying small nucleotide variation and structural variations in DNA that impact phenotypes of interest, progress has not been as dramatic regarding epigenetic changes and base-level damage to DNA, largely due to technological limitations in assaying all known and unknown types of modifications at genome scale. Recently, single-molecule real time (SMRT) sequencing has been reported to identify kinetic variation (KV) events that have been demonstrated to reflect epigenetic changes of every known type, providing a path forward for detecting base modifications as a routine part of sequencing. However, to date no statistical framework has been proposed to enhance the power to detect these events while also controlling for false-positive events. By modeling enzyme kinetics in the neighborhood of an arbitrary location in a genomic region of interest as a conditional random field, we provide a statistical framework for incorporating kinetic information at a test position of interest as well as at neighboring sites that help enhance the power to detect KV events. The performance of this and related models is explored, with the best-performing model applied to plasmid DNA isolated from Escherichia coli and mitochondrial DNA isolated from human brain tissue. We highlight widespread kinetic variation events, some of which strongly associate with known modification events, while others represent putative chemically modified sites of unknown types. PMID:23093720
Schadt, Eric E; Banerjee, Onureena; Fang, Gang; Feng, Zhixing; Wong, Wing H; Zhang, Xuegong; Kislyuk, Andrey; Clark, Tyson A; Luong, Khai; Keren-Paz, Alona; Chess, Andrew; Kumar, Vipin; Chen-Plotkin, Alice; Sondheimer, Neal; Korlach, Jonas; Kasarskis, Andrew
2013-01-01
Current generation DNA sequencing instruments are moving closer to seamlessly sequencing genomes of entire populations as a routine part of scientific investigation. However, while significant inroads have been made identifying small nucleotide variation and structural variations in DNA that impact phenotypes of interest, progress has not been as dramatic regarding epigenetic changes and base-level damage to DNA, largely due to technological limitations in assaying all known and unknown types of modifications at genome scale. Recently, single-molecule real time (SMRT) sequencing has been reported to identify kinetic variation (KV) events that have been demonstrated to reflect epigenetic changes of every known type, providing a path forward for detecting base modifications as a routine part of sequencing. However, to date no statistical framework has been proposed to enhance the power to detect these events while also controlling for false-positive events. By modeling enzyme kinetics in the neighborhood of an arbitrary location in a genomic region of interest as a conditional random field, we provide a statistical framework for incorporating kinetic information at a test position of interest as well as at neighboring sites that help enhance the power to detect KV events. The performance of this and related models is explored, with the best-performing model applied to plasmid DNA isolated from Escherichia coli and mitochondrial DNA isolated from human brain tissue. We highlight widespread kinetic variation events, some of which strongly associate with known modification events, while others represent putative chemically modified sites of unknown types.
DNA barcoding for species identification in deep-sea clams (Mollusca: Bivalvia: Vesicomyidae).
Liu, Jun; Zhang, Haibin
2018-01-15
Deep-sea clams (Bivalvia: Vesicomyidae) have been found in reduced environments over the world oceans, but taxonomy of this group remains confusing at species and supraspecific levels due to their high-morphological similarity and plasticity. In the present study, we collected mitochondrial COI sequences to evaluate the utility of DNA barcoding on identifying vesicomyid species. COI dataset identified 56 well-supported putative species/operational taxonomic units (OTUs), approximately covering half of the extant vesicomyid species. One species (OTU2) was first detected, and may represent a new species. Average distances between species ranged from 1.65 to 29.64%, generally higher than average intraspecific distances (0-1.41%) when excluding Pliocardia sp.10 cf. venusta (average intraspecific distance 1.91%). Local barcoding gap existed in 33 of the 35 species when comparing distances of maximum interspecific and minimum interspecific distances with two exceptions (Abyssogena southwardae and Calyptogena rectimargo-starobogatovi). The barcode index number (BIN) system determined 41 of the 56 species/OTUs, each with a unique BIN, indicating their validity. Three species were found to have two BINs, together with their high level of intraspecific variation, implying cryptic diversity within them. Although fewer 16 S sequences were collected, similar results were obtained. Nineteen putative species were determined and no overlap observed between intra- and inter-specific variation. Implications of DNA barcoding for the Vesicomyidae taxonomy were then discussed. Findings of this study will provide important evidence for taxonomic revision in this problematic clam group, and accelerate the discovery of new vesicomyid species in the future.
Eraso, Jesus M.; Olsen, Randall J.; Beres, Stephen B.; Kachroo, Priyanka; Porter, Adeline R.; Nasser, Waleed; Bernard, Paul E.; DeLeo, Frank R.
2016-01-01
To obtain new information about Streptococcus pyogenes intrahost genetic variation during invasive infection, we sequenced the genomes of 2,954 serotype M1 strains recovered from a nonhuman primate experimental model of necrotizing fasciitis. A total of 644 strains (21.8%) acquired polymorphisms relative to the input parental strain. The fabT gene, encoding a transcriptional regulator of fatty acid biosynthesis genes, contained 54.5% of these changes. The great majority of polymorphisms were predicted to deleteriously alter FabT function. Transcriptome-sequencing (RNA-seq) analysis of a wild-type strain and an isogenic fabT deletion mutant strain found that between 3.7 and 28.5% of the S. pyogenes transcripts were differentially expressed, depending on the growth temperature (35°C or 40°C) and growth phase (mid-exponential or stationary phase). Genes implicated in fatty acid synthesis and lipid metabolism were significantly upregulated in the fabT deletion mutant strain. FabT also directly or indirectly regulated central carbon metabolism genes, including pyruvate hub enzymes and fermentation pathways and virulence genes. Deletion of fabT decreased virulence in a nonhuman primate model of necrotizing fasciitis. In addition, the fabT deletion strain had significantly decreased survival in human whole blood and during phagocytic interaction with polymorphonuclear leukocytes ex vivo. We conclude that FabT mutant progeny arise during infection, constitute a metabolically distinct subpopulation, and are less virulent in the experimental models used here. PMID:27600505
Host genetic variation in mucosal immunity pathways influences the upper airway microbiome.
Igartua, Catherine; Davenport, Emily R; Gilad, Yoav; Nicolae, Dan L; Pinto, Jayant; Ober, Carole
2017-02-01
The degree to which host genetic variation can modulate microbial communities in humans remains an open question. Here, we performed a genetic mapping study of the microbiome in two accessible upper airway sites, the nasopharynx and the nasal vestibule, during two seasons in 144 adult members of a founder population of European decent. We estimated the relative abundances (RAs) of genus level bacteria from 16S rRNA gene sequences and examined associations with 148,653 genetic variants (linkage disequilibrium [LD] r 2 < 0.5) selected from among all common variants discovered in genome sequences in this population. We identified 37 microbiome quantitative trait loci (mbQTLs) that showed evidence of association with the RAs of 22 genera (q < 0.05) and were enriched for genes in mucosal immunity pathways. The most significant association was between the RA of Dermacoccus (phylum Actinobacteria) and a variant 8 kb upstream of TINCR (rs117042385; p = 1.61 × 10 -8 ; q = 0.002), a long non-coding RNA that binds to peptidoglycan recognition protein 3 (PGLYRP3) mRNA, a gene encoding a known antimicrobial protein. A second association was between a missense variant in PGLYRP4 (rs3006458) and the RA of an unclassified genus of family Micrococcaceae (phylum Actinobacteria) (p = 5.10 × 10 -7 ; q = 0.032). Our findings provide evidence of host genetic influences on upper airway microbial composition in humans and implicate mucosal immunity genes in this relationship.
Sudha, Dhandayuthapani; Patric, Irene Rosita Pia; Ganapathy, Aparna; Agarwal, Smitha; Krishna, Shuba; Neriyanuri, Srividya; Sripriya, Sarangapani; Sen, Parveen; Chidambaram, Subbulakshmi; Arunachalam, Jayamuruga Pandian
2017-01-01
In this study, we present a juvenile retinoschisis patient with developmental delay, sensorineural hearing loss, and reduced axial tone. X-linked juvenile retinoschisis (XLRS) is a retinal dystrophy, most often not associated with systemic anomalies and also not showing any locus heterogeneity. Therefore it was of interest to understand the genetic basis of the condition in this patient. RS1 gene screening for XLRS was performed by Sanger sequencing. Whole genome SNP 6.0 array analysis was carried out to investigate gross chromosomal aberrations that could result in systemic phenotype. In addition, targeted next generation sequencing (NGS) was employed to determine any possible involvement of X-linked syndromic and non-syndromic mental retardation genes. This NGS panel consisted of 550 genes implicated in several other rare inherited diseases. RS1 gene screening revealed a pathogenic hemizygous splice site mutation (c.78+1G>T), inherited from the mother. SNP 6.0 array analysis did not indicate any significant chromosomal aberrations that could be disease-associated. Targeted resequencing did not identify any mutations in the X-linked mental retardation genes. However, variations in three other genes (NSD1, LARGE, and POLG) were detected, which were all inherited from the patient's unaffected father. Taken together, RS1 mutation was found to segregate with retinoschisis phenotype while none of the other identified variations were co-segregating with the systemic defects. Hereby, we infer that the multisystemic defects harbored by the patient are a rare coexistence of XLRS, developmental delay, sensorineural hearing loss, and reduced axial tone reported for the first time in the literature.
Artificial mismatch hybridization
Guo, Zhen; Smith, Lloyd M.
1998-01-01
An improved nucleic acid hybridization process is provided which employs a modified oligonucleotide and improves the ability to discriminate a control nucleic acid target from a variant nucleic acid target containing a sequence variation. The modified probe contains at least one artificial mismatch relative to the control nucleic acid target in addition to any mismatch(es) arising from the sequence variation. The invention has direct and advantageous application to numerous existing hybridization methods, including, applications that employ, for example, the Polymerase Chain Reaction, allele-specific nucleic acid sequencing methods, and diagnostic hybridization methods.
Tandemly repeated sequences in mtDNA control region of whitefish, Coregonus lavaretus.
Brzuzan, P
2000-06-01
Length variation of the mitochondrial DNA control region was observed with PCR amplification of a sample of 138 whitefish (Coregonus lavaretus). Nucleotide sequences of representative PCR products showed that the variation was due to the presence of an approximately 100-bp motif tandemly repeated two, three, or five times in the region between the conserved sequence block-3 (CSB-3) and the gene for phenylalanine tRNA. This is the first report on the tandem array composed of long repeat units in mitochondrial DNA of salmonids.
Genetic and Epigenetic Variations Induced by Wheat-Rye 2R and 5R Monosomic Addition Lines
Fu, Shulan; Sun, Chuanfei; Yang, Manyu; Fei, Yunyan; Tan, Feiqun; Yan, Benju; Ren, Zhenglong; Tang, Zongxiang
2013-01-01
Background Monosomic alien addition lines (MAALs) can easily induce structural variation of chromosomes and have been used in crop breeding; however, it is unclear whether MAALs will induce drastic genetic and epigenetic alterations. Methodology/Principal Findings In the present study, wheat-rye 2R and 5R MAALs together with their selfed progeny and parental common wheat were investigated through amplified fragment length polymorphism (AFLP) and methylation-sensitive amplification polymorphism (MSAP) analyses. The MAALs in different generations displayed different genetic variations. Some progeny that only contained 42 wheat chromosomes showed great genetic/epigenetic alterations. Cryptic rye chromatin has introgressed into the wheat genome. However, one of the progeny that contained cryptic rye chromatin did not display outstanding genetic/epigenetic variation. 78 and 49 sequences were cloned from changed AFLP and MSAP bands, respectively. Blastn search indicated that almost half of them showed no significant similarity to known sequences. Retrotransposons were mainly involved in genetic and epigenetic variations. Genetic variations basically affected Gypsy-like retrotransposons, whereas epigenetic alterations affected Copia-like and Gypsy-like retrotransposons equally. Genetic and epigenetic variations seldom affected low-copy coding DNA sequences. Conclusions/Significance The results in the present study provided direct evidence to illustrate that monosomic wheat-rye addition lines could induce different and drastic genetic/epigenetic variations and these variations might not be caused by introgression of rye chromatins into wheat. Therefore, MAALs may be directly used as an effective means to broaden the genetic diversity of common wheat. PMID:23342073
Genetic and epigenetic variations induced by wheat-rye 2R and 5R monosomic addition lines.
Fu, Shulan; Sun, Chuanfei; Yang, Manyu; Fei, Yunyan; Tan, Feiqun; Yan, Benju; Ren, Zhenglong; Tang, Zongxiang
2013-01-01
Monosomic alien addition lines (MAALs) can easily induce structural variation of chromosomes and have been used in crop breeding; however, it is unclear whether MAALs will induce drastic genetic and epigenetic alterations. In the present study, wheat-rye 2R and 5R MAALs together with their selfed progeny and parental common wheat were investigated through amplified fragment length polymorphism (AFLP) and methylation-sensitive amplification polymorphism (MSAP) analyses. The MAALs in different generations displayed different genetic variations. Some progeny that only contained 42 wheat chromosomes showed great genetic/epigenetic alterations. Cryptic rye chromatin has introgressed into the wheat genome. However, one of the progeny that contained cryptic rye chromatin did not display outstanding genetic/epigenetic variation. 78 and 49 sequences were cloned from changed AFLP and MSAP bands, respectively. Blastn search indicated that almost half of them showed no significant similarity to known sequences. Retrotransposons were mainly involved in genetic and epigenetic variations. Genetic variations basically affected Gypsy-like retrotransposons, whereas epigenetic alterations affected Copia-like and Gypsy-like retrotransposons equally. Genetic and epigenetic variations seldom affected low-copy coding DNA sequences. The results in the present study provided direct evidence to illustrate that monosomic wheat-rye addition lines could induce different and drastic genetic/epigenetic variations and these variations might not be caused by introgression of rye chromatins into wheat. Therefore, MAALs may be directly used as an effective means to broaden the genetic diversity of common wheat.
Harrisson, Katherine A; Amish, Stephen J; Pavlova, Alexandra; Narum, Shawn R; Telonis-Scott, Marina; Rourke, Meaghan L; Lyon, Jarod; Tonkin, Zeb; Gilligan, Dean M; Ingram, Brett A; Lintermans, Mark; Gan, Han Ming; Austin, Christopher M; Luikart, Gordon; Sunnucks, Paul
2017-11-01
Adaptive differences across species' ranges can have important implications for population persistence and conservation management decisions. Despite advances in genomic technologies, detecting adaptive variation in natural populations remains challenging. Key challenges in gene-environment association studies involve distinguishing the effects of drift from those of selection and identifying subtle signatures of polygenic adaptation. We used paired-end restriction site-associated DNA sequencing data (6,605 biallelic single nucleotide polymorphisms; SNPs) to examine population structure and test for signatures of adaptation across the geographic range of an iconic Australian endemic freshwater fish species, the Murray cod Maccullochella peelii. Two univariate gene-association methods identified 61 genomic regions associated with climate variation. We also tested for subtle signatures of polygenic adaptation using a multivariate method (redundancy analysis; RDA). The RDA analysis suggested that climate (temperature- and precipitation-related variables) and geography had similar magnitudes of effect in shaping the distribution of SNP genotypes across the sampled range of Murray cod. Although there was poor agreement among the candidate SNPs identified by the univariate methods, the top 5% of SNPs contributing to significant RDA axes included 67% of the SNPs identified by univariate methods. We discuss the potential implications of our findings for the management of Murray cod and other species generally, particularly in relation to informing conservation actions such as translocations to improve evolutionary resilience of natural populations. Our results highlight the value of using a combination of different approaches, including polygenic methods, when testing for signatures of adaptation in landscape genomic studies. © 2017 John Wiley & Sons Ltd.
Is Homo sapiens polytypic? Human taxonomic diversity and its implications.
Woodley, Michael A
2010-01-01
The term race is a traditional synonym for subspecies, however it is frequently asserted that Homo sapiens is monotypic and that what are termed races are nothing more than biological illusions. In this manuscript a case is made for the hypothesis that H. sapiens is polytypic, and in this way is no different from other species exhibiting similar levels of genetic and morphological diversity. First it is demonstrated that the four major definitions of race/subspecies can be shown to be synonymous within the context of the framework of race as a correlation structure of traits. Next the issue of taxonomic classification is considered where it is demonstrated that H. sapiens possesses high levels morphological diversity, genetic heterozygosity and differentiation (F(ST)) compared to many species that are acknowledged to be polytypic with respect to subspecies. Racial variation is then evaluated in light of the phylogenetic species concept, where it is suggested that the least inclusive monophyletic units exist below the level of species within H. sapiens indicating the existence of a number of potential human phylogenetic species; and the biological species concept, where it is determined that racial variation is too small to represent differentiation at the level of biological species. Finally the implications of this are discussed in the context of anthropology where an accurate picture of the sequence and timing of events during the evolution of human taxa are required for a complete picture of human evolution, and medicine, where a greater appreciation of the role played by human taxonomic differences in disease susceptibility and treatment responsiveness will save lives in the future.
NASA Astrophysics Data System (ADS)
Antoine, Pierre; Rousseau, Denis-Didier; Degeai, Jean-Philippe; Moine, Olivier; Lagroix, France; kreutzer, Sebastian; Fuchs, Markus; Hatté, Christine; Gauthier, Caroline; Svoboda, Jiri; Lisá, Lenka
2013-05-01
High-resolution multidisciplinary investigation of key European loess-palaeosols profiles have demonstrated that loess sequences result from rapid and cyclic aeolian sedimentation which is reflected in variations of loess grain size indexes and correlated with Greenland ice-core dust records. This correlation suggests a global connection between North Atlantic and west-European air masses. Herein, we present a revised stratigraphy and a continuous high-resolution record of grain-size, magnetic susceptibility and organic carbon δ13C of the famous of Dolní Vestonice (DV) loess sequence in the Moravian region of the Czech Republic. A new set of quartz OSL ages provides a reliable and accurate chronology of the sequence's main pedosedimentary events. The grain size record shows strongly contrasting variations with numerous abrupt coarse-grained events, especially in the upper part of the sequence between ca 20-30 ka. This time period is also characterised by a progressive coarsening of the loess deposits as already observed in other western European sequences. The base of the DV sequence exhibits an exceptionally well-preserved soil complex composed of three chernozem soil horizons and 5 aeolian silt layers (marker silts). This complex is, at present, the most complete record of environmental variations and dust deposition in the European loess belt for the Weichselian Early-glacial period spanning about 110 to 70 ka, allowing correlations with various global palaeoclimatic records. OSL ages combined with sedimentological and palaeopedological observations lead to the conclusion that this soil complex recorded all of the main climatic events expressed in the North GRIP record from Greenland Interstadials (GIS) 25 to 19.
A Laboratory Exercise for Genotyping Two Human Single Nucleotide Polymorphisms
ERIC Educational Resources Information Center
Fernando, James; Carlson, Bradley; LeBard, Timothy; McCarthy, Michael; Umali, Finianne; Ashton, Bryce; Rose, Ferrill F., Jr.
2016-01-01
The dramatic decrease in the cost of sequencing a human genome is leading to an era in which a wide range of students will benefit from having an understanding of human genetic variation. Since over 90% of sequence variation between humans is in the form of single nucleotide polymorphisms (SNPs), a laboratory exercise has been devised in order to…
RSAT 2018: regulatory sequence analysis tools 20th anniversary.
Nguyen, Nga Thi Thuy; Contreras-Moreira, Bruno; Castro-Mondragon, Jaime A; Santana-Garcia, Walter; Ossio, Raul; Robles-Espinoza, Carla Daniela; Bahin, Mathieu; Collombet, Samuel; Vincens, Pierre; Thieffry, Denis; van Helden, Jacques; Medina-Rivera, Alejandra; Thomas-Chollier, Morgane
2018-05-02
RSAT (Regulatory Sequence Analysis Tools) is a suite of modular tools for the detection and the analysis of cis-regulatory elements in genome sequences. Its main applications are (i) motif discovery, including from genome-wide datasets like ChIP-seq/ATAC-seq, (ii) motif scanning, (iii) motif analysis (quality assessment, comparisons and clustering), (iv) analysis of regulatory variations, (v) comparative genomics. Six public servers jointly support 10 000 genomes from all kingdoms. Six novel or refactored programs have been added since the 2015 NAR Web Software Issue, including updated programs to analyse regulatory variants (retrieve-variation-seq, variation-scan, convert-variations), along with tools to extract sequences from a list of coordinates (retrieve-seq-bed), to select motifs from motif collections (retrieve-matrix), and to extract orthologs based on Ensembl Compara (get-orthologs-compara). Three use cases illustrate the integration of new and refactored tools to the suite. This Anniversary update gives a 20-year perspective on the software suite. RSAT is well-documented and available through Web sites, SOAP/WSDL (Simple Object Access Protocol/Web Services Description Language) web services, virtual machines and stand-alone programs at http://www.rsat.eu/.
Sequence-length variation of mtDNA HVS-I C-stretch in Chinese ethnic groups.
Chen, Feng; Dang, Yong-hui; Yan, Chun-xia; Liu, Yan-ling; Deng, Ya-jun; Fulton, David J R; Chen, Teng
2009-10-01
The purpose of this study was to investigate mitochondrial DNA (mtDNA) hypervariable segment-I (HVS-I) C-stretch variations and explore the significance of these variations in forensic and population genetics studies. The C-stretch sequence variation was studied in 919 unrelated individuals from 8 Chinese ethnic groups using both direct and clone sequencing approaches. Thirty eight C-stretch haplotypes were identified, and some novel and population specific haplotypes were also detected. The C-stretch genetic diversity (GD) values were relatively high, and probability (P) values were low. Additionally, C-stretch length heteroplasmy was observed in approximately 9% of individuals studied. There was a significant correlation (r=-0.961, P<0.01) between the expansion of the cytosine sequence length in the C-stretch of HVS-I and a reduction in the number of upstream adenines. These results indicate that the C-stretch could be a useful genetic maker in forensic identification of Chinese populations. The results from the Fst and dA genetic distance matrix, neighbor-joining tree, and principal component map also suggest that C-stretch could be used as a reliable genetic marker in population genetics.
VARiD: a variation detection framework for color-space and letter-space platforms.
Dalca, Adrian V; Rumble, Stephen M; Levy, Samuel; Brudno, Michael
2010-06-15
High-throughput sequencing (HTS) technologies are transforming the study of genomic variation. The various HTS technologies have different sequencing biases and error rates, and while most HTS technologies sequence the residues of the genome directly, generating base calls for each position, the Applied Biosystem's SOLiD platform generates dibase-coded (color space) sequences. While combining data from the various platforms should increase the accuracy of variation detection, to date there are only a few tools that can identify variants from color space data, and none that can analyze color space and regular (letter space) data together. We present VARiD--a probabilistic method for variation detection from both letter- and color-space reads simultaneously. VARiD is based on a hidden Markov model and uses the forward-backward algorithm to accurately identify heterozygous, homozygous and tri-allelic SNPs, as well as micro-indels. Our analysis shows that VARiD performs better than the AB SOLiD toolset at detecting variants from color-space data alone, and improves the calls dramatically when letter- and color-space reads are combined. The toolset is freely available at http://compbio.cs.utoronto.ca/varid.
McRobie, Helen R; King, Linda M; Fanutti, Cristina; Coussons, Peter J; Moncrief, Nancy D; Thomas, Alison P M
2014-01-01
Sequence variations in the melanocortin 1 receptor (MC1R) gene are associated with melanism in many different species of mammals, birds, and reptiles. The gray squirrel (Sciurus carolinensis), found in the British Isles, was introduced from North America in the late 19th century. Melanism in the British gray squirrel is associated with a 24-bp deletion in the MC1R. To investigate the origin of this mutation, we sequenced the MC1R of 95 individuals including 44 melanic gray squirrels from both the British Isles and North America. Melanic gray squirrels of both populations had the same 24-bp deletion associated with melanism. Given the significant deletion associated with melanism in the gray squirrel, we sequenced the MC1R of both wild-type and melanic fox squirrels (Sciurus niger) (9 individuals) and red squirrels (Sciurus vulgaris) (39 individuals). Unlike the gray squirrel, no association between sequence variation in the MC1R and melanism was found in these 2 species. We conclude that the melanic gray squirrel found in the British Isles originated from one or more introductions of melanic gray squirrels from North America. We also conclude that variations in the MC1R are not associated with melanism in the fox and red squirrels.
NASA Astrophysics Data System (ADS)
Brown, A. G.; Basell, L. S.; Toms, P. S.
2015-05-01
The current model of mid-latitude late Quaternary terrace sequences, is that they are uplift-driven but climatically controlled terrace staircases, relating to both regional-scale crustal and tectonic factors, and palaeohydrological variations forced by quasi-cyclic climatic conditions in the 100 K world (post Mid Pleistocene Transition). This model appears to hold for the majority of the river valleys draining into the English Channel which exhibit 8-15 terrace levels over approximately 60-100 m of altitudinal elevation. However, one valley, the Axe, has only one major morphological terrace and has long-been regarded as anomalous. This paper uses both conventional and novel stratigraphical methods (digital granulometry and terrestrial laser scanning) to show that this terrace is a stacked sedimentary sequence of 20-30 m thickness with a quasi-continuous (i.e. with hiatuses) pulsed, record of fluvial and periglacial sedimentation over at least the last 300-400 K yrs as determined principally by OSL dating of the upper two thirds of the sequence. Since uplift has been regional, there is no evidence of anomalous neotectonics, and climatic history must be comparable to the adjacent catchments (both of which have staircase sequences) a catchment-specific mechanism is required. The Axe is the only valley in North West Europe incised entirely into the near-horizontally bedded chert (crypto-crystalline quartz) and sand-rich Lower Cretaceous rocks creating a buried valley. Mapping of the valley slopes has identified many large landslide scars associated with past and present springs. It is proposed that these are thaw-slump scars and represent large hill-slope failures caused by Vauclausian water pressures and hydraulic fracturing of the chert during rapid permafrost melting. A simple 1D model of this thermokarstic process is used to explore this mechanism, and it is proposed that the resultant anomalously high input of chert and sand into the valley during terminations caused pulsed aggradation until the last termination. It is also proposed that interglacial and interstadial incision may have been prevented by the over-sized and interlocking nature of the sub-angular chert clasts until the Lateglacial when confinement of the river overcame this immobility threshold. One result of this hydrogeologically mediated valley evolution was to provide a sequence of proximal Palaeolithic archaeology over two MIS cycles. This study demonstrates that uplift tectonics and climate alone do not fully determine Quaternary valley evolution and that lithological and hydrogeological conditions are a fundamental cause of variation in terrestrial Quaternary records and landform evolution.
NASA Astrophysics Data System (ADS)
Qiang, Xiaoke; Xu, Xinwen; Zhao, Hui; Fu, Chaofeng
2018-05-01
The ferrimagnetic iron sulfide greigite (Fe3S4) occurs widely in sulfidic lacustrine and marine sedimentary environments. Knowledge of its formation and persistence is important for both magnetostratigraphic and paleoenvironmental studies. Although the formation mechanism of greigite has been widely demonstrated, the sedimentary environments associated with greigite formation in lakes, especially on relatively long timescales, are poorly understood. A long and continuous sequence of Pleistocene lacustrine sediments was recovered in the Heqing drill core from southwestern China, which provides an outstanding record of continental climate and environment. Integrated magnetic, geochemical, and paleoclimatic analysis of the lacustrine sequence provides an opportunity to improve our understanding of the environmental controls on greigite formation. Rock magnetic and scanning electron microscope analyses of selected samples from the core reveal that greigite is present in the lower part of the core (part 1, 665.8-372.5 m). Greigite occurs throughout this interval and is the dominant magnetic mineral, irrespective of the climatic state. The magnetic susceptibility (χ) record, which is mainly controlled by the concentration of greigite, matches well with variations in the Indian Summer Monsoon (ISM) index and total organic carbon (TOC) content, with no significant time lag. This indicates that the greigite formed during early diagenesis. In greigite-bearing intervals, with the χ increase, Bc value increase and tends to be stable at about 50 mT. Therefore, we suggest that χ values could estimate the variation of greigite concentration approximately in the Heqing core. Greigite favored more abundant in terrigenous-rich and organic-poor layers associated with weak summer monsoon which are characterized by high χ values, high Fe content, high Rb/Sr ratio and low TOC content. Greigite enhancement can be explained by variations in terrigenous inputs. Our studies demonstrate that, not only the greigite formation, but also its concentration changes could be useful for studying climatic and environmental variability in sulfidic environments.
Karan, Ratna; DeLeon, Teresa; Biradar, Hanamareddy; Subudhi, Prasanta K.
2012-01-01
Background Salinity is a major environmental factor limiting productivity of crop plants including rice in which wide range of natural variability exists. Although recent evidences implicate epigenetic mechanisms for modulating the gene expression in plants under environmental stresses, epigenetic changes and their functional consequences under salinity stress in rice are underexplored. DNA methylation is one of the epigenetic mechanisms regulating gene expression in plant’s responses to environmental stresses. Better understanding of epigenetic regulation of plant growth and response to environmental stresses may create novel heritable variation for crop improvement. Methodology/Principal Findings Methylation sensitive amplification polymorphism (MSAP) technique was used to assess the effect of salt stress on extent and patterns of DNA methylation in four genotypes of rice differing in the degree of salinity tolerance. Overall, the amount of DNA methylation was more in shoot compared to root and the contribution of fully methylated loci was always more than hemi-methylated loci. Sequencing of ten randomly selected MSAP fragments indicated gene-body specific DNA methylation of retrotransposons, stress responsive genes, and chromatin modification genes, distributed on different rice chromosomes. Bisulphite sequencing and quantitative RT-PCR analysis of selected MSAP loci showed that cytosine methylation changes under salinity as well as gene expression varied with genotypes and tissue types irrespective of the level of salinity tolerance of rice genotypes. Conclusions/Significance The gene body methylation may have an important role in regulating gene expression in organ and genotype specific manner under salinity stress. Association between salt tolerance and methylation changes observed in some cases suggested that many methylation changes are not “directed”. The natural genetic variation for salt tolerance observed in rice germplasm may be independent of the extent and pattern of DNA methylation which may have been induced by abiotic stress followed by accumulation through the natural selection process. PMID:22761959
Elbers, Jean P; Brown, Mary B; Taylor, Sabrina S
2018-01-19
Infectious disease is the single greatest threat to taxa such as amphibians (chytrid fungus), bats (white nose syndrome), Tasmanian devils (devil facial tumor disease), and black-footed ferrets (canine distemper virus, plague). Although understanding the genetic basis to disease susceptibility is important for the long-term persistence of these groups, most research has been limited to major-histocompatibility and Toll-like receptor genes. To better understand the genetic basis of infectious disease susceptibility in a species of conservation concern, we sequenced all known/predicted immune response genes (i.e., the immunomes) in 16 Florida gopher tortoises, Gopherus polyphemus. All tortoises produced antibodies against Mycoplasma agassizii (an etiologic agent of infectious upper respiratory tract disease; URTD) and, at the time of sampling, either had (n = 10) or lacked (n = 6) clinical signs. We found several variants associated with URTD clinical status in complement and lectin genes, which may play a role in Mycoplasma immunity. Thirty-five genes deviated from neutrality according to Tajima's D. These genes were enriched in functions relating to macromolecule and protein modifications, which are vital to immune system functioning. These results are suggestive of genetic differences that might contribute to disease severity, a finding that is consistent with other mycoplasmal diseases. This has implications for management because tortoises across their range may possess genetic variation associated with a more severe response to URTD. More generally: 1) this approach demonstrates that a broader consideration of immune genes is better able to identify important variants, and; 2) this data pipeline can be adopted to identify alleles associated with disease susceptibility or resistance in other taxa, and therefore provide information on a population's risk of succumbing to disease, inform translocations to increase genetic variation for disease resistance, and help to identify potential treatments.
Lefty Blocks a Subset of TGFβ Signals by Antagonizing EGF-CFC Coreceptors
Cheng, Simon K; Olale, Felix; Brivanlou, Ali H
2004-01-01
Members of the EGF-CFC family play essential roles in embryonic development and have been implicated in tumorigenesis. The TGFβ signals Nodal and Vg1/GDF1, but not Activin, require EGF-CFC coreceptors to activate Activin receptors. We report that the TGFβ signaling antagonist Lefty also acts through an EGF-CFC-dependent mechanism. Lefty inhibits Nodal and Vg1 signaling, but not Activin signaling. Lefty genetically interacts with EGF-CFC proteins and competes with Nodal for binding to these coreceptors. Chimeras between Activin and Nodal or Vg1 identify a 14 amino acid region that confers independence from EGF-CFC coreceptors and resistance to Lefty. These results indicate that coreceptors are targets for both TGFβ agonists and antagonists and suggest that subtle sequence variations in TGFβ signals result in greater ligand diversity. PMID:14966532
Giovannetti, Elisa; Ugrasena, Dewa G; Supriyadi, Eddy; Vroling, Laura; Azzarello, Antonino; de Lange, Desiree; Peters, Godefridus J; Veerman, Anjo J P; Cloos, Jacqueline
2008-01-01
Genetic variations in the polymorphic tandem repeat sequence of the enhancer region of the thymidylate synthase promoter (TSER), as well as in methylenetetrahydrofolate reductase (MTHFR) C677T polymorphism, influence methotrexate sensitivity. We studied these polymorphisms in children with acute lymphoblastic leukaemia (ALL) and in subjects without malignancy in Indonesia and Holland. The frequencies of TT and CT genotypes were two-fold higher in Dutch children. The TSER 3R/3R repeat was three-fold more frequent in the Indonesian children, while the 2R/2R repeat was only 1% compared to 21% in the Dutch children. No differences of these polymorphisms were found between ALL cells and normal blood cells, indicating an ethnic rather than leukemic origin. These results may have implications for treatment of Indonesian children with ALL.
NASA Astrophysics Data System (ADS)
Frau, Camille; Bulot, Luc G.; Wimbledon, William A. P.; Ifrim, Christina
2016-06-01
This study describes ammonite taxa of the Perisphinctoidea in the Jurassic/Cretaceous boundary interval at Le Chouet (Drôme, France). Emphasis is placed on new and poorly known Himalayitidae, Neocomitidae and Olcostephanidae from the lower part of the Jacobi Zone auctorum. Significant results relate the introduction of Lopeziceras gen. nov., grouping himalayitid-like forms with two rows of tubercles, and Praedalmasiceras gen. nov., grouping the early Berriasian Dalmasiceras taxa. Study of the ontogenetic sequences of both genera show that they were derived from late Tithonian Himalayitidae. This supports the distinction between the subfamilies Himalayitinae and Dalmasiceratidae subfam. nov. Content, variation, dimorphism and vertical range of the Neocomitidae Berriasella, Pseudoneocomites, Elenaella and Delphinella are discussed. A conservative use of the Olcostephanidae Proniceras is followed herein.
What Advances Are Being Made in DNA Sequencing?
... to identify genetic variations; both methods rely on new technologies that allow rapid sequencing of large amounts of ... describes the different sequencing technologies and what the new technologies have meant for the study of the genetic ...
Bashir, Ali; Bansal, Vikas; Bafna, Vineet
2010-06-18
Massively parallel DNA sequencing technologies have enabled the sequencing of several individual human genomes. These technologies are also being used in novel ways for mRNA expression profiling, genome-wide discovery of transcription-factor binding sites, small RNA discovery, etc. The multitude of sequencing platforms, each with their unique characteristics, pose a number of design challenges, regarding the technology to be used and the depth of sequencing required for a particular sequencing application. Here we describe a number of analytical and empirical results to address design questions for two applications: detection of structural variations from paired-end sequencing and estimating mRNA transcript abundance. For structural variation, our results provide explicit trade-offs between the detection and resolution of rearrangement breakpoints, and the optimal mix of paired-read insert lengths. Specifically, we prove that optimal detection and resolution of breakpoints is achieved using a mix of exactly two insert library lengths. Furthermore, we derive explicit formulae to determine these insert length combinations, enabling a 15% improvement in breakpoint detection at the same experimental cost. On empirical short read data, these predictions show good concordance with Illumina 200 bp and 2 Kbp insert length libraries. For transcriptome sequencing, we determine the sequencing depth needed to detect rare transcripts from a small pilot study. With only 1 Million reads, we derive corrections that enable almost perfect prediction of the underlying expression probability distribution, and use this to predict the sequencing depth required to detect low expressed genes with greater than 95% probability. Together, our results form a generic framework for many design considerations related to high-throughput sequencing. We provide software tools http://bix.ucsd.edu/projects/NGS-DesignTools to derive platform independent guidelines for designing sequencing experiments (amount of sequencing, choice of insert length, mix of libraries) for novel applications of next generation sequencing.
Wang, Daxi; Korhonen, Pasi K; Gasser, Robin B; Young, Neil D
Clonorchis sinensis (family Opisthorchiidae) is an important foodborne parasite that has a major socioeconomic impact on ~35 million people predominantly in China, Vietnam, Korea and the Russian Far East. In humans, infection with C. sinensis causes clonorchiasis, a complex hepatobiliary disease that can induce cholangiocarcinoma (CCA), a malignant cancer of the bile ducts. Central to understanding the epidemiology of this disease is knowledge of genetic variation within and among populations of this parasite. Although most published molecular studies seem to suggest that C. sinensis represents a single species, evidence of karyotypic variation within C. sinensis and cryptic species within a related opisthorchiid fluke (Opisthorchis viverrini) emphasise the importance of studying and comparing the genes and genomes of geographically distinct isolates of C. sinensis. Recently, we sequenced, assembled and characterised a draft nuclear genome of a C. sinensis isolate from Korea and compared it with a published draft genome of a Chinese isolate of this species using a bioinformatic workflow established for comparing draft genome assemblies and their gene annotations. We identified that 50.6% and 51.3% of the Korean and Chinese C. sinensis genomic scaffolds were syntenic, respectively. Within aligned syntenic blocks, the genomes had a high level of nucleotide identity (99.1%) and encoded 15 variable proteins likely to be involved in diverse biological processes. Here, we review current technical challenges of using draft genome assemblies to undertake comparative genomic analyses to quantify genetic variation between isolates of the same species. Using a workflow that overcomes these challenges, we report on a high-quality draft genome for C. sinensis from Korea and comparative genomic analyses, as a basis for future investigations of the genetic structures of C. sinensis populations, and discuss the biotechnological implications of these explorations. Copyright © 2018 Elsevier Inc. All rights reserved.
Vinner, Lasse; Mourier, Tobias; Friis-Nielsen, Jens; Gniadecki, Robert; Dybkaer, Karen; Rosenberg, Jacob; Langhoff, Jill Levin; Cruz, David Flores Santa; Fonager, Jannik; Izarzugaza, Jose M G; Gupta, Ramneek; Sicheritz-Ponten, Thomas; Brunak, Søren; Willerslev, Eske; Nielsen, Lars Peter; Hansen, Anders Johannes
2015-08-19
Although nearly one fifth of all human cancers have an infectious aetiology, the causes for the majority of cancers remain unexplained. Despite the enormous data output from high-throughput shotgun sequencing, viral DNA in a clinical sample typically constitutes a proportion of host DNA that is too small to be detected. Sequence variation among virus genomes complicates application of sequence-specific, and highly sensitive, PCR methods. Therefore, we aimed to develop and characterize a method that permits sensitive detection of sequences despite considerable variation. We demonstrate that our low-stringency in-solution hybridization method enables detection of <100 viral copies. Furthermore, distantly related proviral sequences may be enriched by orders of magnitude, enabling discovery of hitherto unknown viral sequences by high-throughput sequencing. The sensitivity was sufficient to detect retroviral sequences in clinical samples. We used this method to conduct an investigation for novel retrovirus in samples from three cancer types. In accordance with recent studies our investigation revealed no retroviral infections in human B-cell lymphoma cells, cutaneous T-cell lymphoma or colorectal cancer biopsies. Nonetheless, our generally applicable method makes sensitive detection possible and permits sequencing of distantly related sequences from complex material.
2013-01-01
Background Genetic variation at the melanocortin-1 receptor (MC1R) gene is correlated with melanin color variation in many birds. Feral pigeons (Columba livia) show two major melanin-based colorations: a red coloration due to pheomelanic pigment and a black coloration due to eumelanic pigment. Furthermore, within each color type, feral pigeons display continuous variation in the amount of melanin pigment present in the feathers, with individuals varying from pure white to a full dark melanic color. Coloration is highly heritable and it has been suggested that it is under natural or sexual selection, or both. Our objective was to investigate whether MC1R allelic variants are associated with plumage color in feral pigeons. Findings We sequenced 888 bp of the coding sequence of MC1R among pigeons varying both in the type, eumelanin or pheomelanin, and the amount of melanin in their feathers. We detected 10 non-synonymous substitutions and 2 synonymous substitution but none of them were associated with a plumage type. It remains possible that non-synonymous substitutions that influence coloration are present in the short MC1R fragment that we did not sequence but this seems unlikely because we analyzed the entire functionally important region of the gene. Conclusions Our results show that color differences among feral pigeons are probably not attributable to amino acid variation at the MC1R locus. Therefore, variation in regulatory regions of MC1R or variation in other genes may be responsible for the color polymorphism of feral pigeons. PMID:23915680
Genome Sequence of Stachybotrys chartarum Strain 51-11
Kim, Jean; Levy, Josh
2015-01-01
The Stachybotrys chartarum strain 51-11 genome was sequenced by shotgun sequencing utilizing Illumina HiSeq 2000 and PacBio technologies. Since S. chartarum has been implicated as having health impacts within water-damaged buildings, any information extracted from the genomic sequence data relating to toxins or the metabolism of the fungus might be useful. PMID:26430036
Gritsun, T S; Venugopal, K; Zanotto, P M; Mikhailov, M V; Sall, A A; Holmes, E C; Polkinghorne, I; Frolova, T V; Pogodina, V V; Lashkevich, V A; Gould, E A
1997-05-01
The complete nucleotide sequence of two tick-transmitted flaviviruses, Vasilchenko (Vs) from Siberia and louping ill (LI) from the UK, have been determined. The genomes were respectively, 10928 and 10871 nucleotides (nt) in length. The coding strategy and functional protein sequence motifs of tick-borne flaviviruses are presented in both Vs and LI viruses. The phylogenies based on maximum likelihood, maximum parsimony and distance analysis of the polyproteins, identified Vs virus as a member of the tick-borne encephalitis virus subgroup within the tick-borne serocomplex, genus Flavivirus, family Flaviviridae. Comparative alignment of the 3'-untranslated regions revealed deletions of different lengths essentially at the same position downstream of the stop codon for all tick-borne viruses. Two direct 27 nucleotide repeats at the 3'-end were found only for Vs and LI virus. Immediately following the deletions a region of 332-334 nt with relatively conserved primary structure (67-94% identity) was observed at the 3'-non-coding end of the virus genome. Pairwise comparisons of the nucleotide sequence data revealed similar levels of variation between the coding region, and the 5' and 3'-termini of the genome, implying an equivalent strong selective control for translated and untranslated regions. Indeed the predicted folding of the 5' and 3'-untranslated regions revealed patterns of stem and loop structures conserved for all tick-borne flaviviruses suggesting a purifying selection for preservation of essential RNA secondary structures which could be involved in translational control and replication. The possible implications of these findings are discussed.
Fleck, R.J.; Mattinson, J.M.; Busby, C.J.; Carr, M.D.; Davis, G.A.; Burchfiel, B.C.
1994-01-01
Combined U-Pb zircon, Rb-Sr, 40Ar/39Ar laser-fusion, and conventional K-Ar geochronology establish a late Early Cretaceous age for the Delfonte volcanic rocks. U-Pb zircon analyses define a lower intercept age of 100.5 ± 2 Ma that is interpreted as the crystallization age of the Delfonte sequence. Argon studies document both xenocrystic contamination and postemplacement Ar loss. Rb-Sr results from mafic lavas at the base of the sequence demonstrate compositionally correlated variations in initial 87Sr/86Sr ratios (Sri) from 0.706 for basalts to 0.716 for andesitic compositions. This covariation indicates substantial mixing of subcontinental lithosphere with Proterozoic upper crust. Correlations between Rb/Sr and Sri may result not only in pseudoisochrons approaching the age of the crustal component, but also in reasonable but incorrect apparent ages approaching the true age.Ages obtained in this study require that at least some of the thrust faulting in the Mescal Range-Clark Mountain portion of the foreland fold-and-thrust belt occurred later than ca. 100 Ma and was broadly contemporaneous with emplacement of the Keystone thrust plate in the Spring Mountains to the northeast. Comparison of the age and Rb-Sr systematics of ash-flow tuff boulders in the synorogenic Lavinia Wash sequence near Goodsprings, Nevada, with those of the Delfonte volcanic rocks supports a Delfonte source for the boulders. The 99 Ma age of the Lavinia Wash sequence is nearly identical to the Delfonte age, requiring rapid erosion, transport, and deposition following Delfonte volcanism.
Cuadrado, A; Cardoso, M; Jouve, N
2008-01-01
A significant fraction of the nuclear DNA of all eukaryotes is occupied by simple sequence repeats (SSRs) or microsatellites. This type of sequence has sparked great interest as a means of studying genetic variation, linkage mapping, gene tagging and evolution. Although SSRs at different positions in a gene help determine the regulation of expression and the function of the protein produced, little attention has been paid to the chromosomal organisation and distribution of these sequences, even in model species. This review discusses the main achievements in the characterisation of long-range SSR organisation in the chromosomes of Triticum aestivum L., Secale cereale L., and Hordeum vulgare L. (all members of Triticeae). We have detected SSRs using an improved FISH technique based on the random primer labelling of synthetic oligonucleotides (15-24 bases) in multi-colour experiments. Detailed information on the presence and distribution of AC, AG and all the possible classes of trinucleotide repeats has been acquired. These data have revealed the motif-dependent and non-random chromosome distributions of SSRs in the different genomes, and allowed the correlation of particular SSRs with chromosome areas characterised by specific features (e.g., heterochromatin, euchromatin and centromeres) in all three species. The present review provides a detailed comparative study of the distribution of these SSRs in each of the seven chromosomes of the genomes A, B and D of wheat, H of barley and R of rye. The importance of SSRs in plant breeding and their possible role in chromosome structure, function and evolution is discussed. 2008 S. Karger AG, Basel
New genetic variants of LATS1 detected in urinary bladder and colon cancer.
Saadeldin, Mona K; Shawer, Heba; Mostafa, Ahmed; Kassem, Neemat M; Amleh, Asma; Siam, Rania
2014-01-01
LATS1, the large tumor suppressor 1 gene, encodes for a serine/threonine kinase protein and is implicated in cell cycle progression. LATS1 is down-regulated in various human cancers, such as breast cancer, and astrocytoma. Point mutations in LATS1 were reported in human sarcomas. Additionally, loss of heterozygosity of LATS1 chromosomal region predisposes to breast, ovarian, and cervical tumors. In the current study, we investigated LATS1 genetic variations including single nucleotide polymorphisms (SNPs), in 28 Egyptian patients with either urinary bladder or colon cancers. The LATS1 gene was amplified and sequenced and the expression of LATS1 at the RNA level was assessed in 12 urinary bladder cancer samples. We report, the identification of a total of 29 variants including previously identified SNPs within LATS1 coding and non-coding sequences. A total of 18 variants were novel. Majority of the novel variants, 13, were mapped to intronic sequences and un-translated regions of the gene. Four of the five novel variants located in the coding region of the gene, represented missense mutations within the serine/threonine kinase catalytic domain. Interestingly, LATS1 RNA steady state levels was lost in urinary bladder cancerous tissue harboring four specific SNPs (16045 + 41736 + 34614 + 56177) positioned in the 5'UTR, intron 6, and two silent mutations within exon 4 and exon 8, respectively. This study identifies novel single-base-sequence alterations in the LATS1 gene. These newly identified variants could potentially be used as novel diagnostic or prognostic tools in cancer.
Fluorescent signatures for variable DNA sequences
Rice, John E.; Reis, Arthur H.; Rice, Lisa M.; Carver-Brown, Rachel K.; Wangh, Lawrence J.
2012-01-01
Life abounds with genetic variations writ in sequences that are often only a few hundred nucleotides long. Rapid detection of these variations for identification of genetic diseases, pathogens and organisms has become the mainstay of molecular science and medicine. This report describes a new, highly informative closed-tube polymerase chain reaction (PCR) strategy for analysis of both known and unknown sequence variations. It combines efficient quantitative amplification of single-stranded DNA targets through LATE-PCR with sets of Lights-On/Lights-Off probes that hybridize to their target sequences over a broad temperature range. Contiguous pairs of Lights-On/Lights-Off probes of the same fluorescent color are used to scan hundreds of nucleotides for the presence of mutations. Sets of probes in different colors can be combined in the same tube to analyze even longer single-stranded targets. Each set of hybridized Lights-On/Lights-Off probes generates a composite fluorescent contour, which is mathematically converted to a sequence-specific fluorescent signature. The versatility and broad utility of this new technology is illustrated in this report by characterization of variant sequences in three different DNA targets: the rpoB gene of Mycobacterium tuberculosis, a sequence in the mitochondrial cytochrome C oxidase subunit 1 gene of nematodes and the V3 hypervariable region of the bacterial 16 s ribosomal RNA gene. We anticipate widespread use of these technologies for diagnostics, species identification and basic research. PMID:22879378
Complex multifractal nature in Mycobacterium tuberculosis genome
Mandal, Saurav; Roychowdhury, Tanmoy; Chirom, Keilash; Bhattacharya, Alok; Brojen Singh, R. K.
2017-01-01
The mutifractal and long range correlation (C(r)) properties of strings, such as nucleotide sequence can be a useful parameter for identification of underlying patterns and variations. In this study C(r) and multifractal singularity function f(α) have been used to study variations in the genomes of a pathogenic bacteria Mycobacterium tuberculosis. Genomic sequences of M. tuberculosis isolates displayed significant variations in C(r) and f(α) reflecting inherent differences in sequences among isolates. M. tuberculosis isolates can be categorised into different subgroups based on sensitivity to drugs, these are DS (drug sensitive isolates), MDR (multi-drug resistant isolates) and XDR (extremely drug resistant isolates). C(r) follows significantly different scaling rules in different subgroups of isolates, but all the isolates follow one parameter scaling law. The richness in complexity of each subgroup can be quantified by the measures of multifractal parameters displaying a pattern in which XDR isolates have highest value and lowest for drug sensitive isolates. Therefore C(r) and multifractal functions can be useful parameters for analysis of genomic sequences. PMID:28440326
Rare variants and autoimmune disease.
Massey, Jonathan; Eyre, Steve
2014-09-01
The study of rare variants in monogenic forms of autoimmune disease has offered insight into the aetiology of more complex pathologies. Research in complex autoimmune disease initially focused on sequencing candidate genes, with some early successes, notably in uncovering low-frequency variation associated with Type 1 diabetes mellitus. However, other early examples have proved difficult to replicate, and a recent study across six autoimmune diseases, re-sequencing 25 autoimmune disease-associated genes in large sample sizes, failed to find any associated rare variants. The study of rare and low-frequency variation in autoimmune diseases has been made accessible by the inclusion of such variants on custom genotyping arrays (e.g. Immunochip and Exome arrays). Whole-exome sequencing approaches are now also being utilised to uncover the contribution of rare coding variants to disease susceptibility, severity and treatment response. Other sequencing strategies are starting to uncover the role of regulatory rare variation. © The Author 2014. Published by Oxford University Press. All rights reserved. For permissions, please email: journals.permissions@oup.com.
Hart, Reece K; Rico, Rudolph; Hare, Emily; Garcia, John; Westbrook, Jody; Fusaro, Vincent A
2015-01-15
Biological sequence variants are commonly represented in scientific literature, clinical reports and databases of variation using the mutation nomenclature guidelines endorsed by the Human Genome Variation Society (HGVS). Despite the widespread use of the standard, no freely available and comprehensive programming libraries are available. Here we report an open-source and easy-to-use Python library that facilitates the parsing, manipulation, formatting and validation of variants according to the HGVS specification. The current implementation focuses on the subset of the HGVS recommendations that precisely describe sequence-level variation relevant to the application of high-throughput sequencing to clinical diagnostics. The package is released under the Apache 2.0 open-source license. Source code, documentation and issue tracking are available at http://bitbucket.org/hgvs/hgvs/. Python packages are available at PyPI (https://pypi.python.org/pypi/hgvs). Supplementary data are available at Bioinformatics online. © The Author 2014. Published by Oxford University Press.
Complex multifractal nature in Mycobacterium tuberculosis genome
NASA Astrophysics Data System (ADS)
Mandal, Saurav; Roychowdhury, Tanmoy; Chirom, Keilash; Bhattacharya, Alok; Brojen Singh, R. K.
2017-04-01
The mutifractal and long range correlation (C(r)) properties of strings, such as nucleotide sequence can be a useful parameter for identification of underlying patterns and variations. In this study C(r) and multifractal singularity function f(α) have been used to study variations in the genomes of a pathogenic bacteria Mycobacterium tuberculosis. Genomic sequences of M. tuberculosis isolates displayed significant variations in C(r) and f(α) reflecting inherent differences in sequences among isolates. M. tuberculosis isolates can be categorised into different subgroups based on sensitivity to drugs, these are DS (drug sensitive isolates), MDR (multi-drug resistant isolates) and XDR (extremely drug resistant isolates). C(r) follows significantly different scaling rules in different subgroups of isolates, but all the isolates follow one parameter scaling law. The richness in complexity of each subgroup can be quantified by the measures of multifractal parameters displaying a pattern in which XDR isolates have highest value and lowest for drug sensitive isolates. Therefore C(r) and multifractal functions can be useful parameters for analysis of genomic sequences.
Hart, Reece K.; Rico, Rudolph; Hare, Emily; Garcia, John; Westbrook, Jody; Fusaro, Vincent A.
2015-01-01
Summary: Biological sequence variants are commonly represented in scientific literature, clinical reports and databases of variation using the mutation nomenclature guidelines endorsed by the Human Genome Variation Society (HGVS). Despite the widespread use of the standard, no freely available and comprehensive programming libraries are available. Here we report an open-source and easy-to-use Python library that facilitates the parsing, manipulation, formatting and validation of variants according to the HGVS specification. The current implementation focuses on the subset of the HGVS recommendations that precisely describe sequence-level variation relevant to the application of high-throughput sequencing to clinical diagnostics. Availability and implementation: The package is released under the Apache 2.0 open-source license. Source code, documentation and issue tracking are available at http://bitbucket.org/hgvs/hgvs/. Python packages are available at PyPI (https://pypi.python.org/pypi/hgvs). Contact: reecehart@gmail.com Supplementary information: Supplementary data are available at Bioinformatics online. PMID:25273102
Zhang, Zhenying; Liu, Xiaoming; Lv, Xuelian; Lin, Jingrong
2011-12-01
Sporotrichosis is usually a localized, lymphocutaneous disease, but its disseminated type was rarely reported. The main objective of this study was to identify specific DNA sequence variation and virulence of a strain of Sporothrix schenckii isolated from the lesion of disseminated cutaneous sporotrichosis. We confirmed this strain to be S. schenckii by(®) tubulin and chitin synthase gene sequence analysis in addition to the routine mycological and partial ITS and NTS sequencing. We found a 10-bp deletion in the ribosomal NTS region of this strain, in reference to the sequence of control strains isolated from fixed cutaneous sporotrichosis. After inoculated into immunosuppressed mice, this strain caused more extensive system involvement and showed stronger virulence than the control strain isolated from a fixed cutaneous sporotrichosis. Our study thus suggests that different clinical manifestation of sporotrichosis may be associated with variation in genotype and virulence of the strain, independent of effects due to the immune status of the host.
On The Sfr-M* Main Sequence Archetypal Star-Formation History And Analytical Models
NASA Astrophysics Data System (ADS)
Ciesla, Laure; Elbaz, David; Fensch, Jeremy
2017-06-01
From the evolution of the main sequence we can build the star formation history (SFH) of MS galaxies, assuming that they follow this relation all their life. We show that this SFH is not only a function of cosmic time but also involve the seed mass of the galaxy. We discuss the implications of this MS SFH on the stellar mass growth, and the entry in the passive region of the UVJ diagram, while the galaxy is still forming stars. We test the ability of different analytical SFH forms found in the literature to probe the SFR of all type of galaxies. Using a sample of GOODS-South galaxies, we show that these SFHs artificially enhance or create a gradient of age, parallel to the MS. A simple model of a MS galaxy, such as those expected from compaction or variation in gas accretion, undergoing some fluctuations provide does not predict such a gradient, that we show is due to SFH assumptions. We propose an improved analytical form, taking into account a flexibility in the recent SFH that we calibrate as a diagnostic to identify rapidly quenched galaxies from large photometric survey.
NASA Astrophysics Data System (ADS)
Lee, Sang-Rae; Song, Eun Hye; Lee, Tongsup
2018-03-01
Organisms entering the East Sea (Sea of Japan) through the Korea Strait, together with water, salt, and energy, affect the East Sea ecosystem. In this study, we report on the biodiversity of eukaryotic plankton found in the Western Channel of the Korea Strait for the first time using small subunit ribosomal RNA gene (18S rDNA) sequences. We also discuss the characteristics of water masses and their physicochemical factors. Diverse taxonomic groups were recovered from 18S rDNA clone libraries, including putative novel, higher taxonomic entities affiliated with Cercozoa, Raphidophyceae, Picozoa, and novel marine Stramenopiles. We also found that there was cryptic genetic variation at both the intraspecific and interspecific levels among arthropods, diatoms, and green algae. Specific plankton assemblages were identified at different sampling depths and they may provide useful information that could be used to interpret the origin and the subsequent mixing history of the water masses that contribute to the Tsushima Warm Current waters. Furthermore, the biological information highlighted in this study may help improve our understanding about the complex water mass interactions that were highlighted in the Korea Strait.
Catmur, Caroline; Walsh, Vincent; Heyes, Cecilia
2009-01-01
A core requirement for imitation is a capacity to solve the correspondence problem; to map observed onto executed actions, even when observation and execution yield sensory inputs in different modalities and coordinate frames. Until recently, it was assumed that the human capacity to solve the correspondence problem is innate. However, it is now becoming apparent that, as predicted by the associative sequence learning model, experience, and especially sensorimotor experience, plays a critical role in the development of imitation. We review evidence from studies of non-human animals, children and adults, focusing on research in cognitive neuroscience that uses training and naturally occurring variations in expertise to examine the role of experience in the formation of the mirror system. The relevance of this research depends on the widely held assumption that the mirror system plays a causal role in generating imitative behaviour. We also report original data supporting this assumption. These data show that theta-burst transcranial magnetic stimulation of the inferior frontal gyrus, a classical mirror system area, disrupts automatic imitation of finger movements. We discuss the implications of the evidence reviewed for the evolution, development and intentional control of imitation. PMID:19620108
Evolution viewed from physics, physiology and medicine.
Noble, Denis
2017-10-06
Stochasticity is harnessed by organisms to generate functionality. Randomness does not, therefore, necessarily imply lack of function or 'blind chance' at higher levels. In this respect, biology must resemble physics in generating order from disorder. This fact is contrary to Schrödinger's idea of biology generating phenotypic order from molecular- level order, which inspired the central dogma of molecular biology. The order originates at higher levels, which constrain the components at lower levels. We now know that this includes the genome, which is controlled by patterns of transcription factors and various epigenetic and reorganization mechanisms. These processes can occur in response to environmental stress, so that the genome becomes 'a highly sensitive organ of the cell' (McClintock). Organisms have evolved to be able to cope with many variations at the molecular level. Organisms also make use of physical processes in evolution and development when it is possible to arrive at functional development without the necessity to store all information in DNA sequences. This view of development and evolution differs radically from that of neo-Darwinism with its emphasis on blind chance as the origin of variation. Blind chance is necessary, but the origin of functional variation is not at the molecular level. These observations derive from and reinforce the principle of biological relativity, which holds that there is no privileged level of causation. They also have important implications for medical science.
2014-01-01
Background Variation in seed oil composition and content among soybean varieties is largely attributed to differences in transcript sequences and/or transcript accumulation of oil production related genes in seeds. Discovery and analysis of sequence and expression variations in these genes will accelerate soybean oil quality improvement. Results In an effort to identify these variations, we sequenced the transcriptomes of soybean seeds from nine lines varying in oil composition and/or total oil content. Our results showed that 69,338 distinct transcripts from 32,885 annotated genes were expressed in seeds. A total of 8,037 transcript expression polymorphisms and 50,485 transcript sequence polymorphisms (48,792 SNPs and 1,693 small Indels) were identified among the lines. Effects of the transcript polymorphisms on their encoded protein sequences and functions were predicted. The studies also provided independent evidence that the lack of FAD2-1A gene activity and a non-synonymous SNP in the coding sequence of FAB2C caused elevated oleic acid and stearic acid levels in soybean lines M23 and FAM94-41, respectively. Conclusions As a proof-of-concept, we developed an integrated RNA-seq and bioinformatics approach to identify and functionally annotate transcript polymorphisms, and demonstrated its high effectiveness for discovery of genetic and transcript variations that result in altered oil quality traits. The collection of transcript polymorphisms coupled with their predicted functional effects will be a valuable asset for further discovery of genes, gene variants, and functional markers to improve soybean oil quality. PMID:24755115
Zapata, Luis; Ding, Jia; Willing, Eva-Maria; Hartwig, Benjamin; Bezdan, Daniela; Jiao, Wen-Biao; Patel, Vipul; Velikkakam James, Geo; Koornneef, Maarten; Ossowski, Stephan; Schneeberger, Korbinian
2016-07-12
Resequencing or reference-based assemblies reveal large parts of the small-scale sequence variation. However, they typically fail to separate such local variation into colinear and rearranged variation, because they usually do not recover the complement of large-scale rearrangements, including transpositions and inversions. Besides the availability of hundreds of genomes of diverse Arabidopsis thaliana accessions, there is so far only one full-length assembled genome: the reference sequence. We have assembled 117 Mb of the A. thaliana Landsberg erecta (Ler) genome into five chromosome-equivalent sequences using a combination of short Illumina reads, long PacBio reads, and linkage information. Whole-genome comparison against the reference sequence revealed 564 transpositions and 47 inversions comprising ∼3.6 Mb, in addition to 4.1 Mb of nonreference sequence, mostly originating from duplications. Although rearranged regions are not different in local divergence from colinear regions, they are drastically depleted for meiotic recombination in heterozygotes. Using a 1.2-Mb inversion as an example, we show that such rearrangement-mediated reduction of meiotic recombination can lead to genetically isolated haplotypes in the worldwide population of A. thaliana Moreover, we found 105 single-copy genes, which were only present in the reference sequence or the Ler assembly, and 334 single-copy orthologs, which showed an additional copy in only one of the genomes. To our knowledge, this work gives first insights into the degree and type of variation, which will be revealed once complete assemblies will replace resequencing or other reference-dependent methods.
A reference human genome dataset of the BGISEQ-500 sequencer.
Huang, Jie; Liang, Xinming; Xuan, Yuankai; Geng, Chunyu; Li, Yuxiang; Lu, Haorong; Qu, Shoufang; Mei, Xianglin; Chen, Hongbo; Yu, Ting; Sun, Nan; Rao, Junhua; Wang, Jiahao; Zhang, Wenwei; Chen, Ying; Liao, Sha; Jiang, Hui; Liu, Xin; Yang, Zhaopeng; Mu, Feng; Gao, Shangxian
2017-05-01
BGISEQ-500 is a new desktop sequencer developed by BGI. Using DNA nanoball and combinational probe anchor synthesis developed from Complete Genomics™ sequencing technologies, it generates short reads at a large scale. Here, we present the first human whole-genome sequencing dataset of BGISEQ-500. The dataset was generated by sequencing the widely used cell line HG001 (NA12878) in two sequencing runs of paired-end 50 bp (PE50) and two sequencing runs of paired-end 100 bp (PE100). We also include examples of the raw images from the sequencer for reference. Finally, we identified variations using this dataset, estimated the accuracy of the variations, and compared to that of the variations identified from similar amounts of publicly available HiSeq2500 data. We found similar single nucleotide polymorphism (SNP) detection accuracy for the BGISEQ-500 PE100 data (false positive rate [FPR] = 0.00020%, sensitivity = 96.20%) compared to the PE150 HiSeq2500 data (FPR = 0.00017%, sensitivity = 96.60%) better SNP detection accuracy than the PE50 data (FPR = 0.0006%, sensitivity = 94.15%). But for insertions and deletions (indels), we found lower accuracy for BGISEQ-500 data (FPR = 0.00069% and 0.00067% for PE100 and PE50 respectively, sensitivity = 88.52% and 70.93%) than the HiSeq2500 data (FPR = 0.00032%, sensitivity = 96.28%). Our dataset can serve as the reference dataset, providing basic information not just for future development, but also for all research and applications based on the new sequencing platform. © The Authors 2017. Published by Oxford University Press.
Cho, Anna; Seong, Moon-Woo; Lim, Byung Chan; Lee, Hwa Jeen; Byeon, Jung Hye; Kim, Seung Soo; Kim, Soo Yeon; Choi, Sun Ah; Wong, Ai-Lynn; Lee, Jeongho; Kim, Jon Soo; Ryu, Hye Won; Lee, Jin Sook; Kim, Hunmin; Hwang, Hee; Choi, Ji Eun; Kim, Ki Joong; Hwang, Young Seung; Hong, Ki Ho; Park, Seungman; Cho, Sung Im; Lee, Seung Jun; Park, Hyunwoong; Seo, Soo Hyun; Park, Sung Sup; Chae, Jong Hee
2017-05-01
Duchenne and Becker muscular dystrophies (DMD and BMD) are allelic X-linked recessive muscle diseases caused by mutations in the large and complex dystrophin gene. We analyzed the dystrophin gene in 507 Korean DMD/BMD patients by multiple ligation-dependent probe amplification and direct sequencing. Overall, 117 different deletions, 48 duplications, and 90 pathogenic sequence variations, including 30 novel variations, were identified. Deletions and duplications accounted for 65.4% and 13.3% of Korean dystrophinopathy, respectively, suggesting that the incidence of large rearrangements in dystrophin is similar among different ethnic groups. We also detected sequence variations in >100 probands. The small variations were dispersed across the whole gene, and 12.3% were nonsense mutations. Precise genetic characterization in patients with DMD/BMD is timely and important for implementing nationwide registration systems and future molecular therapeutic trials in Korea and globally. Muscle Nerve 55: 727-734, 2017. © 2016 Wiley Periodicals, Inc.
An experimental phylogeny to benchmark ancestral sequence reconstruction
Randall, Ryan N.; Radford, Caelan E.; Roof, Kelsey A.; Natarajan, Divya K.; Gaucher, Eric A.
2016-01-01
Ancestral sequence reconstruction (ASR) is a still-burgeoning method that has revealed many key mechanisms of molecular evolution. One criticism of the approach is an inability to validate its algorithms within a biological context as opposed to a computer simulation. Here we build an experimental phylogeny using the gene of a single red fluorescent protein to address this criticism. The evolved phylogeny consists of 19 operational taxonomic units (leaves) and 17 ancestral bifurcations (nodes) that display a wide variety of fluorescent phenotypes. The 19 leaves then serve as ‘modern' sequences that we subject to ASR analyses using various algorithms and to benchmark against the known ancestral genotypes and ancestral phenotypes. We confirm computer simulations that show all algorithms infer ancient sequences with high accuracy, yet we also reveal wide variation in the phenotypes encoded by incorrectly inferred sequences. Specifically, Bayesian methods incorporating rate variation significantly outperform the maximum parsimony criterion in phenotypic accuracy. Subsampling of extant sequences had minor effect on the inference of ancestral sequences. PMID:27628687
Wallace, Douglas C
2013-07-19
Two major inconsistencies exist in the current neo-Darwinian evolutionary theory that random chromosomal mutations acted on by natural selection generate new species. First, natural selection does not require the evolution of ever increasing complexity, yet this is the hallmark of biology. Second, human chromosomal DNA sequence variation is predominantly either neutral or deleterious and is insufficient to provide the variation required for speciation or for predilection to common diseases. Complexity is explained by the continuous flow of energy through the biosphere that drives the accumulation of nucleic acids and information. Information then encodes complex forms. In animals, energy flow is primarily mediated by mitochondria whose maternally inherited mitochondrial DNA (mtDNA) codes for key genes for energy metabolism. In mammals, the mtDNA has a very high mutation rate, but the deleterious mutations are removed by an ovarian selection system. Hence, new mutations that subtly alter energy metabolism are continuously introduced into the species, permitting adaptation to regional differences in energy environments. Therefore, the most phenotypically significant gene variants arise in the mtDNA, are regional, and permit animals to occupy peripheral energy environments where rarer nuclear DNA (nDNA) variants can accumulate, leading to speciation. The neutralist-selectionist debate is then a consequence of mammals having two different evolutionary strategies: a fast mtDNA strategy for intra-specific radiation and a slow nDNA strategy for speciation. Furthermore, the missing genetic variation for common human diseases is primarily mtDNA variation plus regional nDNA variants, both of which have been missed by large, inter-population association studies.
Tseng, Shu-Ping; Li, Shou-Hsien; Hsieh, Chia-Hung; Wang, Hurng-Yi; Lin, Si-Min
2014-10-01
Dating the time of divergence and understanding speciation processes are central to the study of the evolutionary history of organisms but are notoriously difficult. The difficulty is largely rooted in variations in the ancestral population size or in the genealogy variation across loci. To depict the speciation processes and divergence histories of three monophyletic Takydromus species endemic to Taiwan, we sequenced 20 nuclear loci and combined with one mitochondrial locus published in GenBank. They were analysed by a multispecies coalescent approach within a Bayesian framework. Divergence dating based on the gene tree approach showed high variation among loci, and the divergence was estimated at an earlier date than when derived by the species-tree approach. To test whether variations in the ancestral population size accounted for the majority of this variation, we conducted computer inferences using isolation-with-migration (IM) and approximate Bayesian computation (ABC) frameworks. The results revealed that gene flow during the early stage of speciation was strongly favoured over the isolation model, and the initiation of the speciation process was far earlier than the dates estimated by gene- and species-based divergence dating. Due to their limited dispersal ability, it is suggested that geographical isolation may have played a major role in the divergence of these Takydromus species. Nevertheless, this study reveals a more complex situation and demonstrates that gene flow during the speciation process cannot be overlooked and may have a great impact on divergence dating. By using multilocus data and incorporating Bayesian coalescence approaches, we provide a more biologically realistic framework for delineating the divergence history of Takydromus. © 2014 John Wiley & Sons Ltd.
Shakhssalim, Nasser; Houshmand, Massoud; Kamalidehghan, Behnam; Faraji, Abolfazl; Sarhangnejad, Reza; Dadgar, Sepideh; Mobaraki, Maryam; Rosli, Rozita; Sanati, Mohammad Hossein
2013-12-05
Bladder cancer is a relatively common and potentially life-threatening neoplasm that ranks ninth in terms of worldwide cancer incidence. The aim of this study was to determine deletions and sequence variations in the mitochondrial displacement loop (D-loop) region from the blood specimens and tumoral tissues of patients with bladder cancer, compared to adjacent non-tumoral tissues. The DNA from blood, tumoral tissues and adjacent non-tumoral tissues of twenty-six patients with bladder cancer and DNA from blood of 504 healthy controls from different ethnicities were investigated to determine sequence variation in the mitochondrial D-loop region using multiplex polymerase chain reaction (PCR), DNA sequencing and southern blotting analysis. From a total of 110 variations, 48 were reported as new mutations. No deletions were detected in tumoral tissues, adjacent non-tumoral tissues and blood samples from patients. Although the polymorphisms at loci 16189, 16261 and 16311 were not significantly correlated with bladder cancer, the C16069T variation was significantly present in patient samples compared to control samples (p < 0.05). Interestingly, there was no significant difference (p > 0.05) of C variations, including C7TC6, C8TC6, C9TC6 and C10TC6, in D310 mitochondrial DNA between patients and control samples. Our study suggests that 16069 mitochondrial DNA D-Loop mutations may play a significant role in the etiology of bladder cancer and facilitate the definition of carcinogenesis-related mutations in human cancer.
Wu, Tsung-Jung; Shamsaddini, Amirhossein; Pan, Yang; Smith, Krista; Crichton, Daniel J; Simonyan, Vahan; Mazumder, Raja
2014-01-01
Years of sequence feature curation by UniProtKB/Swiss-Prot, PIR-PSD, NCBI-CDD, RefSeq and other database biocurators has led to a rich repository of information on functional sites of genes and proteins. This information along with variation-related annotation can be used to scan human short sequence reads from next-generation sequencing (NGS) pipelines for presence of non-synonymous single-nucleotide variations (nsSNVs) that affect functional sites. This and similar workflows are becoming more important because thousands of NGS data sets are being made available through projects such as The Cancer Genome Atlas (TCGA), and researchers want to evaluate their biomarkers in genomic data. BioMuta, an integrated sequence feature database, provides a framework for automated and manual curation and integration of cancer-related sequence features so that they can be used in NGS analysis pipelines. Sequence feature information in BioMuta is collected from the Catalogue of Somatic Mutations in Cancer (COSMIC), ClinVar, UniProtKB and through biocuration of information available from publications. Additionally, nsSNVs identified through automated analysis of NGS data from TCGA are also included in the database. Because of the petabytes of data and information present in NGS primary repositories, a platform HIVE (High-performance Integrated Virtual Environment) for storing, analyzing, computing and curating NGS data and associated metadata has been developed. Using HIVE, 31 979 nsSNVs were identified in TCGA-derived NGS data from breast cancer patients. All variations identified through this process are stored in a Curated Short Read archive, and the nsSNVs from the tumor samples are included in BioMuta. Currently, BioMuta has 26 cancer types with 13 896 small-scale and 308 986 large-scale study-derived variations. Integration of variation data allows identifications of novel or common nsSNVs that can be prioritized in validation studies. Database URL: BioMuta: http://hive.biochemistry.gwu.edu/tools/biomuta/index.php; CSR: http://hive.biochemistry.gwu.edu/dna.cgi?cmd=csr; HIVE: http://hive.biochemistry.gwu.edu.
Singh, Satyendra K; Prasad, Kashi N; Singh, Aloukick K; Gupta, Kamlesh K; Chauhan, Ranjeet S; Singh, Amrita; Singh, Avinash; Rai, Ravi P; Pati, Binod K
2016-10-01
Taenia solium is the major cause of taeniasis and cysticercosis/neurocysticercosis (NCC) in the developing countries including India, but the existence of other Taenia species and genetic variation have not been studied in India. So, we studied the existence of different Taenia species, and sequence variation in Taenia isolates from human (proglottids and cysticerci) and swine (cysticerci) in North India. Amplification of cytochrome c oxidase subunit 1 gene (cox1) was done by polymerase chain reaction (PCR) followed by sequencing and phylogenetic analysis. We identified two species of Taenia i.e. T. solium and Taenia asiatica in our isolates. T. solium isolates showed similarity with Asian genotype and nucleotide variations from 0.25 to 1.01 %, whereas T. asiatica displayed nucleotide variations ranged from 0.25 to 0.5 %. These findings displayed the minimal genetic variations in North Indian isolates of T. solium and T. asiatica.
Jakupciak, John P; Wells, Jeffrey M; Karalus, Richard J; Pawlowski, David R; Lin, Jeffrey S; Feldman, Andrew B
2013-01-01
Large-scale genomics projects are identifying biomarkers to detect human disease. B. pseudomallei and B. mallei are two closely related select agents that cause melioidosis and glanders. Accurate characterization of metagenomic samples is dependent on accurate measurements of genetic variation between isolates with resolution down to strain level. Often single biomarker sensitivity is augmented by use of multiple or panels of biomarkers. In parallel with single biomarker validation, advances in DNA sequencing enable analysis of entire genomes in a single run: population-sequencing. Potentially, direct sequencing could be used to analyze an entire genome to serve as the biomarker for genome identification. However, genome variation and population diversity complicate use of direct sequencing, as well as differences caused by sample preparation protocols including sequencing artifacts and mistakes. As part of a Department of Homeland Security program in bacterial forensics, we examined how to implement whole genome sequencing (WGS) analysis as a judicially defensible forensic method for attributing microbial sample relatedness; and also to determine the strengths and limitations of whole genome sequence analysis in a forensics context. Herein, we demonstrate use of sequencing to provide genetic characterization of populations: direct sequencing of populations.
Jakupciak, John P.; Wells, Jeffrey M.; Karalus, Richard J.; Pawlowski, David R.; Lin, Jeffrey S.; Feldman, Andrew B.
2013-01-01
Large-scale genomics projects are identifying biomarkers to detect human disease. B. pseudomallei and B. mallei are two closely related select agents that cause melioidosis and glanders. Accurate characterization of metagenomic samples is dependent on accurate measurements of genetic variation between isolates with resolution down to strain level. Often single biomarker sensitivity is augmented by use of multiple or panels of biomarkers. In parallel with single biomarker validation, advances in DNA sequencing enable analysis of entire genomes in a single run: population-sequencing. Potentially, direct sequencing could be used to analyze an entire genome to serve as the biomarker for genome identification. However, genome variation and population diversity complicate use of direct sequencing, as well as differences caused by sample preparation protocols including sequencing artifacts and mistakes. As part of a Department of Homeland Security program in bacterial forensics, we examined how to implement whole genome sequencing (WGS) analysis as a judicially defensible forensic method for attributing microbial sample relatedness; and also to determine the strengths and limitations of whole genome sequence analysis in a forensics context. Herein, we demonstrate use of sequencing to provide genetic characterization of populations: direct sequencing of populations. PMID:24455204
Sequence investigation of 34 forensic autosomal STRs with massively parallel sequencing.
Zhang, Suhua; Niu, Yong; Bian, Yingnan; Dong, Rixia; Liu, Xiling; Bao, Yun; Jin, Chao; Zheng, Hancheng; Li, Chengtao
2018-05-01
STRs vary not only in the length of the repeat units and the number of repeats but also in the region with which they conform to an incremental repeat pattern. Massively parallel sequencing (MPS) offers new possibilities in the analysis of STRs since they can simultaneously sequence multiple targets in a single reaction and capture potential internal sequence variations. Here, we sequenced 34 STRs applied in the forensic community of China with a custom-designed panel. MPS performance were evaluated from sequencing reads analysis, concordance study and sensitivity testing. High coverage sequencing data were obtained to determine the constitute ratios and heterozygous balance. No actual inconsistent genotypes were observed between capillary electrophoresis (CE) and MPS, demonstrating the reliability of the panel and the MPS technology. With the sequencing data from the 200 investigated individuals, 346 and 418 alleles were obtained via CE and MPS technologies at the 34 STRs, indicating MPS technology provides higher discrimination than CE detection. The whole study demonstrated that STR genotyping with the custom panel and MPS technology has the potential not only to reveal length and sequence variations but also to satisfy the demands of high throughput and high multiplexing with acceptable sensitivity.
Küpper, Clemens; Burke, Terry; Lank, David B.
2015-01-01
Sequence variation in the melanocortin-1 receptor (MC1R) gene explains color morph variation in several species of birds and mammals. Ruffs (Philomachus pugnax) exhibit major dark/light color differences in melanin-based male breeding plumage which is closely associated with alternative reproductive behavior. A previous study identified a microsatellite marker (Ppu020) near the MC1R locus associated with the presence/absence of ornamental plumage. We investigated whether coding sequence variation in the MC1R gene explains major dark/light plumage color variation and/or the presence/absence of ornamental plumage in ruffs. Among 821bp of the MC1R coding region from 44 male ruffs we found 3 single nucleotide polymorphisms, representing 1 nonsynonymous and 2 synonymous amino acid substitutions. None were associated with major dark/light color differences or the presence/absence of ornamental plumage. At all amino acid sites known to be functionally important in other avian species with dark/light plumage color variation, ruffs were either monomorphic or the shared polymorphism did not coincide with color morph. Neither ornamental plumage color differences nor the presence/absence of ornamental plumage in ruffs are likely to be caused entirely by amino acid variation within the coding regions of the MC1R locus. Regulatory elements and structural variation at other loci may be involved in melanin expression and contribute to the extreme plumage polymorphism observed in this species. PMID:25534935
Fornage, Myriam; Mosley, Thomas H; Jack, Clifford R; de Andrade, Mariza; Kardia, Sharon L R; Boerwinkle, Eric; Turner, Stephen T
2007-01-01
Susceptibility to ischemic damage to the subcortical white matter of the brain has a strong genetic basis. Dysregulation of matrix metalloproteinases (MMPs) contributes to loss of cerebrovascular integrity and white matter injury. We investigated whether sequence variation in the genes encoding MMP3 and MMP9 is associated with variation in leukoaraiosis volume, determined by magnetic resonance imaging, in non-Hispanic whites and African-Americans using family-based association tests. Seven hundred and fifty-six white and 671 African-American individuals from sibships ascertained through two or more siblings with hypertension were genotyped for 7 and 8 haplotype-tagging polymorphisms in the MMP3 and MMP9 genes, respectively. MMP3 sequence variation was significantly associated with variation in leukoaraiosis volume in Whites. Two common haplotypes with opposing relationships to leukoaraiosis volume were identified. MMP9 sequence variation was also significantly associated with variation in leukoaraiosis volume in both African-Americans and Whites. Different haplotypes contributed to these associations in the two racial groups. These findings add to the growing body of evidence from animal models and human clinical studies suggesting a role of MMPs in ischemic white matter injury. They provide the basis for further investigation of the role of these genes in susceptibility and/or progression to clinical disease.
Genetics of Lipid and Lipoprotein Disorders and Traits.
Dron, Jacqueline S; Hegele, Robert A
2016-01-01
Plasma lipids, namely cholesterol and triglyceride, and lipoproteins, such as low-density lipoprotein (LDL) and high-density lipoprotein, serve numerous physiological roles. Perturbed levels of these traits underlie monogenic dyslipidemias, a diverse group of multisystem disorders. We are on the verge of having a relatively complete picture of the human dyslipidemias and their components. Recent advances in genetics of plasma lipids and lipoproteins include the following: (1) expanding the range of genes causing monogenic dyslipidemias, particularly elevated LDL cholesterol; (2) appreciating the role of polygenic effects in such traits as familial hypercholesterolemia and combined hyperlipidemia; (3) accumulating a list of common variants that determine plasma lipids and lipoproteins; (4) applying exome sequencing to identify collections of rare variants determining plasma lipids and lipoproteins that via Mendelian randomization have also implicated gene products such as NPC1L1 , APOC3 , LDLR , APOA5 , and ANGPTL4 as causal for atherosclerotic cardiovascular disease; and (5) using naturally occurring genetic variation to identify new drug targets, including inhibitors of apolipoprotein (apo) C-III, apo(a), ANGPTL3, and ANGPTL4. Here, we compile this disparate range of data linking human genetic variation to plasma lipids and lipoproteins, providing a "one stop shop" for the interested reader.
Honnavar, Prasanna; Prasad, Gandham S.; Ghosh, Anup; Dogra, Sunil; Handa, Sanjeev
2016-01-01
The majority of species within the genus Malassezia are lipophilic yeasts that colonize the skin of warm-blooded animals. Two species, Malassezia globosa and Malassezia restricta, are implicated in the causation of seborrheic dermatitis/dandruff (SD/D). During our survey of SD/D cases, we isolated several species of Malassezia and noticed vast variations within a few lipid-dependent species. Variations observed in the phenotypic characteristics (colony morphology, absence of catalase activity, growth at 37°C, and precipitation surrounding wells containing Tween 20 or Cremophor EL) suggested the possible presence of a novel species. Sequence divergence observed in the internal transcribed spacer (ITS) region, the D1/D2 domain, and the intergenic spacer 1 (IGS1) region of rDNA and the TEF1 gene, PCR-restriction fragment length polymorphism (RFLP) analysis of the ITS2 region, and fluorescent amplified fragment length polymorphism analysis support the existence of a novel species. Based on phenotypic and molecular characterization of these strains, we propose a new species, namely, M. arunalokei sp. nov., and we designate NCCPF 127130 (= MTCC 12054 = CBS 13387) as the type strain. PMID:27147721
Honnavar, Prasanna; Prasad, Gandham S; Ghosh, Anup; Dogra, Sunil; Handa, Sanjeev; Rudramurthy, Shivaprakash M
2016-07-01
The majority of species within the genus Malassezia are lipophilic yeasts that colonize the skin of warm-blooded animals. Two species, Malassezia globosa and Malassezia restricta, are implicated in the causation of seborrheic dermatitis/dandruff (SD/D). During our survey of SD/D cases, we isolated several species of Malassezia and noticed vast variations within a few lipid-dependent species. Variations observed in the phenotypic characteristics (colony morphology, absence of catalase activity, growth at 37°C, and precipitation surrounding wells containing Tween 20 or Cremophor EL) suggested the possible presence of a novel species. Sequence divergence observed in the internal transcribed spacer (ITS) region, the D1/D2 domain, and the intergenic spacer 1 (IGS1) region of rDNA and the TEF1 gene, PCR-restriction fragment length polymorphism (RFLP) analysis of the ITS2 region, and fluorescent amplified fragment length polymorphism analysis support the existence of a novel species. Based on phenotypic and molecular characterization of these strains, we propose a new species, namely, M. arunalokei sp. nov., and we designate NCCPF 127130 (= MTCC 12054 = CBS 13387) as the type strain. Copyright © 2016, American Society for Microbiology. All Rights Reserved.
Miller, Mark P.; Haig, Susan M.; Wagner, R.S.
2006-01-01
The Southern torrent salamander (Rhyacotriton variegatus) was recently found not warranted for listing under the US Endangered Species Act due to lack of information regarding population fragmentation and gene flow. Found in small-order streams associated with late-successional coniferous forests of the US Pacific Northwest, threats to their persistence include disturbance related to timber harvest activities. We conducted a study of genetic diversity throughout this species' range to 1) identify major phylogenetic lineages and phylogeographic barriers and 2) elucidate regional patterns of population genetic and spatial phylogeographic structure. Cytochrome b sequence variation was examined for 189 individuals from 72 localities. We identified 3 major lineages corresponding to nonoverlapping geographic regions: a northern California clade, a central Oregon clade, and a northern Oregon clade. The Yaquina River may be a phylogeographic barrier between the northern Oregon and central Oregon clades, whereas the Smith River in northern California appears to correspond to the discontinuity between the central Oregon and northern California clades. Spatial analyses of genetic variation within regions encompassing major clades indicated that the extent of genetic structure is comparable among regions. We discuss our results in the context of conservation efforts for Southern torrent salamanders.
Conservation genetics of the Far Eastern leopard (Panthera pardus orientalis).
Uphyrkina, O; Miquelle, D; Quigley, H; Driscoll, C; O'Brien, S J
2002-01-01
The Far Eastern or Amur leopard (Panthera pardus orientalis) survives today as a tiny relict population of 25-40 individuals in the Russian Far East. The population descends from a 19th-century northeastern Asian subspecies whose range extended over southeastern Russia, the Korean peninsula, and northeastern China. A molecular genetic survey of nuclear microsatellite and mitochondrial DNA (mtDNA) sequence variation validates subspecies distinctiveness but also reveals a markedly reduced level of genetic variation. The amount of genetic diversity measured is the lowest among leopard subspecies and is comparable to the genetically depleted Florida panther and Asiatic lion populations. When considered in the context of nonphysiological perils that threaten small populations (e.g., chance mortality, poaching, climatic extremes, and infectious disease), the genetic and demographic data indicate a critically diminished wild population under severe threat of extinction. An established captive population of P. p. orientalis displays much higher diversity than the wild population sample, but nearly all captive individuals are derived from a history of genetic admixture with the adjacent Chinese subspecies, P. p. japonensis. The conservation management implications of potential restoration/augmentation of the wild population with immigrants from the captive population are discussed.
Sources of Variation in the Gut Microbial Community of Lycaeides melissa Caterpillars.
Chaturvedi, Samridhi; Rego, Alexandre; Lucas, Lauren K; Gompert, Zachariah
2017-09-12
Microbes can mediate insect-plant interactions and have been implicated in major evolutionary transitions to herbivory. Whether microbes also play a role in more modest host shifts or expansions in herbivorous insects is less clear. Here we evaluate the potential for gut microbial communities to constrain or facilitate host plant use in the Melissa blue butterfly (Lycaeides melissa). We conducted a larval rearing experiment where caterpillars from two populations were fed plant tissue from two hosts. We used 16S rRNA sequencing to quantify the relative effects of sample type (frass versus whole caterpillar), diet (plant species), butterfly population and development (caterpillar age) on the composition and diversity of the caterpillar gut microbial communities, and secondly, to test for a relationship between microbial community and larval performance. Gut microbial communities varied over time (that is, with caterpillar age) and differed between frass and whole caterpillar samples. Diet (host plant) and butterfly population had much more limited effects on microbial communities. We found no evidence that gut microbe community composition was associated with caterpillar weight, and thus, our results provide no support for the hypothesis that variation in microbial community affects performance in L. melissa.
Hypermethylation in the ZBTB20 gene is associated with major depressive disorder
2014-01-01
Background Although genetic variation is believed to contribute to an individual’s susceptibility to major depressive disorder, genome-wide association studies have not yet identified associations that could explain the full etiology of the disease. Epigenetics is increasingly believed to play a major role in the development of common clinical phenotypes, including major depressive disorder. Results Genome-wide MeDIP-Sequencing was carried out on a total of 50 monozygotic twin pairs from the UK and Australia that are discordant for depression. We show that major depressive disorder is associated with significant hypermethylation within the coding region of ZBTB20, and is replicated in an independent cohort of 356 unrelated case-control individuals. The twins with major depressive disorder also show increased global variation in methylation in comparison with their unaffected co-twins. ZBTB20 plays an essential role in the specification of the Cornu Ammonis-1 field identity in the developing hippocampus, a region previously implicated in the development of major depressive disorder. Conclusions Our results suggest that aberrant methylation profiles affecting the hippocampus are associated with major depressive disorder and show the potential of the epigenetic twin model in neuro-psychiatric disease. PMID:24694013
Donaldson, Michael E; Rico, Yessica; Hueffer, Karsten; Rando, Halie M; Kukekova, Anna V; Kyle, Christopher J
2018-01-01
Pathogens are recognized as major drivers of local adaptation in wildlife systems. By determining which gene variants are favored in local interactions among populations with and without disease, spatially explicit adaptive responses to pathogens can be elucidated. Much of our current understanding of host responses to disease comes from a small number of genes associated with an immune response. High-throughput sequencing (HTS) technologies, such as genotype-by-sequencing (GBS), facilitate expanded explorations of genomic variation among populations. Hybridization-based GBS techniques can be leveraged in systems not well characterized for specific variants associated with disease outcome to "capture" specific genes and regulatory regions known to influence expression and disease outcome. We developed a multiplexed, sequence capture assay for red foxes to simultaneously assess ~300-kbp of genomic sequence from 116 adaptive, intrinsic, and innate immunity genes of predicted adaptive significance and their putative upstream regulatory regions along with 23 neutral microsatellite regions to control for demographic effects. The assay was applied to 45 fox DNA samples from Alaska, where three arctic rabies strains are geographically restricted and endemic to coastal tundra regions, yet absent from the boreal interior. The assay provided 61.5% on-target enrichment with relatively even sequence coverage across all targeted loci and samples (mean = 50×), which allowed us to elucidate genetic variation across introns, exons, and potential regulatory regions (4,819 SNPs). Challenges remained in accurately describing microsatellite variation using this technique; however, longer-read HTS technologies should overcome these issues. We used these data to conduct preliminary analyses and detected genetic structure in a subset of red fox immune-related genes between regions with and without endemic arctic rabies. This assay provides a template to assess immunogenetic variation in wildlife disease systems.
2013-01-01
Background SNPs&GO is a method for the prediction of deleterious Single Amino acid Polymorphisms (SAPs) using protein functional annotation. In this work, we present the web server implementation of SNPs&GO (WS-SNPs&GO). The server is based on Support Vector Machines (SVM) and for a given protein, its input comprises: the sequence and/or its three-dimensional structure (when available), a set of target variations and its functional Gene Ontology (GO) terms. The output of the server provides, for each protein variation, the probabilities to be associated to human diseases. Results The server consists of two main components, including updated versions of the sequence-based SNPs&GO (recently scored as one of the best algorithms for predicting deleterious SAPs) and of the structure-based SNPs&GO3d programs. Sequence and structure based algorithms are extensively tested on a large set of annotated variations extracted from the SwissVar database. Selecting a balanced dataset with more than 38,000 SAPs, the sequence-based approach achieves 81% overall accuracy, 0.61 correlation coefficient and an Area Under the Curve (AUC) of the Receiver Operating Characteristic (ROC) curve of 0.88. For the subset of ~6,600 variations mapped on protein structures available at the Protein Data Bank (PDB), the structure-based method scores with 84% overall accuracy, 0.68 correlation coefficient, and 0.91 AUC. When tested on a new blind set of variations, the results of the server are 79% and 83% overall accuracy for the sequence-based and structure-based inputs, respectively. Conclusions WS-SNPs&GO is a valuable tool that includes in a unique framework information derived from protein sequence, structure, evolutionary profile, and protein function. WS-SNPs&GO is freely available at http://snps.biofold.org/snps-and-go. PMID:23819482
Sands, Chester J; Convey, Peter; Linse, Katrin; McInnes, Sandra J
2008-04-30
Meiofauna - multicellular animals captured between sieve size 45 mum and 1000 mum - are a fundamental component of terrestrial, and marine benthic ecosystems, forming an integral element of food webs, and playing a critical roll in nutrient recycling. Most phyla have meiofaunal representatives and studies of these taxa impact on a wide variety of sub-disciplines as well as having social and economic implications. However, studies of variation in meiofauna are presented with several important challenges. Isolating individuals from a sample substrate is a time consuming process, and identification requires increasingly scarce taxonomic expertise. Finding suitable morphological characters in many of these organisms is often difficult even for experts. Molecular markers are extremely useful for identifying variation in morphologically conserved organisms. However, for many species markers need to be developed de novo, while DNA can often only be extracted from pooled samples in order to obtain sufficient quantity and quality. Importantly, multiple independent markers are required to reconcile gene evolution with species evolution. In this primarily methodological paper we provide a proof of principle of a novel and effective protocol for the isolation of meiofauna from an environmental sample. We also go on to illustrate examples of the implications arising from subsequent screening for genetic variation at the level of the individual using ribosomal, mitochondrial and single copy nuclear markers. To isolate individual tardigrades from their habitat substrate we used a non-toxic density gradient media that did not interfere with downstream biochemical processes. Using a simple DNA release technique and nested polymerase chain reaction with universal primers we were able amplify multi-copy and, to some extent, single copy genes from individual tardigrades. Maximum likelihood trees from ribosomal 18S, mitochondrial cytochrome oxidase subunit 1, and the single copy nuclear gene Wingless support a recent study indicating that the family Hypsibiidae is a non-monophyletic group. From these sequences we were able to detect variation between individuals at each locus that allowed us to identify the presence of cryptic taxa that would otherwise have been overlooked. Molecular results obtained from individuals, rather than pooled samples, are a prerequisite to enable levels of variation to be placed into context. In this study we have provided a proof of principle of this approach for meiofaunal tardigrades, an important group of soil biota previously not considered amenable to such studies, thereby paving the way for more comprehensive phylogenetic studies using multiple nuclear markers, and population genetic studies.
Sands, Chester J; Convey, Peter; Linse, Katrin; McInnes, Sandra J
2008-01-01
Background Meiofauna – multicellular animals captured between sieve size 45 μm and 1000 μm – are a fundamental component of terrestrial, and marine benthic ecosystems, forming an integral element of food webs, and playing a critical roll in nutrient recycling. Most phyla have meiofaunal representatives and studies of these taxa impact on a wide variety of sub-disciplines as well as having social and economic implications. However, studies of variation in meiofauna are presented with several important challenges. Isolating individuals from a sample substrate is a time consuming process, and identification requires increasingly scarce taxonomic expertise. Finding suitable morphological characters in many of these organisms is often difficult even for experts. Molecular markers are extremely useful for identifying variation in morphologically conserved organisms. However, for many species markers need to be developed de novo, while DNA can often only be extracted from pooled samples in order to obtain sufficient quantity and quality. Importantly, multiple independent markers are required to reconcile gene evolution with species evolution. In this primarily methodological paper we provide a proof of principle of a novel and effective protocol for the isolation of meiofauna from an environmental sample. We also go on to illustrate examples of the implications arising from subsequent screening for genetic variation at the level of the individual using ribosomal, mitochondrial and single copy nuclear markers. Results To isolate individual tardigrades from their habitat substrate we used a non-toxic density gradient media that did not interfere with downstream biochemical processes. Using a simple DNA release technique and nested polymerase chain reaction with universal primers we were able amplify multi-copy and, to some extent, single copy genes from individual tardigrades. Maximum likelihood trees from ribosomal 18S, mitochondrial cytochrome oxidase subunit 1, and the single copy nuclear gene Wingless support a recent study indicating that the family Hypsibiidae is a non-monophyletic group. From these sequences we were able to detect variation between individuals at each locus that allowed us to identify the presence of cryptic taxa that would otherwise have been overlooked. Conclusion Molecular results obtained from individuals, rather than pooled samples, are a prerequisite to enable levels of variation to be placed into context. In this study we have provided a proof of principle of this approach for meiofaunal tardigrades, an important group of soil biota previously not considered amenable to such studies, thereby paving the way for more comprehensive phylogenetic studies using multiple nuclear markers, and population genetic studies. PMID:18447908
Machczyńska, Joanna; Zimny, Janusz; Bednarek, Piotr Tomasz
2015-10-01
Plant regeneration via in vitro culture can induce genetic and epigenetic variation; however, the extent of such changes in triticale is not yet understood. In the present study, metAFLP, a variation of methylation-sensitive amplified fragment length polymorphism analysis, was used to investigate tissue culture-induced variation in triticale regenerants derived from four distinct genotypes using androgenesis and somatic embryogenesis. The metAFLP technique enabled identification of both sequence and DNA methylation pattern changes in a single experiment. Moreover, it was possible to quantify subtle effects such as sequence variation, demethylation, and de novo methylation, which affected 19, 5.5, 4.5% of sites, respectively. Comparison of variation in different genotypes and with different in vitro regeneration approaches demonstrated that both the culture technique and genetic background of donor plants affected tissue culture-induced variation. The results showed that the metAFLP approach could be used for quantification of tissue culture-induced variation and provided direct evidence that in vitro plant regeneration could cause genetic and epigenetic variation.
Wu, Gary D; Lewis, James D; Hoffmann, Christian; Chen, Ying-Yu; Knight, Rob; Bittinger, Kyle; Hwang, Jennifer; Chen, Jun; Berkowsky, Ronald; Nessel, Lisa; Li, Hongzhe; Bushman, Frederic D
2010-07-30
Intense interest centers on the role of the human gut microbiome in health and disease, but optimal methods for analysis are still under development. Here we present a study of methods for surveying bacterial communities in human feces using 454/Roche pyrosequencing of 16S rRNA gene tags. We analyzed fecal samples from 10 individuals and compared methods for storage, DNA purification and sequence acquisition. To assess reproducibility, we compared samples one cm apart on a single stool specimen for each individual. To analyze storage methods, we compared 1) immediate freezing at -80 degrees C, 2) storage on ice for 24 or 3) 48 hours. For DNA purification methods, we tested three commercial kits and bead beating in hot phenol. Variations due to the different methodologies were compared to variation among individuals using two approaches--one based on presence-absence information for bacterial taxa (unweighted UniFrac) and the other taking into account their relative abundance (weighted UniFrac). In the unweighted analysis relatively little variation was associated with the different analytical procedures, and variation between individuals predominated. In the weighted analysis considerable variation was associated with the purification methods. Particularly notable was improved recovery of Firmicutes sequences using the hot phenol method. We also carried out surveys of the effects of different 454 sequencing methods (FLX versus Titanium) and amplification of different 16S rRNA variable gene segments. Based on our findings we present recommendations for protocols to collect, process and sequence bacterial 16S rDNA from fecal samples--some major points are 1) if feasible, bead-beating in hot phenol or use of the PSP kit improves recovery; 2) storage methods can be adjusted based on experimental convenience; 3) unweighted (presence-absence) comparisons are less affected by lysis method.
Wyllie, David H; Sanderson, Nicholas; Myers, Richard; Peto, Tim; Robinson, Esther; Crook, Derrick W; Smith, E Grace; Walker, A Sarah
2018-06-06
Contact tracing requires reliable identification of closely related bacterial isolates. When we noticed the reporting of artefactual variation between M. tuberculosis isolates during routine next generation sequencing of Mycobacterium spp, we investigated its basis in 2,018 consecutive M. tuberculosis isolates. In the routine process used, clinical samples were decontaminated and inoculated into broth cultures; from positive broth cultures DNA was extracted, sequenced, reads mapped, and consensus sequences determined. We investigated the process of consensus sequence determination, which selects the most common nucleotide at each position. Having determined the high-quality read depth and depth of minor variants across 8,006 M. tuberculosis genomic regions, we quantified the relationship between the minor variant depth and the amount of non-Mycobacterial bacterial DNA, which originates from commensal microbes killed during sample decontamination. In the presence of non-Mycobacterial bacterial DNA, we found significant increases in minor variant frequencies of more than 1.5 fold in 242 regions covering 5.1% of the M. tuberculosis genome. Included within these were four high variation regions strongly influenced by the amount of non-Mycobacterial bacterial DNA. Excluding these four regions from pairwise distance comparisons reduced biologically implausible variation from 5.2% to 0% in an independent validation set derived from 226 individuals. Thus, we have demonstrated an approach identifying critical genomic regions contributing to clinically relevant artefactual variation in bacterial similarity searches. The approach described monitors the outputs of the complex multi-step laboratory and bioinformatics process, allows periodic process adjustments, and will have application to quality control of routine bacterial genomics. Copyright © 2018 Wyllie et al.
Salleh, Mohd Zaki; Teh, Lay Kek; Lee, Lian Shien; Ismet, Rose Iszati; Patowary, Ashok; Joshi, Kandarp; Pasha, Ayesha; Ahmed, Azni Zain; Janor, Roziah Mohd; Hamzah, Ahmad Sazali; Adam, Aishah; Yusoff, Khalid; Hoh, Boon Peng; Hatta, Fazleen Haslinda Mohd; Ismail, Mohamad Izwan; Scaria, Vinod; Sivasubbu, Sridhar
2013-01-01
With a higher throughput and lower cost in sequencing, second generation sequencing technology has immense potential for translation into clinical practice and in the realization of pharmacogenomics based patient care. The systematic analysis of whole genome sequences to assess patient to patient variability in pharmacokinetics and pharmacodynamics responses towards drugs would be the next step in future medicine in line with the vision of personalizing medicine. Genomic DNA obtained from a 55 years old, self-declared healthy, anonymous male of Malay descent was sequenced. The subject's mother died of lung cancer and the father had a history of schizophrenia and deceased at the age of 65 years old. A systematic, intuitive computational workflow/pipeline integrating custom algorithm in tandem with large datasets of variant annotations and gene functions for genetic variations with pharmacogenomics impact was developed. A comprehensive pathway map of drug transport, metabolism and action was used as a template to map non-synonymous variations with potential functional consequences. Over 3 million known variations and 100,898 novel variations in the Malay genome were identified. Further in-depth pharmacogenetics analysis revealed a total of 607 unique variants in 563 proteins, with the eventual identification of 4 drug transport genes, 2 drug metabolizing enzyme genes and 33 target genes harboring deleterious SNVs involved in pharmacological pathways, which could have a potential role in clinical settings. The current study successfully unravels the potential of personal genome sequencing in understanding the functionally relevant variations with potential influence on drug transport, metabolism and differential therapeutic outcomes. These will be essential for realizing personalized medicine through the use of comprehensive computational pipeline for systematic data mining and analysis.
Genome Sequence of Stachybotrys chartarum Strain 51-11.
Betancourt, Doris A; Dean, Timothy R; Kim, Jean; Levy, Josh
2015-10-01
The Stachybotrys chartarum strain 51-11 genome was sequenced by shotgun sequencing utilizing Illumina HiSeq 2000 and PacBio technologies. Since S. chartarum has been implicated as having health impacts within water-damaged buildings, any information extracted from the genomic sequence data relating to toxins or the metabolism of the fungus might be useful. Copyright © 2015 Betancourt et al.
NASA Astrophysics Data System (ADS)
Dominguez, L. A.; Taira, T.; Hjorleifsdottir, V.; Santoyo, M. A.
2015-12-01
Repeating earthquake sequences are sets of events that are thought to rupture the same area on the plate interface and thus provide nearly identical waveforms. We systematically analyzed seismic records from 2001 through 2014 to identify repeating earthquakes with highly correlated waveforms occurring along the subduction zone of the Cocos plate. Using the correlation coefficient (cc) and spectral coherency (coh) of the vertical components as selection criteria, we found a set of 214 sequences whose waveforms exceed cc≥95% and coh≥95%. Spatial clustering along the trench shows large variations in repeating earthquakes activity. Particularly, the rupture zone of the M8.1, 1985 earthquake shows an almost absence of characteristic repeating earthquakes, whereas the Guerrero Gap zone and the segment of the trench close to the Guerrero-Oaxaca border shows a significantly larger number of repeating earthquakes sequences. Furthermore, temporal variations associated to stress changes due to major shows episodes of unlocking and healing of the interface. Understanding the different components that control the location and recurrence time of characteristic repeating sequences is a key factor to pinpoint areas where large megathrust earthquakes may nucleate and consequently to improve the seismic hazard assessment.
Hoy, Marshal S.; Rodriguez, Rusty J.
2013-01-01
Molecular genetic analysis was conducted on two populations of the invasive non-native New Zealand mud snail (Potamopyrgus antipodarum), one from a freshwater ecosystem in Devil's Lake (Oregon, USA) and the other from an ecosystem of higher salinity in the Columbia River estuary (Hammond Harbor, Oregon, USA). To elucidate potential genetic differences between the two populations, three segments of nuclear ribosomal DNA (rDNA), the ITS1-ITS2 regions and the 18S and 28S rDNA genes were cloned and sequenced. Variant sequences within each individual were found in all three rDNA segments. Folding models were utilized for secondary structure analysis and results indicated that there were many sequences which contained structure-altering polymorphisms, which suggests they could be nonfunctional pseudogenes. In addition, analysis of molecular variance (AMOVA) was used for hierarchical analysis of genetic variance to estimate variation within and among populations and within individuals. AMOVA revealed significant variation in the ITS region between the populations and among clones within individuals, while in the 5.8S rDNA significant variation was revealed among individuals within the two populations. High levels of intragenomic variation were found in the ITS regions, which are known to be highly variable in many organisms. More interestingly, intragenomic variation was also found in the 18S and 28S rDNA, which has rarely been observed in animals and is so far unreported in Mollusca. We postulate that in these P. antipodarum populations the effects of concerted evolution are diminished due to the fact that not all of the rDNA genes in their polyploid genome should be essential for sustaining cellular function. This could lead to a lessening of selection pressures, allowing mutations to accumulate in some copies, changing them into variant sequences.
Microfluidic droplet enrichment for targeted sequencing
Eastburn, Dennis J.; Huang, Yong; Pellegrino, Maurizio; Sciambi, Adam; Ptáček, Louis J.; Abate, Adam R.
2015-01-01
Targeted sequence enrichment enables better identification of genetic variation by providing increased sequencing coverage for genomic regions of interest. Here, we report the development of a new target enrichment technology that is highly differentiated from other approaches currently in use. Our method, MESA (Microfluidic droplet Enrichment for Sequence Analysis), isolates genomic DNA fragments in microfluidic droplets and performs TaqMan PCR reactions to identify droplets containing a desired target sequence. The TaqMan positive droplets are subsequently recovered via dielectrophoretic sorting, and the TaqMan amplicons are removed enzymatically prior to sequencing. We demonstrated the utility of this approach by generating an average 31.6-fold sequence enrichment across 250 kb of targeted genomic DNA from five unique genomic loci. Significantly, this enrichment enabled a more comprehensive identification of genetic polymorphisms within the targeted loci. MESA requires low amounts of input DNA, minimal prior locus sequence information and enriches the target region without PCR bias or artifacts. These features make it well suited for the study of genetic variation in a number of research and diagnostic applications. PMID:25873629
Typing Clostridium difficile strains based on tandem repeat sequences
2009-01-01
Background Genotyping of epidemic Clostridium difficile strains is necessary to track their emergence and spread. Portability of genotyping data is desirable to facilitate inter-laboratory comparisons and epidemiological studies. Results This report presents results from a systematic screen for variation in repetitive DNA in the genome of C. difficile. We describe two tandem repeat loci, designated 'TR6' and 'TR10', which display extensive sequence variation that may be useful for sequence-based strain typing. Based on an investigation of 154 C. difficile isolates comprising 75 ribotypes, tandem repeat sequencing demonstrated excellent concordance with widely used PCR ribotyping and equal discriminatory power. Moreover, tandem repeat sequences enabled the reconstruction of the isolates' largely clonal population structure and evolutionary history. Conclusion We conclude that sequence analysis of the two repetitive loci introduced here may be highly useful for routine typing of C. difficile. Tandem repeat sequence typing resolves phylogenetic diversity to a level equivalent to PCR ribotypes. DNA sequences may be stored in databases accessible over the internet, obviating the need for the exchange of reference strains. PMID:19133124
Gomez-Smith, C Kimloi; LaPara, Timothy M; Hozalski, Raymond M
2015-07-21
The quantity and composition of bacterial biofilms growing on 10 water mains from a full-scale chloraminated water distribution system were analyzed using real-time PCR targeting the 16S rRNA gene and next-generation, high-throughput Illumina sequencing. Water mains with corrosion tubercles supported the greatest amount of bacterial biomass (n = 25; geometric mean = 2.5 × 10(7) copies cm(-2)), which was significantly higher (P = 0.04) than cement-lined cast-iron mains (n = 6; geometric mean = 2.0 × 10(6) copies cm(-2)). Despite spatial variation of community composition and bacterial abundance in water main biofilms, the communities on the interior main surfaces were surprisingly similar, containing a core group of operational taxonomic units (OTUs) assigned to only 17 different genera. Bacteria from the genus Mycobacterium dominated all communities at the main wall-bulk water interface (25-78% of the community), regardless of main age, estimated water age, main material, and the presence of corrosion products. Further sequencing of the mycobacterial heat shock protein gene (hsp65) provided species-level taxonomic resolution of mycobacteria. The two dominant Mycobacteria present, M. frederiksbergense (arithmetic mean = 85.7% of hsp65 sequences) and M. aurum (arithmetic mean = 6.5% of hsp65 sequences), are generally considered to be nonpathogenic. Two opportunistic pathogens, however, were detected at low numbers: M. hemophilum (arithmetic mean = 1.5% of hsp65 sequences) and M. abscessus (arithmetic mean = 0.006% of hsp65 sequences). Sulfate-reducing bacteria from the genus Desulfovibrio, which have been implicated in microbially influenced corrosion, dominated all communities located underneath corrosion tubercules (arithmetic mean = 67.5% of the community). This research provides novel insights into the quantity and composition of biofilms in full-scale drinking water distribution systems, which is critical for assessing the risks to public health and to the water supply infrastructure.
Positive Selection Underlies Faster-Z Evolution of Gene Expression in Birds.
Dean, Rebecca; Harrison, Peter W; Wright, Alison E; Zimmer, Fabian; Mank, Judith E
2015-10-01
The elevated rate of evolution for genes on sex chromosomes compared with autosomes (Fast-X or Fast-Z evolution) can result either from positive selection in the heterogametic sex or from nonadaptive consequences of reduced relative effective population size. Recent work in birds suggests that Fast-Z of coding sequence is primarily due to relaxed purifying selection resulting from reduced relative effective population size. However, gene sequence and gene expression are often subject to distinct evolutionary pressures; therefore, we tested for Fast-Z in gene expression using next-generation RNA-sequencing data from multiple avian species. Similar to studies of Fast-Z in coding sequence, we recover clear signatures of Fast-Z in gene expression; however, in contrast to coding sequence, our data indicate that Fast-Z in expression is due to positive selection acting primarily in females. In the soma, where gene expression is highly correlated between the sexes, we detected Fast-Z in both sexes, although at a higher rate in females, suggesting that many positively selected expression changes in females are also expressed in males. In the gonad, where intersexual correlations in expression are much lower, we detected Fast-Z for female gene expression, but crucially, not males. This suggests that a large amount of expression variation is sex-specific in its effects within the gonad. Taken together, our results indicate that Fast-Z evolution of gene expression is the product of positive selection acting on recessive beneficial alleles in the heterogametic sex. More broadly, our analysis suggests that the adaptive potential of Z chromosome gene expression may be much greater than that of gene sequence, results which have important implications for the role of sex chromosomes in speciation and sexual selection. © The Author 2015. Published by Oxford University Press on behalf of the Society for Molecular Biology and Evolution.
DOE Office of Scientific and Technical Information (OSTI.GOV)
Fitchen, W.M.; Bebout, D.G.; Hoffman, C.L.
1994-12-31
Core descriptions and regional log correlation/interpretation of Ferry Lake-Upper Glen Rose strata in the East Texas Basin exhibit the uniformity of cyclicity in these shelf units. The cyclicity is defined by an upward decrease in shale content within each cycle accompanied by an upward increase in anhydrite (Ferry Lake) or carbonate (Upper Glen Rose). Core-to-log calibration of facies indicates that formation resistivity is inversely proportional to shale content and thus is a potential proxy for facies identification beyond core control. Cycles (delineated by resistivity log patterns) were correlated for 90 mi across the shelf; they show little change in logmore » signature despite significant updip thinning due to the regional subsidence gradient. The Ferry-Lake-Upper Glen Rose intervals is interpreted as a composite sequence composed of 13 high-frequency sequences (4 in the Ferry Lake and 9 in the Upper Glen Rose). High-frequency sequences contain approximately 20 ({+-}5) cycles; in the Upper Glen Rose, successive cycles exhibit decreasing proportions of shale and increasing proportions of grain-rich carbonate. High-frequency sequences were terminated by terrigenous inundation, possibly preceded by subaerial exposure. Cycle and high-frequency sequence composition is interpreted to reflect composite, periodic(?) fluctuations is terrigeneous dilution from nearby source areas. Grainstones typically occur (stratigraphically) within the upper cycles of high-frequency sequences, where terrigeneous dilution and turbidity were least and potential for carbonate production and shoaling was greatest. Published mid-Cretaceous geographic reconstructions and climate models suggest that precipitation and runoff in the area were controlled by the seasonal amplitude in solar insolation. In this model, orbital variations, combined with subsidence, hydrography, and bathymetry, were in primary controls on Ferry Lake-Upper Glen Rose facies architecture and stratigraphic development.« less
West, Claire; James, Stephen A; Davey, Robert P; Dicks, Jo; Roberts, Ian N
2014-07-01
The ribosomal RNA encapsulates a wealth of evolutionary information, including genetic variation that can be used to discriminate between organisms at a wide range of taxonomic levels. For example, the prokaryotic 16S rDNA sequence is very widely used both in phylogenetic studies and as a marker in metagenomic surveys and the internal transcribed spacer region, frequently used in plant phylogenetics, is now recognized as a fungal DNA barcode. However, this widespread use does not escape criticism, principally due to issues such as difficulties in classification of paralogous versus orthologous rDNA units and intragenomic variation, both of which may be significant barriers to accurate phylogenetic inference. We recently analyzed data sets from the Saccharomyces Genome Resequencing Project, characterizing rDNA sequence variation within multiple strains of the baker's yeast Saccharomyces cerevisiae and its nearest wild relative Saccharomyces paradoxus in unprecedented detail. Notably, both species possess single locus rDNA systems. Here, we use these new variation datasets to assess whether a more detailed characterization of the rDNA locus can alleviate the second of these phylogenetic issues, sequence heterogeneity, while controlling for the first. We demonstrate that a strong phylogenetic signal exists within both datasets and illustrate how they can be used, with existing methodology, to estimate intraspecies phylogenies of yeast strains consistent with those derived from whole-genome approaches. We also describe the use of partial Single Nucleotide Polymorphisms, a type of sequence variation found only in repetitive genomic regions, in identifying key evolutionary features such as genome hybridization events and show their consistency with whole-genome Structure analyses. We conclude that our approach can transform rDNA sequence heterogeneity from a problem to a useful source of evolutionary information, enabling the estimation of highly accurate phylogenies of closely related organisms, and discuss how it could be extended to future studies of multilocus rDNA systems. [concerted evolution; genome hydridisation; phylogenetic analysis; ribosomal DNA; whole genome sequencing; yeast]. © The Author(s) 2014. Published by Oxford University Press, on behalf of the Society of Systematic Biologists.
Schmidt, DJ; Pickett, BE; Camacho, D; Comach, G; Xhaja, K; Lennon, NJ; Rizzolo, K; de Bosch, N; Becerra, A; Nogueira, ML; Mondini, A; da Silva, EV; Vasconcelos, PF; Muñoz-Jordán, JL; Santiago, GA; Ocazionez, R; Gehrke, L; Lefkowitz, EJ; Birren, BW; Henn, MR; Bosch, I
2013-01-01
Dengue virus currently causes 50-100 million infections annually. Comprehensive knowledge about the evolution of Dengue in response to selection pressure is currently unavailable, but would greatly enhance vaccine design efforts. In the current study, we sequenced 187 new dengue virus serotype 3(DENV-3) genotype III whole genomes isolated from Asia and the Americas. We analyzed them together with previously-sequenced isolates to gain a more detailed understanding of the evolutionary adaptations existing in this prevalent American serotype. In order to analyze the phylogenetic dynamics of DENV-3 during outbreak periods; we incorporated datasets of 48 and 11 sequences spanning two major outbreaks in Venezuela during 2001 and 2007-2008 respectively. Our phylogenetic analysis of newly sequenced viruses shows that subsets of genomes cluster primarily by geographic location, and secondarily by time of virus isolation. DENV-3 genotype III sequences from Asia are significantly divergent from those from the Americas due to their geographical separation and subsequent speciation. We measured amino acid variation for the E protein by calculating the Shannon entropy at each position between Asian and American genomes. We found a cluster of 7 amino acid substitutions having high variability within E protein domain III, which has previously been implicated in serotype-specific neutralization escape mutants. No novel mutations were found in the E protein of sequences isolated during either Venezuelan outbreak. Shannon entropy analysis of the NS5 polymerase mature protein revealed that a G374E mutation, in a region that contributes to interferon resistance in other flaviviruses by interfering with JAK-STAT signaling was present in both the Asian and American sequences from the 2007-2008 Venezuelan outbreak, but was absent in the sequences from the 2001 Venezuelan outbreak. In addition to E, several NS5 amino acid changes were unique to the 2007-2008 epidemic in Venezuela and may give additional insight into the adaptive response of DENV-3 at the population level. PMID:21964598
Can Yeast (S. cerevisiae) Metabolic Volatiles Provide Polymorphic Signaling?
Arguello, J. Roman; Sellanes, Carolina; Lou, Yann Ru; Raguso, Robert A.
2013-01-01
Chemical signaling between organisms is a ubiquitous and evolutionarily dynamic process that helps to ensure mate recognition, location of nutrients, avoidance of toxins, and social cooperation. Evolutionary changes in chemical communication systems progress through natural variation within the organism generating the signal as well as the responding individuals. A promising yet poorly understood system with which to probe the importance of this variation exists between D. melanogaster and S. cerevisiae. D. melanogaster relies on yeast for nutrients, while also serving as a vector for yeast cell dispersal. Both are outstanding genetic and genomic models, with Drosophila also serving as a preeminent model for sensory neurobiology. To help develop these two genetic models as an ecological model, we have tested if - and to what extent - S. cerevisiae is capable of producing polymorphic signaling through variation in metabolic volatiles. We have carried out a chemical phenotyping experiment for 14 diverse accessions within a common garden random block design. Leveraging genomic sequences for 11 of the accessions, we ensured a genetically broad sample and tested for phylogenetic signal arising from phenotypic dataset. Our results demonstrate that significant quantitative differences for volatile blends do exist among S. cerevisiae accessions. Of particular ecological relevance, the compounds driving the blend differences (acetoin, 2-phenyl ethanol and 3-methyl-1-butanol) are known ligands for D. melanogasters chemosensory receptors, and are related to sensory behaviors. Though unable to correlate the genetic and volatile measurements, our data point clear ways forward for behavioral assays aimed at understanding the implications of this variation. PMID:23990899
VarDetect: a nucleotide sequence variation exploratory tool
Ngamphiw, Chumpol; Kulawonganunchai, Supasak; Assawamakin, Anunchai; Jenwitheesuk, Ekachai; Tongsima, Sissades
2008-01-01
Background Single nucleotide polymorphisms (SNPs) are the most commonly studied units of genetic variation. The discovery of such variation may help to identify causative gene mutations in monogenic diseases and SNPs associated with predisposing genes in complex diseases. Accurate detection of SNPs requires software that can correctly interpret chromatogram signals to nucleotides. Results We present VarDetect, a stand-alone nucleotide variation exploratory tool that automatically detects nucleotide variation from fluorescence based chromatogram traces. Accurate SNP base-calling is achieved using pre-calculated peak content ratios, and is enhanced by rules which account for common sequence reading artifacts. The proposed software tool is benchmarked against four other well-known SNP discovery software tools (PolyPhred, novoSNP, Genalys and Mutation Surveyor) using fluorescence based chromatograms from 15 human genes. These chromatograms were obtained from sequencing 16 two-pooled DNA samples; a total of 32 individual DNA samples. In this comparison of automatic SNP detection tools, VarDetect achieved the highest detection efficiency. Availability VarDetect is compatible with most major operating systems such as Microsoft Windows, Linux, and Mac OSX. The current version of VarDetect is freely available at . PMID:19091032
Model-based quality assessment and base-calling for second-generation sequencing data.
Bravo, Héctor Corrada; Irizarry, Rafael A
2010-09-01
Second-generation sequencing (sec-gen) technology can sequence millions of short fragments of DNA in parallel, making it capable of assembling complex genomes for a small fraction of the price and time of previous technologies. In fact, a recently formed international consortium, the 1000 Genomes Project, plans to fully sequence the genomes of approximately 1200 people. The prospect of comparative analysis at the sequence level of a large number of samples across multiple populations may be achieved within the next five years. These data present unprecedented challenges in statistical analysis. For instance, analysis operates on millions of short nucleotide sequences, or reads-strings of A,C,G, or T's, between 30 and 100 characters long-which are the result of complex processing of noisy continuous fluorescence intensity measurements known as base-calling. The complexity of the base-calling discretization process results in reads of widely varying quality within and across sequence samples. This variation in processing quality results in infrequent but systematic errors that we have found to mislead downstream analysis of the discretized sequence read data. For instance, a central goal of the 1000 Genomes Project is to quantify across-sample variation at the single nucleotide level. At this resolution, small error rates in sequencing prove significant, especially for rare variants. Sec-gen sequencing is a relatively new technology for which potential biases and sources of obscuring variation are not yet fully understood. Therefore, modeling and quantifying the uncertainty inherent in the generation of sequence reads is of utmost importance. In this article, we present a simple model to capture uncertainty arising in the base-calling procedure of the Illumina/Solexa GA platform. Model parameters have a straightforward interpretation in terms of the chemistry of base-calling allowing for informative and easily interpretable metrics that capture the variability in sequencing quality. Our model provides these informative estimates readily usable in quality assessment tools while significantly improving base-calling performance. © 2009, The International Biometric Society.
Intraspecific variation in Cryptocaryon irritans.
Diggles, B K; Adlard, R D
1997-01-01
Intraspecific variation in the ciliate Cryptocaryon irritans was examined using sequences of the first internal transcribed spacer region (ITS-1) of ribosomal DNA (rDNA) combined with developmental and morphological characters. Amplified rDNA sequences consisting of 151 bases of the flanking 18 S and 5.8 S regions, and the entire ITS-1 region (169 or 170 bases), were determined and compared for 16 isolates of C. irritans from Australia, Israel and the USA. There was one variable base between isolates in the 18 S region and 11 variable bases in the ITS-1 region. Despite their similar morphology, significant sequence variation (4.1% divergence) and developmental differences indicate that Australian C. irritans isolates from estuarine (Moreton Bay) and coral reef (Heron Island) environments are distinct. The Heron Island isolate was genetically closer to morphologically dissimilar isolates from Israel (1.8% divergence) and the USA (2.3% divergence) than it was to the Moreton Bay isolates. Three isolates maintained in our laboratory since February 1994 differed in sequence from earlier laboratory isolates (2.9% to 3.5% divergence), even though all were similar morphologically and originated from the same source. During this time the sequence of the isolates from wild fish in Moreton Bay remained unchanged. These genetic differences indicate the existence of a founder effect in laboratory populations of C. irritans. The genetic variation found here, combined with known morphological and developmental differences, is used to characterise four strains of C. irritans.
2010-01-01
Background The maturing field of genomics is rapidly increasing the number of sequenced genomes and producing more information from those previously sequenced. Much of this additional information is variation data derived from sampling multiple individuals of a given species with the goal of discovering new variants and characterising the population frequencies of the variants that are already known. These data have immense value for many studies, including those designed to understand evolution and connect genotype to phenotype. Maximising the utility of the data requires that it be stored in an accessible manner that facilitates the integration of variation data with other genome resources such as gene annotation and comparative genomics. Description The Ensembl project provides comprehensive and integrated variation resources for a wide variety of chordate genomes. This paper provides a detailed description of the sources of data and the methods for creating the Ensembl variation databases. It also explores the utility of the information by explaining the range of query options available, from using interactive web displays, to online data mining tools and connecting directly to the data servers programmatically. It gives a good overview of the variation resources and future plans for expanding the variation data within Ensembl. Conclusions Variation data is an important key to understanding the functional and phenotypic differences between individuals. The development of new sequencing and genotyping technologies is greatly increasing the amount of variation data known for almost all genomes. The Ensembl variation resources are integrated into the Ensembl genome browser and provide a comprehensive way to access this data in the context of a widely used genome bioinformatics system. All Ensembl data is freely available at http://www.ensembl.org and from the public MySQL database server at ensembldb.ensembl.org. PMID:20459805
Hermes Transposon Distribution and Structure in Musca domestica
Subramanian, Ramanand A.; Cathcart, Laura A.; Krafsur, Elliot S.; Atkinson, Peter W.
2009-01-01
Hermes are hAT transposons from Musca domestica that are very closely related to the hobo transposons from Drosophila melanogaster and are useful as gene vectors in a wide variety of organisms including insects, planaria, and yeast. hobo elements show distinct length variations in a rapidly evolving region of the transposase-coding region as a result of expansions and contractions of a simple repeat sequence encoding 3 amino acids threonine, proline, and glutamic acid (TPE). These variations in length may influence the function of the protein and the movement of hobo transposons in natural populations. Here, we determine the distribution of Hermes in populations of M. domestica as well as whether Hermes transposase has undergone similar sequence expansions and contractions during its evolution in this species. Hermes transposons were found in all M. domestica individuals sampled from 14 populations collected from 4 continents. All individuals with Hermes transposons had evidence for the presence of intact transposase open reading frames, and little sequence variation was observed among Hermes elements. A systematic analysis of the TPE-homologous region of the Hermes transposase-coding region revealed no evidence for length variation. The simple sequence repeat found in hobo elements is a feature of this transposon that evolved since the divergence of hobo and Hermes. PMID:19366812
Species conservation and natural variation among populations [Chapter 5
Leonard F. Ruggiero; Michael K. Schwartz; Keith B. Aubry; Charles J. Krebs; Amanda Stanley; Steven W. Buskirk
2000-01-01
In conservation planning, the importance of natural variation is often given inadequate consideration. However, ignoring the implications of variation within species may result in conservation strategies that jeopardize, rather than conserve, target species (see Grieg 1979; Turcek 1951; Storfer 1999). Natural variation in the traits of individuals and populations is...
USDA-ARS?s Scientific Manuscript database
1. Plant functional traits provide a mechanistic basis for understanding ecological variation among plant species and the implications of this variation for species distribution, community assembly and restoration. 2. The bulk of our functional trait understanding, however, is centered on traits rel...
Variation analysis and gene annotation of eight MHC haplotypes: The MHC Haplotype Project
Horton, Roger; Gibson, Richard; Coggill, Penny; Miretti, Marcos; Allcock, Richard J.; Almeida, Jeff; Forbes, Simon; Gilbert, James G. R.; Halls, Karen; Harrow, Jennifer L.; Hart, Elizabeth; Howe, Kevin; Jackson, David K.; Palmer, Sophie; Roberts, Anne N.; Sims, Sarah; Stewart, C. Andrew; Traherne, James A.; Trevanion, Steve; Wilming, Laurens; Rogers, Jane; de Jong, Pieter J.; Elliott, John F.; Sawcer, Stephen; Todd, John A.; Trowsdale, John
2008-01-01
The human major histocompatibility complex (MHC) is contained within about 4 Mb on the short arm of chromosome 6 and is recognised as the most variable region in the human genome. The primary aim of the MHC Haplotype Project was to provide a comprehensively annotated reference sequence of a single, human leukocyte antigen-homozygous MHC haplotype and to use it as a basis against which variations could be assessed from seven other similarly homozygous cell lines, representative of the most common MHC haplotypes in the European population. Comparison of the haplotype sequences, including four haplotypes not previously analysed, resulted in the identification of >44,000 variations, both substitutions and indels (insertions and deletions), which have been submitted to the dbSNP database. The gene annotation uncovered haplotype-specific differences and confirmed the presence of more than 300 loci, including over 160 protein-coding genes. Combined analysis of the variation and annotation datasets revealed 122 gene loci with coding substitutions of which 97 were non-synonymous. The haplotype (A3-B7-DR15; PGF cell line) designated as the new MHC reference sequence, has been incorporated into the human genome assembly (NCBI35 and subsequent builds), and constitutes the largest single-haplotype sequence of the human genome to date. The extensive variation and annotation data derived from the analysis of seven further haplotypes have been made publicly available and provide a framework and resource for future association studies of all MHC-associated diseases and transplant medicine. PMID:18193213
Natarajan, Sathishkumar; Kim, Hoy-Taek; Thamilarasan, Senthil Kumar; Veerappan, Karpagam; Park, Jong-In; Nou, Ill-Sup
2016-01-01
Powdery mildew is one of the most common fungal diseases in the world. This disease frequently affects melon (Cucumis melo L.) and other Cucurbitaceous family crops in both open field and greenhouse cultivation. One of the goals of genomics is to identify the polymorphic loci responsible for variation in phenotypic traits. In this study, powdery mildew disease assessment scores were calculated for four melon accessions, 'SCNU1154', 'Edisto47', 'MR-1', and 'PMR5'. To investigate the genetic variation of these accessions, whole genome re-sequencing using the Illumina HiSeq 2000 platform was performed. A total of 754,759,704 quality-filtered reads were generated, with an average of 82.64% coverage relative to the reference genome. Comparisons of the sequences for the melon accessions revealed around 7.4 million single nucleotide polymorphisms (SNPs), 1.9 million InDels, and 182,398 putative structural variations (SVs). Functional enrichment analysis of detected variations classified them into biological process, cellular component and molecular function categories. Further, a disease-associated QTL map was constructed for 390 SNPs and 45 InDels identified as related to defense-response genes. Among them 112 SNPs and 12 InDels were observed in powdery mildew responsive chromosomes. Accordingly, this whole genome re-sequencing study identified SNPs and InDels associated with defense genes that will serve as candidate polymorphisms in the search for sources of resistance against powdery mildew disease and could accelerate marker-assisted breeding in melon.
Doddapaneni, Harshavardhan; Yao, Jiqiang; Lin, Hong; Walker, M Andrew; Civerolo, Edwin L
2006-01-01
Background The Gram-negative, xylem-limited phytopathogenic bacterium Xylella fastidiosa is responsible for causing economically important diseases in grapevine, citrus and many other plant species. Despite its economic impact, relatively little is known about the genomic variations among strains isolated from different hosts and their influence on the population genetics of this pathogen. With the availability of genome sequence information for four strains, it is now possible to perform genome-wide analyses to identify and categorize such DNA variations and to understand their influence on strain functional divergence. Results There are 1,579 genes and 194 non-coding homologous sequences present in the genomes of all four strains, representing a 76. 2% conservation of the sequenced genome. About 60% of the X. fastidiosa unique sequences exist as tandem gene clusters of 6 or more genes. Multiple alignments identified 12,754 SNPs and 14,449 INDELs in the 1528 common genes and 20,779 SNPs and 10,075 INDELs in the 194 non-coding sequences. The average SNP frequency was 1.08 × 10-2 per base pair of DNA and the average INDEL frequency was 2.06 × 10-2 per base pair of DNA. On an average, 60.33% of the SNPs were synonymous type while 39.67% were non-synonymous type. The mutation frequency, primarily in the form of external INDELs was the main type of sequence variation. The relative similarity between the strains was discussed according to the INDEL and SNP differences. The number of genes unique to each strain were 60 (9a5c), 54 (Dixon), 83 (Ann1) and 9 (Temecula-1). A sub-set of the strain specific genes showed significant differences in terms of their codon usage and GC composition from the native genes suggesting their xenologous origin. Tandem repeat analysis of the genomic sequences of the four strains identified associations of repeat sequences with hypothetical and phage related functions. Conclusion INDELs and strain specific genes have been identified as the main source of variations among strains, with individual strains showing different rates of genome evolution. Based on these genome comparisons, it appears that the Pierce's disease strain Temecula-1 genome represents the ancestral genome of the X. fastidiosa. Results of this analysis are publicly available in the form of a web database. PMID:16948851
Jelokhani-Niaraki, Saber; Tahmoorespur, Mojtaba; Bitaraf-Sani, Morteza
2015-01-01
Very little is known about LHR and FSHR genes of domestic dromedary camels. The main objective of this study was to determine and analyze partial genomic regions of FSHR and LHR genes in dromedary camels for the first time. To this end, a total of50 DNA samples belonging to dromedary camels raised in Iran were sent for sequencing (25 samples of each gene). We compared the nucleotide sequences of Camelus dromedarius with corresponding sequences of previously published FSHR and LHR genes in bactrian camels and other species. According to the data, the same nucleotide variation was identified in both regions of the two camel species. The alignment of deduced protein sequences of the two different species revealed an amino acid variation at the FSHR region. No evidence of amino acid variation was observed, however, in LHR sequences. Phylogenetic analysis indicated that both camel species had a close relationship and clustered together in a separate branch. This was further confirmed by genetic distance values illustrating significant sequence identity between Camelus dromedarius and Camelus bactrianus. Interestingly, sequence comparisons revealed heterozygote patterns in FSHR sequences isolated from dromedary camels of Iran. In comparison to other species, this camel contains three amino acid substitutions at 5, 67, and 105 positions in the FSHR coding region. These positions are found exclusively in camels and can be considered as species specific. The results of our study can be used for hormone functionality research (FSHR and LHR) as well as reproduction-linked polymorphisms and breeding programs. PMID:27844002
Jelokhani-Niaraki, Saber; Tahmoorespur, Mojtaba; Bitaraf-Sani, Morteza
2015-06-01
Very little is known about LHR and FSHR genes of domestic dromedary camels. The main objective of this study was to determine and analyze partial genomic regions of FSHR and LHR genes in dromedary camels for the first time. To this end, a total of50 DNA samples belonging to dromedary camels raised in Iran were sent for sequencing (25 samples of each gene). We compared the nucleotide sequences of Camelus dromedarius with corresponding sequences of previously published FSHR and LHR genes in bactrian camels and other species. According to the data, the same nucleotide variation was identified in both regions of the two camel species. The alignment of deduced protein sequences of the two different species revealed an amino acid variation at the FSHR region. No evidence of amino acid variation was observed, however, in LHR sequences. Phylogenetic analysis indicated that both camel species had a close relationship and clustered together in a separate branch. This was further confirmed by genetic distance values illustrating significant sequence identity between Camelus dromedarius and Camelus bactrianus . Interestingly, sequence comparisons revealed heterozygote patterns in FSHR sequences isolated from dromedary camels of Iran. In comparison to other species, this camel contains three amino acid substitutions at 5, 67, and 105 positions in the FSHR coding region. These positions are found exclusively in camels and can be considered as species specific. The results of our study can be used for hormone functionality research ( FSHR and LHR ) as well as reproduction-linked polymorphisms and breeding programs.
Karched, Maribasappa; Furgang, David; Planet, Paul J; DeSalle, Rob; Fine, Daniel H
2012-03-01
Aggregatibacter actinomycetemcomitans is implicated in localized aggressive periodontitis. We report the first genome sequence of an A. actinomycetemcomitans strain isolated from an Old World primate.
Aokic, Jun-ya; Kawase, Junya; Hamada, Kazuhisa; Fujimoto, Hiroshi; Yamamoto, Ikki; Usuki, Hironori
2018-01-01
Greater amberjack (Seriola dumerili) is distributed in tropical and temperate waters worldwide and is an important aquaculture fish. We carried out de novo sequencing of the greater amberjack genome to construct a reference genome sequence to identify single nucleotide polymorphisms (SNPs) for breeding amberjack by marker-assisted or gene-assisted selection as well as to identify functional genes for biological traits. We obtained 200 times coverage and constructed a high-quality genome assembly using next generation sequencing technology. The assembled sequences were aligned onto a yellowtail (Seriola quinqueradiata) radiation hybrid (RH) physical map by sequence homology. A total of 215 of the longest amberjack sequences, with a total length of 622.8 Mbp (92% of the total length of the genome scaffolds), were lined up on the yellowtail RH map. We resequenced the whole genomes of 20 greater amberjacks and mapped the resulting sequences onto the reference genome sequence. About 186,000 nonredundant SNPs were successfully ordered on the reference genome. Further, we found differences in the genome structural variations between two greater amberjack populations using BreakDancer. We also analyzed the greater amberjack transcriptome and mapped the annotated sequences onto the reference genome sequence. PMID:29785397
Wong, Gerard; Leckie, Christopher; Gorringe, Kylie L; Haviv, Izhak; Campbell, Ian G; Kowalczyk, Adam
2010-04-15
High-density single nucleotide polymorphism (SNP) genotyping arrays are efficient and cost effective platforms for the detection of copy number variation (CNV). To ensure accuracy in probe synthesis and to minimize production costs, short oligonucleotide probe sequences are used. The use of short probe sequences limits the specificity of binding targets in the human genome. The specificity of these short probeset sequences has yet to be fully analysed against a normal reference human genome. Sequence similarity can artificially elevate or suppress copy number measurements, and hence reduce the reliability of affected probe readings. For the purpose of detecting narrow CNVs reliably down to the width of a single probeset, sequence similarity is an important issue that needs to be addressed. We surveyed the Affymetrix Human Mapping SNP arrays for probeset sequence similarity against the reference human genome. Utilizing sequence similarity results, we identified a collection of fine-scaled putative CNVs between gender from autosomal probesets whose sequence matches various loci on the sex chromosomes. To detect these variations, we utilized our statistical approach, Detecting REcurrent Copy number change using rank-order Statistics (DRECS), and showed that its performance was superior and more stable than the t-test in detecting CNVs. Through the application of DRECS on the HapMap population datasets with multi-matching probesets filtered, we identified biologically relevant SNPs in aberrant regions across populations with known association to physical traits, such as height, covered by the span of a single probe. This provided empirical confirmation of the existence of naturally occurring narrow CNVs as well as the sensitivity of the Affymetrix SNP array technology in detecting them. The MATLAB implementation of DRECS is available at http://ww2.cs.mu.oz.au/ approximately gwong/DRECS/index.html.
Analysis of human herpesvirus-6 IE1 sequence variation in clinical samples.
Stanton, Richard; Wilkinson, Gavin W G; Fox, Julie D
2003-12-01
Herpesvirus immediate early (IE) proteins are known to play key roles in establishing productive infections, regulating reactivation from latency, and creating a cellular environment favourable to viral replication. Human herpesvirus-6 (HHV-6) IE genes have not been studied as intensively as their homologues in the prototype betaherpesvirus human cytomegalovirus (HCMV). Whilst the HCMV IE1 gene is relatively conserved, early studies indicated that HHV-6 IE1 exhibited a high level of sequence variation between HHV-6A and HHV-6B isolates, although the observation was based primarily on virus stocks that had been isolated and propagated in vitro. In this study, we investigated the level of HHV-6 IE1 sequence variation in vivo by direct sequencing of circulating virus in clinical samples without prior in vitro culture. Sequences exactly matching those reported for reference HHV-6 isolates were identified in clinical samples, thus the HHV-6 laboratory strains used in the majority of in vitro studies appear to be representative of virus circulating in vivo with respect to the IE1 gene. The HHV-6 IE1 sequence is also conserved in reference strains that had been passaged extensively in vitro. The high degree of divergence between variant A and B type IE1 sequences was confirmed, but interestingly HHV-6B IE1 sequences were observed to further segregate into two distinct subgroups, with the laboratory strains Z29 and HST representative of these two subgroups. Within each HHV-6B subgroup, a remarkably high level of homology was observed. Thus the HHV-6 IE1 sequence appears highly stable, underlining its potential importance to the viral life cycle. Copyright 2003 Wiley-Liss, Inc.
LenVarDB: database of length-variant protein domains.
Mutt, Eshita; Mathew, Oommen K; Sowdhamini, Ramanathan
2014-01-01
Protein domains are functionally and structurally independent modules, which add to the functional variety of proteins. This array of functional diversity has been enabled by evolutionary changes, such as amino acid substitutions or insertions or deletions, occurring in these protein domains. Length variations (indels) can introduce changes at structural, functional and interaction levels. LenVarDB (freely available at http://caps.ncbs.res.in/lenvardb/) traces these length variations, starting from structure-based sequence alignments in our Protein Alignments organized as Structural Superfamilies (PASS2) database, across 731 structural classification of proteins (SCOP)-based protein domain superfamilies connected to 2 730 625 sequence homologues. Alignment of sequence homologues corresponding to a structural domain is available, starting from a structure-based sequence alignment of the superfamily. Orientation of the length-variant (indel) regions in protein domains can be visualized by mapping them on the structure and on the alignment. Knowledge about location of length variations within protein domains and their visual representation will be useful in predicting changes within structurally or functionally relevant sites, which may ultimately regulate protein function. Non-technical summary: Evolutionary changes bring about natural changes to proteins that may be found in many organisms. Such changes could be reflected as amino acid substitutions or insertions-deletions (indels) in protein sequences. LenVarDB is a database that provides an early overview of observed length variations that were set among 731 protein families and after examining >2 million sequences. Indels are followed up to observe if they are close to the active site such that they can affect the activity of proteins. Inclusion of such information can aid the design of bioengineering experiments.
Ginther, C; Corach, D; Penacino, G A; Rey, J A; Carnese, F R; Hutz, M H; Anderson, A; Just, J; Salzano, F M; King, M C
1993-01-01
DNA samples from 60 Mapuche Indians, representing 39 maternal lineages, were genetically characterized for (1) nucleotide sequences of the mtDNA control region; (2) presence or absence of a nine base duplication in mtDNA region V; (3) HLA loci DRB1 and DQA1; (4) variation at three nuclear genes with short tandem repeats; and (5) variation at the polymorphic marker D2S44. The genetic profile of the Mapuche population was compared to other Amerinds and to worldwide populations. Two highly polymorphic portions of the mtDNA control region, comprising 650 nucleotides, were amplified by the polymerase chain reaction (PCR) and directly sequenced. The 39 maternal lineages were defined by two or three generation families identified by the Mapuches. These 39 lineages included 19 different mtDNA sequences that could be grouped into four classes. The same classes of sequences appear in other Amerinds from North, Central, and South American populations separated by thousands of miles, suggesting that the origin of the mtDNA patterns predates the migration to the Americas. The mtDNA sequence similarity between Amerind populations suggests that the migration throughout the Americas occurred rapidly relative to the mtDNA mutation rate. HLA DRB1 alleles 1602 and 1402 were frequent among the Mapuches. These alleles also occur at high frequency among other Amerinds in North and South America, but not among Spanish, Chinese or African-American populations. The high frequency of these alleles throughout the Americas, and their specificity to the Americas, supports the hypothesis that Mapuches and other Amerind groups are closely related.(ABSTRACT TRUNCATED AT 250 WORDS)
Combelas, Nicolas; Holmblat, Barbara; Joffret, Marie-Line; Colbère-Garapin, Florence; Delpeyroux, Francis
2011-01-01
Genetic recombination in RNA viruses was discovered many years ago for poliovirus (PV), an enterovirus of the Picornaviridae family, and studied using PV or other picornaviruses as models. Recently, recombination was shown to be a general phenomenon between different types of enteroviruses of the same species. In particular, the interest for this mechanism of genetic plasticity was renewed with the emergence of pathogenic recombinant circulating vaccine-derived polioviruses (cVDPVs), which were implicated in poliomyelitis outbreaks in several regions of the world with insufficient vaccination coverage. Most of these cVDPVs had mosaic genomes constituted of mutated poliovaccine capsid sequences and part or all of the non-structural sequences from other human enteroviruses of species C (HEV-C), in particular coxsackie A viruses. A study in Madagascar showed that recombinant cVDPVs had been co-circulating in a small population of children with many different HEV-C types. This viral ecosystem showed a surprising and extensive biodiversity associated to several types and recombinant genotypes, indicating that intertypic genetic recombination was not only a mechanism of evolution for HEV-C, but an usual mode of genetic plasticity shaping viral diversity. Results suggested that recombination may be, in conjunction with mutations, implicated in the phenotypic diversity of enterovirus strains and in the emergence of new pathogenic strains. Nevertheless, little is known about the rules and mechanisms which govern genetic exchanges between HEV-C types, as well as about the importance of intertypic recombination in generating phenotypic variation. This review summarizes our current knowledge of the mechanisms of evolution of PV, in particular recombination events leading to the emergence of recombinant cVDPVs. PMID:21994791
Musunuru, Kiran; Bernstein, Daniel; Cole, F Sessions; Khokha, Mustafa K; Lee, Frank S; Lin, Shin; McDonald, Thomas V; Moskowitz, Ivan P; Quertermous, Thomas; Sankaran, Vijay G; Schwartz, David A; Silverman, Edwin K; Zhou, Xiaobo; Hasan, Ahmed A K; Luo, Xiao-Zhong James
2018-04-01
The National Institutes of Health have made substantial investments in genomic studies and technologies to identify DNA sequence variants associated with human disease phenotypes. The National Heart, Lung, and Blood Institute has been at the forefront of these commitments to ascertain genetic variation associated with heart, lung, blood, and sleep diseases and related clinical traits. Genome-wide association studies, exome- and genome-sequencing studies, and exome-genotyping studies of the National Heart, Lung, and Blood Institute-funded epidemiological and clinical case-control studies are identifying large numbers of genetic variants associated with heart, lung, blood, and sleep phenotypes. However, investigators face challenges in identification of genomic variants that are functionally disruptive among the myriad of computationally implicated variants. Studies to define mechanisms of genetic disruption encoded by computationally identified genomic variants require reproducible, adaptable, and inexpensive methods to screen candidate variant and gene function. High-throughput strategies will permit a tiered variant discovery and genetic mechanism approach that begins with rapid functional screening of a large number of computationally implicated variants and genes for discovery of those that merit mechanistic investigation. As such, improved variant-to-gene and gene-to-function screens-and adequate support for such studies-are critical to accelerating the translation of genomic findings. In this White Paper, we outline the variety of novel technologies, assays, and model systems that are making such screens faster, cheaper, and more accurate, referencing published work and ongoing work supported by the National Heart, Lung, and Blood Institute's R21/R33 Functional Assays to Screen Genomic Hits program. We discuss priorities that can accelerate the impressive but incomplete progress represented by big data genomic research. © 2018 American Heart Association, Inc.
Combelas, Nicolas; Holmblat, Barbara; Joffret, Marie-Line; Colbère-Garapin, Florence; Delpeyroux, Francis
2011-08-01
Genetic recombination in RNA viruses was discovered many years ago for poliovirus (PV), an enterovirus of the Picornaviridae family, and studied using PV or other picornaviruses as models. Recently, recombination was shown to be a general phenomenon between different types of enteroviruses of the same species. In particular, the interest for this mechanism of genetic plasticity was renewed with the emergence of pathogenic recombinant circulating vaccine-derived polioviruses (cVDPVs), which were implicated in poliomyelitis outbreaks in several regions of the world with insufficient vaccination coverage. Most of these cVDPVs had mosaic genomes constituted of mutated poliovaccine capsid sequences and part or all of the non-structural sequences from other human enteroviruses of species C (HEV-C), in particular coxsackie A viruses. A study in Madagascar showed that recombinant cVDPVs had been co-circulating in a small population of children with many different HEV-C types. This viral ecosystem showed a surprising and extensive biodiversity associated to several types and recombinant genotypes, indicating that intertypic genetic recombination was not only a mechanism of evolution for HEV-C, but an usual mode of genetic plasticity shaping viral diversity. Results suggested that recombination may be, in conjunction with mutations, implicated in the phenotypic diversity of enterovirus strains and in the emergence of new pathogenic strains. Nevertheless, little is known about the rules and mechanisms which govern genetic exchanges between HEV-C types, as well as about the importance of intertypic recombination in generating phenotypic variation. This review summarizes our current knowledge of the mechanisms of evolution of PV, in particular recombination events leading to the emergence of recombinant cVDPVs.
Bolin, Lisa L; Ahmad, Shamim; Levy, Laura S
2011-10-15
Feline leukemia virus (FeLV) is a natural retrovirus of domestic cats associated with degenerative, proliferative and malignant diseases. Studies of FeLV infection in a cohort of naturally infected cats were undertaken to examine FeLV variation, the selective pressures operative in FeLV infection that lead to predominance of natural variants, and the consequences for infection and disease progression. A unique variant, designated FeLV-945, was identified as the predominant isolate in the cohort and was associated with non-T-cell diseases including multicentric lymphoma. FeLV-945 was assigned to the FeLV-A subgroup based on sequence analysis and receptor utilization, but was shown to differ in sequence from a prototype member of FeLV-A, designated FeLV-A/61E, in the long terminal repeat (LTR) and the surface glycoprotein gene (SU). A unique sequence motif in the FeLV-945 LTR was shown to function as a transcriptional enhancer and to confer a replicative advantage. The FeLV-945 SU protein was observed to differ in sequence as compared to FeLV-A/61E within functional domains known to determine receptor selection and binding. Experimental infection of newborn cats was performed using wild type FeLV-A/61E or recombinant FeLV-A/61E in which the LTR (61E/945L) or LTR and SU (61E/945SL) were exchanged for that of FeLV-945. Infection with either FeLV-A/61E or 61E/945L resulted in T-cell lymphoma of the thymus, although 61E/945L caused disease significantly more rapidly. In contrast, infection with 61E/945SL resulted in the rapid induction of a multicentric lymphoma of B-cell origin, thus recapitulating the outcome of natural infection and implicating FeLV-945 SU as a determinant of disease outcome. Recombinant FeLV-B was detected infrequently and at low levels in multicentric lymphomas, and was thereby not implicated in disease induction. Preliminary studies of receptor interaction indicated that virus particles bearing FeLV-945 SU bind to the FeLV-A receptor more efficiently than do particles bearing FeLV-A/61E SU, and that soluble SU proteins expressed from the viruses demonstrate the same differential binding phenotype. Preliminary mutational analysis of FeLV-945 was performed by exchanging regions containing either the primary receptor binding determinant, VRA, the secondary determinant, VRB, or a proline-rich region, PRR, with that of FeLV-A/61E. Results implicated a region containing VRA as a minor contributor, while a region containing VRB largely conferred increased binding efficiency. Copyright © 2011 Elsevier B.V. All rights reserved.
Sampathkumar, Raghavan; Shadabi, Elnaz; Luo, Ma
2012-01-01
As of February 2012, 50 circulating recombinant forms (CRFs) have been reported for HIV-1 while one CRF for HIV-2. Also according to HIV sequence compendium 2011, the HIV sequence database is replete with 414,398 sequences. The fact that there are CRFs, which are an amalgamation of sequences derived from six or more subtypes (CRF27_cpx (cpx refers to complex) is a mosaic with sequences from 6 different subtypes besides an unclassified fragment), serves as a testimony to the continual divergent evolution of the virus with its approximate 1% per year rate of evolution, and this phenomena per se poses tremendous challenge for vaccine development against HIV/AIDS, a devastating disease that has killed 1.8 million patients in 2010. Here, we explore the interaction between HIV-1 and host genetic variation in the context of HIV/AIDS and antiretroviral therapy response. PMID:22666249
Liu, Qing; Zhu, Shenghua; Mizuno, Sahoko; Kimura, Masatsugu; Liu, Peina; Isomura, Shin; Wang, Xingzhen; Kawamoto, Fumihiko
1998-01-01
By two PCR-based diagnostic methods, Plasmodium malariae infections have been rediscovered at two foci in the Sichuan province of China, a region where no cases of P. malariae have been officially reported for the last 2 decades. In addition, a variant form of P. malariae which has a deletion of 19 bp and seven substitutions of base pairs in the target sequence of the small-subunit (SSU) rRNA gene was detected with high frequency. Alignment analysis of Plasmodium sp. SSU rRNA gene sequences revealed that the 5′ region of the variant sequence is identical to that of P. vivax or P. knowlesi and its 3′ region is identical to that of P. malariae. The same sequence variations were also found in P. malariae isolates collected along the Thai-Myanmar border, suggesting a wide distribution of this variant form from southern China to Southeast Asia. PMID:9774600
A survey of tools for variant analysis of next-generation genome sequencing data
Pabinger, Stephan; Dander, Andreas; Fischer, Maria; Snajder, Rene; Sperk, Michael; Efremova, Mirjana; Krabichler, Birgit; Speicher, Michael R.; Zschocke, Johannes
2014-01-01
Recent advances in genome sequencing technologies provide unprecedented opportunities to characterize individual genomic landscapes and identify mutations relevant for diagnosis and therapy. Specifically, whole-exome sequencing using next-generation sequencing (NGS) technologies is gaining popularity in the human genetics community due to the moderate costs, manageable data amounts and straightforward interpretation of analysis results. While whole-exome and, in the near future, whole-genome sequencing are becoming commodities, data analysis still poses significant challenges and led to the development of a plethora of tools supporting specific parts of the analysis workflow or providing a complete solution. Here, we surveyed 205 tools for whole-genome/whole-exome sequencing data analysis supporting five distinct analytical steps: quality assessment, alignment, variant identification, variant annotation and visualization. We report an overview of the functionality, features and specific requirements of the individual tools. We then selected 32 programs for variant identification, variant annotation and visualization, which were subjected to hands-on evaluation using four data sets: one set of exome data from two patients with a rare disease for testing identification of germline mutations, two cancer data sets for testing variant callers for somatic mutations, copy number variations and structural variations, and one semi-synthetic data set for testing identification of copy number variations. Our comprehensive survey and evaluation of NGS tools provides a valuable guideline for human geneticists working on Mendelian disorders, complex diseases and cancers. PMID:23341494
Ntumngia, Francis B.; McHenry, Amy M.; Barnwel, John W.; Cole-Tobian, Jennifer; King, Christopher L.; Adams, John H.
2009-01-01
Plasmodium vivax Duffy binding protein (DBP) is vital for parasite development, thereby making this molecule a good vaccine candidate. Preclinical development of a P. vivax vaccine often involves use of primate models prior to testing efficacy in humans, but primate isolates are poorly characterized. We analyzed the complete gene coding for the DBP in several P. vivax isolates that are used for experimental primate infections and compared these sequences with the Salvador I DBP isolate, which is being used for vaccine development. Our results affirm that primate-adapted isolates are genetically similar to P. vivax circulating in humans, but variability is greatest in the putative target of protective antibodies. In addition, some P. vivax isolates contain multiple genetically different clones. Testing a DBP vaccine may therefore be complicated by heterogeneity and diversity of the P. vivax isolates available for in vivo challenge. PMID:19190217
Contrasting evolutionary genome dynamics between domesticated and wild yeasts
Yue, Jia-Xing; Li, Jing; Aigrain, Louise; Hallin, Johan; Persson, Karl; Oliver, Karen; Bergström, Anders; Coupland, Paul; Warringer, Jonas; Lagomarsino, Marco Consentino; Fischer, Gilles; Durbin, Richard; Liti, Gianni
2017-01-01
Structural rearrangements have long been recognized as an important source of genetic variation with implications in phenotypic diversity and disease, yet their detailed evolutionary dynamics remain elusive. Here, we use long-read sequencing to generate end-to-end genome assemblies for 12 strains representing major subpopulations of the partially domesticated yeast Saccharomyces cerevisiae and its wild relative Saccharomyces paradoxus. These population-level high-quality genomes with comprehensive annotation allow for the first time a precise definition of chromosomal boundaries between cores and subtelomeres and a high-resolution view of evolutionary genome dynamics. In chromosomal cores, S. paradoxus exhibits faster accumulation of balanced rearrangements (inversions, reciprocal translocations and transpositions) whereas S. cerevisiae accumulates unbalanced rearrangements (novel insertions, deletions and duplications) more rapidly. In subtelomeres, both species show extensive interchromosomal reshuffling, with a higher tempo in S. cerevisiae. Such striking contrasts between wild and domesticated yeasts likely reflect the influence of human activities on structural genome evolution. PMID:28416820
Androgen Receptor Gene Polymorphisms and Alterations in Prostate Cancer: Of Humanized Mice and Men
Robins, Diane M.
2011-01-01
Germline polymorphisms and somatic mutations of the androgen receptor (AR) have been intensely investigated in prostate cancer but even with genomic approaches their impact remains controversial. To assess the functional significance of AR genetic variation, we converted the mouse gene to the human sequence by germline recombination and engineered alleles to query the role of a polymorphic glutamine (Q) tract implicated in cancer risk. In a prostate cancer model, AR Q tract length influences progression and castration response. Mutation profiling in mice provides direct evidence that somatic AR variants are selected by therapy, a finding validated in human metastases from distinct treatment groups. Mutant ARs exploit multiple mechanisms to resist hormone ablation, including alterations in ligand specificity, target gene selectivity, chaperone interaction and nuclear localization. Regardless of their frequency, these variants permute normal function to reveal novel means to target wild type AR and its key interacting partners. PMID:21689727
NASA Technical Reports Server (NTRS)
Hoffman, P. F.
1986-01-01
A prograding (direction unspecified) trench-arc system is favored as a simple yet comprehensive model for crustal generation in a 250,000 sq km granite-greenstone terrain. The model accounts for the evolutionary sequence of volcanism, sedimentation, deformation, metamorphism and plutonism, observed througout the Slave province. Both unconformable (trench inner slope) and subconformable (trench outer slope) relations between the volcanics and overlying turbidities; and the existence of relatively minor amounts of pre-greenstone basement (microcontinents) and syn-greenstone plutons (accreted arc roots) are explained. Predictions include: a varaiable gap between greenstone volcanism and trench turbidite sedimentation (accompanied by minor volcanism) and systematic regional variations in age span of volcanism and plutonism. Implications of the model will be illustrated with reference to a 1:1 million scale geological map of the Slave Province (and its bounding 1.0 Ga orogens).
The genetics of monarch butterfly migration and warning colouration.
Zhan, Shuai; Zhang, Wei; Niitepõld, Kristjan; Hsu, Jeremy; Haeger, Juan Fernández; Zalucki, Myron P; Altizer, Sonia; de Roode, Jacobus C; Reppert, Steven M; Kronforst, Marcus R
2014-10-16
The monarch butterfly, Danaus plexippus, is famous for its spectacular annual migration across North America, recent worldwide dispersal, and orange warning colouration. Despite decades of study and broad public interest, we know little about the genetic basis of these hallmark traits. Here we uncover the history of the monarch's evolutionary origin and global dispersal, characterize the genes and pathways associated with migratory behaviour, and identify the discrete genetic basis of warning colouration by sequencing 101 Danaus genomes from around the globe. The results rewrite our understanding of this classic system, showing that D. plexippus was ancestrally migratory and dispersed out of North America to occupy its broad distribution. We find the strongest signatures of selection associated with migration centre on flight muscle function, resulting in greater flight efficiency among migratory monarchs, and that variation in monarch warning colouration is controlled by a single myosin gene not previously implicated in insect pigmentation.
Regional comparisons of on-site solar potential in the residential and industrial sectors
NASA Astrophysics Data System (ADS)
Gatzke, A. E.; Skewes-Cox, A. O.
1980-10-01
Regional and subregional differences in the potential development of decentralized solar technologies are studied. Two sectors of the economy were selected for intensive analysis: the residential and industrial sectors. The sequence of analysis follows the same general steps: (1) selection of appropriate prototypes within each land use sector disaggregated by census region; (2) characterization of the end-use energy demand of each prototype in order to match an appropriate decentralized solar technology to the energy demand; (3) assessment of the energy conservation potential within each prototype limited by land use patterns, technology efficiency, and variation in solar insolation; and (4) evaluation of the regional and subregional differences in the land use implications of decentralized energy supply technologies that result from the combination of energy demand, energy supply potential, and the subsequent addition of increasingly more restrictive policies to increase the percent contribution of on-site solar energy.
Karched, Maribasappa; Furgang, David; Planet, Paul J.; DeSalle, Rob
2012-01-01
Aggregatibacter actinomycetemcomitans is implicated in localized aggressive periodontitis. We report the first genome sequence of an A. actinomycetemcomitans strain isolated from an Old World primate. PMID:22328766
Liu, Mei; Li, Mijie; Wang, Shaoqiang; Xu, Yao; Lan, Xianyong; Li, Zhuanjian; Lei, Chuzhao; Yang, Dongying; Jia, Yutang; Chen, Hong
2014-02-25
Forkhead box A2 (Foxa2) has been recognized as one of the most potent transcriptional activators that is implicated in the control of feeding behavior and energy homeostasis. However, similar researches about the effects of genetic variations of Foxa2 gene on growth traits are lacking. Therefore, this study detected Foxa2 gene polymorphisms by DNA pool sequencing, PCR-RFLP and PCR-ACRS methods in 822 individuals from three Chinese cattle breeds. The results showed that four sequence variants (SVs) were screened, including two mutations (SV1, g. 7005 C>T and SV2, g. 7044 C>G) in intron 4, one mutation (SV3, g. 8449 A>G) in exon 5 and one mutation (SV4, g. 8537 T>C) in the 3'UTR. Notably, association analysis of the single mutations with growth traits in total individuals (at 24months) revealed that significant statistical difference was found in four SVs, and SV4 locus was highly significantly associated with growth traits throughout all three breeds (P<0.05 or P<0.01). Meanwhile, haplotype combination CCCCAGTC also indicated remarkably associated to better chest girth and body weight in Jiaxian Red cattle (P<0.05). We herein described a comprehensive study on the variability of bovine Foxa2 gene that was predictive of molecular markers in cattle breeding for the first time. Copyright © 2013 Elsevier B.V. All rights reserved.
Sarno, Stefania; Sevini, Federica; Vianello, Dario; Tamm, Erika; Metspalu, Ene; van Oven, Mannis; Hübner, Alexander; Sazzini, Marco; Franceschi, Claudio; Pettener, Davide; Luiselli, Donata
2015-01-01
Genetic signatures from the Paleolithic inhabitants of Eurasia can be traced from the early divergent mitochondrial DNA lineages still present in contemporary human populations. Previous studies already suggested a pre-Neolithic diffusion of mitochondrial haplogroup HV*(xH,V) lineages, a relatively rare class of mtDNA types that includes parallel branches mainly distributed across Europe and West Asia with a certain degree of structure. Up till now, variation within haplogroup HV was addressed mainly by analyzing sequence data from the mtDNA control region, except for specific sub-branches, such as HV4 or the widely distributed haplogroups H and V. In this study, we present a revised HV topology based on full mtDNA genome data, and we include a comprehensive dataset consisting of 316 complete mtDNA sequences including 60 new samples from the Italian peninsula, a previously underrepresented geographic area. We highlight points of instability in the particular topology of this haplogroup, reconstructed with BEAST-generated trees and networks. We also confirm a major lineage expansion that probably followed the Late Glacial Maximum and preceded Neolithic population movements. We finally observe that Italy harbors a reservoir of mtDNA diversity, with deep-rooting HV lineages often related to sequences present in the Caucasus and the Middle East. The resulting hypothesis of a glacial refugium in Southern Italy has implications for the understanding of late Paleolithic population movements and is discussed within the archaeological cultural shifts occurred over the entire continent. PMID:26640946
Das, Rahul K; Crick, Scott L; Pappu, Rohit V
2012-02-17
Basic region leucine zippers (bZIPs) are modular transcription factors that play key roles in eukaryotic gene regulation. The basic regions of bZIPs (bZIP-bRs) are necessary and sufficient for DNA binding and specificity. Bioinformatic predictions and spectroscopic studies suggest that unbound monomeric bZIP-bRs are uniformly disordered as isolated domains. Here, we test this assumption through a comparative characterization of conformational ensembles for 15 different bZIP-bRs using a combination of atomistic simulations and circular dichroism measurements. We find that bZIP-bRs have quantifiable preferences for α-helical conformations in their unbound monomeric forms. This helicity varies from one bZIP-bR to another despite a significant sequence similarity of the DNA binding motifs (DBMs). Our analysis reveals that intramolecular interactions between DBMs and eight-residue segments directly N-terminal to DBMs are the primary modulators of bZIP-bR helicities. We test the accuracy of this inference by designing chimeras of bZIP-bRs to have either increased or decreased overall helicities. Our results yield quantitative insights regarding the relationship between sequence and the degree of intrinsic disorder within bZIP-bRs, and might have general implications for other intrinsically disordered proteins. Understanding how natural sequence variations lead to modulation of disorder is likely to be important for understanding the evolution of specificity in molecular recognition through intrinsically disordered regions (IDRs). Copyright © 2011 Elsevier Ltd. All rights reserved.
Barr, Norman B; Ledezma, Lisa A; Leblanc, Luc; San Jose, Michael; Rubinoff, Daniel; Geib, Scott M; Fujita, Brian; Bartels, David W; Garza, Daniel; Kerr, Peter; Hauser, Martin; Gaimari, Stephen
2014-10-01
Population genetic diversity of the oriental fruit fly, Bactrocera dorsalis (Hendel), on the Hawaiian islands of Oahu, Maui, Kauai, and Hawaii (the Big Island) was estimated using DNA sequences of the mitochondrial cytochrome c oxidase subunit I gene. In total, 932 flies representing 36 sampled sites across the four islands were sequenced for a 1,500-bp fragment of the gene named the C1500 marker. Genetic variation was low on the Hawaiian Islands with >96% of flies having just two haplotypes: C1500-Haplotype 1 (63.2%) or C1500-Haplotype 2 (33.3%). The other 33 flies (3.5%) had haplotypes similar to the two dominant haplotypes. No population structure was detected among the islands or within islands. The two haplotypes were present at similar frequencies at each sample site, suggesting that flies on the various islands can be considered one population. Comparison of the Hawaiian data set to DNA sequences of 165 flies from outbreaks in California between 2006 and 2012 indicates that a single-source introduction pathway of Hawaiian origin cannot explain many of the flies in California. Hawaii, however, could not be excluded as a maternal source for 69 flies. There was no clear geographic association for Hawaiian or non-Hawaiian haplotypes in the Bay Area or Los Angeles Basin over time. This suggests that California experienced multiple, independent introductions from different sources. © 2014 Entomological Society of America.
Bavykin, Sergei G.; Mirzabekova, legal representative, Natalia V.; Mirzabekov, deceased, Andrei D.
2007-12-04
The present invention relates to methods and compositions for using nucleotide sequence variations of 16S and 23S rRNA within the B. cereus group to discriminate a highly infectious bacterium B. anthracis from closely related microorganisms. Sequence variations in the 16S and 23S rRNA of the B. cereus subgroup including B. anthracis are utilized to construct an array that can detect these sequence variations through selective hybridizations and discriminate B. cereus group that includes B. anthracis. Discrimination of single base differences in rRNA was achieved with a microchip during analysis of B. cereus group isolates from both single and in mixed samples, as well as identification of polymorphic sites. Successful use of a microchip to determine the appropriate subgroup classification using eight reference microorganisms from the B. cereus group as a study set, was demonstrated.