Science.gov

Sample records for acid sequences conserved

  1. Dynamic behavior of an intrinsically unstructured linker domain is conserved in the face of negligible amino acid sequence conservation.

    PubMed

    Daughdrill, Gary W; Narayanaswami, Pranesh; Gilmore, Sara H; Belczyk, Agniezka; Brown, Celeste J

    2007-09-01

    Proteins or regions of proteins that do not form compact globular structures are classified as intrinsically unstructured proteins (IUPs). IUPs are common in nature and have essential molecular functions, but even a limited understanding of the evolution of their dynamic behavior is lacking. The primary objective of this work was to test the evolutionary conservation of dynamic behavior for a particular class of IUPs that form intrinsically unstructured linker domains (IULD) that tether flanking folded domains. This objective was accomplished by measuring the backbone flexibility of several IULD homologues using nuclear magnetic resonance (NMR) spectroscopy. The backbone flexibility of five IULDs, representing three kingdoms, was measured and analyzed. Two IULDs from animals, one IULD from fungi, and two IULDs from plants showed similar levels of backbone flexibility that were consistent with the absence of a compact globular structure. In contrast, the amino acid sequences of the IULDs from these three taxa showed no significant similarity. To investigate how the dynamic behavior of the IULDs could be conserved in the absence of detectable sequence conservation, evolutionary rate studies were performed on a set of nine mammalian IULDs. The results of this analysis showed that many sites in the IULD are evolving neutrally, suggesting that dynamic behavior can be maintained in the absence of natural selection. This work represents the first experimental test of the evolutionary conservation of dynamic behavior and demonstrates that amino acid sequence conservation is not required for the conservation of dynamic behavior and presumably molecular function. PMID:17721672

  2. The Chinese hamster Alu-equivalent sequence: a conserved highly repetitious, interspersed deoxyribonucleic acid sequence in mammals has a structure suggestive of a transposable element.

    PubMed Central

    Haynes, S R; Toomey, T P; Leinwand, L; Jelinek, W R

    1981-01-01

    A consensus sequence has been determined for a major interspersed deoxyribonucleic acid repeat in the genome of Chinese hamster ovary cells (CHO cells). This sequence is extensively homologous to (i) the human Alu sequence (P. L. Deininger et al., J. Mol. Biol., in press), (ii) the mouse B1 interspersed repetitious sequence (Krayev et al., Nucleic Acids Res. 8:1201-1215, 1980) (iii) an interspersed repetitious sequence from African green monkey deoxyribonucleic acid (Dhruva et al., Proc. Natl. Acad. Sci. U.S.A. 77:4514-4518, 1980) and (iv) the CHO and mouse 4.5S ribonucleic acid (this report; F. Harada and N. Kato, Nucleic Acids Res. 8:1273-1285, 1980). Because the CHO consensus sequence shows significant homology to the human Alu sequence it is termed the CHO Alu-equivalent sequence. A conserved structure surrounding CHO Alu-equivalent family members can be recognized. It is similar to that surrounding the human Alu and the mouse B1 sequences, and is represented as follows: direct repeat-CHO-Alu-A-rich sequence-direct repeat. A composite interspersed repetitious sequence has been identified. Its structure is represented as follows: direct repeat-residue 47 to 107 of CHO-Alu-non-Alu repetitious sequence-A-rich sequence-direct repeat. Because the Alu flanking sequences resemble those that flank known transposable elements, we think it likely that the Alu sequence dispersed throughout the mammalian genome by transposition. Images PMID:9279371

  3. Conservation of Shannon's redundancy for proteins. [information theory applied to amino acid sequences

    NASA Technical Reports Server (NTRS)

    Gatlin, L. L.

    1974-01-01

    Concepts of information theory are applied to examine various proteins in terms of their redundancy in natural originators such as animals and plants. The Monte Carlo method is used to derive information parameters for random protein sequences. Real protein sequence parameters are compared with the standard parameters of protein sequences having a specific length. The tendency of a chain to contain some amino acids more frequently than others and the tendency of a chain to contain certain amino acid pairs more frequently than other pairs are used as randomness measures of individual protein sequences. Non-periodic proteins are generally found to have random Shannon redundancies except in cases of constraints due to short chain length and genetic codes. Redundant characteristics of highly periodic proteins are discussed. A degree of periodicity parameter is derived.

  4. Evolution of a "conserved" amino acid sequence: a model study of an in silico investigation of the phylogenesis of some immune receptors.

    PubMed

    Panaro, M A; Acquafredda, A; Sisto, M; Lisi, S; Saccia, M; Mitolo, V

    2006-01-01

    In this paper we analyze a 55-amino acid (aa) sequence which is relatively well conserved in several seven-transmembrane receptor families (from Insects to Mammals) and in some Viruses. This sequence, which covers the second transmembrane domain, the first extracellular loop and the third transmembrane domain, appears in its complete configuration in most of the seven-transmembrane receptor families, as well as in the protein products of some viruses. Other seven-transmembrane receptors and viruses exhibit reduced configurations of the conserved sequence, lacking either aa 31 or aa 30-31. 53-aa configurations are typically found in most chemokine receptor (CKR) subfamilies, as well as in some viral protein products. However, the CCR1, CCR3, and CCR6 subfamilies comprise a 54-aa configuration and the CKR-related protein products, ChemR23 and RDC1, include the complete 55-aa sequence. For each CKR subfamily the "modal sequence" of the conserved segment was constructed by selecting the most frequently occurring aa at each position. Then, pairwise alignments were made between: (i) the modal CKR sequences, and (ii) the sequence (53-aa) of the Yaba-like disease virus - 7L protein. From the alignments two consensus matrices were derived: (i) the consensus 1 matrix with reference to the whole conserved segment, and (ii) the consensus 2 matrix with reference to aa 22-29, which appear to be the most variable segment of the sequence. Based on the obtained consensus values and with reference to this specific conserved segment, the following conclusions are proposed: (1) ChemR23 and RDC1 are probably the more primitive CKR forms; (2) CCR1 and CCR3 may be grouped in a single cluster; (3) CCRs 2, 4, and 5 are closely related to each other and may be grouped in a cluster; CCR7 is likely to be evolutionarily related to this cluster; (4) CXCRs 2, 3, and 4 and CCX CKR appear to be evolutionarily related to each other and very likely derived from an CCR6-like gene; (5) CCR2/4/5 and

  5. Data in support of the discovery of alternative splicing variants of quail LEPR and the evolutionary conservation of qLEPRl by nucleotide and amino acid sequences alignment.

    PubMed

    Wang, Dandan; Xu, Chunlin; Wang, Taian; Li, Hong; Li, Yanmin; Ren, Junxiao; Tian, Yadong; Li, Zhuanjian; Jiao, Yuping; Kang, Xiangtao; Liu, Xiaojun

    2016-03-01

    Leptin receptor (LEPR) belongs to the class I cytokine receptor superfamily which share common structural features and signal transduction pathways. Although multiple LEPR isoforms, which are derived from one gene, were identified in mammals, they were rarely found in avian except the long LEPR. Four alternative splicing variants of quail LEPR (qLEPR) had been cloned and sequenced for the first time (Wang et al., 2015 [1]). To define patterns of the four splicing variants (qLEPRl, qLEPR-a, qLEPR-b and qLEPR-c) and locate the conserved regions of qLEPRl, this data article provides nucleotide sequence alignment of qLEPR and amino acid sequence alignment of representative vertebrate LEPR. The detailed analysis was shown in [1]. PMID:26759819

  6. Data in support of the discovery of alternative splicing variants of quail LEPR and the evolutionary conservation of qLEPRl by nucleotide and amino acid sequences alignment

    PubMed Central

    Wang, Dandan; Xu, Chunlin; Wang, Taian; Li, Hong; Li, Yanmin; Ren, Junxiao; Tian, Yadong; Li, Zhuanjian; Jiao, Yuping; Kang, Xiangtao; Liu, Xiaojun

    2015-01-01

    Leptin receptor (LEPR) belongs to the class I cytokine receptor superfamily which share common structural features and signal transduction pathways. Although multiple LEPR isoforms, which are derived from one gene, were identified in mammals, they were rarely found in avian except the long LEPR. Four alternative splicing variants of quail LEPR (qLEPR) had been cloned and sequenced for the first time (Wang et al., 2015 [1]). To define patterns of the four splicing variants (qLEPRl, qLEPR-a, qLEPR-b and qLEPR-c) and locate the conserved regions of qLEPRl, this data article provides nucleotide sequence alignment of qLEPR and amino acid sequence alignment of representative vertebrate LEPR. The detailed analysis was shown in [1]. PMID:26759819

  7. Creation of a data base for sequences of ribosomal nucleic acids and detection of conserved restriction endonucleases sites through computerized processing.

    PubMed Central

    Patarca, R; Dorta, B; Ramirez, J L

    1982-01-01

    As part of a project pertaining the organization of ribosomal genes in Kinetoplastidae, we have created a data base for published sequences of ribosomal nucleic acids, with information in Spanish. As a first step in their processing, we have written a computer program which introduces the new feature of determining the length of the fragments produced after single or multiple digestion with any of the known restriction enzymes. With this information we have detected conserved SAU 3A sites: (i) at the 5' end of the 5.8S rRNA and at the 3' end of the small subunit rRNA, both included in similar larger sequences; (ii) in the 5.8S rRNA of vertebrates (a second one), which is not present in lower eukaryotes, showing a clear evolutive divergence; and, (iii) at the 5' terminal of the small subunit rRNA, included in a larger conserved sequence. The possible biological importance of these sequences is discussed. PMID:6278402

  8. Amino acid binding by the class I aminoacyl-tRNA synthetases: role for a conserved proline in the signature sequence.

    PubMed Central

    Burbaum, J. J.; Schimmel, P.

    1992-01-01

    Although partial or complete three-dimensional structures are known for three Class I aminoacyl-tRNA synthetases, the amino acid-binding sites in these proteins remain poorly characterized. To explore the methionine binding site of Escherichia coli methionyl-tRNA synthetase, we chose to study a specific, randomly generated methionine auxotroph that contains a mutant methionyl-tRNA synthetase whose defect is manifested in an elevated Km for methionine (Barker, D.G., Ebel, J.-P., Jakes, R.C., & Bruton, C.J., 1982, Eur. J. Biochem. 127, 449-457), and employed the polymerase chain reaction to sequence this mutant synthetase directly. We identified a Pro 14 to Ser replacement (P14S), which accounts for a greater than 300-fold elevation in Km for methionine and has little effect on either the Km for ATP or the kcat of the amino acid activation reaction. This mutation destabilizes the protein in vivo, which may partly account for the observed auxotrophy. The altered proline is found in the "signature sequence" of the Class I synthetases and is conserved. This sequence motif is 1 of 2 found in the 10 Class I aminoacyl-tRNA synthetases and, in the known structures, it is in the nucleotide-binding fold as part of a loop between the end of a beta-strand and the start of an alpha-helix. The phenotype of the mutant and the stability and affinity for methionine of the wild-type and mutant enzymes are influenced by the amino acid that is 25 residues beyond the C-terminus of the signature sequence.(ABSTRACT TRUNCATED AT 250 WORDS) PMID:1304356

  9. Acid rain and electricity conservation

    SciTech Connect

    Geller, H.; Miller, E.; Ledbetter, M.; Miller, P.

    1987-01-01

    Conservation directly lowers the emissions of SO/sub 2/ and other pollutants by reducing the amount of coal and other fuels that must be burned to meet electricity demand. This book is the first report to provide an integrated analysis of electricity supply, acid raid abatement, and conservation opportunities. The authors use a utility simulation model to examine SO/sub 2/ emissions, electric rates, and overall costs to consumers for different load growth and emissions control scenarios. The study also suggests how acid rain legislation can be designed to encourage electricity conservation.

  10. Amplification of human papillomavirus DNA sequences by using conserved primers.

    PubMed Central

    Gregoire, L; Arella, M; Campione-Piccardo, J; Lancaster, W D

    1989-01-01

    The polymerase chain reaction has potential for use in the detection of small amounts of human papillomavirus (HPV) viral nucleic acids present in clinical specimens. However, new HPV types for which no probes exist would remain undetected by using type-specific primers for the polymerase chain reaction before hybridization. Primers corresponding to highly conserved HPV sequences may be useful for detecting low amounts of known HPV DNA as well as new HPV types. Here we analyze a pair of primers derived from conserved sequences within the E1 open reading frame for HPV sequence amplification by using the polymerase chain reaction. The longest perfect homology among HPV sequences is a 12-mer within the first exon of E1M. A region of conserved amino acids coded by the E1 open reading frame allowed the detection of another highly conserved region about 850 base pairs downstream. Two 21-mers derived from these conserved regions were used to amplify sequences from all HPV DNAs used as templates. The amplified DNA was shown to be specific for HPV sequences within the E1 open reading frame. DNA from HPVs whose sequences were not available were amplified by using these two primers. HPV DNA sequences in clinical specimens could also be amplified with the primers. Images PMID:2556429

  11. Evolutionarily conserved sequences on human chromosome 21

    SciTech Connect

    Frazer, Kelly A.; Sheehan, John B.; Stokowski, Renee P.; Chen, Xiyin; Hosseini, Roya; Cheng, Jan-Fang; Fodor, Stephen P.A.; Cox, David R.; Patil, Nila

    2001-09-01

    Comparison of human sequences with the DNA of other mammals is an excellent means of identifying functional elements in the human genome. Here we describe the utility of high-density oligonucleotide arrays as a rapid approach for comparing human sequences with the DNA of multiple species whose sequences are not presently available. High-density arrays representing approximately 22.5 Mb of nonrepetitive human chromosome 21 sequence were synthesized and then hybridized with mouse and dog DNA to identify sequences conserved between humans and mice (human-mouse elements) and between humans and dogs (human-dog elements). Our data show that sequence comparison of multiple species provides a powerful empiric method for identifying actively conserved elements in the human genome. A large fraction of these evolutionarily conserved elements are present in regions on chromosome 21 that do not encode known genes.

  12. Conserved sequence pattern in a wide variety of phosphoesterases.

    PubMed Central

    Koonin, E. V.

    1994-01-01

    A unique sequence pattern, designated the GD/GNH signature, was shown to be conserved in a wide variety of phosphoesterases. The enzymes containing this signature cleave phosphoester bonds in such different substrates as (1) phosphoserine and phosphothreonine in polypeptides; (2) bis(5'-nucleosidyl)-tetraphosphates; (3) nucleoside 5' phosphates; (4) 2',3'-cyclic nucleotide phosphates; (5) polynucleotides; (6) 2'-5' phosphodiesters in RNA (intron) lariats; (7) sphingomyelin; and (7) various phosphomonoesters. Two conserved acidic amino acid residues and a conserved histidine residue may be directly involved in phosphoester bond cleavage. PMID:8003970

  13. Sequence conservation on the Y chromosome

    SciTech Connect

    Gibson, L.H.; Yang-Feng, L.; Lau, C.

    1994-09-01

    The Y chromosome is present in all mammals and is considered to be essential to sex determination. Despite intense genomic research, only a few genes have been identified and mapped to this chromosome in humans. Several of them, such as SRY and ZFY, have been demonstrated to be conserved and Y-located in other mammals. In order to address the issue of sequence conservation on the Y chromosome, we performed fluorescence in situ hybridization (FISH) with DNA from a human Y cosmid library as a probe to study the Y chromosomes from other mammalian species. Total DNA from 3,000-4,500 cosmid pools were labeled with biotinylated-dUTP and hybridized to metaphase chromosomes. For human and primate preparations, human cot1 DNA was included in the hybridization mixture to suppress the hybridization from repeat sequences. FISH signals were detected on the Y chromosomes of human, gorilla, orangutan and baboon (Old World monkey) and were absent on those of squirrel monkey (New World monkey), Indian munjac, wood lemming, Chinese hamster, rat and mouse. Since sequence analysis suggested that specific genes, e.g. SRY and ZFY, are conserved between these two groups, the lack of detectable hybridization in the latter group implies either that conservation of the human Y sequences is limited to the Y chromosomes of the great apes and Old World monkeys, or that the size of the syntenic segment is too small to be detected under the resolution of FISH, or that homologeous sequences have undergone considerable divergence. Further studies with reduced hybridization stringency are currently being conducted. Our results provide some clues as to Y-sequence conservation across species and demonstrate the limitations of FISH across species with total DNA sequences from a particular chromosome.

  14. Conserved noncoding sequences (CNSs) in higher plants.

    PubMed

    Freeling, Michael; Subramaniam, Shabarinath

    2009-04-01

    Plant conserved noncoding sequences (CNSs)--a specific category of phylogenetic footprint--have been shown experimentally to function. No plant CNS is conserved to the extent that ultraconserved noncoding sequences are conserved in vertebrates. Plant CNSs are enriched in known transcription factor or other cis-acting binding sites, and are usually clustered around genes. Genes that encode transcription factors and/or those that respond to stimuli are particularly CNS-rich. Only rarely could this function involve small RNA binding. Some transcribed CNSs encode short translation products as a form of negative control. Approximately 4% of Arabidopsis gene content is estimated to be both CNS-rich and occupies a relatively long stretch of chromosome: Bigfoot genes (long phylogenetic footprints). We discuss a 'DNA-templated protein assembly' idea that might help explain Bigfoot gene CNSs. PMID:19249238

  15. Conserved Sequence Preferences Contribute to Substrate Recognition by the Proteasome*

    PubMed Central

    Yu, Houqing; Singh Gautam, Amit K.; Wilmington, Shameika R.; Wylie, Dennis; Martinez-Fonts, Kirby; Kago, Grace; Warburton, Marie; Chavali, Sreenivas; Inobe, Tomonao; Finkelstein, Ilya J.; Babu, M. Madan

    2016-01-01

    The proteasome has pronounced preferences for the amino acid sequence of its substrates at the site where it initiates degradation. Here, we report that modulating these sequences can tune the steady-state abundance of proteins over 2 orders of magnitude in cells. This is the same dynamic range as seen for inducing ubiquitination through a classic N-end rule degron. The stability and abundance of His3 constructs dictated by the initiation site affect survival of yeast cells and show that variation in proteasomal initiation can affect fitness. The proteasome's sequence preferences are linked directly to the affinity of the initiation sites to their receptor on the proteasome and are conserved between Saccharomyces cerevisiae, Schizosaccharomyces pombe, and human cells. These findings establish that the sequence composition of unstructured initiation sites influences protein abundance in vivo in an evolutionarily conserved manner and can affect phenotype and fitness. PMID:27226608

  16. Composition for nucleic acid sequencing

    DOEpatents

    Korlach, Jonas; Webb, Watt W.; Levene, Michael; Turner, Stephen; Craighead, Harold G.; Foquet, Mathieu

    2008-08-26

    The present invention is directed to a method of sequencing a target nucleic acid molecule having a plurality of bases. In its principle, the temporal order of base additions during the polymerization reaction is measured on a molecule of nucleic acid, i.e. the activity of a nucleic acid polymerizing enzyme on the template nucleic acid molecule to be sequenced is followed in real time. The sequence is deduced by identifying which base is being incorporated into the growing complementary strand of the target nucleic acid by the catalytic activity of the nucleic acid polymerizing enzyme at each step in the sequence of base additions. A polymerase on the target nucleic acid molecule complex is provided in a position suitable to move along the target nucleic acid molecule and extend the oligonucleotide primer at an active site. A plurality of labelled types of nucleotide analogs are provided proximate to the active site, with each distinguishable type of nucleotide analog being complementary to a different nucleotide in the target nucleic acid sequence. The growing nucleic acid strand is extended by using the polymerase to add a nucleotide analog to the nucleic acid strand at the active site, where the nucleotide analog being added is complementary to the nucleotide of the target nucleic acid at the active site. The nucleotide analog added to the oligonucleotide primer as a result of the polymerizing step is identified. The steps of providing labelled nucleotide analogs, polymerizing the growing nucleic acid strand, and identifying the added nucleotide analog are repeated so that the nucleic acid strand is further extended and the sequence of the target nucleic acid is determined.

  17. The highly conserved amino acid sequence motif Tyr-Gly-Asp-Thr-Asp-Ser in alpha-like DNA polymerases is required by phage phi 29 DNA polymerase for protein-primed initiation and polymerization.

    PubMed Central

    Bernad, A; Lázaro, J M; Salas, M; Blanco, L

    1990-01-01

    The alpha-like DNA polymerases from bacteriophage phi 29 and other viruses, prokaryotes and eukaryotes contain an amino acid consensus sequence that has been proposed to form part of the dNTP binding site. We have used site-directed mutants to study five of the six highly conserved consecutive amino acids corresponding to the most conserved C-terminal segment (Tyr-Gly-Asp-Thr-Asp-Ser). Our results indicate that in phi 29 DNA polymerase this consensus sequence, although irrelevant for the 3'----5' exonuclease activity, is essential for initiation and elongation. Based on these results and on its homology with known or putative metal-binding amino acid sequences, we propose that in phi 29 DNA polymerase the Tyr-Gly-Asp-Thr-Asp-Ser consensus motif is part of the dNTP binding site, involved in the synthetic activities of the polymerase (i.e., initiation and polymerization), and that it is involved particularly in the metal binding associated with the dNTP site. Images PMID:2191296

  18. High speed nucleic acid sequencing

    DOEpatents

    Korlach, Jonas; Webb, Watt W.; Levene, Michael; Turner, Stephen; Craighead, Harold G.; Foquet, Mathieu

    2011-05-17

    The present invention is directed to a method of sequencing a target nucleic acid molecule having a plurality of bases. In its principle, the temporal order of base additions during the polymerization reaction is measured on a molecule of nucleic acid. Each type of labeled nucleotide comprises an acceptor fluorophore attached to a phosphate portion of the nucleotide such that the fluorophore is removed upon incorporation into a growing strand. Fluorescent signal is emitted via fluorescent resonance energy transfer between the donor fluorophore and the acceptor fluorophore as each nucleotide is incorporated into the growing strand. The sequence is deduced by identifying which base is being incorporated into the growing strand.

  19. Functionally conserved enhancers with divergent sequences in distant vertebrates

    SciTech Connect

    Yang, Song; Oksenberg, Nir; Takayama, Sachiko; Heo, Seok -Jin; Poliakov, Alexander; Ahituv, Nadav; Dubchak, Inna; Boffelli, Dario

    2015-10-30

    To examine the contributions of sequence and function conservation in the evolution of enhancers, we systematically identified enhancers whose sequences are not conserved among distant groups of vertebrate species, but have homologous function and are likely to be derived from a common ancestral sequence. In conclusion, our approach combined comparative genomics and epigenomics to identify potential enhancer sequences in the genomes of three groups of distantly related vertebrate species.

  20. The tryptophan repressor sequence is highly conserved among the Enterobacteriaceae.

    PubMed Central

    Arvidson, D N; Arvidson, C G; Lawson, C L; Miner, J; Adams, C; Youderian, P

    1994-01-01

    Tryptophan biosynthesis in Escherichia coli is regulated by the product of the trpR gene, the tryptophan (Trp) repressor. Trp aporepressor binds the corepressor, L-tryptophan, to form a holorepressor complex, which binds trp operator DNA tightly, and inhibits transcription of the tryptophan biosynthetic operon. The conservation of trp operator sequences among enteric Gram-negative bacteria suggests that trpR genes from other bacterial species can be cloned by complementation in E. coli. To clone trpR homologues, a deletion of the E. coli trpR gene, delta trpR504, was made on a plasmid by site-directed mutagenesis, then crossed onto the E. coli genome. Plasmid clones of the trpR genes of Enterobacter aerogenes and Enterobacter cloacae were isolated by complementation of the delta trpR504 allele, scored as the ability to repress beta-galactosidase synthesis from a prophage-borne trpE-lacZ gene fusion. The predicted amino acid sequences of four enteric TrpR proteins show differences, clustered on the backside of the folded repressor, opposite the DNA-binding helix-turn-helix substructures. These differences are predicted to have little effect on the interactions of the aporepressor with tryptophan, holorepressor with operator DNA, or tandemly bound holorepressor dimers with one another. Although there is some variation observed at the dimer interface, interactions predicted to stabilize the interface are conserved. The phylogenetic relationships revealed by the TrpR amino acid sequence alignment agree with the results of others. PMID:8208606

  1. Bioinformatic Identification of Conserved Cis-Sequences in Coregulated Genes.

    PubMed

    Bülow, Lorenz; Hehl, Reinhard

    2016-01-01

    Bioinformatics tools can be employed to identify conserved cis-sequences in sets of coregulated plant genes because more and more gene expression and genomic sequence data become available. Knowledge on the specific cis-sequences, their enrichment and arrangement within promoters, facilitates the design of functional synthetic plant promoters that are responsive to specific stresses. The present chapter illustrates an example for the bioinformatic identification of conserved Arabidopsis thaliana cis-sequences enriched in drought stress-responsive genes. This workflow can be applied for the identification of cis-sequences in any sets of coregulated genes. The workflow includes detailed protocols to determine sets of coregulated genes, to extract the corresponding promoter sequences, and how to install and run a software package to identify overrepresented motifs. Further bioinformatic analyses that can be performed with the results are discussed. PMID:27557771

  2. Chip-based sequencing nucleic acids

    DOEpatents

    Beer, Neil Reginald

    2014-08-26

    A system for fast DNA sequencing by amplification of genetic material within microreactors, denaturing, demulsifying, and then sequencing the material, while retaining it in a PCR/sequencing zone by a magnetic field. One embodiment includes sequencing nucleic acids on a microchip that includes a microchannel flow channel in the microchip. The nucleic acids are isolated and hybridized to magnetic nanoparticles or to magnetic polystyrene-coated beads. Microreactor droplets are formed in the microchannel flow channel. The microreactor droplets containing the nucleic acids and the magnetic nanoparticles are retained in a magnetic trap in the microchannel flow channel and sequenced.

  3. Relative von Neumann entropy for evaluating amino acid conservation.

    PubMed

    Johansson, Fredrik; Toh, Hiroyuki

    2010-10-01

    The Shannon entropy is a common way of measuring conservation of sites in multiple sequence alignments, and has also been extended with the relative Shannon entropy to account for background frequencies. The von Neumann entropy is another extension of the Shannon entropy, adapted from quantum mechanics in order to account for amino acid similarities. However, there is yet no relative von Neumann entropy defined for sequence analysis. We introduce a new definition of the von Neumann entropy for use in sequence analysis, which we found to perform better than the previous definition. We also introduce the relative von Neumann entropy and a way of parametrizing this in order to obtain the Shannon entropy, the relative Shannon entropy and the von Neumann entropy at special parameter values. We performed an exhaustive search of this parameter space and found better predictions of catalytic sites compared to any of the previously used entropies. PMID:20981889

  4. Distinguishing Proteins From Arbitrary Amino Acid Sequences

    PubMed Central

    Yau, Stephen S.-T.; Mao, Wei-Guang; Benson, Max; He, Rong Lucy

    2015-01-01

    What kinds of amino acid sequences could possibly be protein sequences? From all existing databases that we can find, known proteins are only a small fraction of all possible combinations of amino acids. Beginning with Sanger's first detailed determination of a protein sequence in 1952, previous studies have focused on describing the structure of existing protein sequences in order to construct the protein universe. No one, however, has developed a criteria for determining whether an arbitrary amino acid sequence can be a protein. Here we show that when the collection of arbitrary amino acid sequences is viewed in an appropriate geometric context, the protein sequences cluster together. This leads to a new computational test, described here, that has proved to be remarkably accurate at determining whether an arbitrary amino acid sequence can be a protein. Even more, if the results of this test indicate that the sequence can be a protein, and it is indeed a protein sequence, then its identity as a protein sequence is uniquely defined. We anticipate our computational test will be useful for those who are attempting to complete the job of discovering all proteins, or constructing the protein universe. PMID:25609314

  5. The BsaHI restriction-modification system: Cloning, sequencing and analysis of conserved motifs

    PubMed Central

    Neely, Robert K; Roberts, Richard J

    2008-01-01

    Background Restriction and modification enzymes typically recognise short DNA sequences of between two and eight bases in length. Understanding the mechanism of this recognition represents a significant challenge that we begin to address for the BsaHI restriction-modification system, which recognises the six base sequence GRCGYC. Results The DNA sequences of the genes for the BsaHI methyltransferase, bsaHIM, and restriction endonuclease, bsaHIR, have been determined (GenBank accession #EU386360), cloned and expressed in E. coli. Both the restriction endonuclease and methyltransferase enzymes share significant similarity with a group of 6 other enzymes comprising the restriction-modification systems HgiDI and HgiGI and the putative HindVP, NlaCORFDP, NpuORFC228P and SplZORFNP restriction-modification systems. A sequence alignment of these homologues shows that their amino acid sequences are largely conserved and highlights several motifs of interest. We target one such conserved motif, reading SPERRFD, at the C-terminal end of the bsaHIR gene. A mutational analysis of these amino acids indicates that the motif is crucial for enzymatic activity. Sequence alignment of the methyltransferase gene reveals a short motif within the target recognition domain that is conserved among enzymes recognising the same sequences. Thus, this motif may be used as a diagnostic tool to define the recognition sequences of the cytosine C5 methyltransferases. Conclusion We have cloned and sequenced the BsaHI restriction and modification enzymes. We have identified a region of the R. BsaHI enzyme that is crucial for its activity. Analysis of the amino acid sequence of the BsaHI methyltransferase enzyme led us to propose two new motifs that can be used in the diagnosis of the recognition sequence of the cytosine C5-methyltransferases. PMID:18479503

  6. Sequence conservation of an avian centromeric repeated DNA component.

    PubMed

    Madsen, C S; Brooks, J E; de Kloet, E; de Kloet, S R

    1994-06-01

    The approximately 190-bp centromeric repeat monomers of the spur-winged lapwing (Vanellus spinosus, Charadriidae), the Chilean flamingo (Phoenicopterus chilensis, Phoenicopteridae), the sarus crane (Grus antigone, Gruidae), parrots (Psittacidae), waterfowl (Anatidae), and the merlin (Falco columbarius, Falconidae) contain elements that are interspecifically highly variable, as well as elements (trinucleotides and higher order oligonucleotides) that are highly conserved in sequence and relative location within the repeat. Such conservation suggests that the centromeric repeats of these avian species have evolved from a common ancestral sequence that may date from very early stages of avian radiation. PMID:8034177

  7. Method for sequencing nucleic acid molecules

    DOEpatents

    Korlach, Jonas; Webb, Watt W.; Levene, Michael; Turner, Stephen; Craighead, Harold G.; Foquet, Mathieu

    2006-05-30

    The present invention is directed to a method of sequencing a target nucleic acid molecule having a plurality of bases. In its principle, the temporal order of base additions during the polymerization reaction is measured on a molecule of nucleic acid, i.e. the activity of a nucleic acid polymerizing enzyme on the template nucleic acid molecule to be sequenced is followed in real time. The sequence is deduced by identifying which base is being incorporated into the growing complementary strand of the target nucleic acid by the catalytic activity of the nucleic acid polymerizing enzyme at each step in the sequence of base additions. A polymerase on the target nucleic acid molecule complex is provided in a position suitable to move along the target nucleic acid molecule and extend the oligonucleotide primer at an active site. A plurality of labelled types of nucleotide analogs are provided proximate to the active site, with each distinguishable type of nucleotide analog being complementary to a different nucleotide in the target nucleic acid sequence. The growing nucleic acid strand is extended by using the polymerase to add a nucleotide analog to the nucleic acid strand at the active site, where the nucleotide analog being added is complementary to the nucleotide of the target nucleic acid at the active site. The nucleotide analog added to the oligonucleotide primer as a result of the polymerizing step is identified. The steps of providing labelled nucleotide analogs, polymerizing the growing nucleic acid strand, and identifying the added nucleotide analog are repeated so that the nucleic acid strand is further extended and the sequence of the target nucleic acid is determined.

  8. Method for sequencing nucleic acid molecules

    DOEpatents

    Korlach, Jonas; Webb, Watt W.; Levene, Michael; Turner, Stephen; Craighead, Harold G.; Foquet, Mathieu

    2006-06-06

    The present invention is directed to a method of sequencing a target nucleic acid molecule having a plurality of bases. In its principle, the temporal order of base additions during the polymerization reaction is measured on a molecule of nucleic acid, i.e. the activity of a nucleic acid polymerizing enzyme on the template nucleic acid molecule to be sequenced is followed in real time. The sequence is deduced by identifying which base is being incorporated into the growing complementary strand of the target nucleic acid by the catalytic activity of the nucleic acid polymerizing enzyme at each step in the sequence of base additions. A polymerase on the target nucleic acid molecule complex is provided in a position suitable to move along the target nucleic acid molecule and extend the oligonucleotide primer at an active site. A plurality of labelled types of nucleotide analogs are provided proximate to the active site, with each distinguishable type of nucleotide analog being complementary to a different nucleotide in the target nucleic acid sequence. The growing nucleic acid strand is extended by using the polymerase to add a nucleotide analog to the nucleic acid strand at the active site, where the nucleotide analog being added is complementary to the nucleotide of the target nucleic acid at the active site. The nucleotide analog added to the oligonucleotide primer as a result of the polymerizing step is identified. The steps of providing labelled nucleotide analogs, polymerizing the growing nucleic acid strand, and identifying the added nucleotide analog are repeated so that the nucleic acid strand is further extended and the sequence of the target nucleic acid is determined.

  9. Proteome-Wide Discovery of Evolutionary Conserved Sequences in Disordered Regions

    PubMed Central

    Nguyen Ba, Alex N.; Yeh, Brian J.; van Dyk, Dewald; Davidson, Alan R.; Andrews, Brenda J.; Weiss, Eric L.; Moses, Alan M.

    2016-01-01

    At least 30% of human proteins are thought to contain intrinsically disordered regions, which lack stable structural conformation. Despite lacking enzymatic functions and having few protein domains, disordered regions are functionally important for protein regulation and contain short linear motifs (short peptide sequences involved in protein-protein interactions), but in most disordered regions, the functional amino acid residues remain unknown. We searched for evolutionarily conserved sequences within disordered regions according to the hypothesis that conservation would indicate functional residues. Using a phylogenetic hidden Markov model (phylo-HMM), we made accurate, specific predictions of functional elements in disordered regions even when these elements are only two or three amino acids long. Among the conserved sequences that we identified were previously known and newly identified short linear motifs, and we experimentally verified key examples, including a motif that may mediate interaction between protein kinase Cbk1 and its substrates. We also observed that hub proteins, which interact with many partners in a protein interaction network, are highly enriched in these conserved sequences. Our analysis enabled the systematic identification of the functional residues in disordered regions and suggested that at least 5% of amino acids in disordered regions are important for function. PMID:22416277

  10. Local Function Conservation in Sequence and Structure Space

    PubMed Central

    Weinhold, Nils; Sander, Oliver; Domingues, Francisco S.; Lengauer, Thomas; Sommer, Ingolf

    2008-01-01

    We assess the variability of protein function in protein sequence and structure space. Various regions in this space exhibit considerable difference in the local conservation of molecular function. We analyze and capture local function conservation by means of logistic curves. Based on this analysis, we propose a method for predicting molecular function of a query protein with known structure but unknown function. The prediction method is rigorously assessed and compared with a previously published function predictor. Furthermore, we apply the method to 500 functionally unannotated PDB structures and discuss selected examples. The proposed approach provides a simple yet consistent statistical model for the complex relations between protein sequence, structure, and function. The GOdot method is available online (http://godot.bioinf.mpi-inf.mpg.de). PMID:18604264

  11. Amino acid sequence of Salmonella typhimurium branched-chain amino acid aminotransferase.

    PubMed

    Feild, M J; Nguyen, D C; Armstrong, F B

    1989-06-13

    The complete amino acid sequence of the subunit of branched-chain amino acid aminotransferase (transaminase B, EC 2.6.1.42) of Salmonella typhimurium was determined. An Escherichia coli recombinant containing the ilvGEDAY gene cluster of Salmonella was used as the source of the hexameric enzyme. The peptide fragments used for sequencing were generated by treatment with trypsin, Staphylococcus aureus V8 protease, endoproteinase Lys-C, and cyanogen bromide. The enzyme subunit contains 308 residues and has a molecular weight of 33,920. To determine the coenzyme-binding site, the pyridoxal 5-phosphate containing enzyme was treated with tritiated sodium borohydride prior to trypsin digestion. Peptide map comparisons with an apoenzyme tryptic digest and monitoring radioactivity incorporation allowed identification of the pyridoxylated peptide, which was then isolated and sequenced. The coenzyme-binding site is the lysyl residue at position 159. The amino acid sequence of Salmonella transaminase B is 97.4% identical with that of Escherichia coli, differing in only eight amino acid positions. Sequence comparisons of transaminase B to other known aminotransferase sequences revealed limited sequence similarity (24-33%) when conserved amino acid substitutions are allowed and alignments were forced to occur on the coenzyme-binding site. PMID:2669973

  12. Properties of Sequence Conservation in Upstream Regulatory and Protein Coding Sequences among Paralogs in Arabidopsis thaliana

    NASA Astrophysics Data System (ADS)

    Richardson, Dale N.; Wiehe, Thomas

    Whole genome duplication (WGD) has catalyzed the formation of new species, genes with novel functions, altered expression patterns, complexified signaling pathways and has provided organisms a level of genetic robustness. We studied the long-term evolution and interrelationships of 5’ upstream regulatory sequences (URSs), protein coding sequences (CDSs) and expression correlations (EC) of duplicated gene pairs in Arabidopsis. Three distinct methods revealed significant evolutionary conservation between paralogous URSs and were highly correlated with microarray-based expression correlation of the respective gene pairs. Positional information on exact matches between sequences unveiled the contribution of micro-chromosomal rearrangements on expression divergence. A three-way rank analysis of URS similarity, CDS divergence and EC uncovered specific gene functional biases. Transcription factor activity was associated with gene pairs exhibiting conserved URSs and divergent CDSs, whereas a broad array of metabolic enzymes was found to be associated with gene pairs showing diverged URSs but conserved CDSs.

  13. Sturgeon conservation genomics: SNP discovery and validation using RAD sequencing.

    PubMed

    Ogden, R; Gharbi, K; Mugue, N; Martinsohn, J; Senn, H; Davey, J W; Pourkazemi, M; McEwing, R; Eland, C; Vidotto, M; Sergeev, A; Congiu, L

    2013-06-01

    Caviar-producing sturgeons belonging to the genus Acipenser are considered to be one of the most endangered species groups in the world. Continued overfishing in spite of increasing legislation, zero catch quotas and extensive aquaculture production have led to the collapse of wild stocks across Europe and Asia. The evolutionary relationships among Adriatic, Russian, Persian and Siberian sturgeons are complex because of past introgression events and remain poorly understood. Conservation management, traceability and enforcement suffer a lack of appropriate DNA markers for the genetic identification of sturgeon at the species, population and individual level. This study employed RAD sequencing to discover and characterize single nucleotide polymorphism (SNP) DNA markers for use in sturgeon conservation in these four tetraploid species over three biological levels, using a single sequencing lane. Four population meta-samples and eight individual samples from one family were barcoded separately before sequencing. Analysis of 14.4 Gb of paired-end RAD data focused on the identification of SNPs in the paired-end contig, with subsequent in silico and empirical validation of candidate markers. Thousands of putatively informative markers were identified including, for the first time, SNPs that show population-wide differentiation between Russian and Persian sturgeons, representing an important advance in our ability to manage these cryptic species. The results highlight the challenges of genotyping-by-sequencing in polyploid taxa, while establishing the potential genetic resources for developing a new range of caviar traceability and enforcement tools. PMID:23473098

  14. Conservation patterns in different functional sequence categoriesof divergent Drosophila species

    SciTech Connect

    Papatsenko, Dmitri; Kislyuk, Andrey; Levine, Michael; Dubchak, Inna

    2005-10-01

    We have explored the distributions of fully conservedungapped blocks in genome-wide pairwise alignments of recently completedspecies of Drosophila: D.yakuba, D.ananassae, D.pseudoobscura, D.virilisand D.mojavensis. Based on these distributions we have found that nearlyevery functional sequence category possesses its own distinctiveconservation pattern, sometimes independent of the overall sequenceconservation level. In the coding and regulatory regions, the ungappedblocks were longer than in introns, UTRs and non-functional sequences. Atthe same time, the blocks in the coding regions carried 3N+2 signaturecharacteristic to synonymic substitutions in the 3rd codon positions.Larger block sizes in transcription regulatory regions can be explainedby the presence of conserved arrays of binding sites for transcriptionfactors. We also have shown that the longest ungapped blocks, or'ultraconserved' sequences, are associated with specific gene groups,including those encoding ion channels and components of the cytoskeleton.We discussed how restrained conservation patterns may help in mappingfunctional sequence categories and improving genomeannotation.

  15. Conservation patterns in angiosperm rDNA ITS2 sequences.

    PubMed Central

    Hershkovitz, M A; Zimmer, E A

    1996-01-01

    The two internal transcribed spacers (ITS1 and ITS2) of nuclear ribosomal DNA have become commonly exploited sources of informative variation for interspecific-/intergeneric-level phylogenetic analyses among angiosperms and other eukaryotes. We present an alignment in which one-third to one-half of the ITS2 sequence is alignable above the family level in angiosperms and a phenetic analysis showing that ITS2 contains information sufficient to diagnose lineages at several hierarchical levels. Base compositional analysis shows that angiosperm ITS2 is inherently GC-rich, and that the proportion of T is much more variable than that for other bases. We propose a general model of angiosperm ITS2 secondary structure that shows common pairing relationships for most of the conserved sequence tracts. Variations in our secondary structure predictions for sequences from different taxa indicate that compensatory mutation is not limited to paired positions. PMID:8760866

  16. Conservative Patch Algorithm and Mesh Sequencing for PAB3D

    NASA Technical Reports Server (NTRS)

    Pao, S. P.; Abdol-Hamid, K. S.

    2005-01-01

    A mesh-sequencing algorithm and a conservative patched-grid-interface algorithm (hereafter Patch Algorithm ) have been incorporated into the PAB3D code, which is a computer program that solves the Navier-Stokes equations for the simulation of subsonic, transonic, or supersonic flows surrounding an aircraft or other complex aerodynamic shapes. These algorithms are efficient, flexible, and have added tremendously to the capabilities of PAB3D. The mesh-sequencing algorithm makes it possible to perform preliminary computations using only a fraction of the grid cells (provided the original cell count is divisible by an integer) along any grid coordinate axis, independently of the other axes. The patch algorithm addresses another critical need in multi-block grid situation where the cell faces of adjacent grid blocks may not coincide, leading to errors in calculating fluxes of conserved physical quantities across interfaces between the blocks. The patch algorithm, based on the Stokes integral formulation of the applicable conservation laws, effectively matches each of the interfacial cells on one side of the block interface to the corresponding fractional cell area pieces on the other side. This approach is comprehensive and unified such that all interface topology is automatically processed without user intervention. This algorithm is implemented in a preprocessing code that creates a cell-by-cell database that will maintain flux conservation at any level of full or reduced grid density as the user may choose by way of the mesh-sequencing algorithm. These two algorithms have enhanced the numerical accuracy of the code, reduced the time and effort for grid preprocessing, and provided users with the flexibility of performing computations at any desired full or reduced grid resolution to suit their specific computational requirements.

  17. Conserved Sequences at the Origin of Adenovirus DNA Replication

    PubMed Central

    Stillman, Bruce W.; Topp, William C.; Engler, Jeffrey A.

    1982-01-01

    The origin of adenovirus DNA replication lies within an inverted sequence repetition at either end of the linear, double-stranded viral DNA. Initiation of DNA replication is primed by a deoxynucleoside that is covalently linked to a protein, which remains bound to the newly synthesized DNA. We demonstrate that virion-derived DNA-protein complexes from five human adenovirus serological subgroups (A to E) can act as a template for both the initiation and the elongation of DNA replication in vitro, using nuclear extracts from adenovirus type 2 (Ad2)-infected HeLa cells. The heterologous template DNA-protein complexes were not as active as the homologous Ad2 DNA, most probably due to inefficient initiation by Ad2 replication factors. In an attempt to identify common features which may permit this replication, we have also sequenced the inverted terminal repeated DNA from human adenovirus serotypes Ad4 (group E), Ad9 and Ad10 (group D), and Ad31 (group A), and we have compared these to previously determined sequences from Ad2 and Ad5 (group C), Ad7 (group B), and Ad12 and Ad18 (group A) DNA. In all cases, the sequence around the origin of DNA replication can be divided into two structural domains: a proximal A · T-rich region which is partially conserved among these serotypes, and a distal G · C-rich region which is less well conserved. The G · C-rich region contains sequences similar to sequences present in papovavirus replication origins. The two domains may reflect a dual mechanism for initiation of DNA replication: adenovirus-specific protein priming of replication, and subsequent utilization of this primer by host replication factors for completion of DNA synthesis. Images PMID:7143575

  18. In Vivo Enhancer Analysis Chromosome 16 Conserved NoncodingSequences

    SciTech Connect

    Pennacchio, Len A.; Ahituv, Nadav; Moses, Alan M.; Nobrega,Marcelo; Prabhakar, Shyam; Shoukry, Malak; Minovitsky, Simon; Visel,Axel; Dubchak, Inna; Holt, Amy; Lewis, Keith D.; Plajzer-Frick, Ingrid; Akiyama, Jennifer; De Val, Sarah; Afzal, Veena; Black, Brian L.; Couronne, Olivier; Eisen, Michael B.; Rubin, Edward M.

    2006-02-01

    The identification of enhancers with predicted specificitiesin vertebrate genomes remains a significant challenge that is hampered bya lack of experimentally validated training sets. In this study, weleveraged extreme evolutionary sequence conservation as a filter toidentify putative gene regulatory elements and characterized the in vivoenhancer activity of human-fish conserved and ultraconserved1 noncodingelements on human chromosome 16 as well as such elements from elsewherein the genome. We initially tested 165 of these extremely conservedsequences in a transgenic mouse enhancer assay and observed that 48percent (79/165) functioned reproducibly as tissue-specific enhancers ofgene expression at embryonic day 11.5. While driving expression in abroad range of anatomical structures in the embryo, the majority of the79 enhancers drove expression in various regions of the developingnervous system. Studying a set of DNA elements that specifically droveforebrain expression, we identified DNA signatures specifically enrichedin these elements and used these parameters to rank all ~;3,400human-fugu conserved noncoding elements in the human genome. The testingof the top predictions in transgenic mice resulted in a three-foldenrichment for sequences with forebrain enhancer activity. These datadramatically expand the catalogue of in vivo-characterized human geneenhancers and illustrate the future utility of such training sets for avariety of iological applications including decoding the regulatoryvocabulary of the human genome.

  19. Phenolic acid esterases, coding sequences and methods

    DOEpatents

    Blum, David L.; Kataeva, Irina; Li, Xin-Liang; Ljungdahl, Lars G.

    2002-01-01

    Described herein are four phenolic acid esterases, three of which correspond to domains of previously unknown function within bacterial xylanases, from XynY and XynZ of Clostridium thermocellum and from a xylanase of Ruminococcus. The fourth specifically exemplified xylanase is a protein encoded within the genome of Orpinomyces PC-2. The amino acids of these polypeptides and nucleotide sequences encoding them are provided. Recombinant host cells, expression vectors and methods for the recombinant production of phenolic acid esterases are also provided.

  20. Amino-Acid Sequence of Porcine Pepsin

    PubMed Central

    Tang, J.; Sepulveda, P.; Marciniszyn, J.; Chen, K. C. S.; Huang, W-Y.; Tao, N.; Liu, D.; Lanier, J. P.

    1973-01-01

    As the culmination of several years of experiments, we propose a complete amino-acid sequence for porcine pepsin, an enzyme containing 327 amino-acid residues in a single polypeptide chain. In the sequence determination, the enzyme was treated with cyanogen bromide. Five resulting fragments were purified. The amino-acid sequence of four of the fragments accounted for 290 residues. Because the structure of a 37-residue carboxyl-terminal fragment was already known, it was not studied. The alignment of these fragments was determined from the sequence of methionyl-peptides we had previously reported. We also discovered the locations of activesite aspartyl residues, as well as the pairing of the three disulfide bridges. A minor component of commercial crystalline pepsin was found to contain two extra amino-acid residues, Ala-Leu-, at the amino-terminus of the molecule. This minor component was apparently derived from a different site of cleavage during the activation of porcine pepsinogen. PMID:4587252

  1. Method for identifying and quantifying nucleic acid sequence aberrations

    DOEpatents

    Lucas, Joe N.; Straume, Tore; Bogen, Kenneth T.

    1998-01-01

    A method for detecting nucleic acid sequence aberrations by detecting nucleic acid sequences having both a first and a second nucleic acid sequence type, the presence of the first and second sequence type on the same nucleic acid sequence indicating the presence of a nucleic acid sequence aberration. The method uses a first hybridization probe which includes a nucleic acid sequence that is complementary to a first sequence type and a first complexing agent capable of attaching to a second complexing agent and a second hybridization probe which includes a nucleic acid sequence that selectively hybridizes to the second nucleic acid sequence type over the first sequence type and includes a detectable marker for detecting the second hybridization probe.

  2. Method for identifying and quantifying nucleic acid sequence aberrations

    DOEpatents

    Lucas, J.N.; Straume, T.; Bogen, K.T.

    1998-07-21

    A method is disclosed for detecting nucleic acid sequence aberrations by detecting nucleic acid sequences having both a first and a second nucleic acid sequence type, the presence of the first and second sequence type on the same nucleic acid sequence indicating the presence of a nucleic acid sequence aberration. The method uses a first hybridization probe which includes a nucleic acid sequence that is complementary to a first sequence type and a first complexing agent capable of attaching to a second complexing agent and a second hybridization probe which includes a nucleic acid sequence that selectively hybridizes to the second nucleic acid sequence type over the first sequence type and includes a detectable marker for detecting the second hybridization probe. 11 figs.

  3. Methods for analyzing nucleic acid sequences

    DOEpatents

    Korlach, Jonas; Webb, Watt W.; Levene, Michael; Turner, Stephen; Craighead, Harold G.; Foquet, Mathieu

    2011-05-17

    The present invention is directed to a method of sequencing a target nucleic acid. The method provides a complex comprising a polymerase enzyme, a target nucleic acid molecule, and a primer, wherein the complex is immobilized on a support Fluorescent label is attached to a terminal phosphate group of the nucleotide or nucleotide analog. The growing nucleic acid strand is extended by using the polymerase to add a nucleotide analog to the nucleic acid strand. The nucleotide analog added to the oligonucleotide primer as a result of the polymerizing step is identified. The time duration of the signal from labeled nucleotides or nucleotide analogs that become incorporated is distinguished from freely diffusing labels by a longer retention in the observation volume for the nucleotides or nucleotide analogs that become incorporated than for the freely diffusing labels.

  4. Polymorphism, monomorphism, and sequences in conserved microsatellites in primate species.

    PubMed

    Blanquer-Maumont, A; Crouau-Roy, B

    1995-10-01

    Dimeric short tandem repeats are a source of highly polymorphic markers in the mammalian genome. Genetic variation at these hypervariable loci is extensively used for linkage analysis, for the identification of individuals, and may be useful for interpopulation and interspecies studies. In this paper, we analyze the variability and the sequences of a segment including three microsatellites, first described in man, in several species of primates (chimpanzee, orangutan, gibbon, and macaque) using the heterologous primers (man primers). This region is located on the human chromosome 6p, near the tumor necrosis factor genes, in the major histocompatibility complex. The fact that these primers work in all species studied indicates that they are conserved throughout the different lineages of the two superfamilies, the Hominoidea and the Cercopithecidea, represented by the macaques. However, the intervening sequence displays intraspecific and interspecific variability. The sites of base substitutions and the insertion/deletion events are not evenly distributed within this region. The data suggest that it is necessary to have a minimal number of repeats to increase the rate of mutation sufficiently to allow the development of polymorphism. In some species, the microsatellites present single base variations which reduce the number of contiguous repeats, thus apparently slowing the rate of additional slippage events. Species with such variations or a low number of repeats are monomorphic. These microsatellite sequences are informative in the comparison of closely related species and reflect the phylogeny of the Old World monkeys, apes, and man. PMID:7563137

  5. Human immunodeficiency virus type 1 and 2 envelope glycoproteins oligomerize through conserved sequences.

    PubMed Central

    Center, R J; Kemp, B E; Poumbourios, P

    1997-01-01

    Hetero-oligomerization between human immunodeficiency virus type 2 (HIV-2) envelope glycoprotein (Env) truncation mutants and epitope-tagged gp160 is dependent on the presence of gp41 transmembrane protein (TM) amino acids 552 to 589, a putative amphipathic alpha-helical sequence. HIV-2 Env truncation mutants containing this sequence were also able to form cross-type hetero-oligomers with HIV-1 Env. HIV-2/HIV-1 hetero-oligomerization was, however, more sensitive to disruption by mutagenesis or increased temperature. The conservation of the Env oligomerization function of the HIV-1 and HIV-2 alpha-helical sequences suggests that retroviral TM alpha-helical motifs may have a universal role in oligomerization. PMID:9188654

  6. Quantification of tertiary structural conservation despite primary sequence drift in the globin fold.

    PubMed

    Aronson, H E; Royer, W E; Hendrickson, W A

    1994-10-01

    The globin family of protein structures was the first for which it was recognized that tertiary structure can be highly conserved even when primary sequences have diverged to a virtually undetectable level of similarity. This principle of structural inertia in molecular evolution is now evident for many other protein families. We have performed a systematic comparison of the sequences and structures of 6 representative hemoglobin subunits as diverse in origin as plants, clams, and humans. Our analysis is based on a 97-residue helical core in common to all 6 structures. Amino acid sequence identities range from 12.4% to 42.3% in pairwise comparisons, and, despite these variations, the maximal RMS deviation in alpha-carbon positions is 3.02 A. Overall, sequence similarity and structural deviation are significantly anticorrelated, with a correlation coefficient of -0.71, but for a set of structures having under 20% pairwise identity, this anticorrelation falls to -0.38, which emphasizes the weak connection between a specific sequence and the tertiary fold. There is substantial variability in structure outside the helical core, and functional characteristics of these globins also differ appreciably. Nevertheless, despite variations in detail that the sequence dissimilarities and functional differences imply, the core structures of these globins remain remarkably preserved. PMID:7849587

  7. The amino acid sequence of chymopapain from Carica papaya.

    PubMed Central

    Watson, D C; Yaguchi, M; Lynn, K R

    1990-01-01

    Chymopapain is a polypeptide of 218 amino acid residues. It has considerable structural similarity with papain and papaya proteinase omega, including conservation of the catalytic site and of the disulphide bonding. Chymopapain is like papaya proteinase omega in carrying four extra residues between papain positions 168 and 169, but differs from both papaya proteinases in the composition of its S2 subsite, as well as in having a second thiol group, Cys-117. Some evidence for the amino acid sequence of chymopapain has been deposited as Supplementary Publication SUP 50153 (12 pages) at the British Library Document Supply Centre, Boston Spa., Wetherby, West Yorkshire LS23 7BQ, U.K., from whom copies may be obtained on the terms indicated in Biochem. J. (1990) 265, 5. The information comprises Supplement Tables 1-4, which contain, in order, amino acid compositions of peptides from tryptic, peptic, CNBr and mild acid cleavages, Supplement Fig. 1, showing re-fractionation of selected peaks from Fig. 2 of the main paper. Supplement Fig. 2, showing cation-exchange chromatography of the earliest-eluted peak of Fig. 3 of the main paper, Supplement Fig. 3, showing reverse-phase h.p.l.c. of the later-eluted peak from Fig. 3 of the main paper, and Supplement Fig. 4, showing the separation of peptides after mild acid hydrolysis of CNBr-cleavage fragment CB3. PMID:2106878

  8. The amino acid sequence of chymopapain from Carica papaya.

    PubMed

    Watson, D C; Yaguchi, M; Lynn, K R

    1990-02-15

    Chymopapain is a polypeptide of 218 amino acid residues. It has considerable structural similarity with papain and papaya proteinase omega, including conservation of the catalytic site and of the disulphide bonding. Chymopapain is like papaya proteinase omega in carrying four extra residues between papain positions 168 and 169, but differs from both papaya proteinases in the composition of its S2 subsite, as well as in having a second thiol group, Cys-117. Some evidence for the amino acid sequence of chymopapain has been deposited as Supplementary Publication SUP 50153 (12 pages) at the British Library Document Supply Centre, Boston Spa., Wetherby, West Yorkshire LS23 7BQ, U.K., from whom copies may be obtained on the terms indicated in Biochem. J. (1990) 265, 5. The information comprises Supplement Tables 1-4, which contain, in order, amino acid compositions of peptides from tryptic, peptic, CNBr and mild acid cleavages, Supplement Fig. 1, showing re-fractionation of selected peaks from Fig. 2 of the main paper. Supplement Fig. 2, showing cation-exchange chromatography of the earliest-eluted peak of Fig. 3 of the main paper, Supplement Fig. 3, showing reverse-phase h.p.l.c. of the later-eluted peak from Fig. 3 of the main paper, and Supplement Fig. 4, showing the separation of peptides after mild acid hydrolysis of CNBr-cleavage fragment CB3. PMID:2106878

  9. Detection of nucleic acid sequences by invader-directed cleavage

    DOEpatents

    Brow, Mary Ann D.; Hall, Jeff Steven Grotelueschen; Lyamichev, Victor; Olive, David Michael; Prudent, James Robert

    1999-01-01

    The present invention relates to means for the detection and characterization of nucleic acid sequences, as well as variations in nucleic acid sequences. The present invention also relates to methods for forming a nucleic acid cleavage structure on a target sequence and cleaving the nucleic acid cleavage structure in a site-specific manner. The 5' nuclease activity of a variety of enzymes is used to cleave the target-dependent cleavage structure, thereby indicating the presence of specific nucleic acid sequences or specific variations thereof. The present invention further relates to methods and devices for the separation of nucleic acid molecules based by charge.

  10. Hybridization and sequencing of nucleic acids using base pair mismatches

    DOEpatents

    Fodor, Stephen P. A.; Lipshutz, Robert J.; Huang, Xiaohua

    2001-01-01

    Devices and techniques for hybridization of nucleic acids and for determining the sequence of nucleic acids. Arrays of nucleic acids are formed by techniques, preferably high resolution, light-directed techniques. Positions of hybridization of a target nucleic acid are determined by, e.g., epifluorescence microscopy. Devices and techniques are proposed to determine the sequence of a target nucleic acid more efficiently and more quickly through such synthesis and detection techniques.

  11. Search for conserved amino acid residues of the [Formula: see text]-crystallin proteins of vertebrates.

    PubMed

    Shiliaev, Nikita G; Selivanova, Olga M; Galzitskaya, Oxana V

    2016-04-01

    [Formula: see text]-crystallin is the major eye lens protein and a member of the small heat-shock protein (sHsp) family. [Formula: see text]-crystallins have been shown to support lens clarity by preventing the aggregation of lens proteins. We performed the bioinformatics analysis of [Formula: see text]-crystallin sequences from vertebrates to find conserved amino acid residues as the three-dimensional (3D) structure of [Formula: see text]-crystallin is not identified yet. We are the first who demonstrated that the N-terminal region is conservative along with the central domain for vertebrate organisms. We have found that there is correlation between the conserved and structured regions. Moreover, amyloidogenic regions also correspond to the structured regions. We analyzed the amino acid composition of [Formula: see text]-crystallin A and B chains. Analyzing the occurrence of each individual amino acid residue, we have found that such amino acid residues as leucine, serine, lysine, proline, phenylalanine, histidine, isoleucine, glutamic acid, and valine change their content simultaneously in A and B chains in different classes of vertebrates. Aromatic amino acids occur more often in [Formula: see text]-crystallins from vertebrates than on the average in proteins among 17 animal proteomes. We obtained that the identity between A and B chains in the mammalian group is 0.35, which is lower than the published 0.60. PMID:26972563

  12. 77 FR 65537 - Requirements for Patent Applications Containing Nucleotide Sequence and/or Amino Acid Sequence...

    Federal Register 2010, 2011, 2012, 2013, 2014

    2012-10-29

    ... Amino Acid Sequence Disclosures ACTION: Proposed collection; comment request. SUMMARY: The United States....'' SUPPLEMENTARY INFORMATION: I. Abstract Patent applications that contain nucleotide and/or amino acid sequence disclosures must include a copy of the sequence listing in accordance with the requirements in 37 CFR...

  13. Conservation of acid waterlogged shipwrecks: nanotechnologies for de-acidification

    NASA Astrophysics Data System (ADS)

    Giorgi, R.; Chelazzi, D.; Baglioni, P.

    2006-06-01

    Preservation of waterlogged wooden artifacts, and in particular ancient wrecks, is a challenge in cultural heritage conservation. Samples, from the Swedish warship Vasa, are under investigation in order to develop innovative methods for wood de-acidification and preservation. The Vasa represents a unique case in the study of ancient wrecks. In the past four years the problem of the acidity of wood emerged as a strong threat to its conservation. The production of sulphuric acid inside the ship wood might be the cause of both chemical damage through the acid hydrolysis of cellulose, and of physical damage of the wood’s pore structure, due to the crystallization of sulphate minerals in the wood pores. In this paper we show that wood acidity can be neutralized by the application of nanoparticles of alkaline-earth carbonates and/or hydroxides. The treatment provides an alkaline reservoir inside the wood. Nanoparticles absorbed in the wood from an alcoholic dispersion adhere to the wood wall and release hydroxyl ions leading to the wood neutralization. Oak and pine samples from the Vasa wreck were characterized and treated with alkaline magnesium or calcium nanoparticle dispersions in non-aqueous solvents. De-acidification was monitored by pH changes and thermal analysis, and all the treated samples were submitted to thermal artificial ageing in order to demonstrate the efficacy of the method. The results obtained opened a new perspective in wood conservation.

  14. Matrix genes of measles virus and canine distemper virus: cloning, nucleotide sequences, and deduced amino acid sequences.

    PubMed Central

    Bellini, W J; Englund, G; Richardson, C D; Rozenblatt, S; Lazzarini, R A

    1986-01-01

    The nucleotide sequences encoding the matrix (M) proteins of measles virus (MV) and canine distemper virus (CDV) were determined from cDNA clones containing these genes in their entirety. In both cases, single open reading frames specifying basic proteins of 335 amino acid residues were predicted from the nucleotide sequences. Both viral messages were composed of approximately 1,450 nucleotides and contained 400 nucleotides of presumptive noncoding sequences at their respective 3' ends. MV and CDV M-protein-coding regions were 67% homologous at the nucleotide level and 76% homologous at the amino acid level. Only chance homology was observed in the 400-nucleotide trailer sequences. Comparisons of the M protein sequences of MV and CDV with the sequence reported for Sendai virus (B. M. Blumberg, K. Rose, M. G. Simona, L. Roux, C. Giorgi, and D. Kolakofsky, J. Virol. 52:656-663; Y. Hidaka, T. Kanda, K. Iwasaki, A. Nomoto, T. Shioda, and H. Shibuta, Nucleic Acids Res. 12:7965-7973) indicated the greatest homology among these M proteins in the carboxyterminal third of the molecule. Secondary-structure analyses of this shared region indicated a structurally conserved, hydrophobic sequence which possibly interacted with the lipid bilayer. Images PMID:3754588

  15. Sequence Conservation, Radial Distance and Packing Density in Spherical Viral Capsids

    PubMed Central

    Lee, Chi-Wen; Huang, Tsun-Tsao; Shih, Chung-Shiuan; Hwang, Jenn-Kang

    2015-01-01

    The conservation level of a residue is a useful measure about the importance of that residue in protein structure and function. Much information about sequence conservation comes from aligning homologous sequences. Profiles showing the variation of the conservation level along the sequence are usually interpreted in evolutionary terms and dictated by site similarities of a proper set of homologous sequences. Here, we report that, of the viral icosahedral capsids, the sequence conservation profile can be determined by variations in the distances between residues and the centroid of the capsid – with a direct inverse proportionality between the conservation level and the centroid distance – as well as by the spatial variations in local packing density. Examining both the centroid and the packing density models against a dataset of 51 crystal structures of nonhomologous icosahedral capsids, we found that many global patterns and minor features derived from the viral structures are consistent with those present in the sequence conservation profiles. The quantitative link between the level of conservation and structural features like centroid-distance or packing density allows us to look at residue conservation from a structural viewpoint as well as from an evolutionary viewpoint. PMID:26132081

  16. Comparison of mealybug (Planococcus lilacinus) and fruit fly genomes: isolation and analysis of conserved sequences and their utility in studying synteny in the mealybug.

    PubMed

    Mohan, K N; Rani, B S; Selvam, S; Debarshi, S; Kadandale, J S

    2007-01-01

    By using ligation-mediated PCR products from mealybug DNA as tester and biotinylated fly DNA as driver, we recovered a fraction of the tester that remains hybridized to driver following high-stringency washing conditions. This fraction is expected to contain mealybug sequences conserved in the fly (MCF). Reciprocal experiments enabled the isolation of fly sequences conserved in the mealybug (FCM). Coding sequences among MCF show amino acid identities >40% with fly proteins, allowing a reliable identification of orthologs. Three sequences from the fly cytogenetic positions 98-99 were hybridized onto mealybug chromosomes and the results identified differences in synteny between the two species. Taken together, our results present a method for direct isolation of sequences conserved between an 'orphan' (mealybug) genome and a 'reference' (fly) genome and showed that these sequences can be used to study chromosome synteny in the mealybug. PMID:18253039

  17. Predicting intrinsic disorder from amino acid sequence.

    PubMed

    Obradovic, Zoran; Peng, Kang; Vucetic, Slobodan; Radivojac, Predrag; Brown, Celeste J; Dunker, A Keith

    2003-01-01

    Blind predictions of intrinsic order and disorder were made on 42 proteins subsequently revealed to contain 9,044 ordered residues, 284 disordered residues in 26 segments of length 30 residues or less, and 281 disordered residues in 2 disordered segments of length greater than 30 residues. The accuracies of the six predictors used in this experiment ranged from 77% to 91% for the ordered regions and from 56% to 78% for the disordered segments. The average of the order and disorder predictions ranged from 73% to 77%. The prediction of disorder in the shorter segments was poor, from 25% to 66% correct, while the prediction of disorder in the longer segments was better, from 75% to 95% correct. Four of the predictors were composed of ensembles of neural networks. This enabled them to deal more efficiently with the large asymmetry in the training data through diversified sampling from the significantly larger ordered set and achieve better accuracy on ordered and long disordered regions. The exclusive use of long disordered regions for predictor training likely contributed to the disparity of the predictions on long versus short disordered regions, while averaging the output values over 61-residue windows to eliminate short predictions of order or disorder probably contributed to the even greater disparity for three of the predictors. This experiment supports the predictability of intrinsic disorder from amino acid sequence. PMID:14579347

  18. BlockLogo: visualization of peptide and sequence motif conservation.

    PubMed

    Olsen, Lars Rønn; Kudahl, Ulrich Johan; Simon, Christian; Sun, Jing; Schönbach, Christian; Reinherz, Ellis L; Zhang, Guang Lan; Brusic, Vladimir

    2013-12-31

    BlockLogo is a web-server application for the visualization of protein and nucleotide fragments, continuous protein sequence motifs, and discontinuous sequence motifs using calculation of block entropy from multiple sequence alignments. The user input consists of a multiple sequence alignment, selection of motif positions, type of sequence, and output format definition. The output has BlockLogo along with the sequence logo, and a table of motif frequencies. We deployed BlockLogo as an online application and have demonstrated its utility through examples that show visualization of T-cell epitopes and B-cell epitopes (both continuous and discontinuous). Our additional example shows a visualization and analysis of structural motifs that determine the specificity of peptide binding to HLA-DR molecules. The BlockLogo server also employs selected experimentally validated prediction algorithms to enable on-the-fly prediction of MHC binding affinity to 15 common HLA class I and class II alleles as well as visual analysis of discontinuous epitopes from multiple sequence alignments. It enables the visualization and analysis of structural and functional motifs that are usually described as regular expressions. It provides a compact view of discontinuous motifs composed of distant positions within biological sequences. BlockLogo is available at: http://research4.dfci.harvard.edu/cvc/blocklogo/ and http://met-hilab.bu.edu/blocklogo/. PMID:24001880

  19. Methods and compositions for efficient nucleic acid sequencing

    DOEpatents

    Drmanac, Radoje

    2002-01-01

    Disclosed are novel methods and compositions for rapid and highly efficient nucleic acid sequencing based upon hybridization with two sets of small oligonucleotide probes of known sequences. Extremely large nucleic acid molecules, including chromosomes and non-amplified RNA, may be sequenced without prior cloning or subcloning steps. The methods of the invention also solve various current problems associated with sequencing technology such as, for example, high noise to signal ratios and difficult discrimination, attaching many nucleic acid fragments to a surface, preparing many, longer or more complex probes and labelling more species.

  20. Methods and compositions for efficient nucleic acid sequencing

    DOEpatents

    Drmanac, Radoje

    2006-07-04

    Disclosed are novel methods and compositions for rapid and highly efficient nucleic acid sequencing based upon hybridization with two sets of small oligonucleotide probes of known sequences. Extremely large nucleic acid molecules, including chromosomes and non-amplified RNA, may be sequenced without prior cloning or subcloning steps. The methods of the invention also solve various current problems associated with sequencing technology such as, for example, high noise to signal ratios and difficult discrimination, attaching many nucleic acid fragments to a surface, preparing many, longer or more complex probes and labelling more species.

  1. High Sequence Conservation of Human Immunodeficiency Virus Type 1 Reverse Transcriptase under Drug Pressure despite the Continuous Appearance of Mutations

    PubMed Central

    Ceccherini-Silberstein, Francesca; Gago, Federico; Santoro, Maria; Gori, Caterina; Svicher, Valentina; Rodríguez-Barrios, Fátima; d'Arrigo, Roberta; Ciccozzi, Massimo; Bertoli, Ada; Monforte, Antonella d'Arminio; Balzarini, Jan; Antinori, Andrea; Perno, Carlo-Federico

    2005-01-01

    To define the extent of sequence conservation in human immunodeficiency virus type 1 (HIV-1) reverse transcriptase (RT) in vivo, the first 320 amino acids of RT obtained from 2,236 plasma-derived samples from a well-defined cohort of 1,704 HIV-1-infected individuals (457 drug naïve and 1,247 drug treated) were analyzed and examined in structural terms. In naïve patients, 233 out of these 320 residues (73%) were conserved (<1% variability). The majority of invariant amino acids clustered into defined regions comprising between 5 and 29 consecutive residues. Of the nine longest invariant regions identified, some contained residues and domains critical for enzyme stability and function. In patients treated with RT inhibitors, despite profound drug pressure and the appearance of mutations primarily associated with resistance, 202 amino acids (63%) remained highly conserved and appeared mostly distributed in regions of variable length. This finding suggests that participation of consecutive residues in structural domains is strictly required for cooperative functions and sustainability of HIV-1 RT activity. Besides confirming the conservation of amino acids that are already known to be important for catalytic activity, stability of the heterodimer interface, and/or primer/template binding, the other 62 new invariable residues are now identified and mapped onto the three-dimensional structure of the enzyme. This new knowledge could be of help in the structure-based design of novel resistance-evading drugs. PMID:16051864

  2. Kit for detecting nucleic acid sequences using competitive hybridization probes

    DOEpatents

    Lucas, Joe N.; Straume, Tore; Bogen, Kenneth T.

    2001-01-01

    A kit is provided for detecting a target nucleic acid sequence in a sample, the kit comprising: a first hybridization probe which includes a nucleic acid sequence that is sufficiently complementary to selectively hybridize to a first portion of the target sequence, the first hybridization probe including a first complexing agent for forming a binding pair with a second complexing agent; and a second hybridization probe which includes a nucleic acid sequence that is sufficiently complementary to selectively hybridize to a second portion of the target sequence to which the first hybridization probe does not selectively hybridize, the second hybridization probe including a detectable marker; a third hybridization probe which includes a nucleic acid sequence that is sufficiently complementary to selectively hybridize to a first portion of the target sequence, the third hybridization probe including the same detectable marker as the second hybridization probe; and a fourth hybridization probe which includes a nucleic acid sequence that is sufficiently complementary to selectively hybridize to a second portion of the target sequence to which the third hybridization probe does not selectively hybridize, the fourth hybridization probe including the first complexing agent for forming a binding pair with the second complexing agent; wherein the first and second hybridization probes are capable of simultaneously hybridizing to the target sequence and the third and fourth hybridization probes are capable of simultaneously hybridizing to the target sequence, the detectable marker is not present on the first or fourth hybridization probes and the first, second, third, and fourth hybridization probes each include a competitive nucleic acid sequence which is sufficiently complementary to a third portion of the target sequence that the competitive sequences of the first, second, third, and fourth hybridization probes compete with each other to hybridize to the third portion of the

  3. High sequence conservation among cucumber mosaic virus isolates from lily.

    PubMed

    Chen, Y K; Derks, A F; Langeveld, S; Goldbach, R; Prins, M

    2001-08-01

    For classification of Cucumber mosaic virus (CMV) isolates from ornamental crops of different geographical areas, these were characterized by comparing the nucleotide sequences of RNAs 4 and the encoded coat proteins. Within the ornamental-infecting CMV viruses both subgroups were represented. CMV isolates of Alstroemeria and crocus were classified as subgroup II isolates, whereas 8 other isolates, from lily, gladiolus, amaranthus, larkspur, and lisianthus, were identified as subgroup I members. In general, nucleotide sequence comparisons correlated well with geographic distribution, with one notable exception: the analyzed nucleotide sequences of 5 lily isolates showed remarkably high homology despite different origins. PMID:11676424

  4. Studies on monotreme proteins. VII. Amino acid sequence of myoglobin from the platypus, Ornithoryhynchus anatinus.

    PubMed

    Fisher, W K; Thompson, E O

    1976-03-01

    Myoglobin isolated from skeletal muscle of the platypus contains 153 amino acid residues. The complete amino acid sequence has been determined following cleavage with cyanogen bromide and further digestion of the four fragments with trypsin, chymotrypsin, pepsin and thermolysin. Sequences of the purified peptides were determined by the dansyl-Edman procedure. The amino acid sequence showed 25 differences from human myoglobin and 24 from kangaroo myoglobin. Amino acid sequences in myoglobins are more conserved than sequences in the alpha- and beta-globin chains, and platypus myoglobin shows a similar number of variations in sequence to kangaroo myoglobin when compared with myoglobin of other species. The date of divergence of the platypus from other mammals was estimated at 102 +/- 31 million years, based on the number of amino acid differences between species and allowing for mutations during the evolutionary period. This estimate differs widely from the estimate given by similar treatment of the alpha- and beta-chain sequences and a constant rate of mutation of globin chains is not supported. PMID:962722

  5. Solid phase sequencing of double-stranded nucleic acids

    DOEpatents

    Fu, Dong-Jing; Cantor, Charles R.; Koster, Hubert; Smith, Cassandra L.

    2002-01-01

    This invention relates to methods for detecting and sequencing of target double-stranded nucleic acid sequences, to nucleic acid probes and arrays of probes useful in these methods, and to kits and systems which contain these probes. Useful methods involve hybridizing the nucleic acids or nucleic acids which represent complementary or homologous sequences of the target to an array of nucleic acid probes. These probe comprise a single-stranded portion, an optional double-stranded portion and a variable sequence within the single-stranded portion. The molecular weights of the hybridized nucleic acids of the set can be determined by mass spectroscopy, and the sequence of the target determined from the molecular weights of the fragments. Nucleic acids whose sequences can be determined include nucleic acids in biological samples such as patient biopsies and environmental samples. Probes may be fixed to a solid support such as a hybridization chip to facilitate automated determination of molecular weights and identification of the target sequence.

  6. Analysis and Annotation of Nucleic Acid Sequence

    SciTech Connect

    States, David J.

    2004-07-28

    The aims of this project were to develop improved methods for computational genome annotation and to apply these methods to improve the annotation of genomic sequence data with a specific focus on human genome sequencing. The project resulted in a substantial body of published work. Notable contributions of this project were the identification of basecalling and lane tracking as error processes in genome sequencing and contributions to improved methods for these steps in genome sequencing. This technology improved the accuracy and throughput of genome sequence analysis. Probabilistic methods for physical map construction were developed. Improved methods for sequence alignment, alternative splicing analysis, promoter identification and NF kappa B response gene prediction were also developed.

  7. Analysis and Annotation of Nucleic Acid Sequence

    SciTech Connect

    David J. States

    1998-08-01

    The aims of this project were to develop improved methods for computational genome annotation and to apply these methods to improve the annotation of genomic sequence data with a specific focus on human genome sequencing. The project resulted in a substantial body of published work. Notable contributions of this project were the identification of basecalling and lane tracking as error processes in genome sequencing and contributions to improved methods for these steps in genome sequencing. This technology improved the accuracy and throughput of genome sequence analysis. Probabilistic methods for physical map construction were developed. Improved methods for sequence alignment, alternative splicing analysis, promoter identification and NF kappa B response gene prediction were also developed.

  8. Sequence analysis of the L protein of the Ebola 2014 outbreak: Insight into conserved regions and mutations.

    PubMed

    Ayub, Gohar; Waheed, Yasir

    2016-06-01

    The 2014 Ebola outbreak was one of the largest that have occurred; it started in Guinea and spread to Nigeria, Liberia and Sierra Leone. Phylogenetic analysis of the current virus species indicated that this outbreak is the result of a divergent lineage of the Zaire ebolavirus. The L protein of Ebola virus (EBOV) is the catalytic subunit of the RNA‑dependent RNA polymerase complex, which, with VP35, is key for the replication and transcription of viral RNA. Earlier sequence analysis demonstrated that the L protein of all non‑segmented negative‑sense (NNS) RNA viruses consists of six domains containing conserved functional motifs. The aim of the present study was to analyze the presence of these motifs in 2014 EBOV isolates, highlight their function and how they may contribute to the overall pathogenicity of the isolates. For this purpose, 81 2014 EBOV L protein sequences were aligned with 475 other NNS RNA viruses, including Paramyxoviridae and Rhabdoviridae viruses. Phylogenetic analysis of all EBOV outbreak L protein sequences was also performed. Analysis of the amino acid substitutions in the 2014 EBOV outbreak was conducted using sequence analysis. The alignment demonstrated the presence of previously conserved motifs in the 2014 EBOV isolates and novel residues. Notably, all the mutations identified in the 2014 EBOV isolates were tolerant, they were pathogenic with certain examples occurring within previously determined functional conserved motifs, possibly altering viral pathogenicity, replication and virulence. The phylogenetic analysis demonstrated that all sequences with the exception of the 2014 EBOV sequences were clustered together. The 2014 EBOV outbreak has acquired a great number of mutations, which may explain the reasons behind this unprecedented outbreak. Certain residues critical to the function of the polymerase remain conserved and may be targets for the development of antiviral therapeutic agents. PMID:27082438

  9. From Artificial Amino Acids to Sequence-Defined Targeted Oligoaminoamides.

    PubMed

    Morys, Stephan; Wagner, Ernst; Lächelt, Ulrich

    2016-01-01

    Artificial oligoamino acids with appropriate protecting groups can be used for the sequential assembly of oligoaminoamides on solid-phase. With the help of these oligoamino acids multifunctional nucleic acid (NA) carriers can be designed and produced in highly defined topologies. Here we describe the synthesis of the artificial oligoamino acid Fmoc-Stp(Boc3)-OH, the subsequent assembly into sequence-defined oligomers and the formulation of tumor-targeted plasmid DNA (pDNA) polyplexes. PMID:27436323

  10. Involvement of phylogenetically conserved acidic amino acid residues in catalysis by an oxidative DNA damage enzyme formamidopyrimidine glycosylase.

    PubMed

    Lavrukhin, O V; Lloyd, R S

    2000-12-12

    Formamidopyrimidine glycosylase (Fpg) is an important bacterial base excision repair enzyme, which initiates removal of damaged purines such as the highly mutagenic 8-oxoguanine. Similar to other glycosylase/AP lyases, catalysis by Fpg is known to proceed by a nucleophilic attack by an amino group (the secondary amine of its N-terminal proline) on C1' of the deoxyribose sugar at a damaged base, which results in the departure of the base from the DNA and removal of the sugar ring by beta/delta-elimination. However, in contrast to other enzymes in this class, in which acidic amino acids have been shown to be essential for glycosyl and phosphodiester bond scission, the catalytically essential acidic residues have not been documented for Fpg. Multiple sequence alignments of conserved acidic residues in all known bacterial Fpg-like proteins revealed six conserved glutamic and aspartic acid residues. Site-directed mutagenesis was used to change glutamic and aspartic acid residues to glutamines and asparagines, respectively. While the Asp to Asn mutants had no effect on the incision activity on 8-oxoguanine-containing DNA, several of the substitutions at glutamates reduced Fpg activity on the 8-oxoguanosine DNA, with the E3Q and E174Q mutants being essentially devoid of activity. The AP lyase activity of all of the glutamic acid mutants was slightly reduced as compared to the wild-type enzyme. Sodium borohydride trapping of wild-type Fpg and its E3Q and E174Q mutants on 8-oxoguanosine or AP site containing DNA correlated with the relative activity of the mutants on either of these substrates. PMID:11106507

  11. Complete cDNA and derived amino acid sequence of human factor V

    SciTech Connect

    Jenny, R.J.; Pittman, D.D.; Toole, J.J.; Kriz, R.W.; Aldape, R.A.; Hewick, R.M.; Kaufman, R.J.; Mann, K.G.

    1987-07-01

    cDNA clones encoding human factor V have been isolated from an oligo(dT)-primed human fetal liver cDNA library prepared with vector Charon 21A. The cDNA sequence of factor V from three overlapping clones includes a 6672-base-pair (bp) coding region, a 90-bp 5' untranslated region, and a 163-bp 3' untranslated region within which is a poly(A)tail. The deduced amino acid sequence consists of 2224 amino acids inclusive of a 28-amino acid leader peptide. Direct comparison with human factor VIII reveals considerable homology between proteins in amino acid sequence and domain structure: a triplicated A domain and duplicated C domain show approx. 40% identity with the corresponding domains in factor VIII. As in factor VIII, the A domains of factor V share approx. 40% amino acid-sequence homology with the three highly conserved domains in ceruloplasmin. The B domain of factor V contains 35 tandem and approx. 9 additional semiconserved repeats of nine amino acids of the form Asp-Leu-Ser-Gln-Thr-Thr/Asn-Leu-Ser-Pro and 2 additional semiconserved repeats of 17 amino acids. Factor V contains 37 potential N-linked glycosylation sites, 25 of which are in the B domain, and a total of 19 cysteine residues.

  12. Detecting frame shifts by amino acid sequence comparison.

    PubMed

    Claverie, J M

    1993-12-20

    Various amino acid substitution scoring matrices are used in conjunction with local alignments programs to detect regions of similarity and infer potential common ancestry between proteins. The usual scoring schemes derive from the implicit hypothesis that related proteins evolve from a common ancestor by the accumulation of point mutations and that amino acids tend to be progressively substituted by others with similar properties. However, other frequent single mutation events, like nucleotide insertion or deletion and gene inversion, change the translation reading frame and cause previously encoded amino acid sequences to become unrecognizable at once. Here, I derive five new types of scoring matrix, each capable of detecting a specific frame shift (deletion, insertion and inversion in 3 frames) and use them with a regular local alignments program to detect amino acid sequences that may have derived from alternative reading frames of the same nucleotide sequence. Frame shifts are inferred from the sole comparison of the protein sequences. The five scoring matrices were used with the BLASTP program to compare all the protein sequences in the Swissprot database. Surprisingly, the searches revealed hundreds of highly significant frame shift matches, of which many are likely to represent sequencing errors. Others provide some evidence that frame shift mutations might be used in protein evolution as a way to create new amino acid sequences from pre-existing coding regions. PMID:7903399

  13. Distinct Functional Constraints Partition Sequence Conservation in a cis-Regulatory Element

    PubMed Central

    Ruvinsky, Ilya

    2011-01-01

    Different functional constraints contribute to different evolutionary rates across genomes. To understand why some sequences evolve faster than others in a single cis-regulatory locus, we investigated function and evolutionary dynamics of the promoter of the Caenorhabditis elegans unc-47 gene. We found that this promoter consists of two distinct domains. The proximal promoter is conserved and is largely sufficient to direct appropriate spatial expression. The distal promoter displays little if any conservation between several closely related nematodes. Despite this divergence, sequences from all species confer robustness of expression, arguing that this function does not require substantial sequence conservation. We showed that even unrelated sequences have the ability to promote robust expression. A prominent feature shared by all of these robustness-promoting sequences is an AT-enriched nucleotide composition consistent with nucleosome depletion. Because general sequence composition can be maintained despite sequence turnover, our results explain how different functional constraints can lead to vastly disparate rates of sequence divergence within a promoter. PMID:21655084

  14. Segments of amino acid sequence similarity in beta-amylases.

    PubMed

    Friedberg, F; Rhodes, C

    1988-01-01

    In alpha-amylases from animals, plants and bacteria and in beta-amylases from plants and bacteria a number of segments exhibit amino acid sequence similarity specific to the alpha or to the beta type, respectively. In the case of the beta-amylases the similar sequence regions are extensive and they are disrupted only by short interspersed dissimilar regions. Close to the C terminus, however, no such sequence similarity exist. PMID:2464171

  15. Concentration of Specific Amino Acids at the Catalytic/Active Centers of Highly-Conserved ``Housekeeping'' Enzymes of Central Metabolism in Archaea, Bacteria and Eukaryota: Is There a Widely Conserved Chemical Signal of Prebiotic Assembly?

    NASA Astrophysics Data System (ADS)

    Pollack, J. Dennis; Pan, Xueliang; Pearl, Dennis K.

    2010-06-01

    In alignments of 1969 protein sequences the amino acid glycine and others were found concentrated at most-conserved sites within ˜15 Å of catalytic/active centers (C/AC) of highly conserved kinases, dehydrogenases or lyases of Archaea, Bacteria and Eukaryota. Lysine and glutamic acid were concentrated at least-conserved sites furthest from their C/ACs. Logistic-regression analyses corroborated the “movement” of glycine towards and lysine away from their C/ACs: the odds of a glycine occupying a site were decreased by 19%, while the odds for a lysine were increased by 53%, for every 10 Å moving away from the C/AC. Average conservation of MSA consensus sites was highest surrounding the C/AC and directly decreased in transition toward model’s peripheries. Findings held with statistical confidence using sequences restricted to individual Domains or enzyme classes or to both. Our data describe variability in the rate of mutation and likelihoods for phylogenetic trees based on protein sequence data and endorse the extension of substitution models by incorporating data on conservation and distance to C/ACs rather than only using cumulative levels. The data support the view that in the most-conserved environment immediately surrounding the C/AC of taxonomically distant and highly conserved essential enzymes of central metabolism there are amino acids whose identity and degree of occupancy is similar to a proposed amino acid set and frequency associated with prebiotic evolution.

  16. Accelerated Evolution of Conserved Noncoding Sequences in theHuman Genome

    SciTech Connect

    Prambhakar, Shyam; Noonan, James P.; Paabo, Svante; Rubin, EdwardM.

    2006-07-06

    Genomic comparisons between human and distant, non-primatemammals are commonly used to identify cis-regulatory elements based onconstrained sequence evolution. However, these methods fail to detect"cryptic" functional elements, which are too weakly conserved amongmammals to distinguish from nonfunctional DNA. To address this problem,we explored the potential of deep intra-primate sequence comparisons. Wesequenced the orthologs of 558 kb of human genomic sequence, coveringmultiple loci involved in cholesterol homeostasis, in 6 nonhumanprimates. Our analysis identified 6 noncoding DNA elements displayingsignificant conservation among primates, but undetectable in more distantcomparisons. In vitro and in vivo tests revealed that at least three ofthese 6 elements have regulatory function. Notably, the mouse orthologsof these three functional human sequences had regulatory activity despitetheir lack of significant sequence conservation, indicating that they arecryptic ancestral cis-regulatory elements. These regulatory elementscould still be detected in a smaller set of three primate speciesincluding human, rhesus and marmoset. Since the human and rhesus genomesequences are already available, and the marmoset genome is activelybeing sequenced, the primate-specific conservation analysis describedhere can be applied in the near future on a whole-genome scale, tocomplement the annotation provided by more distant speciescomparisons.

  17. Detection of Weakly Conserved Ancestral Mammalian RegulatorySequences by Primate Comparisons

    SciTech Connect

    Wang, Qian-fei; Prabhakar, Shyam; Chanan, Sumita; Cheng,Jan-Fang; Rubin, Edward M.; Boffelli, Dario

    2006-06-01

    Genomic comparisons between human and distant, non-primatemammals are commonly used to identify cis-regulatory elements based onconstrained sequence evolution. However, these methods fail to detectcryptic functional elements, which are too weakly conserved among mammalsto distinguish from nonfunctional DNA. To address this problem, weexplored the potential of deep intra-primate sequence comparisons. Wesequenced the orthologs of 558 kb of human genomic sequence, coveringmultiple loci involved in cholesterol homeostasis, in 6 nonhumanprimates. Our analysis identified 6 noncoding DNA elements displayingsignificant conservation among primates, but undetectable in more distantcomparisons. In vitro and in vivo tests revealed that at least three ofthese 6 elements have regulatory function. Notably, the mouse orthologsof these three functional human sequences had regulatory activity despitetheir lack of significant sequence conservation, indicating that they arecryptic ancestral cis-regulatory elements. These regulatory elementscould still be detected in a smaller set of three primate speciesincluding human, rhesus and marmoset. Since the human and rhesus genomesequences are already available, and the marmoset genome is activelybeing sequenced, the primate-specific conservation analysis describedhere can be applied in the near future on a whole-genome scale, tocomplement the annotation provided by more distant speciescomparisons.

  18. The role of evolutionarily conserved germ-line DH sequence in B-1 cell development and natural antibody production.

    PubMed

    Vale, Andre M; Nobrega, Alberto; Schroeder, Harry W

    2015-12-01

    Because of N addition and variation in the site of VDJ joining, the third complementarity-determining region of the heavy chain (CDR-H3) is the most diverse component of the initial immunoglobulin antigen-binding site repertoire. A large component of the peritoneal cavity B-1 cell component is the product of fetal and perinatal B cell production. The CDR-H3 repertoire is thus depleted of N addition, which increases dependency on germ-line sequence. Cross-species comparisons have shown that DH gene sequence demonstrates conservation of amino acid preferences by reading frame. Preference for reading frame 1, which is enriched for tyrosine and glycine, is created both by rearrangement patterns and by pre-BCR and BCR selection. In previous studies, we have assessed the role of conserved DH sequence by examining peritoneal cavity B-1 cell numbers and antibody production in BALB/c mice with altered DH loci. Here, we review our finding that changes in the constraints normally imposed by germ-line-encoded amino acids within the CDR-H3 repertoire profoundly affect B-1 cell development, especially B-1a cells, and thus natural antibody immunity. Our studies suggest that both natural and somatic selection operate to create a restricted B-1 cell CDR-H3 repertoire. PMID:26104486

  19. Conservation of the human telomere sequence (TTAGGG)n among vertebrates.

    PubMed Central

    Meyne, J; Ratliff, R L; Moyzis, R K

    1989-01-01

    To determine the evolutionary origin of the human telomere sequence (TTAGGG)n, biotinylated oligodeoxynucleotides of this sequence were hybridized to metaphase spreads from 91 different species, including representative orders of bony fish, reptiles, amphibians, birds, and mammals. Under stringent hybridization conditions, fluorescent signals were detected at the telomeres of all chromosomes, in all 91 species. The conservation of the (TTAGGG)n sequence and its telomeric location, in species thought to share a common ancestor over 400 million years ago, strongly suggest that this sequence is the functional vertebrate telomere. Images PMID:2780561

  20. Nucleotide and predicted amino acid sequences of cloned human and mouse preprocathepsin B cDNAs.

    PubMed Central

    Chan, S J; San Segundo, B; McCormick, M B; Steiner, D F

    1986-01-01

    Cathepsin B is a lysosomal thiol proteinase that may have additional extralysosomal functions. To further our investigations on the structure, mode of biosynthesis, and intracellular sorting of this enzyme, we have determined the complete coding sequences for human and mouse preprocathepsin B by using cDNA clones isolated from human hepatoma and kidney phage libraries. The nucleotide sequences predict that the primary structure of preprocathepsin B contains 339 amino acids organized as follows: a 17-residue NH2-terminal prepeptide sequence followed by a 62-residue propeptide region, 254 residues in mature (single chain) cathepsin B, and a 6-residue extension at the COOH terminus. A comparison of procathepsin B sequences from three species (human, mouse, and rat) reveals that the homology between the propeptides is relatively conserved with a minimum of 68% sequence identity. In particular, two conserved sequences in the propeptide that may be functionally significant include a potential glycosylation site and the presence of a single cysteine at position 59. Comparative analysis of the three sequences also suggests that processing of procathepsin B is a multistep process, during which enzymatically active intermediate forms may be generated. The availability of the cDNA clones will facilitate the identification of possible active or inactive intermediate processive forms as well as studies on the transcriptional regulation of the cathepsin B gene. PMID:3463996

  1. 37 CFR 1.821 - Nucleotide and/or amino acid sequence disclosures in patent applications.

    Code of Federal Regulations, 2013 CFR

    2013-07-01

    ... acids are not intended to be embraced by this definition. Any amino acid sequence that contains post-translationally modified amino acids may be described as the amino acid sequence that is initially translated... sequence of four or more amino acids or an unbranched sequence of ten or more nucleotides....

  2. 37 CFR 1.821 - Nucleotide and/or amino acid sequence disclosures in patent applications.

    Code of Federal Regulations, 2012 CFR

    2012-07-01

    ... acids are not intended to be embraced by this definition. Any amino acid sequence that contains post-translationally modified amino acids may be described as the amino acid sequence that is initially translated... sequence of four or more amino acids or an unbranched sequence of ten or more nucleotides....

  3. 37 CFR 1.821 - Nucleotide and/or amino acid sequence disclosures in patent applications.

    Code of Federal Regulations, 2014 CFR

    2014-07-01

    ... acids are not intended to be embraced by this definition. Any amino acid sequence that contains post-translationally modified amino acids may be described as the amino acid sequence that is initially translated... sequence of four or more amino acids or an unbranched sequence of ten or more nucleotides....

  4. Phylum-Level Conservation of Regulatory Information in Nematodes despite Extensive Non-coding Sequence Divergence

    PubMed Central

    Gordon, Kacy L.; Arthur, Robert K.; Ruvinsky, Ilya

    2015-01-01

    Gene regulatory information guides development and shapes the course of evolution. To test conservation of gene regulation within the phylum Nematoda, we compared the functions of putative cis-regulatory sequences of four sets of orthologs (unc-47, unc-25, mec-3 and elt-2) from distantly-related nematode species. These species, Caenorhabditis elegans, its congeneric C. briggsae, and three parasitic species Meloidogyne hapla, Brugia malayi, and Trichinella spiralis, represent four of the five major clades in the phylum Nematoda. Despite the great phylogenetic distances sampled and the extensive sequence divergence of nematode genomes, all but one of the regulatory elements we tested are able to drive at least a subset of the expected gene expression patterns. We show that functionally conserved cis-regulatory elements have no more extended sequence similarity to their C. elegans orthologs than would be expected by chance, but they do harbor motifs that are important for proper expression of the C. elegans genes. These motifs are too short to be distinguished from the background level of sequence similarity, and while identical in sequence they are not conserved in orientation or position. Functional tests reveal that some of these motifs contribute to proper expression. Our results suggest that conserved regulatory circuitry can persist despite considerable turnover within cis elements. PMID:26020930

  5. A method to find palindromes in nucleic acid sequences.

    PubMed

    Anjana, Ramnath; Shankar, Mani; Vaishnavi, Marthandan Kirti; Sekar, Kanagaraj

    2013-01-01

    Various types of sequences in the human genome are known to play important roles in different aspects of genomic functioning. Among these sequences, palindromic nucleic acid sequences are one such type that have been studied in detail and found to influence a wide variety of genomic characteristics. For a nucleotide sequence to be considered as a palindrome, its complementary strand must read the same in the opposite direction. For example, both the strands i.e the strand going from 5' to 3' and its complementary strand from 3' to 5' must be complementary. A typical nucleotide palindromic sequence would be TATA (5' to 3') and its complimentary sequence from 3' to 5' would be ATAT. Thus, a new method has been developed using dynamic programming to fetch the palindromic nucleic acid sequences. The new method uses less memory and thereby it increases the overall speed and efficiency. The proposed method has been tested using the bacterial (3891 KB bases) and human chromosomal sequences (Chr-18: 74366 kb and Chr-Y: 25554 kb) and the computation time for finding the palindromic sequences is in milli seconds. PMID:23515654

  6. The amino-acid sequence of leghemoglobin component a from Phaseolus vulgaris (kidney bean).

    PubMed

    Lehtovaara, P; Ellfolk, N

    1975-06-01

    1. Leghemoglobin component a from Phaseolus vulgaris (kidney bean) was digested with trypsin; 15 tryptic peptides and free lysine were purified and the amino acid sequences of the peptides determined. 2. The internal order of the tryptic peptides was determined by the bridge peptides obtained from the thermolytic digest and the dilute acid hydrolyzate of kidney bean leghemoglobin a; 12 thermolytic peptides and two acid hydrolysis peptides were purified and the sequences were partially or completely determined. 3. The complete amino acid sequence of kidney bean leghemoglobin a is compared to that of leghemoglobin a from soybean (Glycine max) and to some animal globins. As regards sequence, the kidney bean globin has 79% identity with the soybean globin and 21% identity with human hemoglobin gamma-chain. Seven of the 14 amino acid residues common to most globins are found in the kidney bean globin. Trp-15 and Tyr-145 are evolutionarily conserved in this globin, which confirms the concept of a common origin of animal and plant globins. PMID:809270

  7. In silico comparative analysis of DNA and amino acid sequences for prion protein gene.

    PubMed

    Kim, Y; Lee, J; Lee, C

    2008-01-01

    Genetic variability might contribute to species specificity of prion diseases in various organisms. In this study, structures of the prion protein gene (PRNP) and its amino acids were compared among species of which sequence data were available. Comparisons of PRNP DNA sequences among 12 species including human, chimpanzee, monkey, bovine, ovine, dog, mouse, rat, wallaby, opossum, chicken and zebrafish allowed us to identify candidate regulatory regions in intron 1 and 3'-untranslated region (UTR) in addition to the coding region. Highly conserved putative binding sites for transcription factors, such as heat shock factor 2 (HSF2) and myocite enhancer factor 2 (MEF2), were discovered in the intron 1. In 3'-UTR, the functional sequence (ATTAAA) for nucleus-specific polyadenylation was found in all the analysed species. The functional sequence (TTTTTAT) for maturation-specific polyadenylation was identically observed only in ovine, and one or two nucleotide mismatches in the other species. A comparison of the amino acid sequences in 53 species revealed a large sequence identity. Especially the octapeptide repeat region was observed in all the species but frog and zebrafish. Functional changes and susceptibility to prion diseases with various isoforms of prion protein could be caused by numeric variability and conformational changes discovered in the repeat sequences. PMID:18397498

  8. Amino acid sequence repertoire of the bacterial proteome and the occurrence of untranslatable sequences.

    PubMed

    Navon, Sharon Penias; Kornberg, Guy; Chen, Jin; Schwartzman, Tali; Tsai, Albert; Puglisi, Elisabetta Viani; Puglisi, Joseph D; Adir, Noam

    2016-06-28

    Bioinformatic analysis of Escherichia coli proteomes revealed that all possible amino acid triplet sequences occur at their expected frequencies, with four exceptions. Two of the four underrepresented sequences (URSs) were shown to interfere with translation in vivo and in vitro. Enlarging the URS by a single amino acid resulted in increased translational inhibition. Single-molecule methods revealed stalling of translation at the entrance of the peptide exit tunnel of the ribosome, adjacent to ribosomal nucleotides A2062 and U2585. Interaction with these same ribosomal residues is involved in regulation of translation by longer, naturally occurring protein sequences. The E. coli exit tunnel has evidently evolved to minimize interaction with the exit tunnel and maximize the sequence diversity of the proteome, although allowing some interactions for regulatory purposes. Bioinformatic analysis of the human proteome revealed no underrepresented triplet sequences, possibly reflecting an absence of regulation by interaction with the exit tunnel. PMID:27307442

  9. Mutational Studies on Resurrected Ancestral Proteins Reveal Conservation of Site-Specific Amino Acid Preferences throughout Evolutionary History

    PubMed Central

    Risso, Valeria A.; Manssour-Triedo, Fadia; Delgado-Delgado, Asunción; Arco, Rocio; Barroso-delJesus, Alicia; Ingles-Prieto, Alvaro; Godoy-Ruiz, Raquel; Gavira, Jose A.; Gaucher, Eric A.; Ibarra-Molero, Beatriz; Sanchez-Ruiz, Jose M.

    2015-01-01

    Local protein interactions (“molecular context” effects) dictate amino acid replacements and can be described in terms of site-specific, energetic preferences for any different amino acid. It has been recently debated whether these preferences remain approximately constant during evolution or whether, due to coevolution of sites, they change strongly. Such research highlights an unresolved and fundamental issue with far-reaching implications for phylogenetic analysis and molecular evolution modeling. Here, we take advantage of the recent availability of phenotypically supported laboratory resurrections of Precambrian thioredoxins and β-lactamases to experimentally address the change of site-specific amino acid preferences over long geological timescales. Extensive mutational analyses support the notion that evolutionary adjustment to a new amino acid may occur, but to a large extent this is insufficient to erase the primitive preference for amino acid replacements. Generally, site-specific amino acid preferences appear to remain conserved throughout evolutionary history despite local sequence divergence. We show such preference conservation to be readily understandable in molecular terms and we provide crystallographic evidence for an intriguing structural-switch mechanism: Energetic preference for an ancestral amino acid in a modern protein can be linked to reorganization upon mutation to the ancestral local structure around the mutated site. Finally, we point out that site-specific preference conservation naturally leads to one plausible evolutionary explanation for the existence of intragenic global suppressor mutations. PMID:25392342

  10. Mutational studies on resurrected ancestral proteins reveal conservation of site-specific amino acid preferences throughout evolutionary history.

    PubMed

    Risso, Valeria A; Manssour-Triedo, Fadia; Delgado-Delgado, Asunción; Arco, Rocio; Barroso-delJesus, Alicia; Ingles-Prieto, Alvaro; Godoy-Ruiz, Raquel; Gavira, Jose A; Gaucher, Eric A; Ibarra-Molero, Beatriz; Sanchez-Ruiz, Jose M

    2015-02-01

    Local protein interactions ("molecular context" effects) dictate amino acid replacements and can be described in terms of site-specific, energetic preferences for any different amino acid. It has been recently debated whether these preferences remain approximately constant during evolution or whether, due to coevolution of sites, they change strongly. Such research highlights an unresolved and fundamental issue with far-reaching implications for phylogenetic analysis and molecular evolution modeling. Here, we take advantage of the recent availability of phenotypically supported laboratory resurrections of Precambrian thioredoxins and β-lactamases to experimentally address the change of site-specific amino acid preferences over long geological timescales. Extensive mutational analyses support the notion that evolutionary adjustment to a new amino acid may occur, but to a large extent this is insufficient to erase the primitive preference for amino acid replacements. Generally, site-specific amino acid preferences appear to remain conserved throughout evolutionary history despite local sequence divergence. We show such preference conservation to be readily understandable in molecular terms and we provide crystallographic evidence for an intriguing structural-switch mechanism: Energetic preference for an ancestral amino acid in a modern protein can be linked to reorganization upon mutation to the ancestral local structure around the mutated site. Finally, we point out that site-specific preference conservation naturally leads to one plausible evolutionary explanation for the existence of intragenic global suppressor mutations. PMID:25392342

  11. Unique amino acid signatures that are evolutionarily conserved distinguish simple-type, epidermal and hair keratins

    PubMed Central

    Strnad, Pavel; Usachov, Valentyn; Debes, Cedric; Gräter, Frauke; Parry, David A. D.; Omary, M. Bishr

    2011-01-01

    Keratins (Ks) consist of central α-helical rod domains that are flanked by non-α-helical head and tail domains. The cellular abundance of keratins, coupled with their selective cell expression patterns, suggests that they diversified to fulfill tissue-specific functions although the primary structure differences between them have not been comprehensively compared. We analyzed keratin sequences from many species: K1, K2, K5, K9, K10, K14 were studied as representatives of epidermal keratins, and compared with K7, K8, K18, K19, K20 and K31, K35, K81, K85, K86, which represent simple-type (single-layered or glandular) epithelial and hair keratins, respectively. We show that keratin domains have striking differences in their amino acids. There are many cysteines in hair keratins but only a small number in epidermal keratins and rare or none in simple-type keratins. The heads and/or tails of epidermal keratins are glycine and phenylalanine rich but alanine poor, whereas parallel domains of hair keratins are abundant in prolines, and those of simple-type epithelial keratins are enriched in acidic and/or basic residues. The observed differences between simple-type, epidermal and hair keratins are highly conserved throughout evolution. Cysteines and histidines, which are infrequent keratin amino acids, are involved in de novo mutations that are markedly overrepresented in keratins. Hence, keratins have evolutionarily conserved and domain-selectively enriched amino acids including glycine and phenylalanine (epidermal), cysteine and proline (hair), and basic and acidic (simple-type epithelial), which reflect unique functions related to structural flexibility, rigidity and solubility, respectively. Our findings also support the importance of human keratin ‘mutation hotspot’ residues and their wild-type counterparts. PMID:22215855

  12. Identification of antimicrobial peptides from teleosts and anurans in expressed sequence tag databases using conserved signal sequences.

    PubMed

    Tessera, Valentina; Guida, Filomena; Juretić, Davor; Tossi, Alessandro

    2012-03-01

    The problem of multidrug resistance requires the efficient and accurate identification of new classes of antimicrobial agents. Endogenous antimicrobial peptides produced by most organisms are a promising source of such molecules. We have exploited the high conservation of signal sequences in teleost and anuran antimicrobial peptides to search cDNA (expressed sequence tag) databases for likely candidates. Subject sequences were then analysed for the presence of potential antimicrobial peptides based on physicochemical properties (amphipathic helical structure, cationicity) and use of the D-descriptor model to predict the therapeutic index (relation between the minimum inhibitory concentration and the concentration giving 50% haemolysis). This analysis also suggested mutations to probe the role of the primary structure in determining potency and selectivity. Selected sequences were chemically synthesized and the antimicrobial activity of the peptides was confirmed. In particular, a short (21-residue) sequence, likely of sticklefish origin, showed potent activity and it was possible to tune the spectrum of action and/or selectivity by combining three directed mutations. Membrane permeabilization studies on both bacterial and host cells indicate that the mode of action was prevalently membranolytic. This method opens up the possibility for more effective searching of the vast and continuously growing expressed sequence tag databases for novel antimicrobial peptides, which are likely abundant, and the efficient identification of the most promising candidates among them. PMID:22188679

  13. Molecular characterization of a bovine Y-specific DNA sequence conserved in taurine and zebu breeds.

    PubMed

    Alves, Beatriz C A; Mayer, Mário G; Taber, Anna Paula; Egito, Andréa A; Fagundes, Valéria; McElreavey, Ken; Moreira-Filho, Carlos A

    2006-06-01

    The identification of new bovine male-specific DNA sequences is of great interest because the bovine Y chromosome remains poorly characterized in terms of physical and genetic maps. Since taurine and zebu Y chromosomes are structurally different, the identification of Y-specific sequences present in both sub-species is particularly important: these sequences are of evolutionary significance and can be broadly used for embryo sexing. In this work, we initially used the random amplified polymorphic DNA (RAPD) technique to search for male-specific sequences present as monomorphic markers in genomic DNA from zebu and taurine bulls. A male-specific RAPD band was found to be present and highly conserved in both sub-species, as demonstrated by Southern blotting, fluorescent in situ hybridization (FISH) and DNA sequencing. In a previous work, a pair of primers derived from this marker was successfully used in taurine and zebu embryo sexing. PMID:17286047

  14. An approach to delineate primers for a group of poorly conserved sequences incorporating the common motif region.

    PubMed

    Sahu, Mousumi; Sahu, Jagajjit; Sahoo, Smita; Dehury, Budheswar; Sarma, Kishore; Sarmah, Ranjan; Sen, Priyabrata; Modi, Mahendra Kumar; Barooah, Madhumita

    2012-01-01

    Glutathione synthetase (gshB) has previously been reported to confer tolerance to acidic soil condition in Rhizobium species. Cloning the gene coding for this enzyme necessitates the designing of proper primer sets which in turn depends on the identification of high quality sequence similarity in multiple global alignments. In this experiment, a group of homologous gene sequences related to gshB gene (accession no: gi-86355669:327589-328536) of Rhizobium etli CFN 42, were extracted from NCBI nucleotide sequence databases using BLASTN and were analyzed for designing degenerate primers. However, the T-coffee multiple global alignment results did not show any block of conserved region for the above sequence set to design the primers. Therefore, we attempted to identify the location of common motif region based on multiple local alignments employing the MEME algorithm supported with MAST and Primer3. The results revealed some common motif regions that enabled us to design the primer sets for related gshB gene sequences. The result will be validated in wet lab. PMID:22419837

  15. On Quantum Algorithm for Multiple Alignment of Amino Acid Sequences

    NASA Astrophysics Data System (ADS)

    Iriyama, Satoshi; Ohya, Masanori

    2009-02-01

    The alignment of genome sequences or amino acid sequences is one of fundamental operations for the study of life. Usual computational complexity for the multiple alignment of N sequences with common length L by dynamic programming is O(LN). This alignment is considered as one of the NP problems, so that it is desirable to find a nice algorithm of the multiple alignment. Thus in this paper we propose the quantum algorithm for the multiple alignment based on the works12,1,2 in which the NP complete problem was shown to be the P problem by means of quantum algorithm and chaos information dynamics.

  16. CONSERVED SEQUENCE IN THE AGGRECAN INTERGLOBULAR DOMAIN MODULATES CLEAVAGE BY ADAMTS-4 AND ADAMTS-5

    PubMed Central

    Miwa, Hazuki E; Gerken, Thomas A; Huynh, Tru D; Duesler, Lori R; Cotter, Meghan; Hering, Thomas M.

    2008-01-01

    Background Cleavage of aggrecan by ADAMTS proteinases at specific sites within highly conserved regions may be important to normal physiological enzyme functions, as well as pathological degradation. Methods To examine ADAMTS selectivity, we assayed ADAMTS-4 and -5 cleavage of recombinant bovine aggrecan mutated at amino acids N-terminal or C-terminal to the interglobular domain cleavage site. Results Mutations of conserved amino acids from P18 to P12 to increase hydrophilicity resulted in ADAMTS-4 cleavage inhibition. Mutation of Thr, but not Asn within the conserved N-glycosylation motif Asn-Ile-Thr from P6 to P4 enhanced cleavage. Mutation of conserved Thr residues from P22 to P17 to increase hydrophobicity enhanced ADAMTS-4 cleavage. A P4′ Ser377Gln mutant inhibited cleavage by ADAMTS-4 and -5, while a neutral Ser377Ala mutant and species mimicking mutants Ser377Thr, Ser377Asn, and Arg375Leu were cleaved normally by ADAMTS-4. The Ser377Thr mutant, however, was resistant to cleavage by ADAMTS-5. Conclusion We have identified multiple conserved amino acids within regions N- and C-terminal to the site of scission that may influence enzyme-substrate recognition, and may interact with exosites on ADAMTS-4 and ADAMTS-5. General Significance Inhibition of the binding of ADAMTS-4 and ADAMTS-5 exosites to aggrecan should be explored as a therapeutic intervention for osteoarthritis. PMID:19101611

  17. Prebiotically plausible mechanisms increase compositional diversity of nucleic acid sequences

    PubMed Central

    Derr, Julien; Manapat, Michael L.; Rajamani, Sudha; Leu, Kevin; Xulvi-Brunet, Ramon; Joseph, Isaac; Nowak, Martin A.; Chen, Irene A.

    2012-01-01

    During the origin of life, the biological information of nucleic acid polymers must have increased to encode functional molecules (the RNA world). Ribozymes tend to be compositionally unbiased, as is the vast majority of possible sequence space. However, ribonucleotides vary greatly in synthetic yield, reactivity and degradation rate, and their non-enzymatic polymerization results in compositionally biased sequences. While natural selection could lead to complex sequences, molecules with some activity are required to begin this process. Was the emergence of compositionally diverse sequences a matter of chance, or could prebiotically plausible reactions counter chemical biases to increase the probability of finding a ribozyme? Our in silico simulations using a two-letter alphabet show that template-directed ligation and high concatenation rates counter compositional bias and shift the pool toward longer sequences, permitting greater exploration of sequence space and stable folding. We verified experimentally that unbiased DNA sequences are more efficient templates for ligation, thus increasing the compositional diversity of the pool. Our work suggests that prebiotically plausible chemical mechanisms of nucleic acid polymerization and ligation could predispose toward a diverse pool of longer, potentially structured molecules. Such mechanisms could have set the stage for the appearance of functional activity very early in the emergence of life. PMID:22319215

  18. The amino-acid sequence of kangaroo pancreatic ribonuclease.

    PubMed

    Gaastra, W; Welling, G W; Beintema, J J

    1978-05-01

    Red kangaroo (Macropus rufus) ribonuclease was isolated from pancreatic tissue by affinity chromatography. The amino acid sequence was determined by automatic sequencing of overlapping large fragments and by analysis of shorter peptides obtained by digestion with a number of proteolytic enzymes. The polypeptide chain consists of 122 amino acid residues. Compared to other ribonucleases, the N-terminal residue and residue 114 are deleted. In other pancreatic ribonucleases position 114 is occupied by a cis proline residue in an external loop at the surface of the molecule. Other remarkable substitutions are the presence of a tyrosine residue at position 123 instead of a serine which forms a hydrogen bond with the pyrimidine ring of a nucleotide substrate, and a number of hydrophobichydrophilic interchanges in the sequence 51-55, which forms part of an alpha-helix in bovine ribonuclease and exhibits few substitutions in the placental mammals. Kangaroo ribonuclease contains no carbohydrate, although the enzyme possesses a recognition site for carbohydrate attachment in the sequence Asn-Val-Thr (62-64). The enzyme differs at about 35-40% of the positions from all other mammalian pancreatic ribonucleases sequenced to date, which is in agreement with the early divergence between the marsupials and the placental mammals. From fragmentary data a tentative sequence of red-necked wallaby (Macropus rufogriseus) pancreatic ribonuclease has been derived. Eight differences with the kangaroo sequence were found. PMID:658039

  19. Studying RNA Homology and Conservation with Infernal: From Single Sequences to RNA Families.

    PubMed

    Barquist, Lars; Burge, Sarah W; Gardner, Paul P

    2016-01-01

    Emerging high-throughput technologies have led to a deluge of putative non-coding RNA (ncRNA) sequences identified in a wide variety of organisms. Systematic characterization of these transcripts will be a tremendous challenge. Homology detection is critical to making maximal use of functional information gathered about ncRNAs: identifying homologous sequence allows us to transfer information gathered in one organism to another quickly and with a high degree of confidence. ncRNA presents a challenge for homology detection, as the primary sequence is often poorly conserved and de novo secondary structure prediction and search remain difficult. This unit introduces methods developed by the Rfam database for identifying "families" of homologous ncRNAs starting from single "seed" sequences, using manually curated sequence alignments to build powerful statistical models of sequence and structure conservation known as covariance models (CMs), implemented in the Infernal software package. We provide a step-by-step iterative protocol for identifying ncRNA homologs and then constructing an alignment and corresponding CM. We also work through an example for the bacterial small RNA MicA, discovering a previously unreported family of divergent MicA homologs in genus Xenorhabdus in the process. © 2016 by John Wiley & Sons, Inc. PMID:27322404

  20. ANTICALIgN: visualizing, editing and analyzing combined nucleotide and amino acid sequence alignments for combinatorial protein engineering.

    PubMed

    Jarasch, Alexander; Kopp, Melanie; Eggenstein, Evelyn; Richter, Antonia; Gebauer, Michaela; Skerra, Arne

    2016-07-01

    ANTIC ALIGN: is an interactive software developed to simultaneously visualize, analyze and modify alignments of DNA and/or protein sequences that arise during combinatorial protein engineering, design and selection. ANTIC ALIGN: combines powerful functions known from currently available sequence analysis tools with unique features for protein engineering, in particular the possibility to display and manipulate nucleotide sequences and their translated amino acid sequences at the same time. ANTIC ALIGN: offers both template-based multiple sequence alignment (MSA), using the unmutated protein as reference, and conventional global alignment, to compare sequences that share an evolutionary relationship. The application of similarity-based clustering algorithms facilitates the identification of duplicates or of conserved sequence features among a set of selected clones. Imported nucleotide sequences from DNA sequence analysis are automatically translated into the corresponding amino acid sequences and displayed, offering numerous options for selecting reading frames, highlighting of sequence features and graphical layout of the MSA. The MSA complexity can be reduced by hiding the conserved nucleotide and/or amino acid residues, thus putting emphasis on the relevant mutated positions. ANTIC ALIGN: is also able to handle suppressed stop codons or even to incorporate non-natural amino acids into a coding sequence. We demonstrate crucial functions of ANTIC ALIGN: in an example of Anticalins selected from a lipocalin random library against the fibronectin extradomain B (ED-B), an established marker of tumor vasculature. Apart from engineered protein scaffolds, ANTIC ALIGN: provides a powerful tool in the area of antibody engineering and for directed enzyme evolution. PMID:27261456

  1. In Silico Structure and Sequence Analysis of Bacterial Porins and Specific Diffusion Channels for Hydrophilic Molecules: Conservation, Multimericity and Multifunctionality.

    PubMed

    Vollan, Hilde S; Tannæs, Tone; Vriend, Gert; Bukholm, Geir

    2016-01-01

    Diffusion channels are involved in the selective uptake of nutrients and form the largest outer membrane protein (OMP) family in Gram-negative bacteria. Differences in pore size and amino acid composition contribute to the specificity. Structure-based multiple sequence alignments shed light on the structure-function relations for all eight subclasses. Entropy-variability analysis results are correlated to known structural and functional aspects, such as structural integrity, multimericity, specificity and biological niche adaptation. The high mutation rate in their surface-exposed loops is likely an important mechanism for host immune system evasion. Multiple sequence alignments for each subclass revealed conserved residue positions that are involved in substrate recognition and specificity. An analysis of monomeric protein channels revealed particular sequence patterns of amino acids that were observed in other classes at multimeric interfaces. This adds to the emerging evidence that all members of the family exist in a multimeric state. Our findings are important for understanding the role of members of this family in a wide range of bacterial processes, including bacterial food uptake, survival and adaptation mechanisms. PMID:27110766

  2. In Silico Structure and Sequence Analysis of Bacterial Porins and Specific Diffusion Channels for Hydrophilic Molecules: Conservation, Multimericity and Multifunctionality

    PubMed Central

    Vollan, Hilde S.; Tannæs, Tone; Vriend, Gert; Bukholm, Geir

    2016-01-01

    Diffusion channels are involved in the selective uptake of nutrients and form the largest outer membrane protein (OMP) family in Gram-negative bacteria. Differences in pore size and amino acid composition contribute to the specificity. Structure-based multiple sequence alignments shed light on the structure-function relations for all eight subclasses. Entropy-variability analysis results are correlated to known structural and functional aspects, such as structural integrity, multimericity, specificity and biological niche adaptation. The high mutation rate in their surface-exposed loops is likely an important mechanism for host immune system evasion. Multiple sequence alignments for each subclass revealed conserved residue positions that are involved in substrate recognition and specificity. An analysis of monomeric protein channels revealed particular sequence patterns of amino acids that were observed in other classes at multimeric interfaces. This adds to the emerging evidence that all members of the family exist in a multimeric state. Our findings are important for understanding the role of members of this family in a wide range of bacterial processes, including bacterial food uptake, survival and adaptation mechanisms. PMID:27110766

  3. Missense mutations and evolutionary conservation of amino acids: evidence that many of the amino acids in factor IX function as "spacer" elements.

    PubMed Central

    Bottema, C D; Ketterling, R P; Ii, S; Yoon, H S; Phillips, J A; Sommer, S S

    1991-01-01

    We report 31 point mutations in the factor IX gene and explore the relationship between the level of evolutionary conservation of an amino acid and the probability of a mutation causing hemophilia B. From our total sample of 125 hemophiliacs and from those reported by others, we identify 95 independent missense mutations, 94 of which occur at amino acids that are evolutionarily conserved in the available mammalian factor IX sequences. The likelihood of a missense mutation causing hemophilia B depends on whether the residue is also conserved in the factor IX-related proteases: factor VII, factor X, and protein C. Most of the possible missense mutations in generically conserved residues (i.e., those conserved in factor IX and in all the related proteases) should cause disease. In contrast, missense mutations in factor IX-specific residues (i.e., those conserved in human, cow, dog, and mouse factor IX but not in the related proteases) are sixfold less likely to cause disease. Missense mutations at nonconserved residues are 33-fold less likely to cause disease. At least three models are compatible with these observations. A comparison of sequence alignments from four and nine species of factor IX and an examination of the missense mutations occurring at CpG residues suggest a model in which most residues fall on opposite ends of a spectrum. In about 40% of residues, virtually any missense mutation in a minority of the residues will cause disease, while virtually no missense mutations will cause disease in most of the remaining residues. Thus, many of the residues in factor IX are spacers; that is, the main chains are presumably necessary to keep other amino acid interactions in register, but the nature of the side chain is unimportant. PMID:1680287

  4. Protein engineering of selected residues from conserved sequence regions of a novel Anoxybacillus α-amylase.

    PubMed

    Ranjani, Velayudhan; Janeček, Stefan; Chai, Kian Piaw; Shahir, Shafinaz; Abdul Rahman, Raja Noor Zaliha Raja; Chan, Kok-Gan; Goh, Kian Mau

    2014-01-01

    The α-amylases from Anoxybacillus species (ASKA and ADTA), Bacillus aquimaris (BaqA) and Geobacillus thermoleovorans (GTA, Pizzo and GtamyII) were proposed as a novel group of the α-amylase family GH13. An ASKA yielding a high percentage of maltose upon its reaction on starch was chosen as a model to study the residues responsible for the biochemical properties. Four residues from conserved sequence regions (CSRs) were thus selected, and the mutants F113V (CSR-I), Y187F and L189I (CSR-II) and A161D (CSR-V) were characterised. Few changes in the optimum reaction temperature and pH were observed for all mutants. Whereas the Y187F (t1/2 43 h) and L189I (t1/2 36 h) mutants had a lower thermostability at 65°C than the native ASKA (t1/2 48 h), the mutants F113V and A161D exhibited an improved t1/2 of 51 h and 53 h, respectively. Among the mutants, only the A161D had a specific activity, k(cat) and k(cat)/K(m) higher (1.23-, 1.17- and 2.88-times, respectively) than the values determined for the ASKA. The replacement of the Ala-161 in the CSR-V with an aspartic acid also caused a significant reduction in the ratio of maltose formed. This finding suggests the Ala-161 may contribute to the high maltose production of the ASKA. PMID:25069018

  5. Using a color-coded ambigraphic nucleic acid notation to visualize conserved palindromic motifs within and across genomes

    PubMed Central

    2014-01-01

    Background Ambiscript is a graphically-designed nucleic acid notation that uses symbol symmetries to support sequence complementation, highlight biologically-relevant palindromes, and facilitate the analysis of consensus sequences. Although the original Ambiscript notation was designed to easily represent consensus sequences for multiple sequence alignments, the notation’s black-on-white ambiguity characters are unable to reflect the statistical distribution of nucleotides found at each position. We now propose a color-augmented ambigraphic notation to encode the frequency of positional polymorphisms in these consensus sequences. Results We have implemented this color-coding approach by creating an Adobe Flash® application ( http://www.ambiscript.org) that shades and colors modified Ambiscript characters according to the prevalence of the encoded nucleotide at each position in the alignment. The resulting graphic helps viewers perceive biologically-relevant patterns in multiple sequence alignments by uniquely combining color, shading, and character symmetries to highlight palindromes and inverted repeats in conserved DNA motifs. Conclusion Juxtaposing an intuitive color scheme over the deliberate character symmetries of an ambigraphic nucleic acid notation yields a highly-functional nucleic acid notation that maximizes information content and successfully embodies key principles of graphic excellence put forth by the statistician and graphic design theorist, Edward Tufte. PMID:24447494

  6. A highly conserved N-terminal sequence for teleost vitellogenin with potential value to the biochemistry, molecular biology and pathology of vitellogenesis

    USGS Publications Warehouse

    Folmar, L.D.; Denslow, N.D.; Wallace, R.A.; LaFleur, G.; Gross, T.S.; Bonomelli, S.; Sullivan, C.V.

    1995-01-01

    N-terminal amino acid sequences for vitellogenin (Vtg) from six species of teleost fish (striped bass, mummichog, pinfish, brown bullhead, medaka, yellow perch and the sturgeon) are compared with published N-terminal Vtg sequences for the lamprey, clawed frog and domestic chicken. Striped bass and mummichog had 100% identical amino acids between positions 7 and 21, while pinfish, brown bullhead, sturgeon, lamprey, Xenopus and chicken had 87%, 93%, 60%, 47%, 47-60%) for four transcripts and had 40% identical, respectively, with striped bass for the same positions. Partial sequences obtained for medaka and yellow perch were 100% identical between positions 5 to 10. The potential utility of this conserved sequence for studies on the biochemistry, molecular biology and pathology of vitellogenesis is discussed.

  7. Antibody-specific model of amino acid substitution for immunological inferences from alignments of antibody sequences.

    PubMed

    Mirsky, Alexander; Kazandjian, Linda; Anisimova, Maria

    2015-03-01

    Antibodies are glycoproteins produced by the immune system as a dynamically adaptive line of defense against invading pathogens. Very elegant and specific mutational mechanisms allow B lymphocytes to produce a large and diversified repertoire of antibodies, which is modified and enhanced throughout all adulthood. One of these mechanisms is somatic hypermutation, which stochastically mutates nucleotides in the antibody genes, forming new sequences with different properties and, eventually, higher affinity and selectivity to the pathogenic target. As somatic hypermutation involves fast mutation of antibody sequences, this process can be described using a Markov substitution model of molecular evolution. Here, using large sets of antibody sequences from mice and humans, we infer an empirical amino acid substitution model AB, which is specific to antibody sequences. Compared with existing general amino acid models, we show that the AB model provides significantly better description for the somatic evolution of mice and human antibody sequences, as demonstrated on large next generation sequencing (NGS) antibody data. General amino acid models are reflective of conservation at the protein level due to functional constraints, with most frequent amino acids exchanges taking place between residues with the same or similar physicochemical properties. In contrast, within the variable part of antibody sequences we observed an elevated frequency of exchanges between amino acids with distinct physicochemical properties. This is indicative of a sui generis mutational mechanism, specific to antibody somatic hypermutation. We illustrate this property of antibody sequences by a comparative analysis of the network modularity implied by the AB model and general amino acid substitution models. We recommend using the new model for computational studies of antibody sequence maturation, including inference of alignments and phylogenetic trees describing antibody somatic hypermutation in

  8. Sequence-related human proteins cluster by degree of evolutionary conservation

    NASA Astrophysics Data System (ADS)

    Mrowka, Ralf; Patzak, Andreas; Herzel, Hanspeter; Holste, Dirk

    2004-11-01

    Gene duplication followed by adaptive evolution is thought to be a central mechanism for the emergence of novel genes. To illuminate the contribution of duplicated protein-coding sequences to the complexity of the human genome, we study the connectivity of pairwise sequence-related human proteins and construct a network (N) of linked protein sequences with shared similarities. We find that (i) the connectivity distribution P(k) for k sequence-related proteins decays as a power law P(k)˜k-γ with γ≈1.2 , (ii) the top rank of N consists of a single large cluster of proteins (≈70%) , while bottom ranks consist of multiple isolated clusters, and (iii) structural characteristics of N show both a high degree of clustering and an intermediate connectivity (“small-world” features). We gain further insight into structural properties of N by studying the relationship between the connectivity distribution and the phylogenetic conservation of proteins in bacteria, plants, invertebrates, and vertebrates. We find that (iv) the proportion of sequence-related proteins increases with increasing extent of evolutionary conservation. Our results support that small-world network properties constitute a footprint of an evolutionary mechanism and extend the traditional interpretation of protein families.

  9. The Large Mitochondrial Genome of Symbiodinium minutum Reveals Conserved Noncoding Sequences between Dinoflagellates and Apicomplexans

    PubMed Central

    Shoguchi, Eiichi; Shinzato, Chuya; Hisata, Kanako; Satoh, Nori; Mungpakdee, Sutada

    2015-01-01

    Even though mitochondrial genomes, which characterize eukaryotic cells, were first discovered more than 50 years ago, mitochondrial genomics remains an important topic in molecular biology and genome sciences. The Phylum Alveolata comprises three major groups (ciliates, apicomplexans, and dinoflagellates), the mitochondrial genomes of which have diverged widely. Even though the gene content of dinoflagellate mitochondrial genomes is reportedly comparable to that of apicomplexans, the highly fragmented and rearranged genome structures of dinoflagellates have frustrated whole genomic analysis. Consequently, noncoding sequences and gene arrangements of dinoflagellate mitochondrial genomes have not been well characterized. Here we report that the continuous assembled genome (∼326 kb) of the dinoflagellate, Symbiodinium minutum, is AT-rich (∼64.3%) and that it contains three protein-coding genes. Based upon in silico analysis, the remaining 99% of the genome comprises transcriptomic noncoding sequences. RNA edited sites and unique, possible start and stop codons clarify conserved regions among dinoflagellates. Our massive transcriptome analysis shows that almost all regions of the genome are transcribed, including 27 possible fragmented ribosomal RNA genes and 12 uncharacterized small RNAs that are similar to mitochondrial RNA genes of the malarial parasite, Plasmodium falciparum. Gene map comparisons show that gene order is only slightly conserved between S. minutum and P. falciparum. However, small RNAs and intergenic sequences share sequence similarities with P. falciparum, suggesting that the function of noncoding sequences has been preserved despite development of very different genome structures. PMID:26199191

  10. The Large Mitochondrial Genome of Symbiodinium minutum Reveals Conserved Noncoding Sequences between Dinoflagellates and Apicomplexans.

    PubMed

    Shoguchi, Eiichi; Shinzato, Chuya; Hisata, Kanako; Satoh, Nori; Mungpakdee, Sutada

    2015-08-01

    Even though mitochondrial genomes, which characterize eukaryotic cells, were first discovered more than 50 years ago, mitochondrial genomics remains an important topic in molecular biology and genome sciences. The Phylum Alveolata comprises three major groups (ciliates, apicomplexans, and dinoflagellates), the mitochondrial genomes of which have diverged widely. Even though the gene content of dinoflagellate mitochondrial genomes is reportedly comparable to that of apicomplexans, the highly fragmented and rearranged genome structures of dinoflagellates have frustrated whole genomic analysis. Consequently, noncoding sequences and gene arrangements of dinoflagellate mitochondrial genomes have not been well characterized. Here we report that the continuous assembled genome (∼326 kb) of the dinoflagellate, Symbiodinium minutum, is AT-rich (∼64.3%) and that it contains three protein-coding genes. Based upon in silico analysis, the remaining 99% of the genome comprises transcriptomic noncoding sequences. RNA edited sites and unique, possible start and stop codons clarify conserved regions among dinoflagellates. Our massive transcriptome analysis shows that almost all regions of the genome are transcribed, including 27 possible fragmented ribosomal RNA genes and 12 uncharacterized small RNAs that are similar to mitochondrial RNA genes of the malarial parasite, Plasmodium falciparum. Gene map comparisons show that gene order is only slightly conserved between S. minutum and P. falciparum. However, small RNAs and intergenic sequences share sequence similarities with P. falciparum, suggesting that the function of noncoding sequences has been preserved despite development of very different genome structures. PMID:26199191

  11. Structure is three to ten times more conserved than sequence--a study of structural response in protein cores.

    PubMed

    Illergård, Kristoffer; Ardell, David H; Elofsson, Arne

    2009-11-15

    Protein structures change during evolution in response to mutations. Here, we analyze the mapping between sequence and structure in a set of structurally aligned protein domains. To avoid artifacts, we restricted our attention only to the core components of these structures. We found that on average, using different measures of structural change, protein cores evolve linearly with evolutionary distance (amino acid substitutions per site). This is true irrespective of which measure of structural change we used, whether RMSD or discrete structural descriptors for secondary structure, accessibility, or contacts. This linear response allows us to quantify the claim that structure is more conserved than sequence. Using structural alphabets of similar cardinality to the sequence alphabet, structural cores evolve three to ten times slower than sequences. Although we observed an average linear response, we found a wide variance. Different domain families varied fivefold in structural response to evolution. An attempt to categorically analyze this variance among subgroups by structural and functional category revealed only one statistically significant trend. This trend can be explained by the fact that beta-sheets change faster than alpha-helices, most likely due to that they are shorter and that change occurs at the ends of the secondary structure elements. PMID:19507241

  12. Auditory sequence processing reveals evolutionarily conserved regions of frontal cortex in macaques and humans.

    PubMed

    Wilson, Benjamin; Kikuchi, Yukiko; Sun, Li; Hunter, David; Dick, Frederic; Smith, Kenny; Thiele, Alexander; Griffiths, Timothy D; Marslen-Wilson, William D; Petkov, Christopher I

    2015-01-01

    An evolutionary account of human language as a neurobiological system must distinguish between human-unique neurocognitive processes supporting language and evolutionarily conserved, domain-general processes that can be traced back to our primate ancestors. Neuroimaging studies across species may determine whether candidate neural processes are supported by homologous, functionally conserved brain areas or by different neurobiological substrates. Here we use functional magnetic resonance imaging in Rhesus macaques and humans to examine the brain regions involved in processing the ordering relationships between auditory nonsense words in rule-based sequences. We find that key regions in the human ventral frontal and opercular cortex have functional counterparts in the monkey brain. These regions are also known to be associated with initial stages of human syntactic processing. This study raises the possibility that certain ventral frontal neural systems, which play a significant role in language function in modern humans, originally evolved to support domain-general abilities involved in sequence processing. PMID:26573340

  13. Auditory sequence processing reveals evolutionarily conserved regions of frontal cortex in macaques and humans

    PubMed Central

    Wilson, Benjamin; Kikuchi, Yukiko; Sun, Li; Hunter, David; Dick, Frederic; Smith, Kenny; Thiele, Alexander; Griffiths, Timothy D.; Marslen-Wilson, William D.; Petkov, Christopher I.

    2015-01-01

    An evolutionary account of human language as a neurobiological system must distinguish between human-unique neurocognitive processes supporting language and evolutionarily conserved, domain-general processes that can be traced back to our primate ancestors. Neuroimaging studies across species may determine whether candidate neural processes are supported by homologous, functionally conserved brain areas or by different neurobiological substrates. Here we use functional magnetic resonance imaging in Rhesus macaques and humans to examine the brain regions involved in processing the ordering relationships between auditory nonsense words in rule-based sequences. We find that key regions in the human ventral frontal and opercular cortex have functional counterparts in the monkey brain. These regions are also known to be associated with initial stages of human syntactic processing. This study raises the possibility that certain ventral frontal neural systems, which play a significant role in language function in modern humans, originally evolved to support domain-general abilities involved in sequence processing. PMID:26573340

  14. Amino acid sequence of bovine heart coupling factor 6.

    PubMed Central

    Fang, J K; Jacobs, J W; Kanner, B I; Racker, E; Bradshaw, R A

    1984-01-01

    The amino acid sequence of bovine heart mitochondrial coupling factor 6 (F6) has been determined by automated Edman degradation of the whole protein and derived peptides. Preparations based on heat precipitation and ethanol extraction showed allotypic variation at three positions while material further purified by HPLC yielded only one sequence that also differed by a Phe-Thr replacement at residue 62. The mature protein contains 76 amino acids with a calculated molecular weight of 9006 and a pI of approximately equal to 5, in good agreement with experimentally measured values. The charged amino acids are mainly clustered at the termini and in one section in the middle; these three polar segments are separated by two segments relatively rich in nonpolar residues. Chou-Fasman analysis suggests three stretches of alpha-helix coinciding (or within) the high-charge-density sequences with a single beta-turn at the first polar-nonpolar junction. Comparison of the F6 sequence with those of other proteins did not reveal any homologous structures. PMID:6149548

  15. Genome-wide identification of conserved regulatory function in diverged sequences

    PubMed Central

    Taher, Leila; McGaughey, David M.; Maragh, Samantha; Aneas, Ivy; Bessling, Seneca L.; Miller, Webb; Nobrega, Marcelo A.; McCallion, Andrew S.; Ovcharenko, Ivan

    2011-01-01

    Plasticity of gene regulatory encryption can permit DNA sequence divergence without loss of function. Functional information is preserved through conservation of the composition of transcription factor binding sites (TFBS) in a regulatory element. We have developed a method that can accurately identify pairs of functional noncoding orthologs at evolutionarily diverged loci by searching for conserved TFBS arrangements. With an estimated 5% false-positive rate (FPR) in approximately 3000 human and zebrafish syntenic loci, we detected approximately 300 pairs of diverged elements that are likely to share common ancestry and have similar regulatory activity. By analyzing a pool of experimentally validated human enhancers, we demonstrated that 7/8 (88%) of their predicted functional orthologs retained in vivo regulatory control. Moreover, in 5/7 (71%) of assayed enhancer pairs, we observed concordant expression patterns. We argue that TFBS composition is often necessary to retain and sufficient to predict regulatory function in the absence of overt sequence conservation, revealing an entire class of functionally conserved, evolutionarily diverged regulatory elements that we term “covert.” PMID:21628450

  16. Conservation of uORF repressiveness and sequence features in mouse, human and zebrafish

    PubMed Central

    Chew, Guo-Liang; Pauli, Andrea; Schier, Alexander F.

    2016-01-01

    Upstream open reading frames (uORFs) are ubiquitous repressive genetic elements in vertebrate mRNAs. While much is known about the regulation of individual genes by their uORFs, the range of uORF-mediated translational repression in vertebrate genomes is largely unexplored. Moreover, it is unclear whether the repressive effects of uORFs are conserved across species. To address these questions, we analyse transcript sequences and ribosome profiling data from human, mouse and zebrafish. We find that uORFs are depleted near coding sequences (CDSes) and have initiation contexts that diminish their translation. Linear modelling reveals that sequence features at both uORFs and CDSes modulate the translation of CDSes. Moreover, the ratio of translation over 5′ leaders and CDSes is conserved between human and mouse, and correlates with the number of uORFs. These observations suggest that the prevalence of vertebrate uORFs may be explained by their conserved role in repressing CDS translation. PMID:27216465

  17. Conservation of uORF repressiveness and sequence features in mouse, human and zebrafish.

    PubMed

    Chew, Guo-Liang; Pauli, Andrea; Schier, Alexander F

    2016-01-01

    Upstream open reading frames (uORFs) are ubiquitous repressive genetic elements in vertebrate mRNAs. While much is known about the regulation of individual genes by their uORFs, the range of uORF-mediated translational repression in vertebrate genomes is largely unexplored. Moreover, it is unclear whether the repressive effects of uORFs are conserved across species. To address these questions, we analyse transcript sequences and ribosome profiling data from human, mouse and zebrafish. We find that uORFs are depleted near coding sequences (CDSes) and have initiation contexts that diminish their translation. Linear modelling reveals that sequence features at both uORFs and CDSes modulate the translation of CDSes. Moreover, the ratio of translation over 5' leaders and CDSes is conserved between human and mouse, and correlates with the number of uORFs. These observations suggest that the prevalence of vertebrate uORFs may be explained by their conserved role in repressing CDS translation. PMID:27216465

  18. [Partial sequence homology of FtsZ in phylogenetics analysis of lactic acid bacteria].

    PubMed

    Zhang, Bin; Dong, Xiu-zhu

    2005-10-01

    FtsZ is a structurally conserved protein, which is universal among the prokaryotes. It plays a key role in prokaryote cell division. A partial fragment of the ftsZ gene about 800bp in length was amplified and sequenced and a partial FtsZ protein phylogenetic tree for the lactic acid bacteria was constructed. By comparing the FtsZ phylogenetic tree with the 16S rDNA tree, it was shown that the two trees were similar in topology. Both trees revealed that Pediococcus spp. were closely related with L. casei group of Lactobacillus spp. , but less related with other lactic acid cocci such as Enterococcus and Streptococcus. The results also showed that the discriminative power of FtsZ was higher than that of 16S rDNA for either inter-species or inter-genus and could be a very useful tool in species identification of lactic acid bacteria. PMID:16342751

  19. Molecular cloning and sequence analysis of expansins--a highly conserved, multigene family of proteins that mediate cell wall extension in plants.

    PubMed Central

    Shcherban, T Y; Shi, J; Durachko, D M; Guiltinan, M J; McQueen-Mason, S J; Shieh, M; Cosgrove, D J

    1995-01-01

    Expansins are unusual proteins discovered by virtue of their ability to mediate cell wall extension in plants. We identified cDNA clones for two cucumber expansins on the basis of peptide sequences of proteins purified from cucumber hypocotyls. The expansin cDNAs encode related proteins with signal peptides predicted to direct protein secretion to the cell wall. Northern blot analysis showed moderate transcript abundance in the growing region of the hypocotyl and no detectable transcripts in the nongrowing region. Rice and Arabidopsis expansin cDNAs were identified from collections of anonymous cDNAs (expressed sequence tags). Sequence comparisons indicate at least four distinct expansin cDNAs in rice and at least six in Arabidopsis. Expansins are highly conserved in size and sequence (60-87% amino acid sequence identity and 75-95% similarity between any pairwise comparison), and phylogenetic trees indicate that this multigene family formed before the evolutionary divergence of monocotyledons and dicotyledons. Sequence and motif analyses show no similarities to known functional domains that might account for expansin action on wall extension. A series of highly conserved tryptophans may function in expansin binding to cellulose or other glycans. The high conservation of this multigene family indicates that the mechanism by which expansins promote wall extensin tolerates little variation in protein structure. Images Fig. 2 PMID:7568110

  20. Sequences Of Amino Acids For Human Serum Albumin

    NASA Technical Reports Server (NTRS)

    Carter, Daniel C.

    1992-01-01

    Sequences of amino acids defined for use in making polypeptides one-third to one-sixth as large as parent human serum albumin molecule. Smaller, chemically stable peptides have diverse applications including service as artificial human serum and as active components of biosensors and chromatographic matrices. In applications involving production of artificial sera from new sequences, little or no concern about viral contaminants. Smaller genetically engineered polypeptides more easily expressed and produced in large quantities, making commercial isolation and production more feasible and profitable.

  1. Sequence conservation and functional constraint on intergenic spacers in reduced genomes of the obligate symbiont Buchnera.

    PubMed

    Degnan, Patrick H; Ochman, Howard; Moran, Nancy A

    2011-09-01

    Analyses of genome reduction in obligate bacterial symbionts typically focus on the removal and retention of protein-coding regions, which are subject to ongoing inactivation and deletion. However, these same forces operate on intergenic spacers (IGSs) and affect their contents, maintenance, and rates of evolution. IGSs comprise both non-coding, non-functional regions, including decaying pseudogenes at varying stages of recognizability, as well as functional elements, such as genes for sRNAs and regulatory control elements. The genomes of Buchnera and other small genome symbionts display biased nucleotide compositions and high rates of sequence evolution and contain few recognizable regulatory elements. However, IGS lengths are highly correlated across divergent Buchnera genomes, suggesting the presence of functional elements. To identify functional regions within the IGSs, we sequenced two Buchnera genomes (from aphid species Uroleucon ambrosiae and Acyrthosiphon kondoi) and applied a phylogenetic footprinting approach to alignments of orthologous IGSs from a total of eight Buchnera genomes corresponding to six aphid species. Inclusion of these new genomes allowed comparative analyses at intermediate levels of divergence, enabling the detection of both conserved elements and previously unrecognized pseudogenes. Analyses of these genomes revealed that 232 of 336 IGS alignments over 50 nucleotides in length displayed substantial sequence conservation. Conserved alignment blocks within these IGSs encompassed 88 Shine-Dalgarno sequences, 55 transcriptional terminators, 5 Sigma-32 binding sites, and 12 novel small RNAs. Although pseudogene formation, and thus IGS formation, are ongoing processes in these genomes, a large proportion of intergenic spacers contain functional sequences. PMID:21912528

  2. Nanopores and nucleic acids: prospects for ultrarapid sequencing

    NASA Technical Reports Server (NTRS)

    Deamer, D. W.; Akeson, M.

    2000-01-01

    DNA and RNA molecules can be detected as they are driven through a nanopore by an applied electric field at rates ranging from several hundred microseconds to a few milliseconds per molecule. The nanopore can rapidly discriminate between pyrimidine and purine segments along a single-stranded nucleic acid molecule. Nanopore detection and characterization of single molecules represents a new method for directly reading information encoded in linear polymers. If single-nucleotide resolution can be achieved, it is possible that nucleic acid sequences can be determined at rates exceeding a thousand bases per second.

  3. Variation in conserved non-coding sequences on chromosome 5q andsusceptibility to asthma and atopy

    SciTech Connect

    Donfack, Joseph; Schneider, Daniel H.; Tan, Zheng; Kurz,Thorsten; Dubchak, Inna; Frazer, Kelly A.; Ober, Carole

    2005-09-10

    Background: Evolutionarily conserved sequences likely havebiological function. Methods: To determine whether variation in conservedsequences in non-coding DNA contributes to risk for human disease, westudied six conserved non-coding elements in the Th2 cytokine cluster onhuman chromosome 5q31 in a large Hutterite pedigree and in samples ofoutbred European American and African American asthma cases and controls.Results: Among six conserved non-coding elements (>100 bp,>70percent identity; human-mouse comparison), we identified one singlenucleotide polymorphism (SNP) in each of two conserved elements and sixSNPs in the flanking regions of three conserved elements. We genotypedour samples for four of these SNPs and an additional three SNPs each inthe IL13 and IL4 genes. While there was only modest evidence forassociation with single SNPs in the Hutterite and European Americansamples (P<0.05), there were highly significant associations inEuropean Americans between asthma and haplotypes comprised of SNPs in theIL4 gene (P<0.001), including a SNP in a conserved non-codingelement. Furthermore, variation in the IL13 gene was strongly associatedwith total IgE (P = 0.00022) and allergic sensitization to mold allergens(P = 0.00076) in the Hutterites, and more modestly associated withsensitization to molds in the European Americans and African Americans (P<0.01). Conclusion: These results indicate that there is overalllittle variation in the conserved non-coding elements on 5q31, butvariation in IL4 and IL13, including possibly one SNP in a conservedelement, influence asthma and atopic phenotypes in diversepopulations.

  4. Amino acid sequence of the Amur tiger prion protein.

    PubMed

    Wu, Changde; Pang, Wanyong; Zhao, Deming

    2006-10-01

    Prion diseases are fatal neurodegenerative disorders in human and animal associated with conformational conversion of a cellular prion protein (PrP(C)) into the pathologic isoform (PrP(Sc)). Various data indicate that the polymorphisms within the open reading frame (ORF) of PrP are associated with the susceptibility and control the species barrier in prion diseases. In the present study, partial Prnp from 25 Amur tigers (tPrnp) were cloned and screened for polymorphisms. Four single nucleotide polymorphisms (T423C, A501G, C511A, A610G) were found; the C511A and A610G nucleotide substitutions resulted in the amino acid changes Lysine171Glutamine and Alanine204Threoine, respectively. The tPrnp amino acid sequence is similar to house cat (Felis catus ) and sheep, but differs significantly from other two cat Prnp sequences that were previously deposited in GenBank. PMID:16780982

  5. Highly conserved D-loop-like nuclear mitochondrial sequences (Numts) in tiger (Panthera tigris).

    PubMed

    Zhang, Wenping; Zhang, Zhihe; Shen, Fujun; Hou, Rong; Lv, Xiaoping; Yue, Bisong

    2006-08-01

    Using oligonucleotide primers designed to match hypervariable segments I (HVS-1) of Panthera tigris mitochondrial DNA (mtDNA), we amplified two different PCR products (500 bp and 287 bp) in the tiger (Panthera tigris), but got only one PCR product (287 bp) in the leopard (Panthera pardus). Sequence analyses indicated that the sequence of 287 bp was a D-loop-like nuclear mitochondrial sequence (Numts), indicating a nuclear transfer that occurred approximately 4.8-17 million years ago in the tiger and 4.6-16 million years ago in the leopard. Although the mtDNA D-loop sequence has a rapid rate of evolution, the 287-bp Numts are highly conserved; they are nearly identical in tiger subspecies and only 1.742% different between tiger and leopard. Thus, such sequences represent molecular 'fossils' that can shed light on evolution of the mitochondrial genome and may be the most appropriate outgroup for phylogenetic analysis. This is also proved by comparing the phylogenetic trees reconstructed using the D-loop sequence of snow leopard and the 287-bp Numts as outgroup. PMID:17072079

  6. cDNA cloning and sequencing of human fibrillarin, a conserved nucleolar protein recognized by autoimmune antisera

    SciTech Connect

    Aris, J.P.; Blobel, G. )

    1991-02-01

    The authors have isolated a 1.1-kilobase cDNA clone that encodes human fibrillarin by screening a hepatoma library in parallel with DNA probes derived from the fibrillarin genes of Saccharomyces cerevisiae (NOP1) and Xenopus laevis. RNA blot analysis indicates that the corresponding mRNA is {approximately}1,300 nucleotides in length. Human fibrillarin expressed in vitro migrates on SDS gels as a 36-kDa protein that is specifically immunoprecipitated by antisera from humans with scleroderma autoimmune disease. Human fibrillarin contains an amino-terminal repetitive domain {approximately}75-80 amino acids in length that is rich in glycine and arginine residues and is similar to amino-terminal domains in the yeast and Xenopus fibrillarins. The occurrence of a putative RNA-binding domain and an RNP consensus sequence within the protein is consistent with the association of fibrillarin with small nucleolar RNAs. Protein sequence alignments show that 67% of amino acids from human fibrillarin are identical to those in yeast fibrillarin and that 81% are identical to those in Xenopus fibrillarin. This identity suggests the evolutionary conservation of an important function early in the pathway for ribosome biosynthesis.

  7. Reptiles and Mammals Have Differentially Retained Long Conserved Noncoding Sequences from the Amniote Ancestor

    PubMed Central

    Janes, D.E.; Chapus, C.; Gondo, Y.; Clayton, D.F.; Sinha, S.; Blatti, C.A.; Organ, C.L.; Fujita, M.K.; Balakrishnan, C.N.; Edwards, S.V.

    2011-01-01

    Many noncoding regions of genomes appear to be essential to genome function. Conservation of large numbers of noncoding sequences has been reported repeatedly among mammals but not thus far among birds and reptiles. By searching genomes of chicken (Gallus gallus), zebra finch (Taeniopygia guttata), and green anole (Anolis carolinensis), we quantified the conservation among birds and reptiles and across amniotes of long, conserved noncoding sequences (LCNS), which we define as sequences ≥500 bp in length and exhibiting ≥95% similarity between species. We found 4,294 LCNS shared between chicken and zebra finch and 574 LCNS shared by the two birds and Anolis. The percent of genomes comprised by LCNS in the two birds (0.0024%) is notably higher than the percent in mammals (<0.0003% to <0.001%), differences that we show may be explained in part by differences in genome-wide substitution rates. We reconstruct a large number of LCNS for the amniote ancestor (ca. 8,630) and hypothesize differential loss and substantial turnover of these sites in descendent lineages. By contrast, we estimated a small role for recruitment of LCNS via acquisition of novel functions over time. Across amniotes, LCNS are significantly enriched with transcription factor binding sites for many developmental genes, and 2.9% of LCNS shared between the two birds show evidence of expression in brain expressed sequence tag databases. These results show that the rate of retention of LCNS from the amniote ancestor differs between mammals and Reptilia (including birds) and that this may reflect differing roles and constraints in gene regulation. PMID:21183607

  8. Grouping of amino acids and recognition of protein structurally conserved regions by reduced alphabets of amino acids.

    PubMed

    Li, Jing; Wang, Wei

    2007-06-01

    Sequence alignment is a common method for finding protein structurally conserved/similar regions. However, sequence alignment is often not accurate if sequence identities between to-be-aligned sequences are less than 30%. This is because that for these sequences, different residues may play similar structural roles and they are incorrectly aligned during the sequence alignment using substitution matrix consisting of 20 types of residues. Based on the similarity of physicochemical features, residues can be clustered into a few groups. Using such simplified alphabets, the complexity of protein sequences is reduced and at the same time the key information encoded in the sequences remains. As a result, the accuracy of sequence alignment might be improved if the residues are properly clustered. Here, by using a database of aligned protein structures (DAPS), a new clustering method based on the substitution scores is proposed for the grouping of residues, and substitution matrices of residues at different levels of simplification are constructed. The validity of the reduced alphabets is confirmed by relative entropy analysis. The reduced alphabets are applied to recognition of protein structurally conserved/similar regions by sequence alignment. The results indicate that the accuracy or efficiency of sequence alignment can be improved with the optimal reduced alphabet with N around 9. PMID:17609897

  9. Conserved nucleotide sequences in the open reading frame and 3' untranslated region of selenoprotein P mRNA.

    PubMed Central

    Hill, K E; Lloyd, R S; Burk, R F

    1993-01-01

    Rat liver selenoprotein P contains 10 selenocysteine residues in its primary structure (deduced). It is the only selenoprotein characterized to date that has more than one selenocysteine residue. Selenoprotein P cDNA has been cloned from human liver and heart cDNA libraries and sequenced. The open reading frames are identical and contain a signal peptide, indicating that the protein is secreted by both organs and is therefore not exclusively produced in the liver. Ten selenocysteine residues (deduced) are present. Comparison of the open reading frame of the human cDNA with the rat cDNA reveals a 69% identity of the nucleotide sequence and 72% identity of the deduced amino acid sequence. Two regions in the 3' untranslated portion have high conservation between human and rat. Each of these regions contains a predicted stable stem-loop structure similar to the single stem-loop structures reported in 3' untranslated regions of type I iodothyronine 5'-deiodinase and glutathione peroxidase. The stem-loop structure of type I iodothyronine 5'-deiodinase has been shown to be necessary for incorporation of the selenocysteine residue at the UGA codon. Because only two stem-loop structures are present in the 3' untranslated region of selenoprotein P mRNA, it can be concluded that a separate stem-loop structure is not required for each selenocysteine residue. Images PMID:8421687

  10. Polyclonal antibody against conserved sequences of mce1A protein blocks MTB infection in macrophages.

    PubMed

    Sivagnanam, Sasikala; Namasivayam, Nalini; Chellam, Rajamanickam

    2012-03-01

    The pathogenesis of Mycobacterium tuberculosis is largely due to its ability to enter and survive within human macrophages. It is suggested that a specific protein namely mammalian cell entry protein is involved in the pathogenesis and the specific gene for this protein mce1A has been identified in several pathogenic organisms such as Rickettsia, Shigella, Escherichia coli, Helicobacter, Streptomyces, Klebsiella, Vibrio, Neisseria, Rhodococcus, Nocardioides, Saccharopolyspora erthyrae, and Pseudomonas. Analysis of mce1 operons in the above mentioned organisms through bioinformatics tools has revealed the presence of unique sequences (conserved regions) suggesting that these sequences may be involved in the process of infection. Presently, the mce1A full-length (1,365 bp) region from Mycobacterium bovis and its conserved regions (303 bp) were cloned in to an expression vector and the purified expressed proteins of molecular weight ~47 and ~11 kDa, respectively, were injected to rabbits to raise the polyclonal antibodies. The purified polyclonal antibodies were checked for their ability to inhibit the Mycobacterium infection in cultured human macrophages. In macrophage invasion assay, when antibody added at high concentration, decrease in viable counts was observed in all cell cultures within the first 5 days after infection, where the intracellular bacterial CFU obtained from the infected MTB increased by the 3rd day at low concentration of antibody. The macrophage invasion assay has indicated that the purified antibodies of mce1A conserved region can inhibit the infection of Mycobacterium. PMID:22159737

  11. The Internally Self-fertilizing Hermaphroditic Teleost Rivulus marmoratus (Cyprinodontiformes, Rivulidae) beta-Actin Gene: Amplification and Sequence Analysis with Conserved Primers.

    PubMed

    Lee

    2000-03-01

    To determine the ease and feasibility of amplifying the beta-actin gene in fish by the polymerase chain reaction (PCR), genomic DNAs of several fish (Rivulus, Southern top mouth minnow, common fat minnow, oily bitterling, carp, Far Eastern catfish, medaka, and European flounder) were extracted and used as a template with conserved primers, designed on the basis of high amino acid homology (approximately 98% or more). Among them, the self-fertilizing hermaphroditic fish Rivulus marmoratus was chosen for further characterization. After amplification of the Rivulus beta-actin PCR product with Taq polymerase, PCR product was subcloned to pCRII vector. After restriction enzyme mapping of Rivulus beta-actin gene, the amplified insert was sequenced using ALF Express automatic DNA sequencer with conserved internal primers. The R. marmoratus beta-actin gene consists of 1763 bp encoding 375 amino acids including 5 exons and 4 introns. The splicing and acceptance sites of the exon and intron boundaries of the Rivulus beta-actin gene were highly conserved with consensus sequences (GT/AG). The amino acid homology of R. marmoratus beta-actin to other species was high: 98.93% to human; 98.93%, Atlantic salmon; 98.93%, common carp; 98.93%, grass carp; 98.93%, zebrafish; 98.67%, medaka; and 98.40%, sea bream. To determine the expression of the R. marmoratus beta-actin gene in liver and ovary, reverse transcriptase-polymerase chain reaction was carried out with internal primers. In conclusion, these universal primers are successful in the rapid cloning of the fish beta-actin gene by PCR, based on a high homology of the beta-actin gene conserved through evolution. This approach will be applicable to the isolation of other beta-actin homologues in the investigation of phylogenetic comparisons of fish species, along with a possible application to cloning strategy in other conserved genes. PMID:10811955

  12. Quantum-Sequencing: Biophysics of quantum tunneling through nucleic acids

    NASA Astrophysics Data System (ADS)

    Casamada Ribot, Josep; Chatterjee, Anushree; Nagpal, Prashant

    2014-03-01

    Tunneling microscopy and spectroscopy has extensively been used in physical surface sciences to study quantum tunneling to measure electronic local density of states of nanomaterials and to characterize adsorbed species. Quantum-Sequencing (Q-Seq) is a new method based on tunneling microscopy for electronic sequencing of single molecule of nucleic acids. A major goal of third-generation sequencing technologies is to develop a fast, reliable, enzyme-free single-molecule sequencing method. Here, we present the unique ``electronic fingerprints'' for all nucleotides on DNA and RNA using Q-Seq along their intrinsic biophysical parameters. We have analyzed tunneling spectra for the nucleotides at different pH conditions and analyzed the HOMO, LUMO and energy gap for all of them. In addition we show a number of biophysical parameters to further characterize all nucleobases (electron and hole transition voltage and energy barriers). These results highlight the robustness of Q-Seq as a technique for next-generation sequencing.

  13. Sequence conservation in the Ancylostoma secreted protein-2 of Necator americanus (Na-ASP-2) from hookworm infected individuals in Thailand.

    PubMed

    Ungcharoensuk, Charoenchai; Putaporntip, Chaturong; Pattanawong, Urassaya; Jongwutiwes, Somchai

    2012-12-01

    The Ancylostoma secreted protein-2 of Necator americanus (Na-ASP-2) was one of the promising vaccine candidates against the most prevalent human hookworm species as adverse vaccine reaction has compromised further human vaccine trials. To elucidate the gene structure and the extent of sequence diversity, we determined the complete nucleotide sequence of the Na-asp-2 gene of individual larvae from 32 infected subjects living in 3 different endemic areas of Thailand. Sequence analysis revealed that the gene encoding Na-ASP-2 comprised 8 exons. Of 3 nucleotide substitutions in these exons, only one causes an amino acid change from leucine to methionine. A consensus conserved GT and AG at the 5' and the 3' boundaries of each intron was observed akin to those found in other eukaryotic genes. Introns of Na-asp-2 contained 23 nucleotide substitutions and 0-18 indels. The mean number of nucleotide substitutions per site (d) in introns was not significantly different from the mean number of synonymous substitutions per synonymous site (d(S)) in exons whereas d in introns was significantly exceeded d(N) (the mean number of nonsynonymous substitutions per nonsynonymous site) in exons (p<0.05), suggesting that introns and synonymous sites in exons may evolve at a similar rate whereas functional constraints at the amino acid could limit amino acid substitutions in Na-ASP-2. A recombination site was identified in an intron near the 3' portion of the gene. The positions of introns and the intron phases in the Na-asp-2 gene comparing with those in other pathogenesis-related-1 proteins of Loa loa, Onchocerca volvulus, Heterodera glycines, Caenorhabditis elegans and human were relatively conserved, suggesting evolutionary conservation of these genes. Sequence conservation in Na-ASP-2 may not compromise further vaccine design if adverse vaccine effects could be resolved whereas microheterogeneity in introns of this locus may be useful for population genetics analysis of N. americanus

  14. Cytochrome Oxidase I (COI) sequence conservation and variation patterns in the yellowfin and longtail tunas.

    PubMed

    Kunal, Swaraj Priyaranjan; Kumar, Girish

    2013-01-01

    Tunas are commercially important fishery worldwide. There are at least 13 species of tuna belonging to three genera, out of which genus Thunnus has maximum eight species. On the basis of their availability, they can be characterised as oceanic such as Thunnus albacares (yellowfin tuna) or coastal such as Thunnus tonggol (longtail tuna). Although these two are different species, morphological differentiation can only be seen in mature individuals, hence misidentification may result in erroneous data set, which ultimately affect conservation strategies. The mitochondrial DNA cytochrome oxidase c subunit 1 (COI) gene is one of the most popular markers for population genetic and phylogeographic studies across the animal kingdom. The present study aims to study the sequence conservation and variation in mitochondrial Cytochrome Oxidase I (COI) between these two species of tuna. COI sequence analysis of yellowfin and longtail revealed the close relationship between them in Thunnus genera. The present study is the first direct comparison of mitochondrial COI sequences of these two tuna species. PMID:23649742

  15. Lack of evidence of conserved lentiviral sequences in pigs with post weaning multisystemic wasting syndrome.

    PubMed Central

    Bratanich, A; Lairmore, M; Heneine, W; Konoby, C; Harding, J; West, K; Vasquez, G; Allan, G; Ellis, J

    1999-01-01

    In order to investigate the role of retroviruses in the recently described porcine postweaning multisystemic wasting syndrome (PMWS) serum and leukocytes were screened for reverse transcriptase (RT) activity, and tissues were examined for the presence of conserved lentiviral sequences using degenerate primers in a polymerase chain reaction (PCR). Serum and stimulated leukocytes from the blood and lymph nodes from pigs with PMWS, as well as from control pigs had RT activity that was detected by the sensitive Amp-RT assay. A 257-bp fragment was amplified from DNA from the blood and bone marrow of pigs with PMWS. This fragment was identical in size to conserved lentiviral sequences that were amplified from plasmids containing DNA from several lentiviruses. Cloning and sequencing of the fragment from affected pigs, however, did not reveal homology with the recognized lentiviruses. Together the results of these analyses suggest that the RT activity present in tissues from control and affected pigs is the result of endogenous retrovirus expression, and that a lentivirus is not a primary pathogen in PMWS. Images Figure 1. Figure 2. PMID:10480463

  16. Comparative sequence analysis suggests a conserved gating mechanism for TRP channels

    PubMed Central

    Palovcak, Eugene; Delemotte, Lucie; Klein, Michael L.

    2015-01-01

    The transient receptor potential (TRP) channel superfamily plays a central role in transducing diverse sensory stimuli in eukaryotes. Although dissimilar in sequence and domain organization, all known TRP channels act as polymodal cellular sensors and form tetrameric assemblies similar to those of their distant relatives, the voltage-gated potassium (Kv) channels. Here, we investigated the related questions of whether the allosteric mechanism underlying polymodal gating is common to all TRP channels, and how this mechanism differs from that underpinning Kv channel voltage sensitivity. To provide insight into these questions, we performed comparative sequence analysis on large, comprehensive ensembles of TRP and Kv channel sequences, contextualizing the patterns of conservation and correlation observed in the TRP channel sequences in light of the well-studied Kv channels. We report sequence features that are specific to TRP channels and, based on insight from recent TRPV1 structures, we suggest a model of TRP channel gating that differs substantially from the one mediating voltage sensitivity in Kv channels. The common mechanism underlying polymodal gating involves the displacement of a defect in the H-bond network of S6 that changes the orientation of the pore-lining residues at the hydrophobic gate. PMID:26078053

  17. Correlation between fibroin amino acid sequence and physical silk properties.

    PubMed

    Fedic, Robert; Zurovec, Michal; Sehnal, Frantisek

    2003-09-12

    The fiber properties of lepidopteran silk depend on the amino acid repeats that interact during H-fibroin polymerization. The aim of our research was to relate repeat composition to insect biology and fiber strength. Representative regions of the H-fibroin genes were sequenced and analyzed in three pyralid species: wax moth (Galleria mellonella), European flour moth (Ephestia kuehniella), and Indian meal moth (Plodia interpunctella). The amino acid repeats are species-specific, evidently a diversification of an ancestral region of 43 residues, and include three types of regularly dispersed motifs: modifications of GSSAASAA sequence, stretches of tripeptides GXZ where X and Z represent bulky residues, and sequences similar to PVIVIEE. No concatenations of GX dipeptide or alanine, which are typical for Bombyx silkworms and Antheraea silk moths, respectively, were found. Despite different repeat structure, the silks of G. mellonella and E. kuehniella exhibit similar tensile strength as the Bombyx and Antheraea silks. We suggest that in these latter two species, variations in the repeat length obstruct repeat alignment, but sufficiently long stretches of iterated residues get superposed to interact. In the pyralid H-fibroins, interactions of the widely separated and diverse motifs depend on the precision of repeat matching; silk is strong in G. mellonella and E. kuehniella, with 2-3 types of long homogeneous repeats, and nearly 10 times weaker in P. interpunctella, with seven types of shorter erratic repeats. The high proportion of large amino acids in the H-fibroin of pyralids has probably evolved in connection with the spinning habit of caterpillars that live in protective silk tubes and spin continuously, enlarging the tubes on one end and partly devouring the other one. The silk serves as a depot of energetically rich and essential amino acids that may be scarce in the diet. PMID:12816957

  18. Conserved Noncoding Sequences Highlight Shared Components of Regulatory Networks in Dicotyledonous Plants[W

    PubMed Central

    Baxter, Laura; Jironkin, Aleksey; Hickman, Richard; Moore, Jay; Barrington, Christopher; Krusche, Peter; Dyer, Nigel P.; Buchanan-Wollaston, Vicky; Tiskin, Alexander; Beynon, Jim; Denby, Katherine; Ott, Sascha

    2012-01-01

    Conserved noncoding sequences (CNSs) in DNA are reliable pointers to regulatory elements controlling gene expression. Using a comparative genomics approach with four dicotyledonous plant species (Arabidopsis thaliana, papaya [Carica papaya], poplar [Populus trichocarpa], and grape [Vitis vinifera]), we detected hundreds of CNSs upstream of Arabidopsis genes. Distinct positioning, length, and enrichment for transcription factor binding sites suggest these CNSs play a functional role in transcriptional regulation. The enrichment of transcription factors within the set of genes associated with CNS is consistent with the hypothesis that together they form part of a conserved transcriptional network whose function is to regulate other transcription factors and control development. We identified a set of promoters where regulatory mechanisms are likely to be shared between the model organism Arabidopsis and other dicots, providing areas of focus for further research. PMID:23110901

  19. Amino acid sequence of the nonsecretory ribonuclease of human urine.

    PubMed

    Beintema, J J; Hofsteenge, J; Iwama, M; Morita, T; Ohgi, K; Irie, M; Sugiyama, R H; Schieven, G L; Dekker, C A; Glitz, D G

    1988-06-14

    The amino acid sequence of a nonsecretory ribonuclease isolated from human urine was determined except for the identity of the residue at position 7. Sequence information indicates that the ribonucleases of human liver and spleen and an eosinophil-derived neurotoxin are identical or very closely related gene products. The sequence is identical at about 30% of the amino acid positions with those of all of the secreted mammalian ribonucleases for which information is available. Identical residues include active-site residues histidine-12, histidine-119, and lysine-41, other residues known to be important for substrate binding and catalytic activity, and all eight half-cystine residues common to these enzymes. Major differences include a deletion of six residues in the (so-called) S-peptide loop, insertions of two, and nine residues, respectively, in three other external loops of the molecule, and an addition of three residues at the amino terminus. The sequence shows the human nonsecretory ribonuclease to belong to the same ribonuclease superfamily as the mammalian secretory ribonucleases, turtle pancreatic ribonuclease, and human angiogenin. Sequence data suggest that a gene duplication occurred in an ancient vertebrate ancestor; one branch led to the nonsecretory ribonuclease, while the other branch led to a second duplication, with one line leading to the secretory ribonucleases (in mammals) and the second line leading to pancreatic ribonuclease in turtle and an angiogenic factor in mammals (human angiogenin). The nonsecretory ribonuclease has five short carbohydrate chains attached via asparagine residues at the surface of the molecule; these chains may have been shortened by exoglycosidase action.(ABSTRACT TRUNCATED AT 250 WORDS) PMID:3166997

  20. A Collection of Conserved Noncoding Sequences to Study Gene Regulation in Flowering Plants1[OPEN

    PubMed Central

    2016-01-01

    Transcription factors (TFs) regulate gene expression by binding cis-regulatory elements, of which the identification remains an ongoing challenge owing to the prevalence of large numbers of nonfunctional TF binding sites. Powerful comparative genomics methods, such as phylogenetic footprinting, can be used for the detection of conserved noncoding sequences (CNSs), which are functionally constrained and can greatly help in reducing the number of false-positive elements. In this study, we applied a phylogenetic footprinting approach for the identification of CNSs in 10 dicot plants, yielding 1,032,291 CNSs associated with 243,187 genes. To annotate CNSs with TF binding sites, we made use of binding site information for 642 TFs originating from 35 TF families in Arabidopsis (Arabidopsis thaliana). In three species, the identified CNSs were evaluated using TF chromatin immunoprecipitation sequencing data, resulting in significant overlap for the majority of data sets. To identify ultraconserved CNSs, we included genomes of additional plant families and identified 715 binding sites for 501 genes conserved in dicots, monocots, mosses, and green algae. Additionally, we found that genes that are part of conserved mini-regulons have a higher coherence in their expression profile than other divergent gene pairs. All identified CNSs were integrated in the PLAZA 3.0 Dicots comparative genomics platform (http://bioinformatics.psb.ugent.be/plaza/versions/plaza_v3_dicots/) together with new functionalities facilitating the exploration of conserved cis-regulatory elements and their associated genes. The availability of this data set in a user-friendly platform enables the exploration of functional noncoding DNA to study gene regulation in a variety of plant species, including crops. PMID:27261064

  1. A Collection of Conserved Noncoding Sequences to Study Gene Regulation in Flowering Plants.

    PubMed

    Van de Velde, Jan; Van Bel, Michiel; Vaneechoutte, Dries; Vandepoele, Klaas

    2016-08-01

    Transcription factors (TFs) regulate gene expression by binding cis-regulatory elements, of which the identification remains an ongoing challenge owing to the prevalence of large numbers of nonfunctional TF binding sites. Powerful comparative genomics methods, such as phylogenetic footprinting, can be used for the detection of conserved noncoding sequences (CNSs), which are functionally constrained and can greatly help in reducing the number of false-positive elements. In this study, we applied a phylogenetic footprinting approach for the identification of CNSs in 10 dicot plants, yielding 1,032,291 CNSs associated with 243,187 genes. To annotate CNSs with TF binding sites, we made use of binding site information for 642 TFs originating from 35 TF families in Arabidopsis (Arabidopsis thaliana). In three species, the identified CNSs were evaluated using TF chromatin immunoprecipitation sequencing data, resulting in significant overlap for the majority of data sets. To identify ultraconserved CNSs, we included genomes of additional plant families and identified 715 binding sites for 501 genes conserved in dicots, monocots, mosses, and green algae. Additionally, we found that genes that are part of conserved mini-regulons have a higher coherence in their expression profile than other divergent gene pairs. All identified CNSs were integrated in the PLAZA 3.0 Dicots comparative genomics platform (http://bioinformatics.psb.ugent.be/plaza/versions/plaza_v3_dicots/) together with new functionalities facilitating the exploration of conserved cis-regulatory elements and their associated genes. The availability of this data set in a user-friendly platform enables the exploration of functional noncoding DNA to study gene regulation in a variety of plant species, including crops. PMID:27261064

  2. Characterization and amino acid sequence of a fatty acid-binding protein from human heart.

    PubMed

    Offner, G D; Brecher, P; Sawlivich, W B; Costello, C E; Troxler, R F

    1988-05-15

    The complete amino acid sequence of a fatty acid-binding protein from human heart was determined by automated Edman degradation of CNBr, BNPS-skatole [3'-bromo-3-methyl-2-(2-nitrobenzenesulphenyl)indolenine], hydroxylamine, Staphylococcus aureus V8 proteinase, tryptic and chymotryptic peptides, and by digestion of the protein with carboxypeptidase A. The sequence of the blocked N-terminal tryptic peptide from citraconylated protein was determined by collisionally induced decomposition mass spectrometry. The protein contains 132 amino acid residues, is enriched with respect to threonine and lysine, lacks cysteine, has an acetylated valine residue at the N-terminus, and has an Mr of 14768 and an isoelectric point of 5.25. This protein contains two short internal repeated sequences from residues 48-54 and from residues 114-119 located within regions of predicted beta-structure and decreasing hydrophobicity. These short repeats are contained within two longer repeated regions from residues 48-60 and residues 114-125, which display 62% sequence similarity. These regions could accommodate the charged and uncharged moieties of long-chain fatty acids and may represent fatty acid-binding domains consistent with the finding that human heart fatty acid-binding protein binds 2 mol of oleate or palmitate/mol of protein. Detailed evidence for the amino acid sequences of the peptides has been deposited as Supplementary Publication SUP 50143 (23 pages) at the British Library Lending Division, Boston Spa, Yorkshire LS23 7BQ, U.K., from whom copies may be obtained as indicated in Biochem. J. (1988) 249, 5. PMID:3421901

  3. Molecular cloning and amino acid sequence of human 5-lipoxygenase

    SciTech Connect

    Matsumoto, T.; Funk, C.D.; Radmark, O.; Hoeoeg, J.O.; Joernvall, H.; Samuelsson, B.

    1988-01-01

    5-Lipoxygenase (EC 1.13.11.34), a Ca/sup 2 +/- and ATP-requiring enzyme, catalyzes the first two steps in the biosynthesis of the peptidoleukotrienes and the chemotactic factor leukotriene B/sub 4/. A cDNA clone corresponding to 5-lipoxygenase was isolated from a human lung lambda gt11 expression library by immunoscreening with a polyclonal antibody. Additional clones from a human placenta lambda gt11 cDNA library were obtained by plaque hybridization with the /sup 32/P-labeled lung cDNA clone. Sequence data obtained from several overlapping clones indicate that the composite DNAs contain the complete coding region for the enzyme. From the deduced primary structure, 5-lipoxygenase encodes a 673 amino acid protein with a calculated molecular weight of 77,839. Direct analysis of the native protein and its proteolytic fragments confirmed the deduced composition, the amino-terminal amino acid sequence, and the structure of many internal segments. 5-Lipoxygenase has no apparent sequence homology with leukotriene A/sub 4/ hydrolase or Ca/sup 2 +/-binding proteins. RNA blot analysis indicated substantial amounts of an mRNA species of approx. = 2700 nucleotides in leukocytes, lung, and placenta.

  4. Nucleic acid sequence detection using multiplexed oligonucleotide PCR

    DOEpatents

    Nolan, John P.; White, P. Scott

    2006-12-26

    Methods for rapidly detecting single or multiple sequence alleles in a sample nucleic acid are described. Provided are all of the oligonucleotide pairs capable of annealing specifically to a target allele and discriminating among possible sequences thereof, and ligating to each other to form an oligonucleotide complex when a particular sequence feature is present (or, alternatively, absent) in the sample nucleic acid. The design of each oligonucleotide pair permits the subsequent high-level PCR amplification of a specific amplicon when the oligonucleotide complex is formed, but not when the oligonucleotide complex is not formed. The presence or absence of the specific amplicon is used to detect the allele. Detection of the specific amplicon may be achieved using a variety of methods well known in the art, including without limitation, oligonucleotide capture onto DNA chips or microarrays, oligonucleotide capture onto beads or microspheres, electrophoresis, and mass spectrometry. Various labels and address-capture tags may be employed in the amplicon detection step of multiplexed assays, as further described herein.

  5. The amino acid sequence of rabbit muscle triose phosphate isomerase.

    PubMed Central

    Corran, P H; Waley, S G

    1975-01-01

    The amino acid sequence of rabbit muscle triose phosphate isomerase was deduced by characterizing peptides that overlap the tryptic peptides. Thiol groups were modified by oxidation, carboxymethylation or aminoen. About 50 peptides that provided information about overlaps were isolated; the peptides were mostly characterized by their compositions and N-terminal residues. The peptide chains contain 248 amino acid residues, and no evidence for dissimilarity of the two subunits that comprise the native enzyme was found. The sequence of the rabbit muscle enzyme may be compared with that of the coelacanth enzyme (Kolb et al., 1974): 84% of the residues are in identical positions. Similarly, comparison of the sequence with that inferred for the chicken enzyme (Furth et al., 1974) shows that 87% of the residues are in identical positions. Limited though these comparisons are, they suggest that triose phosphate isomerase has one of the lowest rates of evolutionary change. An extended version of the present paper has been deposited as Supplementary Publication SUP 50040 (42 pages) at the British Library (Lending Division) (formerly the National Lending Library for Science and Technology), Boston Spa, Yorks. LS23 7BQ, U.K., from whom copies can be obtained on the terms given in Biochem. J. (1975) 145, 5. PMID:1171682

  6. New insights into SRY regulation through identification of 5' conserved sequences

    PubMed Central

    Ross, Diana GF; Bowles, Josephine; Koopman, Peter; Lehnert, Sigrid

    2008-01-01

    Background SRY is the pivotal gene initiating male sex determination in most mammals, but how its expression is regulated is still not understood. In this study we derived novel SRY 5' flanking genomic sequence data from bovine and caprine genomic BAC clones. Results We identified four intervals of high homology upstream of SRY by comparison of human, bovine, pig, goat and mouse genomic sequences. These conserved regions contain putative binding sites for a large number of known transcription factor families, including several that have been implicated previously in sex determination and early gonadal development. Conclusion Our results reveal potentially important SRY regulatory elements, mutations in which might underlie cases of idiopathic human XY sex reversal. PMID:18851760

  7. Vertebrate paralogous conserved noncoding sequences may be related to gene expressions in brain.

    PubMed

    Matsunami, Masatoshi; Saitou, Naruya

    2013-01-01

    Vertebrate genomes include gene regulatory elements in protein-noncoding regions. A part of gene regulatory elements are expected to be conserved according to their functional importance, so that evolutionarily conserved noncoding sequences (CNSs) might be good candidates for those elements. In addition, paralogous CNSs, which are highly conserved among both orthologous loci and paralogous loci, have the possibility of controlling overlapping expression patterns of their adjacent paralogous protein-coding genes. The two-round whole-genome duplications (2R WGDs), which most probably occurred in the vertebrate common ancestors, generated large numbers of paralogous protein-coding genes and their regulatory elements. These events could contribute to the emergence of vertebrate features. However, the evolutionary history and influences of the 2R WGDs are still unclear, especially in noncoding regions. To address this issue, we identified paralogous CNSs. Region-focused Basic Local Alignment Search Tool (BLAST) search of each synteny block revealed 7,924 orthologous CNSs and 309 paralogous CNSs conserved among eight high-quality vertebrate genomes. Paralogous CNSs we found contained 115 previously reported ones and newly detected 194 ones. Through comparisons with VISTA Enhancer Browser and available ChIP-seq data, one-third (103) of paralogous CNSs detected in this study showed gene regulatory activity in the brain at several developmental stages. Their genomic locations are highly enriched near the transcription factor-coding regions, which are expressed in brain and neural systems. These results suggest that paralogous CNSs are conserved mainly because of maintaining gene expression in the vertebrate brain. PMID:23267051

  8. Amino acid sequence prerequisites for the formation of cn ions.

    PubMed

    Downard, K M; Biemann, K

    1993-11-01

    Ammo acid sequence prerequisites are described for the formation of c, ions observed in high-energy collision-induced decomposition spectra of peptides. It is shown that the formation of cn ions is promoted by the nature of the amino acid C-terminal to the cleavage site. A propensity for cn cleavage preceding threonine, and to a lesser extent tryptophan, lysine, and serine, is demonstrated where fragmentation is directed N-terminally at these residues. In addition, the nature of the residue N-terminal to the cleavage site is shown to have little effect on cn ion formation. A mechanism for cn ion formation is proposed and its applicability to the results observed is discussed. PMID:24227531

  9. Ultrasensitive nucleic acid sequence detection by single-molecule electrophoresis

    SciTech Connect

    Castro, A; Shera, E.B.

    1996-09-01

    This is the final report of a one-year laboratory-directed research and development project at Los Alamos National Laboratory. There has been considerable interest in the development of very sensitive clinical diagnostic techniques over the last few years. Many pathogenic agents are often present in extremely small concentrations in clinical samples, especially at the initial stages of infection, making their detection very difficult. This project sought to develop a new technique for the detection and accurate quantification of specific bacterial and viral nucleic acid sequences in clinical samples. The scheme involved the use of novel hybridization probes for the detection of nucleic acids combined with our recently developed technique of single-molecule electrophoresis. This project is directly relevant to the DOE`s Defense Programs strategic directions in the area of biological warfare counter-proliferation.

  10. Sequence of Radiotherapy and Chemotherapy in Breast Cancer After Breast-Conserving Surgery

    SciTech Connect

    Jobsen, Jan J.; Palen, Job van der; Brinkhuis, Marieel; Ong, Francisca; Struikmans, Henk

    2012-04-01

    Purpose: The optimal sequence of radiotherapy and chemotherapy in breast-conserving therapy is unknown. Methods and Materials: From 1983 through 2007, a total of 641 patients with 653 instances of breast-conserving therapy (BCT), received both chemotherapy and radiotherapy and are the basis of this analysis. Patients were divided into three groups. Groups A and B comprised patients treated before 2005, Group A radiotherapy first and Group B chemotherapy first. Group C consisted of patients treated from 2005 onward, when we had a fixed sequence of radiotherapy first, followed by chemotherapy. Results: Local control did not show any differences among the three groups. For distant metastasis, no difference was shown between Groups A and B. Group C, when compared with Group A, showed, on univariate and multivariate analyses, a significantly better distant metastasis-free survival. The same was noted for disease-free survival. With respect to disease-specific survival, no differences were shown on multivariate analysis among the three groups. Conclusion: Radiotherapy, as an integral part of the primary treatment of BCT, should be administered first, followed by adjuvant chemotherapy.

  11. Conserved Noncoding Sequences Regulate lhx5 Expression in the Zebrafish Forebrain

    PubMed Central

    Sun, Liu; Chen, Fengjiao; Peng, Gang

    2015-01-01

    The LIM homeobox family protein Lhx5 plays important roles in forebrain development in the vertebrates. The lhx5 gene exhibits complex temporal and spatial expression patterns during early development but its transcriptional regulation mechanisms are not well understood. Here, we have used transgenesis in zebrafish in order to define regulatory elements that drive lhx5 expression in the forebrain. Through comparative genomic analysis we identified 10 non-coding sequences conserved in five teleost species. We next examined the enhancer activities of these conserved non-coding sequences with Tol2 transposon mediated transgenesis. We found a proximately located enhancer gave rise to robust reporter EGFP expression in the forebrain regions. In addition, we identified an enhancer located at approximately 50 kb upstream of lhx5 coding region that is responsible for reporter gene expression in the hypothalamus. We also identify an enhancer located approximately 40 kb upstream of the lhx5 coding region that is required for expression in the prethalamus (ventral thalamus). Together our results suggest discrete enhancer elements control lhx5 expression in different regions of the forebrain. PMID:26147098

  12. Genomes of sequence type 121 Listeria monocytogenes strains harbor highly conserved plasmids and prophages

    PubMed Central

    Schmitz-Esser, Stephan; Müller, Anneliese; Stessl, Beatrix; Wagner, Martin

    2015-01-01

    The food-borne pathogen Listeria (L.) monocytogenes is often found in food production environments. Thus, controlling the occurrence of L. monocytogenes in food production is a great challenge for food safety. Among a great diversity of L. monocytogenes strains from food production, particularly strains belonging to sequence type (ST)121 are prevalent. The molecular reasons for the abundance of ST121 strains are however currently unknown. We therefore determined the genome sequences of three L. monocytogenes ST121 strains: 6179 and 4423, which persisted for up to 8 years in food production plants in Ireland and Austria, and of the strain 3253 and compared them with available L. monocytogenes ST121 genomes. Our results show that the ST121 genomes are highly similar to each other and show a tremendously high degree of conservation among some of their prophages and particularly among their plasmids. This remarkably high level of conservation among prophages and plasmids suggests that strong selective pressure is acting on them. We thus hypothesize that plasmids and prophages are providing important adaptations for survival in food production environments. In addition, the ST121 genomes share common adaptations which might be related to their persistence in food production environments such as the presence of Tn6188, a transposon responsible for increased tolerance against quaternary ammonium compounds, a yet undescribed insertion harboring recombination hotspot (RHS) repeat proteins, which are most likely involved in competition against other bacteria, and presence of homologs of the L. innocua genes lin0464 and lin0465. PMID:25972859

  13. Conservation.

    ERIC Educational Resources Information Center

    National Audubon Society, New York, NY.

    This set of teaching aids consists of seven Audubon Nature Bulletins, providing the teacher and student with informational reading on various topics in conservation. The bulletins have these titles: Plants as Makers of Soil, Water Pollution Control, The Ground Water Table, Conservation--To Keep This Earth Habitable, Our Threatened Air Supply,…

  14. A search for small noncoding RNAs in Staphylococcus aureus reveals a conserved sequence motif for regulation

    PubMed Central

    Geissmann, Thomas; Chevalier, Clément; Cros, Marie-Josée; Boisset, Sandrine; Fechter, Pierre; Noirot, Céline; Schrenzel, Jacques; François, Patrice; Vandenesch, François; Gaspin, Christine; Romby, Pascale

    2009-01-01

    Bioinformatic analysis of the intergenic regions of Staphylococcus aureus predicted multiple regulatory regions. From this analysis, we characterized 11 novel noncoding RNAs (RsaA‐K) that are expressed in several S. aureus strains under different experimental conditions. Many of them accumulate in the late-exponential phase of growth. All ncRNAs are stable and their expression is Hfq-independent. The transcription of several of them is regulated by the alternative sigma B factor (RsaA, D and F) while the expression of RsaE is agrA-dependent. Six of these ncRNAs are specific to S. aureus, four are conserved in other Staphylococci, and RsaE is also present in Bacillaceae. Transcriptomic and proteomic analysis indicated that RsaE regulates the synthesis of proteins involved in various metabolic pathways. Phylogenetic analysis combined with RNA structure probing, searches for RsaE‐mRNA base pairing, and toeprinting assays indicate that a conserved and unpaired UCCC sequence motif of RsaE binds to target mRNAs and prevents the formation of the ribosomal initiation complex. This study unexpectedly shows that most of the novel ncRNAs carry the conserved C−rich motif, suggesting that they are members of a class of ncRNAs that target mRNAs by a shared mechanism. PMID:19786493

  15. cDNA sequence, genomic organization, and evolutionary conservation of a novel gene from the WAGR region

    SciTech Connect

    Schwartz, F.; Eisenman, R.; Knoll, J.; Bruns, G.

    1995-09-20

    A new gene (239FB) with predominant and differential expression in fetal brain has recently been isolated from a chromosome 11p13-p14 boundary area near FSHB. The corresponding mRNA has an open reading frame of 294 amino acids, a 3` untranslated region of 1247 nucleotides, and a highly GC-rich 5` untranslated region. The coding and 3` UT sequence is specified by 6 exons within nearly 87 kb of isolated genomic locus. The 5` end region of the transcript maps adjacent to the only genomically defined CpG island in a chromosomal subregion that may be associated with part of the mental retardation of some WAGR (Wilms tumor, aniridia, genitourinary anomalies, and mental retardation) syndrome patients. In addition to nucleotide and amino acid similarity to an EST from a normalized infant brain cDNA library, the predicted protein has extensive similarity to Caenorhbditis elegans polypeptides of, as yet, unknown function. The 239FB locus is, therefore, likely part of a family of genes with two members expressed in human brain. The extensive conservation of the predicted protein suggests a fundamental function of the gene product and will enable evaluation of the role of the 239FB gene in neurogenesis in model organisms. 48 refs., 4 figs., 1 tab.

  16. Transactivation specificity is conserved among p53 family proteins and depends on a response element sequence code

    PubMed Central

    Ciribilli, Yari; Monti, Paola; Bisio, Alessandra; Nguyen, H. Thien; Ethayathulla, Abdul S.; Ramos, Ana; Foggetti, Giorgia; Menichini, Paola; Menendez, Daniel; Resnick, Michael A.; Viadiu, Hector; Fronza, Gilberto; Inga, Alberto

    2013-01-01

    Structural and biochemical studies have demonstrated that p73, p63 and p53 recognize DNA with identical amino acids and similar binding affinity. Here, measuring transactivation activity for a large number of response elements (REs) in yeast and human cell lines, we show that p53 family proteins also have overlapping transactivation profiles. We identified mutations at conserved amino acids of loops L1 and L3 in the DNA-binding domain that tune the transactivation potential nearly equally in p73, p63 and p53. For example, the mutant S139F in p73 has higher transactivation potential towards selected REs, enhanced DNA-binding cooperativity in vitro and a flexible loop L1 as seen in the crystal structure of the protein–DNA complex. By studying, how variations in the RE sequence affect transactivation specificity, we discovered a RE-transactivation code that predicts enhanced transactivation; this correlation is stronger for promoters of genes associated with apoptosis. PMID:23892287

  17. Evolutionary conservation of sequence and secondary structures inCRISPR repeats

    SciTech Connect

    Kunin, Victor; Sorek, Rotem; Hugenholtz, Philip

    2006-09-01

    Clustered Regularly Interspaced Palindromic Repeats (CRISPRs) are a novel class of direct repeats, separated by unique spacer sequences of similar length, that are present in {approx}40% of bacterial and all archaeal genomes analyzed to date. More than 40 gene families, called CRISPR-associated sequences (CAS), appear in conjunction with these repeats and are thought to be involved in the propagation and functioning of CRISPRs. It has been proposed that the CRISPR/CAS system samples, maintains a record of, and inactivates invasive DNA that the cell has encountered, and therefore constitutes a prokaryotic analog of an immune system. Here we analyze CRISPR repeats identified in 195 microbial genomes and show that they can be organized into multiple clusters based on sequence similarity. All individual repeats in any given cluster were inferred to form characteristic RNA secondary structure, ranging from non-existent to pronounced. Stable secondary structures included G:U base pairs and exhibited multiple compensatory base changes in the stem region, indicating evolutionary conservation and functional importance. We also show that the repeat-based classification corresponds to, and expands upon, a previously reported CAS gene-based classification including specific relationships between CRISPR and CAS subtypes.

  18. Homology analyses of the protein sequences of fatty acid synthases from chicken liver, rat mammary gland, and yeast

    SciTech Connect

    Chang, Soo-Ik ); Hammes, G.G. )

    1989-11-01

    Homology analyses of the protein sequences of chicken liver and rat mammary gland fatty acid synthases were carried out. The amino acid sequences of the chicken and rat enzymes are 67% identical. If conservative substitutions are allowed, 78% of the amino acids are matched. A region of low homologies exists between the functional domains, in particular around amino acid residues 1059-1264 of the chicken enzyme. Homologies between the active sites of chicken and rat and of chicken and yeast enzymes have been analyzed by an alignment method. A high degree of homology exists between the active sites of the chicken and rat enzymes. However, the chicken and yeast enzymes show a lower degree of homology. The DADPH-binding dinucleotide folds of the {beta}-ketoacyl reductase and the enoyl reductase sites were identified by comparison with a known consensus sequence for the DADP- and FAD-binding dinucleotide folds. The active sites of all of the enzymes are primarily in hydrophobic regions of the protein. This study suggests that the genes for the functional domains of fatty acid synthase were originally separated, and these genes were connected to each other by using different connecting nucleotide sequences in different species. An alternative explanation for the differences in rat and chicken is a common ancestry and mutations in the joining regions during evolution.

  19. Characterization of the Role of a Highly Conserved Sequence in ATP Binding Cassette Transporter G (ABCG) Family in ABCG1 Stability, Oligomerization, and Trafficking

    PubMed Central

    2013-01-01

    ATP-binding cassette transporter G1 (ABCG1) mediates cholesterol and oxysterol efflux onto lipidated lipoproteins and plays an important role in macrophage reverse cholesterol transport. Here, we identified a highly conserved sequence present in the five ABCG transporter family members. The conserved sequence is located between the nucleotide binding domain and the transmembrane domain and contains five amino acid residues from Asn at position 316 to Phe at position 320 in ABCG1 (NPADF). We found that cells expressing mutant ABCG1, in which Asn316, Pro317, Asp319, and Phe320 in the conserved sequence were replaced with Ala simultaneously, showed impaired cholesterol efflux activity compared with wild type ABCG1-expressing cells. A more detailed mutagenesis study revealed that mutation of Asn316 or Phe 320 to Ala significantly reduced cellular cholesterol and 7-ketocholesterol efflux conferred by ABCG1, whereas replacement of Pro317 or Asp319 with Ala had no detectable effect. To confirm the important role of Asn316 and Phe320, we mutated Asn316 to Asp (N316D) and Gln (N316Q), and Phe320 to Ile (F320I) and Tyr (F320Y). The mutant F320Y showed the same phenotype as wild type ABCG1. However, the efflux of cholesterol and 7-ketocholesterol was reduced in cells expressing ABCG1 mutant N316D, N316Q, or F320I compared with wild type ABCG1. Further, mutations N316Q and F320I impaired ABCG1 trafficking while having no marked effect on the stability and oligomerization of ABCG1. The mutant N316Q and F320I could not be transported to the cell surface efficiently. Instead, the mutant proteins were mainly localized intracellularly. Thus, these findings indicate that the two highly conserved amino acid residues, Asn and Phe, play an important role in ABCG1-dependent export of cellular cholesterol, mainly through the regulation of ABCG1 trafficking. PMID:24320932

  20. 37 CFR 1.822 - Symbols and format to be used for nucleotide and/or amino acid sequence data.

    Code of Federal Regulations, 2013 CFR

    2013-07-01

    ... in the sequence. (4) The enumeration of amino acids may start at the first amino acid of the first..., counting backwards starting with the amino acid next to number 1. Otherwise, the enumeration of amino acids... sequence every 5 amino acids. The enumeration method for amino acid sequences that is set forth......

  1. 37 CFR 1.822 - Symbols and format to be used for nucleotide and/or amino acid sequence data.

    Code of Federal Regulations, 2014 CFR

    2014-07-01

    ... in the sequence. (4) The enumeration of amino acids may start at the first amino acid of the first..., counting backwards starting with the amino acid next to number 1. Otherwise, the enumeration of amino acids... sequence every 5 amino acids. The enumeration method for amino acid sequences that is set forth......

  2. 37 CFR 1.822 - Symbols and format to be used for nucleotide and/or amino acid sequence data.

    Code of Federal Regulations, 2012 CFR

    2012-07-01

    ... in the sequence. (4) The enumeration of amino acids may start at the first amino acid of the first..., counting backwards starting with the amino acid next to number 1. Otherwise, the enumeration of amino acids... sequence every 5 amino acids. The enumeration method for amino acid sequences that is set forth......

  3. Regulation of SHOOT MERISTEMLESS genes via an upstream-conserved noncoding sequence coordinates leaf development

    PubMed Central

    Uchida, Naoyuki; Townsley, Brad; Chung, Kook-Hyun; Sinha, Neelima

    2007-01-01

    The indeterminate shoot apical meristem of plants is characterized by the expression of the Class 1 KNOTTED1-LIKE HOMEOBOX (KNOX1) genes. KNOX1 genes have been implicated in the acquisition and/or maintenance of meristematic fate. One of the earliest indicators of a switch in fate from indeterminate meristem to determinate leaf primordium is the down-regulation of KNOX1 genes orthologous to SHOOT MERISTEMLESS (STM) in Arabidopsis (hereafter called STM genes) in the initiating primordia. In simple leafed plants, this down-regulation persists during leaf formation. In compound leafed plants, however, KNOX1 gene expression is reestablished later in the developing primordia, creating an indeterminate environment for leaflet formation. Despite this knowledge, most aspects of how STM gene expression is regulated remain largely unknown. Here, we identify two evolutionarily conserved noncoding sequences within the 5′ upstream region of STM genes in both simple and compound leafed species across monocots and dicots. We show that one of these elements is involved in the regulation of the persistent repression and/or the reestablishment of STM expression in the developing leaves but is not involved in the initial down-regulation in the initiating primordia. We also show evidence that this regulation is developmentally significant for leaf formation in the pathway involving ASYMMETRIC LEAVES1/2 (AS1/2) gene expression; these genes are known to function in leaf development. Together, these findings reveal a regulatory point of leaf development mediated through a conserved, noncoding sequence in STM genes. PMID:17898165

  4. Conservation of plasmid DNA sequences in coronatine-producing pathovars of Pseudomonas syringae

    SciTech Connect

    Bender, C.L.; Young, S.A. ); Mitchell, R.E. )

    1991-04-01

    In Pseudomonas syringae pv. tomato PT23.2, plasmid pPT23A (101 kb) is involved in synthesis of the phytotoxin coronatine. The physical characterization of mutations that abolished coronatine production indicated that at least 30 kb of pPT23A DNA are required for toxin synthesis. In the present study, {sup 32}P-labeled DNA fragments from the 30-kb region of pPT23A hybridized to plasmid DNAs from several coronatine-producing pathovars of P. syringae under conditions of high stringency. These experiments indicated that this region of pPT23A was strongly conserved in large plasmids (90 to 105 kb) that reside in P. syringae pv. atropurpurea, glycinea, and morsprunorum. The functional significance of the observed homology was demonstrated in marker-exchange experiments in which Tn5-inactivated sequences from the 30-kb region of pPT23A were used to mutate coronatine synthesis genes in the three heterologous pathovars. Physical characterization of the Tn5 insertions generated by marker exchange indicated that genes controlling coronatine synthesis in P. syringae pv. atropurpurea 1304, glycinea 4180, and morsprunorum 567 and 3714 were located on the large indigenous plasmids where homology was originally detected. Therefore, coronatine biosynthesis genes are strongly conserved in the plasmid DNAs of four producing pathovars, despite their disparate origins (California, Japan, New Zealand, Great Britain, and Italy).

  5. Hemagglutinin Sequence Conservation Guided Stem Immunogen Design from Influenza A H3 Subtype

    PubMed Central

    Mallajosyula, V. Vamsee Aditya; Citron, Michael; Ferrara, Francesca; Temperton, Nigel J.; Liang, Xiaoping; Flynn, Jessica A.; Varadarajan, Raghavan

    2015-01-01

    Seasonal epidemics caused by influenza A (H1 and H3 subtypes) and B viruses are a major global health threat. The traditional, trivalent influenza vaccines have limited efficacy because of rapid antigenic evolution of the circulating viruses. This antigenic variability mediates viral escape from the host immune responses, necessitating annual vaccine updates. Influenza vaccines elicit a protective antibody response, primarily targeting the viral surface glycoprotein hemagglutinin (HA). However, the predominant humoral response is against the hypervariable head domain of HA, thereby restricting the breadth of protection. In contrast, the conserved, subdominant stem domain of HA is a potential “universal” vaccine candidate. We designed an HA stem-fragment immunogen from the 1968 pandemic H3N2 strain (A/Hong Kong/1/68) guided by a comprehensive H3 HA sequence conservation analysis. The biophysical properties of the designed immunogen were further improved by C-terminal fusion of a trimerization motif, “isoleucine-zipper”, or “foldon”. These immunogens elicited cross-reactive, antiviral antibodies and conferred partial protection against a lethal, homologous HK68 virus challenge in vivo. Furthermore, bacterial expression of these immunogens is economical and facilitates rapid scale-up. PMID:26167164

  6. Predicting protein disorder by analyzing amino acid sequence

    PubMed Central

    Yang, Jack Y; Yang, Mary Qu

    2008-01-01

    Background Many protein regions and some entire proteins have no definite tertiary structure, presenting instead as dynamic, disorder ensembles under different physiochemical circumstances. These proteins and regions are known as Intrinsically Unstructured Proteins (IUP). IUP have been associated with a wide range of protein functions, along with roles in diseases characterized by protein misfolding and aggregation. Results Identifying IUP is important task in structural and functional genomics. We exact useful features from sequences and develop machine learning algorithms for the above task. We compare our IUP predictor with PONDRs (mainly neural-network-based predictors), disEMBL (also based on neural networks) and Globplot (based on disorder propensity). Conclusion We find that augmenting features derived from physiochemical properties of amino acids (such as hydrophobicity, complexity etc.) and using ensemble method proved beneficial. The IUP predictor is a viable alternative software tool for identifying IUP protein regions and proteins. PMID:18831799

  7. Snake venom toxins. The amino acid sequence of toxin Vi2, a homologue of pancreatic trypsin inhibitor, from Dendroaspis polylepis polylepis (black mamba) venom.

    PubMed

    Strydom, D J

    1977-04-25

    The amino acid sequence of venom component Vi2, a protein of low toxicity from Dendroaspis polylepis polylepis venom was determined by automatic sequence analysis in combination with sequence studies on tryptic peptides. This protein, the most retarded fraction of this venom on a cation-exchange resin, is a homologue of bovine pancreatic trypsin inhibitor consisting of a single chain of 57 amino acid residues containing six half-cystine residues. The active site lysyl residue of bovine trypsin inhibitor is conserved in Vi2 although large differences are found in the rest of the molecule. PMID:857902

  8. Clostridium sticklandii, a specialist in amino acid degradation:revisiting its metabolism through its genome sequence

    PubMed Central

    2010-01-01

    Background Clostridium sticklandii belongs to a cluster of non-pathogenic proteolytic clostridia which utilize amino acids as carbon and energy sources. Isolated by T.C. Stadtman in 1954, it has been generally regarded as a "gold mine" for novel biochemical reactions and is used as a model organism for studying metabolic aspects such as the Stickland reaction, coenzyme-B12- and selenium-dependent reactions of amino acids. With the goal of revisiting its carbon, nitrogen, and energy metabolism, and comparing studies with other clostridia, its genome has been sequenced and analyzed. Results C. sticklandii is one of the best biochemically studied proteolytic clostridial species. Useful additional information has been obtained from the sequencing and annotation of its genome, which is presented in this paper. Besides, experimental procedures reveal that C. sticklandii degrades amino acids in a preferential and sequential way. The organism prefers threonine, arginine, serine, cysteine, proline, and glycine, whereas glutamate, aspartate and alanine are excreted. Energy conservation is primarily obtained by substrate-level phosphorylation in fermentative pathways. The reactions catalyzed by different ferredoxin oxidoreductases and the exergonic NADH-dependent reduction of crotonyl-CoA point to a possible chemiosmotic energy conservation via the Rnf complex. C. sticklandii possesses both the F-type and V-type ATPases. The discovery of an as yet unrecognized selenoprotein in the D-proline reductase operon suggests a more detailed mechanism for NADH-dependent D-proline reduction. A rather unusual metabolic feature is the presence of genes for all the enzymes involved in two different CO2-fixation pathways: C. sticklandii harbours both the glycine synthase/glycine reductase and the Wood-Ljungdahl pathways. This unusual pathway combination has retrospectively been observed in only four other sequenced microorganisms. Conclusions Analysis of the C. sticklandii genome and

  9. A search for conserved sequences in coding regions reveals that the let-7 microRNA targets Dicer within its coding sequence

    PubMed Central

    Forman, Joshua J.; Legesse-Miller, Aster; Coller, Hilary A.

    2008-01-01

    Recognition sites for microRNAs (miRNAs) have been reported to be located in the 3′ untranslated regions of transcripts. In a computational screen for highly conserved motifs within coding regions, we found an excess of sequences conserved at the nucleotide level within coding regions in the human genome, the highest scoring of which are enriched for miRNA target sequences. To validate our results, we experimentally demonstrated that the let-7 miRNA directly targets the miRNA-processing enzyme Dicer within its coding sequence, thus establishing a mechanism for a miRNA/Dicer autoregulatory negative feedback loop. We also found computational evidence to suggest that miRNA target sites in coding regions and 3′ UTRs may differ in mechanism. This work demonstrates that miRNAs can directly target transcripts within their coding region in animals, and it suggests that a complete search for the regulatory targets of miRNAs should be expanded to include genes with recognition sites within their coding regions. As more genomes are sequenced, the methodological approach that we used for identifying motifs with high sequence conservation will be increasingly valuable for detecting functional sequence motifs within coding regions. PMID:18812516

  10. Morphological tranformation of calcite crystal growth by prismatic "acidic" polypeptide sequences.

    SciTech Connect

    Kim, I; Giocondi, J L; Orme, C A; Collino, J; Evans, J S

    2007-02-13

    Many of the interesting mechanical and materials properties of the mollusk shell are thought to stem from the prismatic calcite crystal assemblies within this composite structure. It is now evident that proteins play a major role in the formation of these assemblies. Recently, a superfamily of 7 conserved prismatic layer-specific mollusk shell proteins, Asprich, were sequenced, and the 42 AA C-terminal sequence region of this protein superfamily was found to introduce surface voids or porosities on calcite crystals in vitro. Using AFM imaging techniques, we further investigate the effect that this 42 AA domain (Fragment-2) and its constituent subdomains, DEAD-17 and Acidic-2, have on the morphology and growth kinetics of calcite dislocation hillocks. We find that Fragment-2 adsorbs on terrace surfaces and pins acute steps, accelerates then decelerates the growth of obtuse steps, forms clusters and voids on terrace surfaces, and transforms calcite hillock morphology from a rhombohedral form to a rounded one. These results mirror yet are distinct from some of the earlier findings obtained for nacreous polypeptides. The subdomains Acidic-2 and DEAD-17 were found to accelerate then decelerate obtuse steps and induce oval rather than rounded hillock morphologies. Unlike DEAD-17, Acidic-2 does form clusters on terrace surfaces and exhibits stronger obtuse velocity inhibition effects than either DEAD-17 or Fragment-2. Interestingly, a 1:1 mixture of both subdomains induces an irregular polygonal morphology to hillocks, and exhibits the highest degree of acute step pinning and obtuse step velocity inhibition. This suggests that there is some interplay between subdomains within an intra (Fragment-2) or intermolecular (1:1 mixture) context, and sequence interplay phenomena may be employed by biomineralization proteins to exert net effects on crystal growth and morphology.

  11. Mammalian mitochondrial D-loop region structural analysis: identification of new conserved sequences and their functional and evolutionary implications.

    PubMed

    Sbisà, E; Tanzariello, F; Reyes, A; Pesole, G; Saccone, C

    1997-12-31

    This paper reports the first comprehensive analysis of Displacement loop (D-loop) region sequences from ten different mammalian orders. It represents a systematic evolutionary study at the molecular level on regulatory homologous regions in organisms belonging to a well defined class, mammalia, which radiated about 150 million years ago (Mya). We have aligned and analyzed 26 complete D-loop region sequences available in the literature and the fat dormouse sequence, recently determined in our laboratory. The novelty of our alignment consists of the extensive manual revision of the preliminary output obtained by computer program to optimize sequence similarity, particularly for the two peripheral domains displaying heterogeneity in length and the presence of repeated sequences. The multialignment is available at the WWW site: http://www.ba.cnr.it/dloop.html. Our comparative study has allowed us to identify new conserved sequence blocks present in all the species under consideration and events of insertion/deletion which have important implications in both functional and evolutionary aspects. In particular we have detected two blocks, about 60 bp long, extended termination associated sequences (ETAS1 and ETAS2) conserved in all the organisms considered. Evaluation against experimental work suggests a possible functional role of ETAS1 and ETAS2 in the regulation of replication and transcription and targeted experimental approaches. The analyses on conserved sequence blocks (CSBs) clearly indicate that CSB1 is the only very essential element, common to all mammalian mt genomes, while CSB2 and CSB3 could be involved in different though related functions, probably species specific, and thus more linked to nuclear mitochondrial coevolutionary processes. Our hypothesis on the different functional implications of the conserved elements, CSBs and TASs, reported so far as main regulatory signals, would explain the different conservation of these elements in evolution. Moreover

  12. Lineage-Specific Conserved Noncoding Sequences of Plant Genomes: Their Possible Role in Nucleosome Positioning

    PubMed Central

    Hettiarachchi, Nilmini; Kryukov, Kirill; Sumiyama, Kenta; Saitou, Naruya

    2014-01-01

    Many studies on conserved noncoding sequences (CNSs) have found that CNSs are enriched significantly in regulatory sequence elements. We conducted whole-genome analysis on plant CNSs to identify lineage-specific CNSs in eudicots, monocots, angiosperms, and vascular plants based on the premise that lineage-specific CNSs define lineage-specific characters and functions in groups of organisms. We identified 27 eudicot, 204 monocot, 6,536 grass, 19 angiosperm, and 2 vascular plant lineage-specific CNSs (lengths range from 16 to 1,517 bp) that presumably originated in their respective common ancestors. A stronger constraint on the CNSs located in the untranslated regions was observed. The CNSs were often flanked by genes involved in transcription regulation. A drop of A+T content near the border of CNSs was observed and CNS regions showed a higher nucleosome occupancy probability. These CNSs are candidate regulatory elements, which are expected to define lineage-specific features of various plant groups. PMID:25364802

  13. G-boxes, bigfoot genes, and environmental response: characterization of intragenomic conserved noncoding sequences in Arabidopsis.

    PubMed

    Freeling, Michael; Rapaka, Lakshmi; Lyons, Eric; Pedersen, Brent; Thomas, Brian C

    2007-05-01

    A tetraploidy left Arabidopsis thaliana with 6358 pairs of homoeologs that, when aligned, generated 14,944 intragenomic conserved noncoding sequences (CNSs). Our previous work assembled these phylogenetic footprints into a database. We show that known transcription factor (TF) binding motifs, including the G-box, are overrepresented in these CNSs. A total of 254 genes spanning long lengths of CNS-rich chromosomes (Bigfoot) dominate this database. Therefore, we made subdatabases: one containing Bigfoot genes and the other containing genes with three to five CNSs (Smallfoot). Bigfoot genes are generally TFs that respond to signals, with their modal CNS positioned 3.1 kb 5' from the ATG. Smallfoot genes encode components of signal transduction machinery, the cytoskeleton, or involve transcription. We queried each subdatabase with each possible 7-nucleotide sequence. Among hundreds of hits, most were purified from CNSs, and almost all of those significantly enriched in CNSs had no experimental history. The 7-mers in CNSs are not 5'- to 3'-oriented in Bigfoot genes but are often oriented in Smallfoot genes. CNSs with one G-box tend to have two G-boxes. CNSs were shared with the homoeolog only and with no other gene, suggesting that binding site turnover impedes detection. Bigfoot genes may function in adaptation to environmental change. PMID:17496117

  14. Discovery of Novel ncRNA Sequences in Multiple Genome Alignments on the Basis of Conserved and Stable Secondary Structures.

    PubMed

    Fu, Yinghan; Xu, Zhenjiang Zech; Lu, Zhi J; Zhao, Shan; Mathews, David H

    2015-01-01

    Recently, non-coding RNAs (ncRNAs) have been discovered with novel functions, and it has been appreciated that there is pervasive transcription of genomes. Moreover, many novel ncRNAs are not conserved on the primary sequence level. Therefore, de novo computational ncRNA detection that is accurate and efficient is desirable. The purpose of this study is to develop a ncRNA detection method based on conservation of structure in more than two genomes. A new method called Multifind, using Multilign, was developed. Multilign predicts the common secondary structure for multiple input sequences. Multifind then uses measures of structure conservation to estimate the probability that the input sequences are a conserved ncRNA using a classification support vector machine. Multilign is based on Dynalign, which folds and aligns two sequences simultaneously using a scoring scheme that does not include sequence identity; its structure prediction quality is therefore not affected by input sequence diversity. Additionally, ensemble defect was introduced to Multifind as an additional discriminating feature that quantifies the compactness of the folding space for a sequence. Benchmarks showed Multifind performs better than RNAz and LocARNATE+RNAz, a method that uses RNAz on structure alignments generated by LocARNATE, on testing sequences extracted from the Rfam database. For de novo ncRNA discovery in three genomes, Multifind and LocARNATE+RNAz had an advantage over RNAz in low similarity regions of genome alignments. Additionally, Multifind and LocARNATE+RNAz found different subsets of known ncRNA sequences, suggesting the two approaches are complementary. PMID:26075601

  15. Whole genome sequencing of Ethiopian highlanders reveals conserved hypoxia tolerance genes

    PubMed Central

    2014-01-01

    Background Although it has long been proposed that genetic factors contribute to adaptation to high altitude, such factors remain largely unverified. Recent advances in high-throughput sequencing have made it feasible to analyze genome-wide patterns of genetic variation in human populations. Since traditionally such studies surveyed only a small fraction of the genome, interpretation of the results was limited. Results We report here the results of the first whole genome resequencing-based analysis identifying genes that likely modulate high altitude adaptation in native Ethiopians residing at 3,500 m above sea level on Bale Plateau or Chennek field in Ethiopia. Using cross-population tests of selection, we identify regions with a significant loss of diversity, indicative of a selective sweep. We focus on a 208 kbp gene-rich region on chromosome 19, which is significant in both of the Ethiopian subpopulations sampled. This region contains eight protein-coding genes and spans 135 SNPs. To elucidate its potential role in hypoxia tolerance, we experimentally tested whether individual genes from the region affect hypoxia tolerance in Drosophila. Three genes significantly impact survival rates in low oxygen: cic, an ortholog of human CIC, Hsl, an ortholog of human LIPE, and Paf-AHα, an ortholog of human PAFAH1B3. Conclusions Our study reveals evolutionarily conserved genes that modulate hypoxia tolerance. In addition, we show that many of our results would likely be unattainable using data from exome sequencing or microarray studies. This highlights the importance of whole genome sequencing for investigating adaptation by natural selection. PMID:24555826

  16. Multiple Amino Acid Sequence Alignment Nitrogenase Component 1: Insights into Phylogenetics and Structure-Function Relationships

    PubMed Central

    Howard, James B.; Kechris, Katerina J.; Rees, Douglas C.; Glazer, Alexander N.

    2013-01-01

    Amino acid residues critical for a protein's structure-function are retained by natural selection and these residues are identified by the level of variance in co-aligned homologous protein sequences. The relevant residues in the nitrogen fixation Component 1 α- and β-subunits were identified by the alignment of 95 protein sequences. Proteins were included from species encompassing multiple microbial phyla and diverse ecological niches as well as the nitrogen fixation genotypes, anf, nif, and vnf, which encode proteins associated with cofactors differing at one metal site. After adjusting for differences in sequence length, insertions, and deletions, the remaining >85% of the sequence co-aligned the subunits from the three genotypes. Six Groups, designated Anf, Vnf , and Nif I-IV, were assigned based upon genetic origin, sequence adjustments, and conserved residues. Both subunits subdivided into the same groups. Invariant and single variant residues were identified and were defined as “core” for nitrogenase function. Three species in Group Nif-III, Candidatus Desulforudis audaxviator, Desulfotomaculum kuznetsovii, and Thermodesulfatator indicus, were found to have a seleno-cysteine that replaces one cysteinyl ligand of the 8Fe:7S, P-cluster. Subsets of invariant residues, limited to individual groups, were identified; these unique residues help identify the gene of origin (anf, nif, or vnf) yet should not be considered diagnostic of the metal content of associated cofactors. Fourteen of the 19 residues that compose the cofactor pocket are invariant or single variant; the other five residues are highly variable but do not correlate with the putative metal content of the cofactor. The variable residues are clustered on one side of the cofactor, away from other functional centers in the three dimensional structure. Many of the invariant and single variant residues were not previously recognized as potentially critical and their identification provides the bases

  17. Bacterial periplasmic sialic acid-binding proteins exhibit a conserved binding site

    SciTech Connect

    Gangi Setty, Thanuja; Cho, Christine; Govindappa, Sowmya; Apicella, Michael A.; Ramaswamy, S.

    2014-07-01

    Structure–function studies of sialic acid-binding proteins from F. nucleatum, P. multocida, V. cholerae and H. influenzae reveal a conserved network of hydrogen bonds involved in conformational change on ligand binding. Sialic acids are a family of related nine-carbon sugar acids that play important roles in both eukaryotes and prokaryotes. These sialic acids are incorporated/decorated onto lipooligosaccharides as terminal sugars in multiple bacteria to evade the host immune system. Many pathogenic bacteria scavenge sialic acids from their host and use them for molecular mimicry. The first step of this process is the transport of sialic acid to the cytoplasm, which often takes place using a tripartite ATP-independent transport system consisting of a periplasmic binding protein and a membrane transporter. In this paper, the structural characterization of periplasmic binding proteins from the pathogenic bacteria Fusobacterium nucleatum, Pasteurella multocida and Vibrio cholerae and their thermodynamic characterization are reported. The binding affinities of several mutations in the Neu5Ac binding site of the Haemophilus influenzae protein are also reported. The structure and the thermodynamics of the binding of sugars suggest that all of these proteins have a very well conserved binding pocket and similar binding affinities. A significant conformational change occurs when these proteins bind the sugar. While the C1 carboxylate has been identified as the primary binding site, a second conserved hydrogen-bonding network is involved in the initiation and stabilization of the conformational states.

  18. Discovery and profiling of novel and conserved microRNAs during flower development in Carya cathayensis via deep sequencing.

    PubMed

    Wang, Zheng Jia; Huang, Jian Qin; Huang, You Jun; Li, Zheng; Zheng, Bing Song

    2012-08-01

    Hickory (Carya cathayensis Sarg.) is an economically important woody plant in China, but its long juvenile phase delays yield. MicroRNAs (miRNAs) are critical regulators of genes and important for normal plant development and physiology, including flower development. We used Solexa technology to sequence two small RNA libraries from two floral differentiation stages in hickory to identify miRNAs related to flower development. We identified 39 conserved miRNA sequences from 114 loci belonging to 23 families as well as two novel and ten potential novel miRNAs belonging to nine families. Moreover, 35 conserved miRNA*s and two novel miRNA*s were detected. Twenty miRNA sequences from 49 loci belonging to 11 families were differentially expressed; all were up-regulated at the later stage of flower development in hickory. Quantitative real-time PCR of 12 conserved miRNA sequences, five novel miRNA families, and two novel miRNA*s validated that all were expressed during hickory flower development, and the expression patterns were similar to those detected with Solexa sequencing. Finally, a total of 146 targets of the novel and conserved miRNAs were predicted. This study identified a diverse set of miRNAs that were closely related to hickory flower development and that could help in plant floral induction. PMID:22481137

  19. [Identification of new conserved and variable regions in the 16S rRNA gene of acetic acid bacteria and acetobacteraceae family].

    PubMed

    Chakravorty, S; Sarkar, S; Gachhui, R

    2015-01-01

    The Acetobacteraceae family of the class Alpha Proteobacteria is comprised of high sugar and acid tolerant bacteria. The Acetic Acid Bacteria are the economically most significant group of this family because of its association with food products like vinegar, wine etc. Acetobacteraceae are often hard to culture in laboratory conditions and they also maintain very low abundances in their natural habitats. Thus identification of the organisms in such environments is greatly dependent on modern tools of molecular biology which require a thorough knowledge of specific conserved gene sequences that may act as primers and or probes. Moreover unconserved domains in genes also become markers for differentiating closely related genera. In bacteria, the 16S rRNA gene is an ideal candidate for such conserved and variable domains. In order to study the conserved and variable domains of the 16S rRNA gene of Acetic Acid Bacteria and the Acetobacteraceae family, sequences from publicly available databases were aligned and compared. Near complete sequences of the gene were also obtained from Kombucha tea biofilm, a known Acetobacteraceae family habitat, in order to corroborate the domains obtained from the alignment studies. The study indicated that the degree of conservation in the gene is significantly higher among the Acetic Acid Bacteria than the whole Acetobacteraceae family. Moreover it was also observed that the previously described hypervariable regions V1, V3, V5, V6 and V7 were more or less conserved in the family and the spans of the variable regions are quite distinct as well. PMID:26510592

  20. Characterization of Protective Epitopes in a Highly Conserved Plasmodium falciparum Antigenic Protein Containing Repeats of Acidic and Basic Residues

    PubMed Central

    Sharma, Pawan; Kumar, Anil; Singh, Balwan; Bharadwaj, Ashima; Sailaja, V. Naga; Adak, T.; Kushwaha, Ashima; Malhotra, Pawan; Chauhan, V. S.

    1998-01-01

    The delineation of putatively protective and immunogenic epitopes in vaccine candidate proteins constitutes a major research effort towards the development of an effective malaria vaccine. By virtue of its role in the formation of the immune clusters of merozoites, its location on the surface of merozoites, and its highly conserved nature both at the nucleotide sequence level and the amino acid sequence level, the antigen which contains repeats of acidic and basic residues (ABRA) of the human malaria parasite Plasmodium falciparum represents such an antigen. Based upon the predicted amino acid sequence of ABRA, we synthesized eight peptides, with six of these (AB-1 to AB-6) ranging from 12 to 18 residues covering the most hydrophilic regions of the protein, and two more peptides (AB-7 and AB-8) representing its repetitive sequences. We found that all eight constructs bound an appreciable amount of antibody in sera from a large proportion of P. falciparum malaria patients; two of these peptides (AB-1 and AB-3) also elicited a strong proliferation response in peripheral blood mononuclear cells from all 11 human subjects recovering from malaria. When used as carrier-free immunogens, six peptides induced a strong, boostable, immunoglobulin G-type antibody response in rabbits, indicating the presence of both B-cell determinants and T-helper-cell epitopes in these six constructs. These antibodies specifically cross-reacted with the parasite protein(s) in an immunoblot and in an immunofluorescence assay. In another immunoblot, rabbit antipeptide sera also recognized recombinant fragments of ABRA expressed in bacteria. More significantly, rabbit antibodies against two constructs (AB-1 and AB-5) inhibited the merozoite reinvasion of human erythrocytes in vitro up to ∼90%. These results favor further studies so as to determine possible inclusion of these two constructs in a multicomponent subunit vaccine against asexual blood stages of P. falciparum. PMID:9596765

  1. Conserved regulators of Rag GTPases orchestrate amino acid-dependent TORC1 signaling

    PubMed Central

    Powis, Katie; De Virgilio, Claudio

    2016-01-01

    The highly conserved target of rapamycin complex 1 (TORC1) is the central component of a signaling network that couples a vast range of internal and external stimuli to cell growth, proliferation and metabolism. TORC1 deregulation is associated with a number of human pathologies, including many cancers and metabolic disorders, underscoring its importance in cellular and organismal growth control. The activity of TORC1 is modulated by multiple inputs; however, the presence of amino acids is a stimulus that is essential for its activation. Amino acid sufficiency is communicated to TORC1 via the highly conserved family of Rag GTPases, which assemble as heterodimeric complexes on lysosomal/vacuolar membranes and are regulated by their guanine nucleotide loading status. Studies in yeast, fly and mammalian model systems have revealed a multitude of conserved Rag GTPase modulators, which have greatly expanded our understanding of amino acid sensing by TORC1. Here we review the major known modulators of the Rag GTPases, focusing on recent mechanistic insights that highlight the evolutionary conservation and divergence of amino acid signaling to TORC1. PMID:27462445

  2. Amino acid sequence analysis and characterization of a ribonuclease from starfish Asterias amurensis.

    PubMed

    Motoyoshi, Naomi; Kobayashi, Hiroko; Itagaki, Tadashi; Inokuchi, Norio

    2016-09-01

    The aim of this study was to phylogenetically characterize the location of the RNase T2 enzyme in the starfish (Asterias amurensis). We isolated an RNase T2 ribonuclease (RNase Aa) from the ovaries of starfish and determined its amino acid sequence by protein chemistry and cloning cDNA encoding RNase Aa. The isolated protein had 231 amino acid residues, a predicted molecular mass of 25,906 Da, and an optimal pH of 5.0. RNase Aa preferentially released guanylic acid from the RNA. The catalytic sites of the RNase T2 family are conserved in RNase Aa; furthermore, the distribution of the cysteine residues in RNase Aa is similar to that in other animal and plant T2 RNases. RNase Aa is cleaved at two points: 21 residues from the N-terminus and 29 residues from the C-terminus; however, both fragments may remain attached to the protein via disulfide bridges, leading to the maintenance of its conformation, as suggested by circular dichroism spectrum analysis. The phylogenetic analysis revealed that starfish RNase Aa is evolutionarily an intermediate between protozoan and oyster RNases. PMID:26920046

  3. Structural gene and complete amino acid sequence of Pseudomonas aeruginosa IFO 3455 elastase.

    PubMed Central

    Fukushima, J; Yamamoto, S; Morihara, K; Atsumi, Y; Takeuchi, H; Kawamoto, S; Okuda, K

    1989-01-01

    The DNA encoding the elastase of Pseudomonas aeruginosa IFO 3455 was cloned, and its complete nucleotide sequence was determined. When the cloned gene was ligated to pUC18, the Escherichia coli expression vector, bacteria carrying the gene exhibited high levels of both elastase activity and elastase antigens. The amino acid sequence, deduced from the nucleotide sequence, revealed that the mature elastase consisted of 301 amino acids with a relative molecular mass of 32,926 daltons. The amino acid composition predicted from the DNA sequence was quite similar to the chemically determined composition of purified elastase reported previously. We also observed nucleotide sequence encoding a signal peptide and "pro" sequence consisting of 197 amino acids upstream from the mature elastase protein gene. The amino acid sequence analysis revealed that both the N-terminal sequence of the purified elastase and the N-terminal side sequences of the C-terminal tryptic peptide as well as the internal lysyl peptide fragment were completely identical to the deduced amino acid sequences. The pattern of identity of amino acid sequences was quite evident in the regions that include structurally and functionally important residues of Bacillus subtilis thermolysin. PMID:2493453

  4. A conserved intronic U1 snRNP-binding sequence promotes trans-splicing in Drosophila

    PubMed Central

    Gao, Jun-Li; Fan, Yu-Jie; Wang, Xiu-Ye; Zhang, Yu; Pu, Jia; Li, Liang; Shao, Wei; Zhan, Shuai; Hao, Jianjiang

    2015-01-01

    Unlike typical cis-splicing, trans-splicing joins exons from two separate transcripts to produce chimeric mRNA and has been detected in most eukaryotes. Trans-splicing in trypanosomes and nematodes has been characterized as a spliced leader RNA-facilitated reaction; in contrast, its mechanism in higher eukaryotes remains unclear. Here we investigate mod(mdg4), a classic trans-spliced gene in Drosophila, and report that two critical RNA sequences in the middle of the last 5′ intron, TSA and TSB, promote trans-splicing of mod(mdg4). In TSA, a 13-nucleotide (nt) core motif is conserved across Drosophila species and is essential and sufficient for trans-splicing, which binds U1 small nuclear RNP (snRNP) through strong base-pairing with U1 snRNA. In TSB, a conserved secondary structure acts as an enhancer. Deletions of TSA and TSB using the CRISPR/Cas9 system result in developmental defects in flies. Although it is not clear how the 5′ intron finds the 3′ introns, compensatory changes in U1 snRNA rescue trans-splicing of TSA mutants, demonstrating that U1 recruitment is critical to promote trans-splicing in vivo. Furthermore, TSA core-like motifs are found in many other trans-spliced Drosophila genes, including lola. These findings represent a novel mechanism of trans-splicing, in which RNA motifs in the 5′ intron are sufficient to bring separate transcripts into close proximity to promote trans-splicing. PMID:25838544

  5. Comparison of SIV and HIV-1 genomic RNA structures reveals impact of sequence evolution on conserved and non-conserved structural motifs.

    PubMed

    Pollom, Elizabeth; Dang, Kristen K; Potter, E Lake; Gorelick, Robert J; Burch, Christina L; Weeks, Kevin M; Swanstrom, Ronald

    2013-01-01

    RNA secondary structure plays a central role in the replication and metabolism of all RNA viruses, including retroviruses like HIV-1. However, structures with known function represent only a fraction of the secondary structure reported for HIV-1(NL4-3). One tool to assess the importance of RNA structures is to examine their conservation over evolutionary time. To this end, we used SHAPE to model the secondary structure of a second primate lentiviral genome, SIVmac239, which shares only 50% sequence identity at the nucleotide level with HIV-1NL4-3. Only about half of the paired nucleotides are paired in both genomic RNAs and, across the genome, just 71 base pairs form with the same pairing partner in both genomes. On average the RNA secondary structure is thus evolving at a much faster rate than the sequence. Structure at the Gag-Pro-Pol frameshift site is maintained but in a significantly altered form, while the impact of selection for maintaining a protein binding interaction can be seen in the conservation of pairing partners in the small RRE stems where Rev binds. Structures that are conserved between SIVmac239 and HIV-1(NL4-3) also occur at the 5' polyadenylation sequence, in the plus strand primer sites, PPT and cPPT, and in the stem-loop structure that includes the first splice acceptor site. The two genomes are adenosine-rich and cytidine-poor. The structured regions are enriched in guanosines, while unpaired regions are enriched in adenosines, and functionaly important structures have stronger base pairing than nonconserved structures. We conclude that much of the secondary structure is the result of fortuitous pairing in a metastable state that reforms during sequence evolution. However, secondary structure elements with important function are stabilized by higher guanosine content that allows regions of structure to persist as sequence evolution proceeds, and, within the confines of selective pressure, allows structures to evolve. PMID:23593004

  6. Human retroviruses and AIDS 1996. A compilation and analysis of nucleic acid and amino acid sequences

    SciTech Connect

    Myers, G.; Foley, B.; Korber, B.; Mellors, J.W.; Jeang, K.T.; Wain-Hobson, S.

    1997-04-01

    This compendium and the accompanying floppy diskettes are the result of an effort to compile and rapidly publish all relevant molecular data concerning the human immunodeficiency viruses (HIV) and related retroviruses. The scope of the compendium and database is best summarized by the five parts that it comprises: (1) Nuclear Acid Alignments and Sequences; (2) Amino Acid Alignments; (3) Analysis; (4) Related Sequences; and (5) Database Communications. Information within all the parts is updated throughout the year on the Web site, http://hiv-web.lanl.gov. While this publication could take the form of a review or sequence monograph, it is not so conceived. Instead, the literature from which the database is derived has simply been summarized and some elementary computational analyses have been performed upon the data. Interpretation and commentary have been avoided insofar as possible so that the reader can form his or her own judgments concerning the complex information. In addition to the general descriptions of the parts of the compendium, the user should read the individual introductions for each part.

  7. Automated conserved non-coding sequence (CNS) discovery reveals differences in gene content and promoter evolution among grasses

    PubMed Central

    Turco, Gina; Schnable, James C.; Pedersen, Brent; Freeling, Michael

    2013-01-01

    Conserved non-coding sequences (CNS) are islands of non-coding sequence that, like protein coding exons, show less divergence in sequence between related species than functionless DNA. Several CNSs have been demonstrated experimentally to function as cis-regulatory regions. However, the specific functions of most CNSs remain unknown. Previous searches for CNS in plants have either anchored on exons and only identified nearby sequences or required years of painstaking manual annotation. Here we present an open source tool that can accurately identify CNSs between any two related species with sequenced genomes, including both those immediately adjacent to exons and distal sequences separated by >12 kb of non-coding sequence. We have used this tool to characterize new motifs, associate CNSs with additional functions, and identify previously undetected genes encoding RNA and protein in the genomes of five grass species. We provide a list of 15,363 orthologous CNSs conserved across all grasses tested. We were also able to identify regulatory sequences present in the common ancestor of grasses that have been lost in one or more extant grass lineages. Lists of orthologous gene pairs and associated CNSs are provided for reference inbred lines of arabidopsis, Japonica rice, foxtail millet, sorghum, brachypodium, and maize. PMID:23874343

  8. A comparative genomics strategy for targeted discovery of single-nucleotide polymorphisms and conserved-noncoding sequences in orphan crops.

    PubMed

    Feltus, F A; Singh, H P; Lohithaswa, H C; Schulze, S R; Silva, T D; Paterson, A H

    2006-04-01

    Completed genome sequences provide templates for the design of genome analysis tools in orphan species lacking sequence information. To demonstrate this principle, we designed 384 PCR primer pairs to conserved exonic regions flanking introns, using Sorghum/Pennisetum expressed sequence tag alignments to the Oryza genome. Conserved-intron scanning primers (CISPs) amplified single-copy loci at 37% to 80% success rates in taxa that sample much of the approximately 50-million years of Poaceae divergence. While the conserved nature of exons fostered cross-taxon amplification, the lesser evolutionary constraints on introns enhanced single-nucleotide polymorphism detection. For example, in eight rice (Oryza sativa) genotypes, polymorphism averaged 12.1 per kb in introns but only 3.6 per kb in exons. Curiously, among 124 CISPs evaluated across Oryza, Sorghum, Pennisetum, Cynodon, Eragrostis, Zea, Triticum, and Hordeum, 23 (18.5%) seemed to be subject to rigid intron size constraints that were independent of per-nucleotide DNA sequence variation. Furthermore, we identified 487 conserved-noncoding sequence motifs in 129 CISP loci. A large CISP set (6,062 primer pairs, amplifying introns from 1,676 genes) designed using an automated pipeline showed generally higher abundance in recombinogenic than in nonrecombinogenic regions of the rice genome, thus providing relatively even distribution along genetic maps. CISPs are an effective means to explore poorly characterized genomes for both DNA polymorphism and noncoding sequence conservation on a genome-wide or candidate gene basis, and also provide anchor points for comparative genomics across a diverse range of species. PMID:16607031

  9. Conserved sequences of sperm-activating peptide and its receptor throughout evolution, despite speciation in the sea star Asterias amurensis and closely related species.

    PubMed

    Nakachi, Mia; Hoshi, Motonori; Matsumoto, Midori; Moriyama, Hideaki

    2008-08-01

    The asteroidal sperm-activating peptides (asterosaps) from the egg jelly bind to their sperm receptor, a membrane-bound guanylate cyclase, on the tail to activate sperm in sea stars. Asterosaps are produced as single peptides and then cleaved into shorter peptides. Sperm activation is followed by the acrosome reaction, which is subfamily specific. In order to investigate the molecular details of the asterosap-receptor interaction, corresponding cDNAs have been cloned, sequenced and analysed from the Asteriinae subfamily including Asterias amurensis, A. rubens, A. forbesi and Aphelasterias japonica, as well as Distolasterias nipon from the Coscinasteriinae subfamily. Averages of 29% and 86% identity were found from the deduced amino acid sequences in asterosap and its receptor extracellular domains, respectively, across all species examined. The phylogenic tree topology for asterosap and its receptor was similar to that of the mitochondrial cytochrome c oxidase subunit I. In spite of a certain homology, the amino acid sequences exhibited speciation. Conservation was found in the asterosap residues involved in disulphide bonding and proteinase-cleaving sites. Conversely, similarities were detected between potential asterosap-binding sites and the structure of the atrial natriuretic peptide receptor. Although the sperm-activating peptide and its receptor share certain common sequences, they may serve as barriers that ensure speciation in the sea star A. amurensis and closely related species. PMID:18578950

  10. Evolution of vertebrate IgM: complete amino acid sequence of the constant region of Ambystoma mexicanum mu chain deduced from cDNA sequence.

    PubMed

    Fellah, J S; Wiles, M V; Charlemagne, J; Schwager, J

    1992-10-01

    cDNA clones coding for the constant region of the Mexican axolotl (Ambystoma mexicanum) mu heavy immunoglobulin chain were selected from total spleen RNA, using a cDNA polymerase chain reaction technique. The specific 5'-end primer was an oligonucleotide homologous to the JH segment of Xenopus laevis mu chain. One of the clones, JHA/3, corresponded to the complete constant region of the axolotl mu chain, consisting of a 1362-nucleotide sequence coding for a polypeptide of 454 amino acids followed in 3' direction by a 179-nucleotide untranslated region and a polyA+ tail. The axolotl C mu is divided into four typical domains (C mu 1-C mu 4) and can be aligned with the Xenopus C mu with an overall identity of 56% at the nucleotide level. Percent identities were particularly high between C mu 1 (59%) and C mu 4 (71%). The C-terminal 20-amino acid segment which constitutes the secretory part of the mu chain is strongly homologous to the equivalent sequences of chondrichthyans and of other tetrapods, including a conserved N-linked oligosaccharide, the penultimate cysteine and the C-terminal lysine. The four C mu domains of 13 vertebrate species ranging from chondrichthyans to mammals were aligned and compared at the amino acid level. The significant number of mu-specific residues which are conserved into each of the four C mu domains argues for a continuous line of evolution of the vertebrate mu chain. This notion was confirmed by the ability to reconstitute a consistent vertebrate evolution tree based on the phylogenic parsimony analysis of the C mu 4 sequences. PMID:1382992

  11. Integrating bioinformatic resources to predict transcription factors interacting with cis-sequences conserved in co-regulated genes

    PubMed Central

    2014-01-01

    Background Using motif detection programs it is fairly straightforward to identify conserved cis-sequences in promoters of co-regulated genes. In contrast, the identification of the transcription factors (TFs) interacting with these cis-sequences is much more elaborate. To facilitate this, we explore the possibility of using several bioinformatic and experimental approaches for TF identification. This starts with the selection of co-regulated gene sets and leads first to the prediction and then to the experimental validation of TFs interacting with cis-sequences conserved in the promoters of these co-regulated genes. Results Using the PathoPlant database, 32 up-regulated gene groups were identified with microarray data for drought-responsive gene expression from Arabidopsis thaliana. Application of the binding site estimation suite of tools (BEST) discovered 179 conserved sequence motifs within the corresponding promoters. Using the STAMP web-server, 49 sequence motifs were classified into 7 motif families for which similarities with known cis-regulatory sequences were identified. All motifs were subjected to a footprintDB analysis to predict interacting DNA binding domains from plant TF families. Predictions were confirmed by using a yeast-one-hybrid approach to select interacting TFs belonging to the predicted TF families. TF-DNA interactions were further experimentally validated in yeast and with a Physcomitrella patens transient expression system, leading to the discovery of several novel TF-DNA interactions. Conclusions The present work demonstrates the successful integration of several bioinformatic resources with experimental approaches to predict and validate TFs interacting with conserved sequence motifs in co-regulated genes. PMID:24773781

  12. THE GRK4 SUBFAMILY OF G PROTEIN-COUPLED RECEPTOR KINASES: ALTERNATIVE SPLICING, GENE ORGANIZATION, AND SEQUENCE CONSERVATION

    EPA Science Inventory

    The GRK4 subfamily of G protein-coupled receptor kinases. Alternative splicing, gene organization, and sequence conservation.

    Premont RT, Macrae AD, Aparicio SA, Kendall HE, Welch JE, Lefkowitz RJ.

    Department of Medicine, Howard Hughes Medical Institute, Duke Univer...

  13. Natural vs. random protein sequences: Discovering combinatorics properties on amino acid words.

    PubMed

    Santoni, Daniele; Felici, Giovanni; Vergni, Davide

    2016-02-21

    Casual mutations and natural selection have driven the evolution of protein amino acid sequences that we observe at present in nature. The question about which is the dominant force of proteins evolution is still lacking of an unambiguous answer. Casual mutations tend to randomize protein sequences while, in order to have the correct functionality, one expects that selection mechanisms impose rigid constraints on amino acid sequences. Moreover, one also has to consider that the space of all possible amino acid sequences is so astonishingly large that it could be reasonable to have a well tuned amino acid sequence indistinguishable from a random one. In order to study the possibility to discriminate between random and natural amino acid sequences, we introduce different measures of association between pairs of amino acids in a sequence, and apply them to a dataset of 1047 natural protein sequences and 10,470 random sequences, carefully generated in order to preserve the relative length and amino acid distribution of the natural proteins. We analyze the multidimensional measures with machine learning techniques and show that, to a reasonable extent, natural protein sequences can be differentiated from random ones. PMID:26656109

  14. Transcriptome Sequencing in Response to Salicylic Acid in Salvia miltiorrhiza

    PubMed Central

    Zhang, Xiaoru; Dong, Juane; Liu, Hailong; Wang, Jiao; Qi, Yuexin; Liang, Zongsuo

    2016-01-01

    Salvia miltiorrhiza is a traditional Chinese herbal medicine, whose quality and yield are often affected by diseases and environmental stresses during its growing season. Salicylic acid (SA) plays a significant role in plants responding to biotic and abiotic stresses, but the involved regulatory factors and their signaling mechanisms are largely unknown. In order to identify the genes involved in SA signaling, the RNA sequencing (RNA-seq) strategy was employed to evaluate the transcriptional profiles in S. miltiorrhiza cell cultures. A total of 50,778 unigenes were assembled, in which 5,316 unigenes were differentially expressed among 0-, 2-, and 8-h SA induction. The up-regulated genes were mainly involved in stimulus response and multi-organism process. A core set of candidate novel genes coding SA signaling component proteins was identified. Many transcription factors (e.g., WRKY, bHLH and GRAS) and genes involved in hormone signal transduction were differentially expressed in response to SA induction. Detailed analysis revealed that genes associated with defense signaling, such as antioxidant system genes, cytochrome P450s and ATP-binding cassette transporters, were significantly overexpressed, which can be used as genetic tools to investigate disease resistance. Our transcriptome analysis will help understand SA signaling and its mechanism of defense systems in S. miltiorrhiza. PMID:26808150

  15. Transcriptome Sequencing in Response to Salicylic Acid in Salvia miltiorrhiza.

    PubMed

    Zhang, Xiaoru; Dong, Juane; Liu, Hailong; Wang, Jiao; Qi, Yuexin; Liang, Zongsuo

    2016-01-01

    Salvia miltiorrhiza is a traditional Chinese herbal medicine, whose quality and yield are often affected by diseases and environmental stresses during its growing season. Salicylic acid (SA) plays a significant role in plants responding to biotic and abiotic stresses, but the involved regulatory factors and their signaling mechanisms are largely unknown. In order to identify the genes involved in SA signaling, the RNA sequencing (RNA-seq) strategy was employed to evaluate the transcriptional profiles in S. miltiorrhiza cell cultures. A total of 50,778 unigenes were assembled, in which 5,316 unigenes were differentially expressed among 0-, 2-, and 8-h SA induction. The up-regulated genes were mainly involved in stimulus response and multi-organism process. A core set of candidate novel genes coding SA signaling component proteins was identified. Many transcription factors (e.g., WRKY, bHLH and GRAS) and genes involved in hormone signal transduction were differentially expressed in response to SA induction. Detailed analysis revealed that genes associated with defense signaling, such as antioxidant system genes, cytochrome P450s and ATP-binding cassette transporters, were significantly overexpressed, which can be used as genetic tools to investigate disease resistance. Our transcriptome analysis will help understand SA signaling and its mechanism of defense systems in S. miltiorrhiza. PMID:26808150

  16. Nucleotide sequence of a cluster of early and late genes in a conserved segment of the vaccinia virus genome.

    PubMed Central

    Plucienniczak, A; Schroeder, E; Zettlmeissl, G; Streeck, R E

    1985-01-01

    The nucleotide sequence of a 7.6 kb vaccinia DNA segment from a genomic region conserved among different orthopox virus has been determined. This segment contains a tight cluster of 12 partly overlapping open reading frames most of which can be correlated with previously identified early and late proteins and mRNAs. Regulatory signals used by vaccinia virus have been studied. Presumptive promoter regions are rich in A, T and carry the consensus sequences TATA and AATAA spaced at 20-24 base pairs. Tandem repeats of a CTATTC consensus sequence are proposed to be involved in the termination of early transcription. PMID:2987815

  17. Using Caenorhabditis elegans to Uncover Conserved Functions of Omega-3 and Omega-6 Fatty Acids

    PubMed Central

    Watts, Jennifer L.

    2016-01-01

    The nematode Caenorhabditis elegans is a powerful model organism to study functions of polyunsaturated fatty acids. The ability to alter fatty acid composition with genetic manipulation and dietary supplementation permits the dissection of the roles of omega-3 and omega-6 fatty acids in many biological process including reproduction, aging and neurobiology. Studies in C. elegans to date have mostly identified overlapping functions of 20-carbon omega-6 and omega-3 fatty acids in reproduction and in neurons, however, specific roles for either omega-3 or omega-6 fatty acids are beginning to emerge. Recent findings with importance to human health include the identification of a conserved Cox-independent prostaglandin synthesis pathway, critical functions for cytochrome P450 derivatives of polyunsaturated fatty acids, the requirements for omega-6 and omega-3 fatty acids in sensory neurons, and the importance of fatty acid desaturation for long lifespan. Furthermore, the ability of C. elegans to interconvert omega-6 to omega-3 fatty acids using the FAT-1 omega-3 desaturase has been exploited in mammalian studies and biotechnology approaches to generate mammals capable of exogenous generation of omega-3 fatty acids. PMID:26848697

  18. A highly conserved G-rich consensus sequence in hepatitis C virus core gene represents a new anti–hepatitis C target

    PubMed Central

    Wang, Shao-Ru; Min, Yuan-Qin; Wang, Jia-Qi; Liu, Chao-Xing; Fu, Bo-Shi; Wu, Fan; Wu, Ling-Yu; Qiao, Zhi-Xian; Song, Yan-Yan; Xu, Guo-Hua; Wu, Zhi-Guo; Huang, Gai; Peng, Nan-Fang; Huang, Rong; Mao, Wu-Xiang; Peng, Shuang; Chen, Yu-Qi; Zhu, Ying; Tian, Tian; Zhang, Xiao-Lian; Zhou, Xiang

    2016-01-01

    G-quadruplex (G4) is one of the most important secondary structures in nucleic acids. Until recently, G4 RNAs have not been reported in any ribovirus, such as the hepatitis C virus. Our bioinformatics analysis reveals highly conserved guanine-rich consensus sequences within the core gene of hepatitis C despite the high genetic variability of this ribovirus; we further show using various methods that such consensus sequences can fold into unimolecular G4 RNA structures, both in vitro and under physiological conditions. Furthermore, we provide direct evidences that small molecules specifically targeting G4 can stabilize this structure to reduce RNA replication and inhibit protein translation of intracellular hepatitis C. Ultimately, the stabilization of G4 RNA in the genome of hepatitis C represents a promising new strategy for anti–hepatitis C drug development. PMID:27051880

  19. A highly conserved G-rich consensus sequence in hepatitis C virus core gene represents a new anti-hepatitis C target.

    PubMed

    Wang, Shao-Ru; Min, Yuan-Qin; Wang, Jia-Qi; Liu, Chao-Xing; Fu, Bo-Shi; Wu, Fan; Wu, Ling-Yu; Qiao, Zhi-Xian; Song, Yan-Yan; Xu, Guo-Hua; Wu, Zhi-Guo; Huang, Gai; Peng, Nan-Fang; Huang, Rong; Mao, Wu-Xiang; Peng, Shuang; Chen, Yu-Qi; Zhu, Ying; Tian, Tian; Zhang, Xiao-Lian; Zhou, Xiang

    2016-04-01

    G-quadruplex (G4) is one of the most important secondary structures in nucleic acids. Until recently, G4 RNAs have not been reported in any ribovirus, such as the hepatitis C virus. Our bioinformatics analysis reveals highly conserved guanine-rich consensus sequences within the core gene of hepatitis C despite the high genetic variability of this ribovirus; we further show using various methods that such consensus sequences can fold into unimolecular G4 RNA structures, both in vitro and under physiological conditions. Furthermore, we provide direct evidences that small molecules specifically targeting G4 can stabilize this structure to reduce RNA replication and inhibit protein translation of intracellular hepatitis C. Ultimately, the stabilization of G4 RNA in the genome of hepatitis C represents a promising new strategy for anti-hepatitis C drug development. PMID:27051880

  20. Genotyping by sequencing resolves shallow population structure to inform conservation of Chinook salmon (Oncorhynchus tshawytscha)

    PubMed Central

    Larson, Wesley A; Seeb, Lisa W; Everett, Meredith V; Waples, Ryan K; Templin, William D; Seeb, James E

    2014-01-01

    Recent advances in population genomics have made it possible to detect previously unidentified structure, obtain more accurate estimates of demographic parameters, and explore adaptive divergence, potentially revolutionizing the way genetic data are used to manage wild populations. Here, we identified 10 944 single-nucleotide polymorphisms using restriction-site-associated DNA (RAD) sequencing to explore population structure, demography, and adaptive divergence in five populations of Chinook salmon (Oncorhynchus tshawytscha) from western Alaska. Patterns of population structure were similar to those of past studies, but our ability to assign individuals back to their region of origin was greatly improved (>90% accuracy for all populations). We also calculated effective size with and without removing physically linked loci identified from a linkage map, a novel method for nonmodel organisms. Estimates of effective size were generally above 1000 and were biased downward when physically linked loci were not removed. Outlier tests based on genetic differentiation identified 733 loci and three genomic regions under putative selection. These markers and genomic regions are excellent candidates for future research and can be used to create high-resolution panels for genetic monitoring and population assignment. This work demonstrates the utility of genomic data to inform conservation in highly exploited species with shallow population structure. PMID:24665338

  1. Genotyping by sequencing resolves shallow population structure to inform conservation of Chinook salmon (Oncorhynchus tshawytscha).

    PubMed

    Larson, Wesley A; Seeb, Lisa W; Everett, Meredith V; Waples, Ryan K; Templin, William D; Seeb, James E

    2014-03-01

    Recent advances in population genomics have made it possible to detect previously unidentified structure, obtain more accurate estimates of demographic parameters, and explore adaptive divergence, potentially revolutionizing the way genetic data are used to manage wild populations. Here, we identified 10 944 single-nucleotide polymorphisms using restriction-site-associated DNA (RAD) sequencing to explore population structure, demography, and adaptive divergence in five populations of Chinook salmon (Oncorhynchus tshawytscha) from western Alaska. Patterns of population structure were similar to those of past studies, but our ability to assign individuals back to their region of origin was greatly improved (>90% accuracy for all populations). We also calculated effective size with and without removing physically linked loci identified from a linkage map, a novel method for nonmodel organisms. Estimates of effective size were generally above 1000 and were biased downward when physically linked loci were not removed. Outlier tests based on genetic differentiation identified 733 loci and three genomic regions under putative selection. These markers and genomic regions are excellent candidates for future research and can be used to create high-resolution panels for genetic monitoring and population assignment. This work demonstrates the utility of genomic data to inform conservation in highly exploited species with shallow population structure. PMID:24665338

  2. Evolutionary divergence and limits of conserved non-coding sequence detection in plant genomes

    PubMed Central

    Reineke, Anna R.; Bornberg-Bauer, Erich; Gu, Jenny

    2011-01-01

    The discovery of regulatory motifs embedded in upstream regions of plants is a particularly challenging bioinformatics task. Previous studies have shown that motifs in plants are short compared with those found in vertebrates. Furthermore, plant genomes have undergone several diversification mechanisms such as genome duplication events which impact the evolution of regulatory motifs. In this article, a systematic phylogenomic comparison of upstream regions is conducted to further identify features of the plant regulatory genomes, the component of genomes regulating gene expression, to enable future de novo discoveries. The findings highlight differences in upstream region properties between major plant groups and the effects of divergence times and duplication events. First, clear differences in upstream region evolution can be detected between monocots and dicots, thus suggesting that a separation of these groups should be made when searching for novel regulatory motifs, particularly since universal motifs such as the TATA box are rare. Second, investigating the decay rate of significantly aligned regions suggests that a divergence time of ∼100 mya sets a limit for reliable conserved non-coding sequence (CNS) detection. Insights presented here will set a framework to help identify embedded motifs of functional relevance by understanding the limits of bioinformatics detection for CNSs. PMID:21470961

  3. Conserved Non-Coding Sequences are Associated with Rates of mRNA Decay in Arabidopsis

    PubMed Central

    Spangler, Jacob B.; Feltus, Frank Alex

    2013-01-01

    Steady-state mRNA levels are tightly regulated through a combination of transcriptional and post-transcriptional control mechanisms. The discovery of cis-acting DNA elements that encode these control mechanisms is of high importance. We have investigated the influence of conserved non-coding sequences (CNSs), DNA patterns retained after an ancient whole genome duplication event, on the breadth of gene expression and the rates of mRNA decay in Arabidopsis thaliana. The absence of CNSs near α duplicate genes was associated with a decrease in breadth of gene expression and slower mRNA decay rates while the presence CNSs near α duplicates was associated with an increase in breadth of gene expression and faster mRNA decay rates. The observed difference in mRNA decay rate was fastest in genes with CNSs in both non-transcribed and transcribed regions, albeit through an unknown mechanism. This study supports the notion that some Arabidopsis CNSs regulate the steady-state mRNA levels through post-transcriptional control mechanisms and that CNSs also play a role in controlling the breadth of gene expression. PMID:23675377

  4. QColors: an algorithm for conservative viral quasispecies reconstruction from short and non-contiguous next generation sequencing reads.

    PubMed

    Huang, Austin; Kantor, Rami; DeLong, Allison; Schreier, Leeann; Istrail, Sorin

    Next generation sequencing technologies have recently been applied to characterize mutational spectra of the heterogeneous population of viral genotypes (known as a quasispecies) within HIV-infected patients. Such information is clinically relevant because minority genetic subpopulations of HIV within patients enable viral escape from selection pressures such as the immune response and antiretroviral therapy. However, methods for quasispecies sequence reconstruction from next generation sequencing reads are not yet widely used and remains an emerging area of research. Furthermore, the majority of research methodology in HIV has focused on 454 sequencing, while many next-generation sequencing platforms used in practice are limited to shorter read lengths relative to 454 sequencing. Little work has been done in determining how best to address the read length limitations of other platforms. The approach described here incorporates graph representations of both read differences and read overlap to conservatively determine the regions of the sequence with sufficient variability to separate quasispecies sequences. Within these tractable regions of quasispecies inference, we use constraint programming to solve for an optimal quasispecies subsequence determination via vertex coloring of the conflict graph, a representation which also lends itself to data with non-contiguous reads such as paired-end sequencing. We demonstrate the utility of the method by applying it to simulations based on actual intra-patient clonal HIV-1 sequencing data. PMID:23202421

  5. Jack bean α-mannosidase: amino acid sequencing and N-glycosylation analysis of a valuable glycomics tool.

    PubMed

    Gnanesh Kumar, B S; Pohlentz, Gottfried; Schulte, Mona; Mormann, Michael; Siva Kumar, Nadimpalli

    2014-03-01

    Jack bean (Canavalia ensiformis) seeds contain several biologically important proteins among which α-mannosidase (EC 3.2.1.24) has been purified, its biochemical properties studied and widely used in glycan analysis. In the present study, we have used the purified enzyme and derived its amino acid sequence covering both the known subunits (molecular mass of ∼66,000 and ∼44,000 Da) hitherto not known in its entirety. Peptide de novo sequencing and structural elucidation of N-glycopeptides obtained either directly from proteolytic digestion or after zwitterionic hydrophilic interaction liquid chromatography solid phase extraction-based separation were performed by use of nanoelectrospray ionization quadrupole time-of-flight mass spectrometry and low-energy collision-induced dissociation experiments. De novo sequencing provided new insights into the disulfide linkage organization, intersection of subunits and complete N-glycan structures along with site specificities. The primary sequence suggests that the enzyme belongs to glycosyl hydrolase family 38 and the N-glycan sequence analysis revealed high-mannose oligosaccharides, which were found to be heterogeneous with varying number of hexoses viz, Man8-9GlcNAc2 and Glc1Man9GlcNAc2 in an evolutionarily conserved N-glycosylation site. This site with two proximal cysteines is present in all the acidic α-mannosidases reported so far in eukaryotes. Further, a truncated paucimannose type was identified to be lacking terminal two mannose, Man1(Xyl)GlcNAc2 (Fuc). PMID:24295789

  6. Detection and isolation of nucleic acid sequences using a bifunctional hybridization probe

    DOEpatents

    Lucas, Joe N.; Straume, Tore; Bogen, Kenneth T.

    2000-01-01

    A method for detecting and isolating a target sequence in a sample of nucleic acids is provided using a bifunctional hybridization probe capable of hybridizing to the target sequence that includes a detectable marker and a first complexing agent capable of forming a binding pair with a second complexing agent. A kit is also provided for detecting a target sequence in a sample of nucleic acids using a bifunctional hybridization probe according to this method.

  7. Amino Acids of Conserved Kinase Motifs of Cytomegalovirus Protein UL97 Are Essential for Autophosphorylation

    PubMed Central

    Michel, Detlef; Kramer, Silke; Höhn, Simone; Schaarschmidt, Peter; Wunderlich, Kirsten; Mertens, Thomas

    1999-01-01

    Thirteen point mutations targeting predicted domains conserved in homologous protein kinases were introduced into the UL97 coding region of the human cytomegalovirus. All mutagenized proteins were expressed in cells infected with recombinant vaccinia viruses (rVV). Several mutations drastically reduced ganciclovir (GCV) phosphorylation. Mutations at amino acids G340, A442, L446, and F523 resulted in a complete loss of pUL97 phosphorylation, which was strictly associated with a loss of GCV phosphorylation. Our results confirm that in rVV-infected cells pUL97 phosphorylation is due to autophosphorylation and show that several amino acids conserved within domains of protein kinases are essential for this pUL97 phosphorylation. GCV phosphorylation is dependent on pUL97 phosphorylation. PMID:10482650

  8. Novel sequences encoding venom C-type lectins are conserved in phylogenetically and geographically distinct Echis and Bitis viper species.

    PubMed

    Harrison, R A; Oliver, J; Hasson, S S; Bharati, K; Theakston, R D G

    2003-10-01

    Envenoming by Echis saw scaled vipers and Bitis arietans puff adders is the leading cause of death and morbidity in Africa due to snake bite. Despite their medical importance, the composition and constituent functionality of venoms from these vipers remains poorly understood. Here, we report the cloning of cDNA sequences encoding seven clusters or isoforms of the haemostasis-disruptive C-type lectin (CTL) proteins from the venom glands of Echis ocellatus, E. pyramidum leakeyi, E. carinatus sochureki and B. arietans. All these CTL sequences encoded the cysteine scaffold that defines the carbohydrate-recognition domain of mammalian CTLs. All but one of the Echis and Bitis CTL sequences showed greater sequence similarity to the beta than alpha CTL subunits in venoms of related Asian and American vipers. Four of the new CTL clusters showed marked inter-cluster sequence conservation across all four viper species which were significantly different from that of previously published viper CTLs. The other three Echis and Bitis CTL clusters showed varying degrees of sequence similarity to published viper venom CTLs. Because viper venom CTLs exhibit a high degree of sequence similarity and yet exert profoundly different effects on the mammalian haemostatic system, no attempt was made to assign functionality to the new Echis and Bitis CTLs on the basis of sequence alone. The extraordinary level of inter-specific and inter-generic sequence conservation exhibited by the Echis and Bitis CTLs leads us to speculate that antibodies to representative molecules should neutralise the biological function of this important group of venom toxins in vipers that are distributed throughout Africa, the Middle East and the Indian subcontinent. PMID:14557069

  9. Two evolutionarily conserved sequence elements for Peg3/Usp29 transcription

    PubMed Central

    Kim, Jeong Do; Yu, Sungryul; Choo, Jung Ha; Kim, Joomyeong

    2008-01-01

    Background Two evolutionarily Conserved Sequence Elements, CSE1 and CSE2 (YY1 binding sites), are found within the 3.8-kb CpG island surrounding the bidirectional promoter of two imprinted genes, Peg3 (Paternally expressed gene 3) and Usp29 (Ubiquitin-specific protease 29). This CpG island is a likely ICR (Imprinting Control Region) that controls transcription of the 500-kb genomic region of the Peg3 imprinted domain. Results The current study investigated the functional roles of CSE1 and CSE2 in the transcriptional control of the two genes, Peg3 and Usp29, using cell line-based promoter assays. The mutation of 6 YY1 binding sites (CSE2) reduced the transcriptional activity of the bidirectional promoter in the Peg3 direction in an orientation-dependent manner, suggesting an activator role for CSE2 (YY1 binding sites). However, the activity in the Usp29 direction was not detectable regardless of the presence/absence of YY1 binding sites. In contrast, mutation of CSE1 increased the transcriptional activity of the promoter in both the Peg3 and Usp29 directions, suggesting a potential repressor role for CSE1. The observed repression by CSE1 was also orientation-dependent. Serial mutational analyses further narrowed down two separate 6-bp-long regions within the 42-bp-long CSE1 which are individually responsible for the repression of Peg3 and Usp29. Conclusion CSE2 (YY1 binding sites) functions as an activator for Peg3 transcription, while CSE1 acts as a repressor for the transcription of both Peg3 and Usp29. PMID:19068137

  10. Use of a Drosophila Genome-Wide Conserved Sequence Database to Identify Functionally Related cis-Regulatory Enhancers

    PubMed Central

    Brody, Thomas; Yavatkar, Amarendra S; Kuzin, Alexander; Kundu, Mukta; Tyson, Leonard J; Ross, Jermaine; Lin, Tzu-Yang; Lee, Chi-Hon; Awasaki, Takeshi; Lee, Tzumin; Odenwald, Ward F

    2012-01-01

    Background: Phylogenetic footprinting has revealed that cis-regulatory enhancers consist of conserved DNA sequence clusters (CSCs). Currently, there is no systematic approach for enhancer discovery and analysis that takes full-advantage of the sequence information within enhancer CSCs. Results: We have generated a Drosophila genome-wide database of conserved DNA consisting of >100,000 CSCs derived from EvoPrints spanning over 90% of the genome. cis-Decoder database search and alignment algorithms enable the discovery of functionally related enhancers. The program first identifies conserved repeat elements within an input enhancer and then searches the database for CSCs that score highly against the input CSC. Scoring is based on shared repeats as well as uniquely shared matches, and includes measures of the balance of shared elements, a diagnostic that has proven to be useful in predicting cis-regulatory function. To demonstrate the utility of these tools, a temporally-restricted CNS neuroblast enhancer was used to identify other functionally related enhancers and analyze their structural organization. Conclusions: cis-Decoder reveals that co-regulating enhancers consist of combinations of overlapping shared sequence elements, providing insights into the mode of integration of multiple regulating transcription factors. The database and accompanying algorithms should prove useful in the discovery and analysis of enhancers involved in any developmental process. Developmental Dynamics 241:169–189, 2012. © 2011 Wiley Periodicals, Inc. Key findings A genome-wide catalog of Drosophila conserved DNA sequence clusters. cis-Decoder discovers functionally related enhancers. Functionally related enhancers share balanced sequence element copy numbers. Many enhancers function during multiple phases of development. PMID:22174086

  11. Sequence conservation of the rad21 Schizosaccharomyces pombe DNA double-strand break repair gene in human and mouse

    SciTech Connect

    McKay, M.J.; Troelstra, C.; Kanaar, R.

    1996-09-01

    The rad21 gene of Schizosaccharomyces pombe is involved in the repair of ionizing radiation-induced DNA double-strand breaks. The isolation of mouse and human putative homologs of rad21 is reported here. Alignment of the predicted amino acid sequence of Rad21 with the mammalian proteins showed that the similarity was distributed across the length of the proteins, with more highly conserved regions at both termini. The mHR21{sup sp} (mouse homolog of Rad21, S. pombe) and hHR21{sup sp} (human homolog of Rad21, S. pombe) predicted proteins were 96% identical, whereas the human and S. pombe proteins were 25% identical and 47% similar. RNA blot analysis showed that mHR21{sup sp} mRNA was abundant in all adult mouse tissues examined, with highest expression in testis and thymus. In addition to a 3.1-kb constitutive mRNA transcript, a 2.2-kb transcript was present at a high level in postmeiotic spermatids, while expression of the 3.1-kb mRNA in testis was confined to the meiotic compartment. hHR21{sup sp} mRNA was cell-cycle regulated in human cells, increasing in late S phase to a peak in G2 phase. The level of hHR21{sup sp} transcripts was not altered by exposure of normal diploid fibroblasts to 10 Gy ionizing radiation. In situ hybridization showed that mHR21{sup sp} resided on chromosome 15D3, whereas hHR21{sup sp} localized to the syntenic 8q24 region. Elevated expression of mHR21{sup sp} in testis and thymus supports a possible role for the rad21 mammalian homologs in V(D)J and meiotic recombination, respectively. Cell cycle regulation of rad21, retained from S. pombe to human, is consistent with a conservation of function between S. pombe and human rad21 genes. 62 refs., 8 figs., 1 tab.

  12. Conservation of the gene for outer membrane protein OprF in the family Pseudomonadaceae: sequence of the Pseudomonas syringae oprF gene.

    PubMed Central

    Ullstrom, C A; Siehnel, R; Woodruff, W; Steinbach, S; Hancock, R E

    1991-01-01

    The conservation of the oprF gene for the major outer membrane protein OprF was determined by restriction mapping and Southern blot hybridization with the Pseudomonas aeruginosa oprF gene as a probe. The restriction map was highly conserved among 16 of the 17 serotype strains and 42 clinical isolates of P. aeruginosa. Only the serotype 12 isolate and one clinical isolate showed small differences in restriction pattern. Southern probing of PstI chromosomal digests of 14 species from the family Pseudomonadaceae revealed that only the nine members of rRNA homology group I hybridized with the oprF gene. To reveal the actual extent of homology, the oprF gene and its product were characterized in Pseudomonas syringae. Nine strains of P. syringae from seven different pathovars hybridized with the P. aeruginosa gene to produce five different but related restriction maps. All produced an OprF protein in their outer membranes with the same apparent molecular weight as that of P.aeruginosa OprF. In each case the protein reacted with monoclonal antibody MA4-10 and was similarly heat and 2-mercaptoethanol modifiable. The purified OprF protein of the type strain P. syringae pv. syringae ATCC 19310 reconstituted small channels in lipid bilayer membranes. The oprF gene from this latter strain was cloned and sequenced. Despite the low level of DNA hybridization between P. aeruginosa and P. syringae DNA, the OprF gene was highly conserved between the species with 72% DNA sequence identity and 68% amino acid sequence identity overall. The carboxy terminus-encoding region of P. syringae oprF showed 85 and 33% identity, respectively, with the same regions of the P. aeruginosa oprF and Escherichia coli ompA genes. Images PMID:1898935

  13. Dominant sequences of human major histocompatibility complex conserved extended haplotypes from HLA-DQA2 to DAXX.

    PubMed

    Larsen, Charles E; Alford, Dennis R; Trautwein, Michael R; Jalloh, Yanoh K; Tarnacki, Jennifer L; Kunnenkeri, Sushruta K; Fici, Dolores A; Yunis, Edmond J; Awdeh, Zuheir L; Alper, Chester A

    2014-10-01

    We resequenced and phased 27 kb of DNA within 580 kb of the MHC class II region in 158 population chromosomes, most of which were conserved extended haplotypes (CEHs) of European descent or contained their centromeric fragments. We determined the single nucleotide polymorphism and deletion-insertion polymorphism alleles of the dominant sequences from HLA-DQA2 to DAXX for these CEHs. Nine of 13 CEHs remained sufficiently intact to possess a dominant sequence extending at least to DAXX, 230 kb centromeric to HLA-DPB1. We identified the regions centromeric to HLA-DQB1 within which single instances of eight "common" European MHC haplotypes previously sequenced by the MHC Haplotype Project (MHP) were representative of those dominant CEH sequences. Only two MHP haplotypes had a dominant CEH sequence throughout the centromeric and extended class II region and one MHP haplotype did not represent a known European CEH anywhere in the region. We identified the centromeric recombination transition points of other MHP sequences from CEH representation to non-representation. Several CEH pairs or groups shared sequence identity in small blocks but had significantly different (although still conserved for each separate CEH) sequences in surrounding regions. These patterns partly explain strong calculated linkage disequilibrium over only short (tens to hundreds of kilobases) distances in the context of a finite number of observed megabase-length CEHs comprising half a population's haplotypes. Our results provide a clearer picture of European CEH class II allelic structure and population haplotype architecture, improved regional CEH markers, and raise questions concerning regional recombination hotspots. PMID:25299700

  14. Dominant Sequences of Human Major Histocompatibility Complex Conserved Extended Haplotypes from HLA-DQA2 to DAXX

    PubMed Central

    Larsen, Charles E.; Alford, Dennis R.; Trautwein, Michael R.; Jalloh, Yanoh K.; Tarnacki, Jennifer L.; Kunnenkeri, Sushruta K.; Fici, Dolores A.; Yunis, Edmond J.; Awdeh, Zuheir L.; Alper, Chester A.

    2014-01-01

    We resequenced and phased 27 kb of DNA within 580 kb of the MHC class II region in 158 population chromosomes, most of which were conserved extended haplotypes (CEHs) of European descent or contained their centromeric fragments. We determined the single nucleotide polymorphism and deletion-insertion polymorphism alleles of the dominant sequences from HLA-DQA2 to DAXX for these CEHs. Nine of 13 CEHs remained sufficiently intact to possess a dominant sequence extending at least to DAXX, 230 kb centromeric to HLA-DPB1. We identified the regions centromeric to HLA-DQB1 within which single instances of eight “common” European MHC haplotypes previously sequenced by the MHC Haplotype Project (MHP) were representative of those dominant CEH sequences. Only two MHP haplotypes had a dominant CEH sequence throughout the centromeric and extended class II region and one MHP haplotype did not represent a known European CEH anywhere in the region. We identified the centromeric recombination transition points of other MHP sequences from CEH representation to non-representation. Several CEH pairs or groups shared sequence identity in small blocks but had significantly different (although still conserved for each separate CEH) sequences in surrounding regions. These patterns partly explain strong calculated linkage disequilibrium over only short (tens to hundreds of kilobases) distances in the context of a finite number of observed megabase-length CEHs comprising half a population's haplotypes. Our results provide a clearer picture of European CEH class II allelic structure and population haplotype architecture, improved regional CEH markers, and raise questions concerning regional recombination hotspots. PMID:25299700

  15. Biochemical Roles for Conserved Residues in the Bacterial Fatty Acid-binding Protein Family.

    PubMed

    Broussard, Tyler C; Miller, Darcie J; Jackson, Pamela; Nourse, Amanda; White, Stephen W; Rock, Charles O

    2016-03-18

    Fatty acid kinase (Fak) is a ubiquitous Gram-positive bacterial enzyme consisting of an ATP-binding protein (FakA) that phosphorylates the fatty acid bound to FakB. In Staphylococcus aureus, Fak is a global regulator of virulence factor transcription and is essential for the activation of exogenous fatty acids for incorporation into phospholipids. The 1.2-Å x-ray structure of S. aureus FakB2, activity assays, solution studies, site-directed mutagenesis, and in vivo complementation were used to define the functions of the five conserved residues that define the FakB protein family (Pfam02645). The fatty acid tail is buried within the protein, and the exposed carboxyl group is bound by a Ser-93-fatty acid carboxyl-Thr-61-His-266 hydrogen bond network. The guanidinium of the invariant Arg-170 is positioned to potentially interact with a bound acylphosphate. The reduced thermal denaturation temperatures of the T61A, S93A, and H266A FakB2 mutants illustrate the importance of the hydrogen bond network in protein stability. The FakB2 T61A, S93A, and H266A mutants are 1000-fold less active in the Fak assay, and the R170A mutant is completely inactive. All FakB2 mutants form FakA(FakB2)2 complexes except FakB2(R202A), which is deficient in FakA binding. Allelic replacement shows that strains expressing FakB2 mutants are defective in fatty acid incorporation into phospholipids and virulence gene transcription. These conserved residues are likely to perform the same critical functions in all bacterial fatty acid-binding proteins. PMID:26774272

  16. Partial amino acid sequence of human factor D:homology with serine proteases.

    PubMed Central

    Volanakis, J E; Bhown, A; Bennett, J C; Mole, J E

    1980-01-01

    Human factor D purified to homogeneity by a modified procedure was subjected to NH2-terminal amino acid sequence analysis by using a modified automated Beckman sequencer. We identified 48 of the first 57 NH2-terminal amino acids in a single sequencer run, using microgram quantities of factor D. The deduced amino acid sequence represents approximately 25% of the primary structure of factor D. This extended NH2-terminal amino acid sequence of factor D was compared to that of other trypsin-related serine proteases. By visual inspection, strong homologies (33--50% identity) were observed with all the serine proteases included in the comparison. Interestingly, factor D showed a higher degree of homology to serine proteases of pancreatic origin than to those of serum origin. Images PMID:6987665

  17. High-throughput genomic sequencing of cassava bacterial blight strains identifies conserved effectors to target for durable resistance.

    PubMed

    Bart, Rebecca; Cohn, Megan; Kassen, Andrew; McCallum, Emily J; Shybut, Mikel; Petriello, Annalise; Krasileva, Ksenia; Dahlbeck, Douglas; Medina, Cesar; Alicai, Titus; Kumar, Lava; Moreira, Leandro M; Rodrigues Neto, Júlio; Verdier, Valerie; Santana, María Angélica; Kositcharoenkul, Nuttima; Vanderschuren, Hervé; Gruissem, Wilhelm; Bernal, Adriana; Staskawicz, Brian J

    2012-07-10

    Cassava bacterial blight (CBB), incited by Xanthomonas axonopodis pv. manihotis (Xam), is the most important bacterial disease of cassava, a staple food source for millions of people in developing countries. Here we present a widely applicable strategy for elucidating the virulence components of a pathogen population. We report Illumina-based draft genomes for 65 Xam strains and deduce the phylogenetic relatedness of Xam across the areas where cassava is grown. Using an extensive database of effector proteins from animal and plant pathogens, we identify the effector repertoire for each sequenced strain and use a comparative sequence analysis to deduce the least polymorphic of the conserved effectors. These highly conserved effectors have been maintained over 11 countries, three continents, and 70 y of evolution and as such represent ideal targets for developing resistance strategies. PMID:22699502

  18. Amino acid sequence of Japanese quail (Coturnix japonica) and northern bobwhite (Colinus virginianus) myoglobin.

    PubMed

    Goodson, John; Beckstead, Robert B; Payne, Jason; Singh, Rakesh K; Mohan, Anand

    2015-08-15

    Myoglobin has an important physiological role in vertebrates, and as the primary sarcoplasmic pigment in meat, influences quality perception and consumer acceptability. In this study, the amino acid sequences of Japanese quail and northern bobwhite myoglobin were deduced by cDNA cloning of the coding sequence from mRNA. Japanese quail myoglobin was isolated from quail cardiac muscles, purified using ammonium sulphate precipitation and gel-filtration, and subjected to multiple enzymatic digestions. Mass spectrometry corroborated the deduced protein amino acid sequence at the protein level. Sequence analysis revealed both species' myoglobin structures consist of 153 amino acids, differing at only three positions. When compared with chicken myoglobin, Japanese quail showed 98% sequence identity, and northern bobwhite 97% sequence identity. The myoglobin in both quail species contained eight histidine residues instead of the nine present in chicken and turkey. PMID:25794748

  19. Identification of conserved genomic regions and variation therein amongst Cetartiodactyla species using next generation sequencing

    Technology Transfer Automated Retrieval System (TEKTRAN)

    Background Next Generation Sequencing has created an opportunity to genetically characterize an individual both inexpensively and comprehensively. In earlier work produced in our collaboration [1], it was demonstrated that, for animals without a reference genome, their Next Generation Sequence data ...

  20. Identification of random nucleic acid sequence aberrations using dual capture probes which hybridize to different chromosome regions

    DOEpatents

    Lucas, Joe N.; Straume, Tore; Bogen, Kenneth T.

    1998-01-01

    A method is provided for detecting nucleic acid sequence aberrations using two immobilization steps. According to the method, a nucleic acid sequence aberration is detected by detecting nucleic acid sequences having both a first nucleic acid sequence type (e.g., from a first chromosome) and a second nucleic acid sequence type (e.g., from a second chromosome), the presence of the first and the second nucleic acid sequence type on the same nucleic acid sequence indicating the presence of a nucleic acid sequence aberration. In the method, immobilization of a first hybridization probe is used to isolate a first set of nucleic acids in the sample which contain the first nucleic acid sequence type. Immobilization of a second hybridization probe is then used to isolate a second set of nucleic acids from within the first set of nucleic acids which contain the second nucleic acid sequence type. The second set of nucleic acids are then detected, their presence indicating the presence of a nucleic acid sequence aberration.

  1. Identification of random nucleic acid sequence aberrations using dual capture probes which hybridize to different chromosome regions

    DOEpatents

    Lucas, J.N.; Straume, T.; Bogen, K.T.

    1998-03-24

    A method is provided for detecting nucleic acid sequence aberrations using two immobilization steps. According to the method, a nucleic acid sequence aberration is detected by detecting nucleic acid sequences having both a first nucleic acid sequence type (e.g., from a first chromosome) and a second nucleic acid sequence type (e.g., from a second chromosome), the presence of the first and the second nucleic acid sequence type on the same nucleic acid sequence indicating the presence of a nucleic acid sequence aberration. In the method, immobilization of a first hybridization probe is used to isolate a first set of nucleic acids in the sample which contain the first nucleic acid sequence type. Immobilization of a second hybridization probe is then used to isolate a second set of nucleic acids from within the first set of nucleic acids which contain the second nucleic acid sequence type. The second set of nucleic acids are then detected, their presence indicating the presence of a nucleic acid sequence aberration. 14 figs.

  2. Conserved sequence-specific lincRNA-steroid receptor interactions drive transcriptional repression and direct cell fate

    PubMed Central

    Hudson, William H.; Pickard, Mark R.; de Vera, Ian Mitchelle S.; Kuiper, Emily G.; Mourtada-Maarabouni, Mirna; Conn, Graeme L.; Kojetin, Douglas J.; Williams, Gwyn T.; Ortlund, Eric A.

    2014-01-01

    The majority of the eukaryotic genome is transcribed, generating a significant number of long intergenic non-coding RNAs (lincRNAs). While lincRNAs represent the most poorly understood product of transcription, recent work has shown lincRNAs fulfill important cellular functions. In addition to low sequence conservation, poor understanding of structural mechanisms driving lincRNA biology hinders systematic prediction of their function. Here, we report the molecular requirements for the recognition of steroid receptors (SRs) by the lincRNA Gas5, which regulates steroid-mediated transcriptional regulation, growth arrest, and apoptosis. We identify the functional Gas5-SR interface and generate point mutations that ablate the SR-Gas5 lincRNA interaction, altering Gas5-driven apoptosis in cancer cell lines. Further, we find that the Gas5 SR-recognition sequence is conserved among haplorhines, with its evolutionary origin as a splice acceptor site. This study demonstrates that lincRNAs can recognize protein targets in a conserved, sequence-specific manner in order to affect critical cell functions. PMID:25377354

  3. Conserved sequence-specific lincRNA-steroid receptor interactions drive transcriptional repression and direct cell fate

    SciTech Connect

    Hudson, William H.; Pickard, Mark R.; de Vera, Ian Mitchelle S.; Kuiper, Emily G.; Mourtada-Maarabouni, Mirna; Conn, Graeme L.; Kojetin, Douglas J.; Williams, Gwyn T.; Ortlund, Eric A.

    2014-12-23

    The majority of the eukaryotic genome is transcribed, generating a significant number of long intergenic noncoding RNAs (lincRNAs). Although lincRNAs represent the most poorly understood product of transcription, recent work has shown lincRNAs fulfill important cellular functions. In addition to low sequence conservation, poor understanding of structural mechanisms driving lincRNA biology hinders systematic prediction of their function. Here we report the molecular requirements for the recognition of steroid receptors (SRs) by the lincRNA growth arrest-specific 5 (Gas5), which regulates steroid-mediated transcriptional regulation, growth arrest and apoptosis. We identify the functional Gas5-SR interface and generate point mutations that ablate the SR-Gas5 lincRNA interaction, altering Gas5-driven apoptosis in cancer cell lines. Further, we find that the Gas5 SR-recognition sequence is conserved among haplorhines, with its evolutionary origin as a splice acceptor site. This study demonstrates that lincRNAs can recognize protein targets in a conserved, sequence-specific manner in order to affect critical cell functions.

  4. Control regions for chromosome replication are conserved with respect to sequence and location among Escherichia coli strains

    PubMed Central

    Frimodt-Møller, Jakob; Charbon, Godefroid; Krogfelt, Karen A.; Løbner-Olesen, Anders

    2015-01-01

    In Escherichia coli, chromosome replication is initiated from oriC by the DnaA initiator protein associated with ATP. Three non-coding regions contribute to the activity of DnaA. The datA locus is instrumental in conversion of DnaAATP to DnaAADP (datA dependent DnaAATP hydrolysis) whereas DnaA rejuvenation sequences 1 and 2 (DARS1 and DARS2) reactivate DnaAADP to DnaAATP. The structural organization of oriC, datA, DARS1, and DARS2 were found conserved among 59 fully sequenced E. coli genomes, with differences primarily in the non-functional spacer regions between key protein binding sites. The relative distances from oriC to datA, DARS1, and DARS2, respectively, was also conserved despite of large variations in genome size, suggesting that the gene dosage of either region is important for bacterial growth. Yet all three regions could be deleted alone or in combination without loss of viability. Competition experiments during balanced growth in rich medium and during mouse colonization indicated roles of datA, DARS1, and DARS2 for bacterial fitness although the relative contribution of each region differed between growth conditions. We suggest that this fitness advantage has contributed to conservation of both sequence and chromosomal location for datA, DARS1, and DARS2. PMID:26441936

  5. tax and rex Sequences of bovine leukaemia virus from globally diverse isolates: rex amino acid sequence more variable than tax.

    PubMed

    McGirr, K M; Buehring, G C

    2005-02-01

    Bovine leukaemia virus (BLV) is an important agricultural problem with high costs to the dairy industry. Here, we examine the variation of the tax and rex genes of BLV. The tax and rex genes share 420 bases and have overlapping reading frames. The tax gene encodes a protein that functions as a transactivator of the BLV promoter, is required for viral replication, acts on cellular promoters, and is responsible for oncogenesis. The rex facilitates the export of viral mRNAs from the nucleus and regulates transcription. We have sequenced five new isolates of the tax/rex gene. We examined the five new and three previously published tax/rex DNA and predicted amino acid sequences of BLV isolates from cattle in representative regions worldwide. The highest variation among nucleic acid sequences for tax and rex was 7% and 5%, respectively; among predicted amino acid sequences for Tax and Rex, 9% and 11%, respectively. Significantly more nucleotide changes resulted in predicted amino acid changes in the rex gene than in the tax gene (P < or = 0.0006). This variability is higher than previously reported for any region of the viral genome. This research may also have implications for the development of Tax-based vaccines. PMID:15702995

  6. The amino acid sequence of protein CM-3 from Dendroaspis polylepis polylepis (black mamba) venom.

    PubMed

    Joubert, F J

    1985-01-01

    Protein CM-3 from Dendroaspis polylepis polylepis venom was purified by gel filtration and ion exchange chromatography. It comprises 65 amino acids including eight half-cystines. The complete amino acid sequence of protein CM-3 has been elucidated. The sequence (residues 1-50) resembles that of the N-terminal sequence of the subunits of a synergistic type protein and residues 51-65 that of the C-terminal sequence of an angusticeps type protein. Mixtures of protein CM-3 and angusticeps type proteins showed no apparent synergistic effect, in that their toxicity in combination was no greater than the sum of their individual toxicities. PMID:4029488

  7. Fad7 gene identification and fatty acids phenotypic variation in an olive collection by EcoTILLING and sequencing approaches.

    PubMed

    Sabetta, Wilma; Blanco, Antonio; Zelasco, Samanta; Lombardo, Luca; Perri, Enzo; Mangini, Giacomo; Montemurro, Cinzia

    2013-08-01

    The ω-3 fatty acid desaturases (FADs) are enzymes responsible for catalyzing the conversion of linoleic acid to α-linolenic acid localized in the plastid or in the endoplasmic reticulum. In this research we report the genotypic and phenotypic variation of Italian Olea europaea L. germoplasm for the fatty acid composition. The phenotypic oil characterization was followed by the molecular analysis of the plastidial-type ω-3 FAD gene (fad7) (EC 1.14.19), whose full-length sequence has been here identified in cultivar Leccino. The gene consisted of 2635 bp with 8 exons and 5'- and 3'-UTRs of 336 and 282 bp respectively, and showed a high level of heterozygousity (1/110 bp). The natural allelic variation was investigated both by a LiCOR EcoTILLING assay and the PCR product direct sequencing. Only three haplotypes were identified among the 96 analysed cultivars, highlighting the strong degree of conservation of this gene. PMID:23685785

  8. Violation of an evolutionarily conserved immunoglobulin diversity gene sequence preference promotes production of dsDNA-specific IgG antibodies.

    PubMed

    Silva-Sanchez, Aaron; Liu, Cun Ren; Vale, Andre M; Khass, Mohamed; Kapoor, Pratibha; Elgavish, Ada; Ivanov, Ivaylo I; Ippolito, Gregory C; Schelonka, Robert L; Schoeb, Trenton R; Burrows, Peter D; Schroeder, Harry W

    2015-01-01

    Variability in the developing antibody repertoire is focused on the third complementarity determining region of the H chain (CDR-H3), which lies at the center of the antigen binding site where it often plays a decisive role in antigen binding. The power of VDJ recombination and N nucleotide addition has led to the common conception that the sequence of CDR-H3 is unrestricted in its variability and random in its composition. Under this view, the immune response is solely controlled by somatic positive and negative clonal selection mechanisms that act on individual B cells to promote production of protective antibodies and prevent the production of self-reactive antibodies. This concept of a repertoire of random antigen binding sites is inconsistent with the observation that diversity (DH) gene segment sequence content by reading frame (RF) is evolutionarily conserved, creating biases in the prevalence and distribution of individual amino acids in CDR-H3. For example, arginine, which is often found in the CDR-H3 of dsDNA binding autoantibodies, is under-represented in the commonly used DH RFs rearranged by deletion, but is a frequent component of rarely used inverted RF1 (iRF1), which is rearranged by inversion. To determine the effect of altering this germline bias in DH gene segment sequence on autoantibody production, we generated mice that by genetic manipulation are forced to utilize an iRF1 sequence encoding two arginines. Over a one year period we collected serial serum samples from these unimmunized, specific pathogen-free mice and found that more than one-fifth of them contained elevated levels of dsDNA-binding IgG, but not IgM; whereas mice with a wild type DH sequence did not. Thus, germline bias against the use of arginine enriched DH sequence helps to reduce the likelihood of producing self-reactive antibodies. PMID:25706374

  9. Genomic Locations of Conserved Noncoding Sequences and Their Proximal Protein-Coding Genes in Mammalian Expression Dynamics.

    PubMed

    Babarinde, Isaac Adeyemi; Saitou, Naruya

    2016-07-01

    Experimental studies have found the involvement of certain conserved noncoding sequences (CNSs) in the regulation of the proximal protein-coding genes in mammals. However, reported cases of long range enhancer activities and inter-chromosomal regulation suggest that proximity of CNSs to protein-coding genes might not be important for regulation. To test the importance of the CNS genomic location, we extracted the CNSs conserved between chicken and four mammalian species (human, mouse, dog, and cattle). These CNSs were confirmed to be under purifying selection. The intergenic CNSs are often found in clusters in gene deserts, where protein-coding genes are in paucity. The distribution pattern, ChIP-Seq, and RNA-Seq data suggested that the CNSs are more likely to be regulatory elements and not corresponding to long intergenic noncoding RNAs. Physical distances between CNS and their nearest protein coding genes were well conserved between human and mouse genomes, and CNS-flanking genes were often found in evolutionarily conserved genomic neighborhoods. ChIP-Seq signal and gene expression patterns also suggested that CNSs regulate nearby genes. Interestingly, genes with more CNSs have more evolutionarily conserved expression than those with fewer CNSs. These computationally obtained results suggest that the genomic locations of CNSs are important for their regulatory functions. In fact, various kinds of evolutionary constraints may be acting to maintain the genomic locations of CNSs and protein-coding genes in mammals to ensure proper regulation. PMID:27017584

  10. PCR-based study of conserved and variable DNA sequences of Tritrichomonas foetus isolates from Saskatchewan, Canada.

    PubMed Central

    Riley, D E; Wagner, B; Polley, L; Krieger, J N

    1995-01-01

    The protozoan parasite Tritrichomonas foetus causes infertility and spontaneous abortion in cattle. In Saskatchewan, Canada, the culture prevalence of trichomonads was 65 of 1,048 (6%) among 1,048 bulls tested within a 1-year period ending in April 1994. Saskatchewan was previously thought to be free of the parasite. To confirm the culture results, possible T. foetus DNA presence was determined by the PCR. All of the 16 culture-positive isolates tested were PCR positive by a single-band test, but one PCR product was weak. DNA fingerprinting by both T17 PCR and randomly amplified polymorphic DNA PCR revealed genetic variation or polymorphism among the T. foetus isolates. T17 PCR also revealed conserved loci that distinguished these T. foetus isolates from Trichomonas vaginalis, from a variety of other protozoa, and from prokaryotes. TCO-1 PCR, a PCR test designed to sample DNA sequence homologous to the 5' flank of a highly conserved cell division control gene, detected genetic polymorphism at low stringency and a conserved, single locus at higher stringency. These findings suggested that T. foetus isolates exhibit both conserved genetic loci and polymorphic loci detectable by independent PCR methods. Both conserved and polymorphic genetic loci may prove useful for improved clinical diagnosis of T. foetus. The polymorphic loci detected by PCR suggested either a long history of infection or multiple lines of T. foetus infection in Saskatchewan. Polymorphic loci detected by PCR may provide data for epidemiologic studies of T. foetus. PMID:7615746

  11. A conserved patch of hydrophobic amino acids modulates Myb activity by mediating protein-protein interactions.

    PubMed

    Dukare, Sandeep; Klempnauer, Karl-Heinz

    2016-07-01

    The transcription factor c-Myb plays a key role in the control of proliferation and differentiation in hematopoietic progenitor cells and has been implicated in the development of leukemia and certain non-hematopoietic tumors. c-Myb activity is highly dependent on the interaction with the coactivator p300 which is mediated by the transactivation domain of c-Myb and the KIX domain of p300. We have previously observed that conservative valine-to-isoleucine amino acid substitutions in a conserved stretch of hydrophobic amino acids have a profound effect on Myb activity. Here, we have explored the function of the hydrophobic region as a mediator of protein-protein interactions. We show that the hydrophobic region facilitates Myb self-interaction and binding of the histone acetyl transferase Tip60, a previously identified Myb interacting protein. We show that these interactions are affected by the valine-to-isoleucine amino acid substitutions and suppress Myb activity by interfering with the interaction of Myb and the KIX domain of p300. Taken together, our work identifies the hydrophobic region in the Myb transactivation domain as a binding site for homo- and heteromeric protein interactions and leads to a picture of the c-Myb transactivation domain as a composite protein binding region that facilitates interdependent protein-protein interactions of Myb with regulatory proteins. PMID:27080133

  12. Structural analysis of complementary DNA and amino acid sequences of human and rat androgen receptors

    SciTech Connect

    Chang, C.; Kokontis, J.; Liao, S. )

    1988-10-01

    Structural analysis of cDNAs for human and rat androgen receptors (ARs) indicates that the amino-terminal regions of ARs are rich in oligo- and poly(amino acid) motifs as in some homeotic genes. The human AR has a long stretch of repeated glycines, whereas rat AR has a long stretch of glutamines. There is a considerable sequence similarity among ARs and the receptors for glucocorticoids, progestins, and mineralocorticoids within the steroid-binding domains. The cysteine-rich DNA-binding domains are well conserved. Translation of mRNA transcribed from AR cDNAs yielded 94- and 76-kDa proteins and smaller forms that bind to DNA and have high affinity toward androgens. These rat or human ARs were recognized by human autoantibodies to natural Ars. Molecular hybridization studies, using AR cDNAs as probes, indicated that the ventral prostate and other male accessory organs are rich in AR mRNA and that the production of AR mRNA in the target organs may be autoregulated by androgens.

  13. Snake venoms. The amino acid sequences of two proteinase inhibitor homologues from Dendroaspis angusticeps venom.

    PubMed

    Joubert, F J; Taljaard, N

    1980-05-01

    Toxins C13S1C3 and C13S2C3 from D. angusticeps venom were purified by gel filtration and ion exchange chromatography. Whereas C13S1C3 contains 57 amino acids, C13S2C3 contains 59 but each include six half-cystine residues. The complete primary structure of the low toxicity proteins have been elucidated. The sequences and the invariant residues of toxins C13S1C3 and C13S2C3 from D. angusticeps venom resemble, respectively, those of the proteinase inhibitor homologues K and I from D. polylepis polylepis venom and they are also homologous to the active proteinase inhibitors from various sources. In C13S1C3 and K the active site lysyl residue of active bovine pancreatic proteinase inhibitor is conserved but the site residue alanine, is replaced by lysine. In C13S2C3 and I the active site residue is replaced by tyrosine. PMID:7429422

  14. Computer Simulation of the Determination of Amino Acid Sequences in Polypeptides

    ERIC Educational Resources Information Center

    Daubert, Stephen D.; Sontum, Stephen F.

    1977-01-01

    Describes a computer program that generates a random string of amino acids and guides the student in determining the correct sequence of a given protein by using experimental analytic data for that protein. (MLH)

  15. Sequence-Based Screening for Rare Enzymes: New Insights into the World of AMDases Reveal a Conserved Motif and 58 Novel Enzymes Clustering in Eight Distinct Families

    PubMed Central

    Maimanakos, Janine; Chow, Jennifer; Gaßmeyer, Sarah K.; Güllert, Simon; Busch, Florian; Kourist, Robert; Streit, Wolfgang R.

    2016-01-01

    Arylmalonate Decarboxylases (AMDases, EC 4.1.1.76) are very rare and mostly underexplored enzymes. Currently only four known and biochemically characterized representatives exist. However, their ability to decarboxylate α-disubstituted malonic acid derivatives to optically pure products without cofactors makes them attractive and promising candidates for the use as biocatalysts in industrial processes. Until now, AMDases could not be separated from other members of the aspartate/glutamate racemase superfamily based on their gene sequences. Within this work, a search algorithm was developed that enables a reliable prediction of AMDase activity for potential candidates. Based on specific sequence patterns and screening methods 58 novel AMDase candidate genes could be identified in this work. Thereby, AMDases with the conserved sequence pattern of Bordetella bronchiseptica’s prototype appeared to be limited to the classes of Alpha-, Beta-, and Gamma-proteobacteria. Amino acid homologies and comparison of gene surrounding sequences enabled the classification of eight enzyme clusters. Particularly striking is the accumulation of genes coding for different transporters of the tripartite tricarboxylate transporters family, TRAP transporters and ABC transporters as well as genes coding for mandelate racemases/muconate lactonizing enzymes that might be involved in substrate uptake or degradation of AMDase products. Further, three novel AMDases were characterized which showed a high enantiomeric excess (>99%) of the (R)-enantiomer of flurbiprofen. These are the recombinant AmdA and AmdV from Variovorax sp. strains HH01 and HH02, originated from soil, and AmdP from Polymorphum gilvum found by a data base search. Altogether our findings give new insights into the class of AMDases and reveal many previously unknown enzyme candidates with high potential for bioindustrial processes. PMID:27610105

  16. Sequence-Based Screening for Rare Enzymes: New Insights into the World of AMDases Reveal a Conserved Motif and 58 Novel Enzymes Clustering in Eight Distinct Families.

    PubMed

    Maimanakos, Janine; Chow, Jennifer; Gaßmeyer, Sarah K; Güllert, Simon; Busch, Florian; Kourist, Robert; Streit, Wolfgang R

    2016-01-01

    Arylmalonate Decarboxylases (AMDases, EC 4.1.1.76) are very rare and mostly underexplored enzymes. Currently only four known and biochemically characterized representatives exist. However, their ability to decarboxylate α-disubstituted malonic acid derivatives to optically pure products without cofactors makes them attractive and promising candidates for the use as biocatalysts in industrial processes. Until now, AMDases could not be separated from other members of the aspartate/glutamate racemase superfamily based on their gene sequences. Within this work, a search algorithm was developed that enables a reliable prediction of AMDase activity for potential candidates. Based on specific sequence patterns and screening methods 58 novel AMDase candidate genes could be identified in this work. Thereby, AMDases with the conserved sequence pattern of Bordetella bronchiseptica's prototype appeared to be limited to the classes of Alpha-, Beta-, and Gamma-proteobacteria. Amino acid homologies and comparison of gene surrounding sequences enabled the classification of eight enzyme clusters. Particularly striking is the accumulation of genes coding for different transporters of the tripartite tricarboxylate transporters family, TRAP transporters and ABC transporters as well as genes coding for mandelate racemases/muconate lactonizing enzymes that might be involved in substrate uptake or degradation of AMDase products. Further, three novel AMDases were characterized which showed a high enantiomeric excess (>99%) of the (R)-enantiomer of flurbiprofen. These are the recombinant AmdA and AmdV from Variovorax sp. strains HH01 and HH02, originated from soil, and AmdP from Polymorphum gilvum found by a data base search. Altogether our findings give new insights into the class of AMDases and reveal many previously unknown enzyme candidates with high potential for bioindustrial processes. PMID:27610105

  17. The Moraxella catarrhalis immunoglobulin D-binding protein MID has conserved sequences and is regulated by a mechanism corresponding to phase variation.

    PubMed

    Möllenkvist, Andrea; Nordström, Therése; Halldén, Christer; Christensen, Jens Jørgen; Forsgren, Arne; Riesbeck, Kristian

    2003-04-01

    The prevalence of the Moraxella catarrhalis immunoglobulin D (IgD)-binding outer membrane protein MID and its gene was determined in 91 clinical isolates and in 7 culture collection strains. Eighty-four percent of the clinical Moraxella strains expressed MID-dependent IgD binding. The mid gene was detected in all strains as revealed by homology of the signal peptide sequence and a conserved area in the 3' end of the gene. When MID proteins from five different strains were compared, an identity of 65.3 to 85.0% and a similarity of 71.2 to 89.1% were detected. Gene analyses showed several amino acid repeat motifs in the open reading frames, and MID could be called a putative autotransport protein. Interestingly, homopolymeric [polyguanine [poly(G)

  18. Accuracy of sequence alignment and fold assessment using reduced amino acid alphabets.

    PubMed

    Melo, Francisco; Marti-Renom, Marc A

    2006-06-01

    Reduced or simplified amino acid alphabets group the 20 naturally occurring amino acids into a smaller number of representative protein residues. To date, several reduced amino acid alphabets have been proposed, which have been derived and optimized by a variety of methods. The resulting reduced amino acid alphabets have been applied to pattern recognition, generation of consensus sequences from multiple alignments, protein folding, and protein structure prediction. In this work, amino acid substitution matrices and statistical potentials were derived based on several reduced amino acid alphabets and their performance assessed in a large benchmark for the tasks of sequence alignment and fold assessment of protein structure models, using as a reference frame the standard alphabet of 20 amino acids. The results showed that a large reduction in the total number of residue types does not necessarily translate into a significant loss of discriminative power for sequence alignment and fold assessment. Therefore, some definitions of a few residue types are able to encode most of the relevant sequence/structure information that is present in the 20 standard amino acids. Based on these results, we suggest that the use of reduced amino acid alphabets may allow to increasing the accuracy of current substitution matrices and statistical potentials for the prediction of protein structure of remote homologs. PMID:16506243

  19. Characterization of mouse cellular deoxyribonucleic acid homologous to Abelson murine leukemia virus-specific sequences.

    PubMed Central

    Dale, B; Ozanne, B

    1981-01-01

    The genome of Abelson murine leukemia virus (A-MuLV) consists of sequences derived from both BALB/c mouse deoxyribonucleic acid and the genome of Moloney murine leukemia virus. Using deoxyribonucleic acid linear intermediates as a source of retroviral deoxyribonucleic acid, we isolated a recombinant plasmid which contained 1.9 kilobases of the 3.5-kilobase mouse-derived sequences found in A-MuLV (A-MuLV-specific sequences). We used this clone, designated pSA-17, as a probe restriction enzyme and Southern blot analyses to examine the arrangement of homologous sequences in BALB/c deoxyribonucleic acid (endogenous Abelson sequences). The endogenous Abelson sequences within the mouse genome were interrupted by noncoding regions, suggesting that a rearrangement of the cell sequences was required to produce the sequence found in the virus. Endogenous Abelson sequences were arranged similarly in mice that were susceptible to A-MuLV tumors and in mice that were resistant to A-MuLV tumors. An examination of three BALB/c plasmacytomas and a BALB/c early B-cell tumor likewise revealed no alteration in the arrangement of the endogenous Abelson sequences. Homology to pSA-17 was also observed in deoxyribonucleic acids prepared from rat, hamster, chicken, and human cells. An isolate of A-MuLV which encoded a 160,000-dalton transforming protein (P160) contained 700 more base pairs of mouse sequences than the standard A-MuLV isolate, which encoded a 120,000-dalton transforming protein (P120). Images PMID:9279386

  20. The amino acid sequence of monal pheasant lysozyme and its activity.

    PubMed

    Araki, T; Matsumoto, T; Torikata, T

    1998-10-01

    The amino acid sequence of monal pheasant lysozyme and its activity were analyzed. Carboxymethylated lysozyme was digested with trypsin and the resulting peptides were sequenced. The established amino acid sequence had one amino acid substitution at position 102 (Arg to Gly) comparing with Indian peafowl lysozyme and four amino acid substitutions at positions 3 (Phe to Tyr), 15 (His to Leu), 41 (Gln to His), and 121 (Gln to His) with chicken lysozyme. Analysis of the time-courses of reaction using N-acetylglucosamine pentamer as a substrate showed a difference of binding free energy change (-0.4 kcal/mol) at subsites A between monal pheasant and Indian peafowl lysozyme. This was assumed to be caused by the amino acid substitution at subsite A with loss of a positive charge at position 102 (Arg102 to Gly). PMID:9836434

  1. cDNA-derived amino acid sequences of myoglobins from nine species of whales and dolphins.

    PubMed

    Iwanami, Kentaro; Mita, Hajime; Yamamoto, Yasuhiko; Fujise, Yoshihiro; Yamada, Tadasu; Suzuki, Tomohiko

    2006-10-01

    We determined the myoglobin (Mb) cDNA sequences of nine cetaceans, of which six are the first reports of Mb sequences: sei whale (Balaenoptera borealis), Bryde's whale (Balaenoptera edeni), pygmy sperm whale (Kogia breviceps), Stejneger's beaked whale (Mesoplodon stejnegeri), Longman's beaked whale (Indopacetus pacificus), and melon-headed whale (Peponocephala electra), and three confirm the previously determined chemical amino acid sequences: sperm whale (Physeter macrocephalus), common minke whale (Balaenoptera acutorostrata) and pantropical spotted dolphin (Stenella attenuata). We found two types of Mb in the skeletal muscle of pantropical spotted dolphin: Mb I with the same amino acid sequence as that deposited in the protein database, and Mb II, which differs at two amino acid residues compared with Mb I. Using an alignment of the amino acid or cDNA sequences of cetacean Mb, we constructed a phylogenetic tree by the NJ method. Clustering of cetacean Mb amino acid and cDNA sequences essentially follows the classical taxonomy of cetaceans, suggesting that Mb sequence data is valid for classification of cetaceans at least to the family level. PMID:16962803

  2. Multi-species sequence comparison reveals conservation of ghrelin gene-derived splice variants encoding a truncated ghrelin peptide.

    PubMed

    Seim, Inge; Jeffery, Penny L; Thomas, Patrick B; Walpole, Carina M; Maugham, Michelle; Fung, Jenny N T; Yap, Pei-Yi; O'Keeffe, Angela J; Lai, John; Whiteside, Eliza J; Herington, Adrian C; Chopin, Lisa K

    2016-06-01

    The peptide hormone ghrelin is a potent orexigen produced predominantly in the stomach. It has a number of other biological actions, including roles in appetite stimulation, energy balance, the stimulation of growth hormone release and the regulation of cell proliferation. Recently, several ghrelin gene splice variants have been described. Here, we attempted to identify conserved alternative splicing of the ghrelin gene by cross-species sequence comparisons. We identified a novel human exon 2-deleted variant and provide preliminary evidence that this splice variant and in1-ghrelin encode a C-terminally truncated form of the ghrelin peptide, termed minighrelin. These variants are expressed in humans and mice, demonstrating conservation of alternative splicing spanning 90 million years. Minighrelin appears to have similar actions to full-length ghrelin, as treatment with exogenous minighrelin peptide stimulates appetite and feeding in mice. Forced expression of the exon 2-deleted preproghrelin variant mirrors the effect of the canonical preproghrelin, stimulating cell proliferation and migration in the PC3 prostate cancer cell line. This is the first study to characterise an exon 2-deleted preproghrelin variant and to demonstrate sequence conservation of ghrelin gene-derived splice variants that encode a truncated ghrelin peptide. This adds further impetus for studies into the alternative splicing of the ghrelin gene and the function of novel ghrelin peptides in vertebrates. PMID:26792793

  3. Genome-wide in Silico Identification of New Conserved and Functional Retinoic Acid Receptor Response Elements (Direct Repeats Separated by 5 bp)*

    PubMed Central

    Lalevée, Sébastien; Anno, Yannick N.; Chatagnon, Amandine; Samarut, Eric; Poch, Olivier; Laudet, Vincent; Benoit, Gerard; Lecompte, Odile; Rochette-Egly, Cécile

    2011-01-01

    The nuclear retinoic acid receptors interact with specific retinoic acid (RA) response elements (RAREs) located in the promoters of target genes to orchestrate transcriptional networks involved in cell growth and differentiation. Here we describe a genome-wide in silico analysis of consensus DR5 RAREs based on the recurrent RGKTSA motifs. More than 15,000 DR5 RAREs were identified and analyzed for their localization and conservation in vertebrates. We selected 138 elements located ±10 kb from transcription start sites and gene ends and conserved across more than 6 species. We also validated the functionality of these RAREs by analyzing their ability to bind retinoic acid receptors (ChIP sequencing experiments) as well as the RA regulation of the corresponding genes (RNA sequencing and quantitative real time PCR experiments). Such a strategy provided a global set of high confidence RAREs expanding the known experimentally validated RAREs repertoire associated to a series of new genes involved in cell signaling, development, and tumor suppression. Finally, the present work provides a valuable knowledge base for the analysis of a wider range of RA-target genes in different species. PMID:21803772

  4. Genome-wide in silico identification of new conserved and functional retinoic acid receptor response elements (direct repeats separated by 5 bp).

    PubMed

    Lalevée, Sébastien; Anno, Yannick N; Chatagnon, Amandine; Samarut, Eric; Poch, Olivier; Laudet, Vincent; Benoit, Gerard; Lecompte, Odile; Rochette-Egly, Cécile

    2011-09-23

    The nuclear retinoic acid receptors interact with specific retinoic acid (RA) response elements (RAREs) located in the promoters of target genes to orchestrate transcriptional networks involved in cell growth and differentiation. Here we describe a genome-wide in silico analysis of consensus DR5 RAREs based on the recurrent RGKTSA motifs. More than 15,000 DR5 RAREs were identified and analyzed for their localization and conservation in vertebrates. We selected 138 elements located ±10 kb from transcription start sites and gene ends and conserved across more than 6 species. We also validated the functionality of these RAREs by analyzing their ability to bind retinoic acid receptors (ChIP sequencing experiments) as well as the RA regulation of the corresponding genes (RNA sequencing and quantitative real time PCR experiments). Such a strategy provided a global set of high confidence RAREs expanding the known experimentally validated RAREs repertoire associated to a series of new genes involved in cell signaling, development, and tumor suppression. Finally, the present work provides a valuable knowledge base for the analysis of a wider range of RA-target genes in different species. PMID:21803772

  5. Draft Genome Sequences of Two Novel Acidimicrobiaceae Members from an Acid Mine Drainage Biofilm Metagenome.

    PubMed

    Pinto, Ameet J; Sharp, Jonathan O; Yoder, Michael J; Almstrand, Robert

    2016-01-01

    Bacteria belonging to the family Acidimicrobiaceae are frequently encountered in heavy metal-contaminated acidic environments. However, their phylogenetic and metabolic diversity is poorly resolved. We present draft genome sequences of two novel and phylogenetically distinct Acidimicrobiaceae members assembled from an acid mine drainage biofilm metagenome. PMID:26769942

  6. Draft Genome Sequences of Two Novel Acidimicrobiaceae Members from an Acid Mine Drainage Biofilm Metagenome

    PubMed Central

    Pinto, Ameet J.; Sharp, Jonathan O.; Yoder, Michael J.

    2016-01-01

    Bacteria belonging to the family Acidimicrobiaceae are frequently encountered in heavy metal-contaminated acidic environments. However, their phylogenetic and metabolic diversity is poorly resolved. We present draft genome sequences of two novel and phylogenetically distinct Acidimicrobiaceae members assembled from an acid mine drainage biofilm metagenome. PMID:26769942

  7. Conserved hypothetical protein Rv1977 in Mycobacterium tuberculosis strains contains sequence polymorphisms and might be involved in ongoing immune evasion

    PubMed Central

    Jiang, Yi; Liu, Haican; Wang, Xuezhi; Li, Guilian; Qiu, Yan; Dou, Xiangfeng; Wan, Kanglin

    2015-01-01

    Host immune pressure and associated parasite immune evasion are key features of host-pathogen co-evolution. A previous study showed that human T cell epitopes of Mycobacterium tuberculosis are evolutionarily hyperconserved and thus it was deduced that M. tuberculosis lacks antigenic variation and immune evasion. Here, we selected 151 clinical Mycobacterium tuberculosis isolates from China, amplified gene encoding Rv1977 and compared the sequences. The results showed that Rv1977, a conserved hypothetical protein, is not conserved in M. tuberculosis strains and there are polymorphisms existed in the protein. Some mutations, especially one frameshift mutation, occurred in the antigen Rv1977, which is uncommon in M.tb strains and may lead to the protein function altering. Mutations and deletion in the gene all affect one of three T cell epitopes and the changed T cell epitope contained more than one variable position, which may suggest ongoing immune evasion. PMID:26261576

  8. Structure-sequence based analysis for identification of conserved regions in proteins

    DOEpatents

    Zemla, Adam T; Zhou, Carol E; Lam, Marisa W; Smith, Jason R; Pardes, Elizabeth

    2013-05-28

    Disclosed are computational methods, and associated hardware and software products for scoring conservation in a protein structure based on a computationally identified family or cluster of protein structures. A method of computationally identifying a family or cluster of protein structures in also disclosed herein.

  9. Species identification using genetic tools: the value of nuclear and mitochondrial gene sequences in whale conservation.

    PubMed

    Palumbi, S R; Cipriano, F

    1998-01-01

    DNA sequence analysis is a powerful tool for identifying the source of samples thought to be derived from threatened or endangered species. Analysis of mitochondrial DNA (mtDNA) from retail whale meat markets has shown consistently that the expected baleen whale in these markets, the minke whale, makes up only about half the products analyzed. The other products are either unregulated small toothed whales like dolphins or are protected baleen whales such as humpback, Bryde's, fin, or blue whales. Independent verification of such mtDNA identifications requires analysis of nuclear genetic loci, but this is technically more difficult than standard mtDNA sequencing. In addition, evolution of species-specific sequences (i.e., fixation of sequence differences to produce reciprocally monophyletic gene trees) is slower in nuclear than in mitochondrial genes primarily because genetic drift is slower at nuclear loci. When will use of nuclear sequences allow forensic DNA identification? Comparison of neutral theories of coalescence of mitochondrial and nuclear loci suggests a simple rule of thumb. The "three-times rule" suggests that phylogenetic sorting at nuclear loci is likely to produce species-specific sequences when mitochondrial alleles are reciprocally monophyletic and the branches leading to the mtDNA sequences of a species are three times longer than the average difference observed within species. A preliminary test of the three-times rule, which depends on many assumptions about the species and genes involved, suggests that blue and fin whales should have species-specific sequences at most neutral nuclear loci, whereas humpback and fin whales should show species-specific sequences at fewer nuclear loci. Partial sequences of actin introns from these species confirm the predictions of the three-times rule and show that blue and fin whales are reciprocally monophyletic at this locus. These intron sequences are thus good tools for the identification of these species

  10. Two distinct ferredoxins from Rhodobacter capsulatus: complete amino acid sequences and molecular evolution.

    PubMed

    Saeki, K; Suetsugu, Y; Yao, Y; Horio, T; Marrs, B L; Matsubara, H

    1990-09-01

    Two distinct ferredoxins were purified from Rhodobacter capsulatus SB1003. Their complete amino acid sequences were determined by a combination of protease digestion, BrCN cleavage and Edman degradation. Ferredoxins I and II were composed of 64 and 111 amino acids, respectively, with molecular weights of 6,728 and 12,549 excluding iron and sulfur atoms. Both contained two Cys clusters in their amino acid sequences. The first cluster of ferredoxin I and the second cluster of ferredoxin II had a sequence, CxxCxxCxxxCP, in common with the ferredoxins found in Clostridia. The second cluster of ferredoxin I had a sequence, CxxCxxxxxxxxCxxxCM, with extra amino acids between the second and third Cys, which has been reported for other photosynthetic bacterial ferredoxins and putative ferredoxins (nif-gene products) from nitrogen-fixing bacteria, and with a unique occurrence of Met. The first cluster of ferredoxin II had a CxxCxxxxCxxxCP sequence, with two additional amino acids between the second and third Cys, a characteristics feature of Azotobacter-[3Fe-4S] [4Fe-4S]-ferredoxin. Ferredoxin II was also similar to Azotobacter-type ferredoxins with an extended carboxyl (C-) terminal sequence compared to the common Clostridium-type. The evolutionary relationship of the two together with a putative one recently found to be encoded in nifENXQ region in this bacterium [Moreno-Vivian et al. (1989) J. Bacteriol. 171, 2591-2598] is discussed. PMID:2277040

  11. Using Chou's pseudo amino acid composition to predict protein quaternary structure: a sequence-segmented PseAAC approach.

    PubMed

    Zhang, Shao-Wu; Chen, Wei; Yang, Feng; Pan, Quan

    2008-10-01

    In the protein universe, many proteins are composed of two or more polypeptide chains, generally referred to as subunits, which associate through noncovalent interactions and, occasionally, disulfide bonds to form protein quaternary structures. It has long been known that the functions of proteins are closely related to their quaternary structures; some examples include enzymes, hemoglobin, DNA polymerase, and ion channels. However, it is extremely labor-expensive and even impossible to quickly determine the structures of hundreds of thousands of protein sequences solely from experiments. Since the number of protein sequences entering databanks is increasing rapidly, it is highly desirable to develop computational methods for classifying the quaternary structures of proteins from their primary sequences. Since the concept of Chou's pseudo amino acid composition (PseAAC) was introduced, a variety of approaches, such as residue conservation scores, von Neumann entropy, multiscale energy, autocorrelation function, moment descriptors, and cellular automata, have been utilized to formulate the PseAAC for predicting different attributes of proteins. Here, in a different approach, a sequence-segmented PseAAC is introduced to represent protein samples. Meanwhile, multiclass SVM classifier modules were adopted to classify protein quaternary structures. As a demonstration, the dataset constructed by Chou and Cai [(2003) Proteins 53:282-289] was adopted as a benchmark dataset. The overall jackknife success rates thus obtained were 88.2-89.1%, indicating that the new approach is quite promising for predicting protein quaternary structure. PMID:18427713

  12. Amino Acid Sequence of Anionic Peroxidase from the Windmill Palm Tree Trachycarpus fortunei

    PubMed Central

    2015-01-01

    Palm peroxidases are extremely stable and have uncommon substrate specificity. This study was designed to fill in the knowledge gap about the structures of a peroxidase from the windmill palm tree Trachycarpus fortunei. The complete amino acid sequence and partial glycosylation were determined by MALDI-top-down sequencing of native windmill palm tree peroxidase (WPTP), MALDI-TOF/TOF MS/MS of WPTP tryptic peptides, and cDNA sequencing. The propeptide of WPTP contained N- and C-terminal signal sequences which contained 21 and 17 amino acid residues, respectively. Mature WPTP was 306 amino acids in length, and its carbohydrate content ranged from 21% to 29%. Comparison to closely related royal palm tree peroxidase revealed structural features that may explain differences in their substrate specificity. The results can be used to guide engineering of WPTP and its novel applications. PMID:25383699

  13. Computer analysis between nucleotide and amino acid sequences of bean golden mosaic virus and those of maize streak, wheat dwarf, chloris striate mosaic, and beet curly top viruses.

    PubMed

    Ikegami, M

    1989-01-01

    Bean golden mosaic virus (BGMV) DNA 1 and 2 have little sequence homology with maize streak virus (MSV), wheat dwarf virus (WDV), and chloris striate mosaic virus (CSMV) DNAs. BGMV DNA 1 and beet curly top virus (BCTV) DNA are closely related, whereas BGMV DNA 2 and BCTV DNA are not related. Direct amino acid homologies of predicted proteins between BGMV ORFs and MSV ORFs, WDV ORFs or CSMV ORFs were 40-50%. BGMV 1L1 and BCTV L1, and BGMV IL3 and BCTV L4 were highly conserved. The sequence TAATATTAC was detected in the loops of hairpin structures of 5 gemini-viruses. PMID:2615677

  14. Protein chemotaxonomy. XIII. Amino acid sequence of ferredoxin from Panax ginseng.

    PubMed

    Mino, Yoshiki

    2006-08-01

    The complete amino acid sequence of [2Fe-2S] ferredoxin from Panax ginseng (Araliaceae) has been determined by automated Edman degradation of the entire S-carboxymethylcysteinyl protein and of the peptides obtained by enzymatic digestion. This ferredoxin has a unique amino acid sequence, which includes an insertion of Tyr at the 3rd position from the amino-terminus and a deletion of two amino acid residues at the carboxyl terminus. This ferredoxin had 18 differences in its amino acid sequence compared to that of Petroselinum sativum (Umbelliferae). In contrast, 23-33 differences were observed compared to other dicotyledonous plants. This suggests that Panax ginseng is related taxonomically to umbelliferous plants. PMID:16880642

  15. Complete amino acid sequence and structure characterization of the taste-modifying protein, miraculin.

    PubMed

    Theerasilp, S; Hitotsuya, H; Nakajo, S; Nakaya, K; Nakamura, Y; Kurihara, Y

    1989-04-25

    The taste-modifying protein, miraculin, has the unusual property of modifying sour taste into sweet taste. The complete amino acid sequence of miraculin purified from miracle fruits by a newly developed method (Theerasilp, S., and Kurihara, Y. (1988) J. Biol. Chem. 263, 11536-11539) was determined by an automatic Edman degradation method. Miraculin was a single polypeptide with 191 amino acid residues. The calculated molecular weight based on the amino acid sequence and the carbohydrate content (13.9%) was 24,600. Asn-42 and Asn-186 were linked N-glycosidically to carbohydrate chains. High homology was found between the amino acid sequences of miraculin and soybean trypsin inhibitor. PMID:2708331

  16. N-terminal sequence of amino acids and some properties of an acid-stable alpha-amylase from citric acid-koji (Aspergillus usamii var.).

    PubMed

    Suganuma, T; Tahara, N; Kitahara, K; Nagahama, T; Inuzuka, K

    1996-01-01

    An acid-stable alpha-amylase (AA) was purified from an acidic extract of citric acid-koji (A. usamii var.). The N-terminal sequence of the first 20 amino acids of the enzyme was identical with that of AA from A. niger, but the two enzymes differed in molecular weight. HPLC analysis for identifying the anomers of products indicated that the AA hydrolyzed maltopentaose (G5) at the third glycoside bond predominantly, which differed from Taka-amylase A and the neutral alpha-amylase (NA) from the citric acid-koji. PMID:8824843

  17. Conservation of nucleotide sequences for molecular diagnosis of Middle East respiratory syndrome coronavirus, 2015.

    PubMed

    Furuse, Yuki; Okamoto, Michiko; Oshitani, Hitoshi

    2015-11-01

    Infection due to the Middle East respiratory syndrome coronavirus (MERS-CoV) is widespread. The present study was performed to assess the protocols used for the molecular diagnosis of MERS-CoV by analyzing the nucleotide sequences of viruses detected between 2012 and 2015, including sequences from the large outbreak in eastern Asia in 2015. Although the diagnostic protocols were established only 2 years ago, mismatches between the sequences of primers/probes and viruses were found for several of the assays. Such mismatches could lead to a lower sensitivity of the assay, thereby leading to false-negative diagnosis. A slight modification in the primer design is suggested. Protocols for the molecular diagnosis of viral infections should be reviewed regularly after they are established, particularly for viruses that pose a great threat to public health such as MERS-CoV. PMID:26432410

  18. Conservation of the C-type lectin fold for massive sequence variation in a Treponema diversity-generating retroelement

    SciTech Connect

    Le Coq, Johanne; Ghosh, Partho

    2012-06-19

    Anticipatory ligand binding through massive protein sequence variation is rare in biological systems, having been observed only in the vertebrate adaptive immune response and in a phage diversity-generating retroelement (DGR). Earlier work has demonstrated that the prototypical DGR variable protein, major tropism determinant (Mtd), meets the demands of anticipatory ligand binding by novel means through the C-type lectin (CLec) fold. However, because of the low sequence identity among DGR variable proteins, it has remained unclear whether the CLec fold is a general solution for DGRs. We have addressed this problem by determining the structure of a second DGR variable protein, TvpA, from the pathogenic oral spirochete Treponema denticola. Despite its weak sequence identity to Mtd ({approx}16%), TvpA was found to also have a CLec fold, with predicted variable residues exposed in a ligand-binding site. However, this site in TvpA was markedly more variable than the one in Mtd, reflecting the unprecedented approximate 10{sup 20} potential variability of TvpA. In addition, similarity between TvpA and Mtd with formylglycine-generating enzymes was detected. These results provide strong evidence for the conservation of the formylglycine-generating enzyme-type CLec fold among DGRs as a means of accommodating massive sequence variation.

  19. Expression of cassini, a murine gamma-satellite sequence conserved in evolution, is regulated in normal and malignant hematopoietic cells

    PubMed Central

    2012-01-01

    Background Acute lymphoblastic leukemia (ALL) cells treated with drugs can become drug-tolerant if co-cultured with protective stromal mouse embryonic fibroblasts (MEFs). Results We performed transcriptional profiling on these stromal fibroblasts to investigate if they were affected by the presence of drug-treated ALL cells. These mitotically inactivated MEFs showed few changes in gene expression, but a family of sequences of which transcription is significantly increased was identified. A sequence related to this family, which we named cassini, was selected for further characterization. We found that cassini was highly upregulated in drug-treated ALL cells. Analysis of RNAs from different normal mouse tissues showed that cassini expression is highest in spleen and thymus, and can be further enhanced in these organs by exposure of mice to bacterial endotoxin. Heat shock, but not other types of stress, significantly induced the transcription of this locus in ALL cells. Transient overexpression of cassini in human 293 embryonic kidney cells did not increase the cytotoxic or cytostatic effects of chemotherapeutic drugs but provided some protection. Database searches revealed that sequences highly homologous to cassini are present in rodents, apicomplexans, flatworms and primates, indicating that they are conserved in evolution. Moreover, CASSINI RNA was induced in human ALL cells treated with vincristine. Surprisingly, cassini belongs to the previously reported murine family of γ-satellite/major satellite DNA sequences, which were not known to be present in other species. Conclusions Our results show that the transcription of at least one member of these sequences is regulated, suggesting that this has a function in normal and transformed immune cells. Expression of these sequences may protect cells when they are exposed to specific stress stimuli. PMID:22916712

  20. Computational identification of riboswitches based on RNA conserved functional sequences and conformations.

    PubMed

    Chang, Tzu-Hao; Huang, Hsien-Da; Wu, Li-Ching; Yeh, Chi-Ta; Liu, Baw-Jhiune; Horng, Jorng-Tzong

    2009-07-01

    Riboswitches are cis-acting genetic regulatory elements within a specific mRNA that can regulate both transcription and translation by interacting with their corresponding metabolites. Recently, an increasing number of riboswitches have been identified in different species and investigated for their roles in regulatory functions. Both the sequence contexts and structural conformations are important characteristics of riboswitches. None of the previously developed tools, such as covariance models (CMs), Riboswitch finder, and RibEx, provide a web server for efficiently searching homologous instances of known riboswitches or considers two crucial characteristics of each riboswitch, such as the structural conformations and sequence contexts of functional regions. Therefore, we developed a systematic method for identifying 12 kinds of riboswitches. The method is implemented and provided as a web server, RiboSW, to efficiently and conveniently identify riboswitches within messenger RNA sequences. The predictive accuracy of the proposed method is comparable with other previous tools. The efficiency of the proposed method for identifying riboswitches was improved in order to achieve a reasonable computational time required for the prediction, which makes it possible to have an accurate and convenient web server for biologists to obtain the results of their analysis of a given mRNA sequence. RiboSW is now available on the web at http://RiboSW.mbc.nctu.edu.tw/. PMID:19460868

  1. Comparison of C. elegans and C. briggsae Genome Sequences Reveals Extensive Conservation of Chromosome Organization and Synteny

    PubMed Central

    Hillier, LaDeana W; Miller, Raymond D; Baird, Scott E; Chinwalla, Asif; Fulton, Lucinda A; Koboldt, Daniel C; Waterston, Robert H

    2007-01-01

    To determine whether the distinctive features of Caenorhabditis elegans chromosomal organization are shared with the C. briggsae genome, we constructed a single nucleotide polymorphism–based genetic map to order and orient the whole genome shotgun assembly along the six C. briggsae chromosomes. Although these species are of the same genus, their most recent common ancestor existed 80–110 million years ago, and thus they are more evolutionarily distant than, for example, human and mouse. We found that, like C. elegans chromosomes, C. briggsae chromosomes exhibit high levels of recombination on the arms along with higher repeat density, a higher fraction of intronic sequence, and a lower fraction of exonic sequence compared with chromosome centers. Despite extensive intrachromosomal rearrangements, 1:1 orthologs tend to remain in the same region of the chromosome, and colinear blocks of orthologs tend to be longer in chromosome centers compared with arms. More strikingly, the two species show an almost complete conservation of synteny, with 1:1 orthologs present on a single chromosome in one species also found on a single chromosome in the other. The conservation of both chromosomal organization and synteny between these two distantly related species suggests roles for chromosome organization in the fitness of an organism that are only poorly understood presently. PMID:17608563

  2. Strong conservation of non-coding sequences during vertebrates evolution: potential involvement in post-transcriptional regulation of gene expression.

    PubMed Central

    Duret, L; Dorkeld, F; Gautier, C

    1993-01-01

    Comparison of nucleotide sequences from different classes of vertebrates that diverged more than 300 million years ago, revealed the existence of highly conserved regions (HCRs) with more than 70% similarity over 100 to 1450 nt in non-coding parts of genes. Such a conservation is unexpected because it is much longer and stronger than what is necessary for specifying the binding of a regulatory protein. HCRs are relatively frequent, particularly in genes that are essential to cell life. In multigene families, conserved regions are specific of each isotype and are probably involved in the control of their specific pattern of expression. Studying HCRs distribution within genes showed that functional constraints are generally much stronger in 3'-non-coding regions than in promoters or introns. The 3'-HCRs are particularly A + T-rich and are always located in the transcribed untranslated regions of genes, which suggests that they are involved in post-transcriptional processes. However, current knowledge of mechanisms that regulate mRNA export, localisation, translation, or degradation is not sufficient to explain the strong functional constraints that we have characterised. PMID:8506129

  3. Detection and isolation of nucleic acid sequences using competitive hybridization probes

    DOEpatents

    Lucas, Joe N.; Straume, Tore; Bogen, Kenneth T.

    1997-01-01

    A method for detecting a target nucleic acid sequence in a sample is provided using hybridization probes which competitively hybridize to a target nucleic acid. According to the method, a target nucleic acid sequence is hybridized to first and second hybridization probes which are complementary to overlapping portions of the target nucleic acid sequence, the first hybridization probe including a first complexing agent capable of forming a binding pair with a second complexing agent and the second hybridization probe including a detectable marker. The first complexing agent attached to the first hybridization probe is contacted with a second complexing agent, the second complexing agent being attached to a solid support such that when the first and second complexing agents are attached, target nucleic acid sequences hybridized to the first hybridization probe become immobilized on to the solid support. The immobilized target nucleic acids are then separated and detected by detecting the detectable marker attached to the second hybridization probe. A kit for performing the method is also provided.

  4. Detection and isolation of nucleic acid sequences using competitive hybridization probes

    DOEpatents

    Lucas, J.N.; Straume, T.; Bogen, K.T.

    1997-04-01

    A method for detecting a target nucleic acid sequence in a sample is provided using hybridization probes which competitively hybridize to a target nucleic acid. According to the method, a target nucleic acid sequence is hybridized to first and second hybridization probes which are complementary to overlapping portions of the target nucleic acid sequence, the first hybridization probe including a first complexing agent capable of forming a binding pair with a second complexing agent and the second hybridization probe including a detectable marker. The first complexing agent attached to the first hybridization probe is contacted with a second complexing agent, the second complexing agent being attached to a solid support such that when the first and second complexing agents are attached, target nucleic acid sequences hybridized to the first hybridization probe become immobilized on to the solid support. The immobilized target nucleic acids are then separated and detected by detecting the detectable marker attached to the second hybridization probe. A kit for performing the method is also provided. 7 figs.

  5. Sequence Divergence and Conservation in Genomes of Helicobacter cetorum Strains from a Dolphin and a Whale

    PubMed Central

    Kersulyte, Dangeruta; Rossi, Mirko; Berg, Douglas E.

    2013-01-01

    Background and Objectives Strains of Helicobacter cetorum have been cultured from several marine mammals and have been found to be closely related in 16 S rDNA sequence to the human gastric pathogen H. pylori, but their genomes were not characterized further. Methods The genomes of H. cetorum strains from a dolphin and a whale were sequenced completely using 454 technology and PCR and capillary sequencing. Results These genomes are 1.8 and 1.95 mb in size, some 7–26% larger than H. pylori genomes, and differ markedly from one another in gene content, and sequences and arrangements of shared genes. However, each strain is more related overall to H. pylori and its descendant H. acinonychis than to other known species. These H. cetorum strains lack cag pathogenicity islands, but contain novel alleles of the virulence-associated vacuolating cytotoxin (vacA) gene. Of particular note are (i) an extra triplet of vacA genes with ≤50% protein-level identity to each other in the 5′ two-thirds of the gene needed for host factor interaction; (ii) divergent sets of outer membrane protein genes; (iii) several metabolic genes distinct from those of H. pylori; (iv) genes for an iron-cofactored urease related to those of Helicobacter species from terrestrial carnivores, in addition to genes for a nickel co-factored urease; and (v) members of the slr multigene family, some of which modulate host responses to infection and improve Helicobacter growth with mammalian cells. Conclusions Our genome sequence data provide a glimpse into the novelty and great genetic diversity of marine helicobacters. These data should aid further analyses of microbial genome diversity and evolution and infection and disease mechanisms in vast and often fragile ocean ecosystems. PMID:24358262

  6. Sequence Similarity of Clostridium difficile Strains by Analysis of Conserved Genes and Genome Content Is Reflected by Their Ribotype Affiliation

    PubMed Central

    Kurka, Hedwig; Ehrenreich, Armin; Ludwig, Wolfgang; Monot, Marc; Rupnik, Maja; Barbut, Frederic; Indra, Alexander; Dupuy, Bruno; Liebl, Wolfgang

    2014-01-01

    PCR-ribotyping is a broadly used method for the classification of isolates of Clostridium difficile, an emerging intestinal pathogen, causing infections with increased disease severity and incidence in several European and North American countries. We have now carried out clustering analysis with selected genes of numerous C. difficile strains as well as gene content comparisons of their genomes in order to broaden our view of the relatedness of strains assigned to different ribotypes. We analyzed the genomic content of 48 C. difficile strains representing 21 different ribotypes. The calculation of distance matrix-based dendrograms using the neighbor joining method for 14 conserved genes (standard phylogenetic marker genes) from the genomes of the C. difficile strains demonstrated that the genes from strains with the same ribotype generally clustered together. Further, certain ribotypes always clustered together and formed ribotype groups, i.e. ribotypes 078, 033 and 126, as well as ribotypes 002 and 017, indicating their relatedness. Comparisons of the gene contents of the genomes of ribotypes that clustered according to the conserved gene analysis revealed that the number of common genes of the ribotypes belonging to each of these three ribotype groups were very similar for the 078/033/126 group (at most 69 specific genes between the different strains with the same ribotype) but less similar for the 002/017 group (86 genes difference). It appears that the ribotype is indicative not only of a specific pattern of the amplified 16S–23S rRNA intergenic spacer but also reflects specific differences in the nucleotide sequences of the conserved genes studied here. It can be anticipated that the sequence deviations of more genes of C. difficile strains are correlated with their PCR-ribotype. In conclusion, the results of this study corroborate and extend the concept of clonal C. difficile lineages, which correlate with ribotypes affiliation. PMID:24482682

  7. Sequence similarity of Clostridium difficile strains by analysis of conserved genes and genome content is reflected by their ribotype affiliation.

    PubMed

    Kurka, Hedwig; Ehrenreich, Armin; Ludwig, Wolfgang; Monot, Marc; Rupnik, Maja; Barbut, Frederic; Indra, Alexander; Dupuy, Bruno; Liebl, Wolfgang

    2014-01-01

    PCR-ribotyping is a broadly used method for the classification of isolates of Clostridium difficile, an emerging intestinal pathogen, causing infections with increased disease severity and incidence in several European and North American countries. We have now carried out clustering analysis with selected genes of numerous C. difficile strains as well as gene content comparisons of their genomes in order to broaden our view of the relatedness of strains assigned to different ribotypes. We analyzed the genomic content of 48 C. difficile strains representing 21 different ribotypes. The calculation of distance matrix-based dendrograms using the neighbor joining method for 14 conserved genes (standard phylogenetic marker genes) from the genomes of the C. difficile strains demonstrated that the genes from strains with the same ribotype generally clustered together. Further, certain ribotypes always clustered together and formed ribotype groups, i.e. ribotypes 078, 033 and 126, as well as ribotypes 002 and 017, indicating their relatedness. Comparisons of the gene contents of the genomes of ribotypes that clustered according to the conserved gene analysis revealed that the number of common genes of the ribotypes belonging to each of these three ribotype groups were very similar for the 078/033/126 group (at most 69 specific genes between the different strains with the same ribotype) but less similar for the 002/017 group (86 genes difference). It appears that the ribotype is indicative not only of a specific pattern of the amplified 16S-23S rRNA intergenic spacer but also reflects specific differences in the nucleotide sequences of the conserved genes studied here. It can be anticipated that the sequence deviations of more genes of C. difficile strains are correlated with their PCR-ribotype. In conclusion, the results of this study corroborate and extend the concept of clonal C. difficile lineages, which correlate with ribotypes affiliation. PMID:24482682

  8. The Most Deeply Conserved Noncoding Sequences in Plants Serve Similar Functions to Those in Vertebrates Despite Large Differences in Evolutionary Rates[W

    PubMed Central

    Burgess, Diane; Freeling, Michael

    2014-01-01

    In vertebrates, conserved noncoding elements (CNEs) are functionally constrained sequences that can show striking conservation over >400 million years of evolutionary distance and frequently are located megabases away from target developmental genes. Conserved noncoding sequences (CNSs) in plants are much shorter, and it has been difficult to detect conservation among distantly related genomes. In this article, we show not only that CNS sequences can be detected throughout the eudicot clade of flowering plants, but also that a subset of 37 CNSs can be found in all flowering plants (diverging ∼170 million years ago). These CNSs are functionally similar to vertebrate CNEs, being highly associated with transcription factor and development genes and enriched in transcription factor binding sites. Some of the most highly conserved sequences occur in genes encoding RNA binding proteins, particularly the RNA splicing–associated SR genes. Differences in sequence conservation between plants and animals are likely to reflect differences in the biology of the organisms, with plants being much more able to tolerate genomic deletions and whole-genome duplication events due, in part, to their far greater fecundity compared with vertebrates. PMID:24681619

  9. A conserved amino acid residue critical for product and substrate specificity in plant triterpene synthases.

    PubMed

    Salmon, Melissa; Thimmappa, Ramesha B; Minto, Robert E; Melton, Rachel E; Hughes, Richard K; O'Maille, Paul E; Hemmings, Andrew M; Osbourn, Anne

    2016-07-26

    Triterpenes are structurally complex plant natural products with numerous medicinal applications. They are synthesized through an origami-like process that involves cyclization of the linear 30 carbon precursor 2,3-oxidosqualene into different triterpene scaffolds. Here, through a forward genetic screen in planta, we identify a conserved amino acid residue that determines product specificity in triterpene synthases from diverse plant species. Mutation of this residue results in a major change in triterpene cyclization, with production of tetracyclic rather than pentacyclic products. The mutated enzymes also use the more highly oxygenated substrate dioxidosqualene in preference to 2,3-oxidosqualene when expressed in yeast. Our discoveries provide new insights into triterpene cyclization, revealing hidden functional diversity within triterpene synthases. They further open up opportunities to engineer novel oxygenated triterpene scaffolds by manipulating the precursor supply. PMID:27412861

  10. A conserved amino acid residue critical for product and substrate specificity in plant triterpene synthases

    PubMed Central

    Salmon, Melissa; Thimmappa, Ramesha B.; Minto, Robert E.; Melton, Rachel E.; O’Maille, Paul E.; Hemmings, Andrew M.; Osbourn, Anne

    2016-01-01

    Triterpenes are structurally complex plant natural products with numerous medicinal applications. They are synthesized through an origami-like process that involves cyclization of the linear 30 carbon precursor 2,3-oxidosqualene into different triterpene scaffolds. Here, through a forward genetic screen in planta, we identify a conserved amino acid residue that determines product specificity in triterpene synthases from diverse plant species. Mutation of this residue results in a major change in triterpene cyclization, with production of tetracyclic rather than pentacyclic products. The mutated enzymes also use the more highly oxygenated substrate dioxidosqualene in preference to 2,3-oxidosqualene when expressed in yeast. Our discoveries provide new insights into triterpene cyclization, revealing hidden functional diversity within triterpene synthases. They further open up opportunities to engineer novel oxygenated triterpene scaffolds by manipulating the precursor supply. PMID:27412861

  11. LEU3 of Saccharomyces cerevisiae activates multiple genes for branched-chain amino acid biosynthesis by binding to a common decanucleotide core sequence

    SciTech Connect

    Friden, P.; Schimmel, P.

    1988-07-01

    LEU3 of Saccharomyces cerevisiae encodes an 886-amino-acid polypeptide that regulates transcription of a group of genes involved in leucine biosynthesis and has been shown to bind specifically to a 114-base-pair DNA fragment of the LEU2 upstream region. The authors show here that, in addition to LEU2, LEU3 binds in vitro to sequences in the promoter regions of LEU1, LEU4, ILV2, and, by inference, ILV5. The largely conserved decanucleotide core sequence shared by the binding sites in these genes is CCGGNNCCGG. Methylation interference footprinting experiements show that LEU 3 makes symmetrical contacts with the conserved bases that lie in the major groove. Synthetic oligonucleides (19 to 29 base pairs) which contain the core decanucleotide and flanking sequences of LEU1, LEU2, LEU4, and ILV2 have individually been placed upstream of a LEU3-insensitive test promoter. The expression of each construction is activated by LEU3, although the degree of activation varies considerably according to the specific oligonucleotide which is introduced. A promoter construction with substitutions in the core sequence remains LEU3 insensitive, however. One of the oligonucleotides (based on a LEU2 sequence) was also tested and shown to confer leucine-sensitive expression on the test promoter. The results demonstrate that only a short sequence element is necessary for LEU3-dependent promoter binding and activation and provide direct evidence for an expanded repertoire of genes that are activated by LEU3.

  12. Co-conservation of rRNA tetraloop sequences and helix length suggests involvement of the tetraloops in higher-order interactions

    NASA Technical Reports Server (NTRS)

    Hedenstierna, K. O.; Siefert, J. L.; Fox, G. E.; Murgola, E. J.

    2000-01-01

    Terminal loops containing four nucleotides (tetraloops) are common in structural RNAs, and they frequently conform to one of three sequence motifs, GNRA, UNCG, or CUUG. Here we compare available sequences and secondary structures for rRNAs from bacteria, and we show that helices capped by phylogenetically conserved GNRA loops display a strong tendency to be of conserved length. The simplest interpretation of this correlation is that the conserved GNRA loops are involved in higher-order interactions, intramolecular or intermolecular, resulting in a selective pressure for maintaining the lengths of these helices. A small number of conserved UNCG loops were also found to be associated with conserved length helices, consistent with the possibility that this type of tetraloop also takes part in higher-order interactions.

  13. The Roles of Four Conserved Basic Amino Acids in a Ferredoxin-Dependent Cyanobacterial Nitrate Reductase

    PubMed Central

    Srivastava, Anurag P.; Hirasawa, Masakazu; Bhalla, Megha; Chung, Jung-Sung; Allen, James P.; Johnson, Michael K.; Tripathy, Jatindra N.; Rubio, Luis M.; Vaccaro, Brian; Subramanian, Sowmya; Flores, Enrique; Zabet-Moghaddam, Masoud; Stitle, Kyle; Knaff, David B.

    2013-01-01

    The roles of four conserved basic amino acids in the reaction catalyzed by the ferredoxin-dependent nitrate reductase from the cyanobacterium Synechococcus sp. PCC 7942 have been investigated using site-directed mutagenesis in combination with measurements of steady-state kinetics, substrate-binding affinities and spectroscopic properties of the enzyme’s two prosthetic groups. Replacement of either Lys58 or Arg70 by glutamine leads to a complete loss of activity, with both the physiological electron donor, reduced ferredoxin and with a non-physiological electron donor, reduced methyl viologen. More conservative, charge-maintaining K58R and R70K variants were also completely inactive. Replacement of Lys130 by glutamine produced a variant that retained 26% of the wild-type activity with methyl viologen as the electron donor and 22% of the wild-type activity with ferredoxin as the electron donor, while replacement by arginine produces a variant that retains a significantly higher percentage of the wild-type activity with both electron donors. In contrast, replacement of Arg146 by glutamine had minimal effect on the activity of the enzyme. These results, along with substrate-binding and spectroscopic measurements, are discussed in terms of an in silico structural model for the enzyme. PMID:23692082

  14. Sorting out relationships among the grouse and ptarmigan using intron, mitochondrial, and ultra-conserved element sequences.

    PubMed

    Persons, Nicholas W; Hosner, Peter A; Meiklejohn, Kelly A; Braun, Edward L; Kimball, Rebecca T

    2016-05-01

    The Holarctic phasianid clade of the grouse and ptarmigan has received substantial attention in areas such as evolution of mating systems, display behavior, and population ecology related to their conservation and management as wild game species. There are multiple molecular phylogenetic studies that focus on grouse and ptarmigan. In spite of this, there is little consensus regarding historical relationships, particularly among genera, which has led to unstable and partial taxonomic revisions. We estimated the phylogeny of all currently recognized species using a combination of novel data from seven nuclear loci (largely intron sequences) and published data from one additional autosomal locus, two W-linked loci, and four mitochondrial regions. To explore relationships among genera and assess paraphyly of one genus more rigorously, we then added over 3000 ultra-conserved element (UCE) loci (over 1.7million bp) gathered using Illumina sequencing. The UCE topology agreed with that of the combined nuclear intron and previously published sequence data with 100% bootstrap support for all relationships. These data strongly support previous studies separating Bonasa from Tetrastes and Dendragapus from Falcipennis. However, the placement of Lagopus differed from previous studies, and we found no support for Falcipennis monophyly. Biogeographic analysis suggests that the ancestors of grouse and ptarmigan were distributed in the New World and subsequently underwent at least four dispersal events between the Old and New Worlds. Divergence time estimates from maternally-inherited and autosomal markers show stark differences across this clade, with divergence time estimates from maternally-inherited markers being nearly half that of the autosomal markers at some nodes, and nearly twice that at other nodes. PMID:26879712

  15. The first myriapod genome sequence reveals conservative arthropod gene content and genome organisation in the centipede Strigamia maritima.

    PubMed

    Chipman, Ariel D; Ferrier, David E K; Brena, Carlo; Qu, Jiaxin; Hughes, Daniel S T; Schröder, Reinhard; Torres-Oliva, Montserrat; Znassi, Nadia; Jiang, Huaiyang; Almeida, Francisca C; Alonso, Claudio R; Apostolou, Zivkos; Aqrawi, Peshtewani; Arthur, Wallace; Barna, Jennifer C J; Blankenburg, Kerstin P; Brites, Daniela; Capella-Gutiérrez, Salvador; Coyle, Marcus; Dearden, Peter K; Du Pasquier, Louis; Duncan, Elizabeth J; Ebert, Dieter; Eibner, Cornelius; Erikson, Galina; Evans, Peter D; Extavour, Cassandra G; Francisco, Liezl; Gabaldón, Toni; Gillis, William J; Goodwin-Horn, Elizabeth A; Green, Jack E; Griffiths-Jones, Sam; Grimmelikhuijzen, Cornelis J P; Gubbala, Sai; Guigó, Roderic; Han, Yi; Hauser, Frank; Havlak, Paul; Hayden, Luke; Helbing, Sophie; Holder, Michael; Hui, Jerome H L; Hunn, Julia P; Hunnekuhl, Vera S; Jackson, LaRonda; Javaid, Mehwish; Jhangiani, Shalini N; Jiggins, Francis M; Jones, Tamsin E; Kaiser, Tobias S; Kalra, Divya; Kenny, Nathan J; Korchina, Viktoriya; Kovar, Christie L; Kraus, F Bernhard; Lapraz, François; Lee, Sandra L; Lv, Jie; Mandapat, Christigale; Manning, Gerard; Mariotti, Marco; Mata, Robert; Mathew, Tittu; Neumann, Tobias; Newsham, Irene; Ngo, Dinh N; Ninova, Maria; Okwuonu, Geoffrey; Ongeri, Fiona; Palmer, William J; Patil, Shobha; Patraquim, Pedro; Pham, Christopher; Pu, Ling-Ling; Putman, Nicholas H; Rabouille, Catherine; Ramos, Olivia Mendivil; Rhodes, Adelaide C; Robertson, Helen E; Robertson, Hugh M; Ronshaugen, Matthew; Rozas, Julio; Saada, Nehad; Sánchez-Gracia, Alejandro; Scherer, Steven E; Schurko, Andrew M; Siggens, Kenneth W; Simmons, DeNard; Stief, Anna; Stolle, Eckart; Telford, Maximilian J; Tessmar-Raible, Kristin; Thornton, Rebecca; van der Zee, Maurijn; von Haeseler, Arndt; Williams, James M; Willis, Judith H; Wu, Yuanqing; Zou, Xiaoyan; Lawson, Daniel; Muzny, Donna M; Worley, Kim C; Gibbs, Richard A; Akam, Michael; Richards, Stephen

    2014-11-01

    Myriapods (e.g., centipedes and millipedes) display a simple homonomous body plan relative to other arthropods. All members of the class are terrestrial, but they attained terrestriality independently of insects. Myriapoda is the only arthropod class not represented by a sequenced genome. We present an analysis of the genome of the centipede Strigamia maritima. It retains a compact genome that has undergone less gene loss and shuffling than previously sequenced arthropods, and many orthologues of genes conserved from the bilaterian ancestor that have been lost in insects. Our analysis locates many genes in conserved macro-synteny contexts, and many small-scale examples of gene clustering. We describe several examples where S. maritima shows different solutions from insects to similar problems. The insect olfactory receptor gene family is absent from S. maritima, and olfaction in air is likely effected by expansion of other receptor gene families. For some genes S. maritima has evolved paralogues to generate coding sequence diversity, where insects use alternate splicing. This is most striking for the Dscam gene, which in Drosophila generates more than 100,000 alternate splice forms, but in S. maritima is encoded by over 100 paralogues. We see an intriguing linkage between the absence of any known photosensory proteins in a blind organism and the additional absence of canonical circadian clock genes. The phylogenetic position of myriapods allows us to identify where in arthropod phylogeny several particular molecular mechanisms and traits emerged. For example, we conclude that juvenile hormone signalling evolved with the emergence of the exoskeleton in the arthropods and that RR-1 containing cuticle proteins evolved in the lineage leading to Mandibulata. We also identify when various gene expansions and losses occurred. The genome of S. maritima offers us a unique glimpse into the ancestral arthropod genome, while also displaying many adaptations to its specific

  16. The First Myriapod Genome Sequence Reveals Conservative Arthropod Gene Content and Genome Organisation in the Centipede Strigamia maritima

    PubMed Central

    Chipman, Ariel D.; Ferrier, David E. K.; Brena, Carlo; Qu, Jiaxin; Hughes, Daniel S. T.; Schröder, Reinhard; Torres-Oliva, Montserrat; Znassi, Nadia; Jiang, Huaiyang; Almeida, Francisca C.; Alonso, Claudio R.; Apostolou, Zivkos; Aqrawi, Peshtewani; Arthur, Wallace; Barna, Jennifer C. J.; Blankenburg, Kerstin P.; Brites, Daniela; Capella-Gutiérrez, Salvador; Coyle, Marcus; Dearden, Peter K.; Du Pasquier, Louis; Duncan, Elizabeth J.; Ebert, Dieter; Eibner, Cornelius; Erikson, Galina; Evans, Peter D.; Extavour, Cassandra G.; Francisco, Liezl; Gabaldón, Toni; Gillis, William J.; Goodwin-Horn, Elizabeth A.; Green, Jack E.; Griffiths-Jones, Sam; Grimmelikhuijzen, Cornelis J. P.; Gubbala, Sai; Guigó, Roderic; Han, Yi; Hauser, Frank; Havlak, Paul; Hayden, Luke; Helbing, Sophie; Holder, Michael; Hui, Jerome H. L.; Hunn, Julia P.; Hunnekuhl, Vera S.; Jackson, LaRonda; Javaid, Mehwish; Jhangiani, Shalini N.; Jiggins, Francis M.; Jones, Tamsin E.; Kaiser, Tobias S.; Kalra, Divya; Kenny, Nathan J.; Korchina, Viktoriya; Kovar, Christie L.; Kraus, F. Bernhard; Lapraz, François; Lee, Sandra L.; Lv, Jie; Mandapat, Christigale; Manning, Gerard; Mariotti, Marco; Mata, Robert; Mathew, Tittu; Neumann, Tobias; Newsham, Irene; Ngo, Dinh N.; Ninova, Maria; Okwuonu, Geoffrey; Ongeri, Fiona; Palmer, William J.; Patil, Shobha; Patraquim, Pedro; Pham, Christopher; Pu, Ling-Ling; Putman, Nicholas H.; Rabouille, Catherine; Ramos, Olivia Mendivil; Rhodes, Adelaide C.; Robertson, Helen E.; Robertson, Hugh M.; Ronshaugen, Matthew; Rozas, Julio; Saada, Nehad; Sánchez-Gracia, Alejandro; Scherer, Steven E.; Schurko, Andrew M.; Siggens, Kenneth W.; Simmons, DeNard; Stief, Anna; Stolle, Eckart; Telford, Maximilian J.; Tessmar-Raible, Kristin; Thornton, Rebecca; van der Zee, Maurijn; von Haeseler, Arndt; Williams, James M.; Willis, Judith H.; Wu, Yuanqing; Zou, Xiaoyan; Lawson, Daniel; Muzny, Donna M.; Worley, Kim C.; Gibbs, Richard A.; Akam, Michael; Richards, Stephen

    2014-01-01

    Myriapods (e.g., centipedes and millipedes) display a simple homonomous body plan relative to other arthropods. All members of the class are terrestrial, but they attained terrestriality independently of insects. Myriapoda is the only arthropod class not represented by a sequenced genome. We present an analysis of the genome of the centipede Strigamia maritima. It retains a compact genome that has undergone less gene loss and shuffling than previously sequenced arthropods, and many orthologues of genes conserved from the bilaterian ancestor that have been lost in insects. Our analysis locates many genes in conserved macro-synteny contexts, and many small-scale examples of gene clustering. We describe several examples where S. maritima shows different solutions from insects to similar problems. The insect olfactory receptor gene family is absent from S. maritima, and olfaction in air is likely effected by expansion of other receptor gene families. For some genes S. maritima has evolved paralogues to generate coding sequence diversity, where insects use alternate splicing. This is most striking for the Dscam gene, which in Drosophila generates more than 100,000 alternate splice forms, but in S. maritima is encoded by over 100 paralogues. We see an intriguing linkage between the absence of any known photosensory proteins in a blind organism and the additional absence of canonical circadian clock genes. The phylogenetic position of myriapods allows us to identify where in arthropod phylogeny several particular molecular mechanisms and traits emerged. For example, we conclude that juvenile hormone signalling evolved with the emergence of the exoskeleton in the arthropods and that RR-1 containing cuticle proteins evolved in the lineage leading to Mandibulata. We also identify when various gene expansions and losses occurred. The genome of S. maritima offers us a unique glimpse into the ancestral arthropod genome, while also displaying many adaptations to its specific

  17. Abscisic acid-induced gene expression in the liverwort Marchantia polymorpha is mediated by evolutionarily conserved promoter elements.

    PubMed

    Ghosh, Totan K; Kaneko, Midori; Akter, Khaleda; Murai, Shuhei; Komatsu, Kenji; Ishizaki, Kimitsune; Yamato, Katsuyuki T; Kohchi, Takayuki; Takezawa, Daisuke

    2016-04-01

    Abscisic acid (ABA) is a phytohormone widely distributed among members of the land plant lineage (Embryophyta), regulating dormancy, stomata closure and tolerance to environmental stresses. In angiosperms (Magnoliophyta), ABA-induced gene expression is mediated by promoter elements such as the G-box-like ACGT-core motifs recognized by bZIP transcription factors. In contrast, the mode of regulation by ABA of gene expression in liverworts (Marchantiophyta), representing one of the earliest diverging land plant groups, has not been elucidated. In this study, we used promoters of the liverwort Marchantia polymorpha dehydrin and the wheat Em genes fused to the β-glucuronidase (GUS) reporter gene to investigate ABA-induced gene expression in liverworts. Transient assays of cultured cells of Marchantia indicated that ACGT-core motifs proximal to the transcription initiation site play a role in the ABA-induced gene expression. The RY sequence recognized by B3 transcriptional regulators was also shown to be responsible for the ABA-induced gene expression. In transgenic Marchantia plants, ABA treatment elicited an increase in GUS expression in young gemmalings, which was abolished by simultaneous disruption of the ACGT-core and RY elements. ABA-induced GUS expression was less obvious in mature thalli than in young gemmalings, associated with reductions in sensitivity to exogenous ABA during gametophyte growth. In contrast, lunularic acid, which had been suggested to function as an ABA-like substance, had no effect on GUS expression. The results demonstrate the presence of ABA-specific response mechanisms mediated by conserved cis-regulatory elements in liverworts, implying that the mechanisms had been acquired in the common ancestors of embryophytes. PMID:26456006

  18. The green-absorbing Drosophila Rh6 visual pigment contains a blue-shifting amino acid substitution that is conserved in vertebrates.

    PubMed

    Salcedo, Ernesto; Farrell, David M; Zheng, Lijun; Phistry, Meridee; Bagg, Eve E; Britt, Steven G

    2009-02-27

    The molecular mechanisms that regulate invertebrate visual pigment absorption are poorly understood. Through sequence analysis and functional investigation of vertebrate visual pigments, numerous amino acid substitutions important for this adaptive process have been identified. Here we describe a serine/alanine (S/A) substitution in long wavelength-absorbing Drosophila visual pigments that occurs at a site corresponding to Ala-292 in bovine rhodopsin. This S/A substitution accounts for a 10-17-nm absorption shift in visual pigments of this class. Additionally, we demonstrate that substitution of a cysteine at the same site, as occurs in the blue-absorbing Rh5 pigment, accounts for a 4-nm shift. Substitutions at this site are the first spectrally significant amino acid changes to be identified for invertebrate pigments sensitive to visible light and are the first evidence of a conserved tuning mechanism in vertebrate and invertebrate pigments of this class. PMID:19126545

  19. Conversion of amino-acid sequence in proteins to classical music: search for auditory patterns

    PubMed Central

    2007-01-01

    We have converted genome-encoded protein sequences into musical notes to reveal auditory patterns without compromising musicality. We derived a reduced range of 13 base notes by pairing similar amino acids and distinguishing them using variations of three-note chords and codon distribution to dictate rhythm. The conversion will help make genomic coding sequences more approachable for the general public, young children, and vision-impaired scientists. PMID:17477882

  20. Conservation and Expression Patterns Divergence of Ascorbic Acid d-mannose/l-galactose Pathway Genes in Brassica rapa.

    PubMed

    Duan, Weike; Ren, Jun; Li, Yan; Liu, Tongkun; Song, Xiaoming; Chen, Zhongwen; Huang, Zhinan; Hou, Xilin; Li, Ying

    2016-01-01

    Ascorbic acid (AsA) participates in diverse biological processes, is regulated by multiple factors and is a potent antioxidant and cellular reductant. The D-Mannose/L-Galactose pathway is a major plant AsA biosynthetic pathway that is highly connected within biosynthetic networks, and generally conserved across plants. Previous work has shown that, although most genes of this pathway are expressed under standard growth conditions in Brassica rapa, some paralogs of these genes are not. We hypothesize that regulatory evolution in duplicate AsA pathway genes has occurred as an adaptation to environmental stressors, and that gene retention has been influenced by polyploidation events in Brassicas. To test these hypotheses, we explored the conservation of these genes in Brassicas and their expression patterns divergence in B. rapa. Similar retention and a high degree of gene sequence similarity were identified in B. rapa (A genome), B. oleracea (C genome) and B. napus (AC genome). However, the number of genes that encode the same type of enzymes varied among the three plant species. With the exception of GMP, which has nine genes, there were one to four genes that encoded the other enzymes. Moreover, we found that expression patterns divergence widely exists among these genes. (i) VTC2 and VTC5 are paralogous genes, but only VTC5 is influenced by FLC. (ii) Under light treatment, PMI1 co-regulates the AsA pool size with other D-Man/L-Gal pathway genes, whereas PMI2 is regulated only by darkness. (iii) Under NaCl, Cu(2+), MeJA and wounding stresses, most of the paralogs exhibit different expression patterns. Additionally, GME and GPP are the key regulatory enzymes that limit AsA biosynthesis in response to these treatments. In conclusion, our data support that the conservative and divergent expression patterns of D-Man/L-Gal pathway genes not only avoid AsA biosynthesis network instability but also allow B. rapa to better adapt to complex environments. PMID:27313597

  1. Conservation and Expression Patterns Divergence of Ascorbic Acid d-mannose/l-galactose Pathway Genes in Brassica rapa

    PubMed Central

    Duan, Weike; Ren, Jun; Li, Yan; Liu, Tongkun; Song, Xiaoming; Chen, Zhongwen; Huang, Zhinan; Hou, Xilin; Li, Ying

    2016-01-01

    Ascorbic acid (AsA) participates in diverse biological processes, is regulated by multiple factors and is a potent antioxidant and cellular reductant. The D-Mannose/L-Galactose pathway is a major plant AsA biosynthetic pathway that is highly connected within biosynthetic networks, and generally conserved across plants. Previous work has shown that, although most genes of this pathway are expressed under standard growth conditions in Brassica rapa, some paralogs of these genes are not. We hypothesize that regulatory evolution in duplicate AsA pathway genes has occurred as an adaptation to environmental stressors, and that gene retention has been influenced by polyploidation events in Brassicas. To test these hypotheses, we explored the conservation of these genes in Brassicas and their expression patterns divergence in B. rapa. Similar retention and a high degree of gene sequence similarity were identified in B. rapa (A genome), B. oleracea (C genome) and B. napus (AC genome). However, the number of genes that encode the same type of enzymes varied among the three plant species. With the exception of GMP, which has nine genes, there were one to four genes that encoded the other enzymes. Moreover, we found that expression patterns divergence widely exists among these genes. (i) VTC2 and VTC5 are paralogous genes, but only VTC5 is influenced by FLC. (ii) Under light treatment, PMI1 co-regulates the AsA pool size with other D-Man/L-Gal pathway genes, whereas PMI2 is regulated only by darkness. (iii) Under NaCl, Cu2+, MeJA and wounding stresses, most of the paralogs exhibit different expression patterns. Additionally, GME and GPP are the key regulatory enzymes that limit AsA biosynthesis in response to these treatments. In conclusion, our data support that the conservative and divergent expression patterns of D-Man/L-Gal pathway genes not only avoid AsA biosynthesis network instability but also allow B. rapa to better adapt to complex environments. PMID:27313597

  2. Protein location prediction using atomic composition and global features of the amino acid sequence

    SciTech Connect

    Cherian, Betsy Sheena; Nair, Achuthsankar S.

    2010-01-22

    Subcellular location of protein is constructive information in determining its function, screening for drug candidates, vaccine design, annotation of gene products and in selecting relevant proteins for further studies. Computational prediction of subcellular localization deals with predicting the location of a protein from its amino acid sequence. For a computational localization prediction method to be more accurate, it should exploit all possible relevant biological features that contribute to the subcellular localization. In this work, we extracted the biological features from the full length protein sequence to incorporate more biological information. A new biological feature, distribution of atomic composition is effectively used with, multiple physiochemical properties, amino acid composition, three part amino acid composition, and sequence similarity for predicting the subcellular location of the protein. Support Vector Machines are designed for four modules and prediction is made by a weighted voting system. Our system makes prediction with an accuracy of 100, 82.47, 88.81 for self-consistency test, jackknife test and independent data test respectively. Our results provide evidence that the prediction based on the biological features derived from the full length amino acid sequence gives better accuracy than those derived from N-terminal alone. Considering the features as a distribution within the entire sequence will bring out underlying property distribution to a greater detail to enhance the prediction accuracy.

  3. Ab initio detection of fuzzy amino acid tandem repeats in protein sequences

    PubMed Central

    2012-01-01

    Background Tandem repetitions within protein amino acid sequences often correspond to regular secondary structures and form multi-repeat 3D assemblies of varied size and function. Developing internal repetitions is one of the evolutionary mechanisms that proteins employ to adapt their structure and function under evolutionary pressure. While there is keen interest in understanding such phenomena, detection of repeating structures based only on sequence analysis is considered an arduous task, since structure and function is often preserved even under considerable sequence divergence (fuzzy tandem repeats). Results In this paper we present PTRStalker, a new algorithm for ab-initio detection of fuzzy tandem repeats in protein amino acid sequences. In the reported results we show that by feeding PTRStalker with amino acid sequences from the UniProtKB/Swiss-Prot database we detect novel tandemly repeated structures not captured by other state-of-the-art tools. Experiments with membrane proteins indicate that PTRStalker can detect global symmetries in the primary structure which are then reflected in the tertiary structure. Conclusions PTRStalker is able to detect fuzzy tandem repeating structures in protein sequences, with performance beyond the current state-of-the art. Such a tool may be a valuable support to investigating protein structural properties when tertiary X-ray data is not available. PMID:22536906

  4. Multimodal phylogeny for taxonomy: integrating information from nucleotide and amino acid sequences.

    PubMed

    Bicego, Manuele; Dellaglio, Franco; Felis, Giovanna E

    2007-10-01

    The crucial role played by the analysis of microbial diversity in biotechnology-based innovations has increased the interest in the microbial taxonomy research area. Phylogenetic sequence analyses have contributed significantly to the advances in this field, also in the view of the large amount of sequence data collected in recent years. Phylogenetic analyses could be realized on the basis of protein-encoding nucleotide sequences or encoded amino acid molecules: these two mechanisms present different peculiarities, still starting from two alternative representations of the same information. This complementarity could be exploited to achieve a multimodal phylogenetic scheme that is able to integrate gene and protein information in order to realize a single final tree. This aspect has been poorly addressed in the literature. In this paper, we propose to integrate the two phylogenetic analyses using basic schemes derived from the multimodality fusion theory (or multiclassifier systems theory), a well-founded and rigorous branch for which its powerfulness has already been demonstrated in other pattern recognition contexts. The proposed approach could be applied to distance matrix-based phylogenetic techniques (like neighbor joining), resulting in a smart and fast method. The proposed methodology has been tested in a real case involving sequences of some species of lactic acid bacteria. With this dataset, both nucleotide sequence- and amino acid sequence-based phylogenetic analyses present some drawbacks, which are overcome with the multimodal analysis. PMID:17933011

  5. Simian immunodeficiency virus (mac 251-32H) transmembrane protein sequence remains conserved throughout the course of infection in macaques.

    PubMed

    Slade, A; Jones, S; Almond, N; Kitchin, P

    1993-02-01

    Two cynomolgus macaques were infected with a genetically complex challenge stock of simian immunodeficiency virus (SIVmac251-32H). The polymerase chain reaction (PCR) was used to amplify the env gp41, rev, and nef overlapping coding sequences from provirus present in the blood of both animals at 1, 6, and 15 months post infection (p.i.). The predominant, env sequences found in both animals at the three time points were very similar to that found in the original 11/88 challenge stock. The functionally important hydrophobic fusion and membrane-spanning domains within gp41 remained conserved throughout the course of infection. Nucleotide variation within the region corresponding to the REV response element (RRE) was limited to four positions, none of which were predicted to cause any significant disruption to the secondary structure of the RRE. Very little genetic variation was observed in and around the cluster of potential glycosylation sites of the external portion of gp41. However, the existence of a previously assigned variable region elsewhere in the cytoplasmic domain of gp41 was confirmed. The three gene loci (env, rev, and nef) examined varied independently. All changes in the predominant protein sequences were brought about by single nucleotide substitutions only. After 15 months of infection with SIV, 1 animal was sick from SIV-induced disease whereas the other remained healthy. In-frame stop codons within the transmembrane protein occurred with a much greater frequency in the healthy animal. PMID:8457380

  6. Comparative Mitogenomics of the Genus Odontobutis (Perciformes: Gobioidei: Odontobutidae) Revealed Conserved Gene Rearrangement and High Sequence Variations

    PubMed Central

    Ma, Zhihong; Yang, Xuefen; Bercsenyi, Miklos; Wu, Junjie; Yu, Yongyao; Wei, Kaijian; Fan, Qixue; Yang, Ruibin

    2015-01-01

    To understand the molecular evolution of mitochondrial genomes (mitogenomes) in the genus Odontobutis, the mitogenome of Odontobutis yaluensis was sequenced and compared with those of another four Odontobutis species. Our results displayed similar mitogenome features among species in genome organization, base composition, codon usage, and gene rearrangement. The identical gene rearrangement of trnS-trnL-trnH tRNA cluster observed in mitogenomes of these five closely related freshwater sleepers suggests that this unique gene order is conserved within Odontobutis. Additionally, the present gene order and the positions of associated intergenic spacers of these Odontobutis mitogenomes indicate that this unusual gene rearrangement results from tandem duplication and random loss of large-scale gene regions. Moreover, these mitogenomes exhibit a high level of sequence variation, mainly due to the differences of corresponding intergenic sequences in gene rearrangement regions and the heterogeneity of tandem repeats in the control regions. Phylogenetic analyses support Odontobutis species with shared gene rearrangement forming a monophyletic group, and the interspecific phylogenetic relationships are associated with structural differences among their mitogenomes. The present study contributes to understanding the evolutionary patterns of Odontobutidae species. PMID:26492246

  7. The complete mitochondrial genome sequence of the tubeworm Lamellibrachia satsuma and structural conservation in the mitochondrial genome control regions of Order Sabellida.

    PubMed

    Patra, Ajit Kumar; Kwon, Yong Min; Kang, Sung Gyun; Fujiwara, Yoshihiro; Kim, Sang-Jin

    2016-04-01

    The control region of the mitochondrial genomes shows high variation in conserved sequence organizations, which follow distinct evolutionary patterns in different species or taxa. In this study, we sequenced the complete mitochondrial genome of Lamellibrachia satsuma from the cold-seep region of Kagoshima Bay, as a part of whole genome study and extensively studied the structural features and patterns of the control region sequences. We obtained 15,037 bp of mitochondrial genome using Illumina sequencing and identified the non-coding AT-rich region or control region (354 bp, AT=83.9%) located between trnH and trnR. We found 7 conserved sequence blocks (CSB), scattered throughout the control region of L. satsuma and other taxa of Annelida. The poly-TA stretches, which commonly form the stem of multiple stem-loop structures, are most conserved in the CSB-I and CSB-II regions. The mitochondrial genome of L. satsuma encodes a unique repetitive sequence in the control region, which forms a unique secondary structure in comparison to Lamellibrachia luymesi. Phylogenetic analyses of all protein-coding genes indicate that L. satsuma forms a monophyletic clade with L. luymesi along with other tubeworms found in cold-seep regions (genera: Lamellibrachia, Escarpia, and Seepiophila). In general, the control region sequences of Annelida could be aligned with certainty within each genus, and to some extent within the family, but with a higher rate of variation in conserved regions. PMID:26776396

  8. Draft genome sequence of the docosahexaenoic acid producing thraustochytrid Aurantiochytrium sp. T66.

    PubMed

    Liu, Bin; Ertesvåg, Helga; Aasen, Inga Marie; Vadstein, Olav; Brautaset, Trygve; Heggeset, Tonje Marita Bjerkan

    2016-06-01

    Thraustochytrids are unicellular, marine protists, and there is a growing industrial interest in these organisms, particularly because some species, including strains belonging to the genus Aurantiochytrium, accumulate high levels of docosahexaenoic acid (DHA). Here, we report the draft genome sequence of Aurantiochytrium sp. T66 (ATCC PRA-276), with a size of 43 Mbp, and 11,683 predicted protein-coding sequences. The data has been deposited at DDBJ/EMBL/Genbank under the accession LNGJ00000000. The genome sequence will contribute new insight into DHA biosynthesis and regulation, providing a basis for metabolic engineering of thraustochytrids. PMID:27222814

  9. The Arachidonic Acid Metabolome Serves as a Conserved Regulator of Cholesterol Metabolism

    PubMed Central

    Demetz, Egon; Schroll, Andrea; Auer, Kristina; Heim, Christiane; Patsch, Josef R.; Eller, Philipp; Theurl, Markus; Theurl, Igor; Theurl, Milan; Seifert, Markus; Lener, Daniela; Stanzl, Ursula; Haschka, David; Asshoff, Malte; Dichtl, Stefanie; Nairz, Manfred; Huber, Eva; Stadlinger, Martin; Moschen, Alexander R.; Li, Xiaorong; Pallweber, Petra; Scharnagl, Hubert; Stojakovic, Tatjana; März, Winfried; Kleber, Marcus E.; Garlaschelli, Katia; Uboldi, Patrizia; Catapano, Alberico L.; Stellaard, Frans; Rudling, Mats; Kuba, Keiji; Imai, Yumiko; Arita, Makoto; Schuetz, John D.; Pramstaller, Peter P.; Tietge, Uwe J.F.; Trauner, Michael; Norata, Giuseppe D.; Claudel, Thierry; Hicks, Andrew A.; Weiss, Guenter; Tancevski, Ivan

    2014-01-01

    Summary Cholesterol metabolism is closely interrelated with cardiovascular disease in humans. Dietary supplementation with omega-6 polyunsaturated fatty acids including arachidonic acid (AA) was shown to favorably affect plasma LDL-C and HDL-C. However, the underlying mechanisms are poorly understood. By combining data from a GWAS screening in >100,000 individuals of European ancestry, mediator lipidomics, and functional validation studies in mice, we identify the AA metabolome as an important regulator of cholesterol homeostasis. Pharmacological modulation of AA metabolism by aspirin induced hepatic generation of leukotrienes (LTs) and lipoxins (LXs), thereby increasing hepatic expression of the bile salt export pump Abcb11. Induction of Abcb11 translated in enhanced reverse cholesterol transport, one key function of HDL. Further characterization of the bioactive AA-derivatives identified LX mimetics to lower plasma LDL-C. Our results define the AA metabolome as conserved regulator of cholesterol metabolism, and identify AA derivatives as promising therapeutics to treat cardiovascular disease in humans. PMID:25444678

  10. Adiponectin receptor 1 conserves docosahexaenoic acid and promotes photoreceptor cell survival

    PubMed Central

    Rice, Dennis S.; Calandria, Jorgelina M.; Gordon, William C.; Jun, Bokkyoo; Zhou, Yongdong; Gelfman, Claire M.; Li, Songhua; Jin, Minghao; Knott, Eric J.; Chang, Bo; Abuin, Alex; Issa, Tawfik; Potter, David; Platt, Kenneth A.; Bazan, Nicolas G.

    2015-01-01

    The identification of pathways necessary for photoreceptor and retinal pigment epithelium (RPE) function is critical to uncover therapies for blindness. Here we report the discovery of adiponectin receptor 1 (AdipoR1) as a regulator of these cells’ functions. Docosahexaenoic acid (DHA) is avidly retained in photoreceptors, while mechanisms controlling DHA uptake and retention are unknown. Thus, we demonstrate that AdipoR1 ablation results in DHA reduction. In situ hybridization reveals photoreceptor and RPE cell AdipoR1 expression, blunted in AdipoR1−/− mice. We also find decreased photoreceptor-specific phosphatidylcholine containing very long-chain polyunsaturated fatty acids and severely attenuated electroretinograms. These changes precede progressive photoreceptor degeneration in AdipoR1−/− mice. RPE-rich eyecup cultures from AdipoR1−/− reveal impaired DHA uptake. AdipoR1 overexpression in RPE cells enhances DHA uptake, whereas AdipoR1 silencing has the opposite effect. These results establish AdipoR1 as a regulatory switch of DHA uptake, retention, conservation and elongation in photoreceptors and RPE, thus preserving photoreceptor cell integrity. PMID:25736573

  11. A classification of glycosyl hydrolases based on amino acid sequence similarities.

    PubMed Central

    Henrissat, B

    1991-01-01

    The amino acid sequences of 301 glycosyl hydrolases and related enzymes have been compared. A total of 291 sequences corresponding to 39 EC entries could be classified into 35 families. Only ten sequences (less than 5% of the sample) could not be assigned to any family. With the sequences available for this analysis, 18 families were found to be monospecific (containing only one EC number) and 17 were found to be polyspecific (containing at least two EC numbers). Implications on the folding characteristics and mechanism of action of these enzymes and on the evolution of carbohydrate metabolism are discussed. With the steady increase in sequence and structural data, it is suggested that the enzyme classification system should perhaps be revised. PMID:1747104

  12. New families in the classification of glycosyl hydrolases based on amino acid sequence similarities.

    PubMed Central

    Henrissat, B; Bairoch, A

    1993-01-01

    301 glycosyl hydrolases and related enzymes corresponding to 39 EC entries of the I.U.B. classification system have been classified into 35 families on the basis of amino-acid-sequence similarities [Henrissat (1991) Biochem. J. 280, 309-316]. Approximately half of the families were found to be monospecific (containing only one EC number), whereas the other half were found to be polyspecific (containing at least two EC numbers). A > 60% increase in sequence data for glycosyl hydrolases (181 additional enzymes or enzyme domains sequences have since become available) allowed us to update the classification not only by the addition of more members to already identified families, but also by the finding of ten new families. On the basis of a comparison of 482 sequences corresponding to 52 EC entries, 45 families, out of which 22 are polyspecific, can now be defined. This classification has been implemented in the SWISS-PROT protein sequence data bank. PMID:8352747

  13. Sequence-specific purification of nucleic acids by PNA-controlled hybrid selection.

    PubMed

    Orum, H; Nielsen, P E; Jørgensen, M; Larsson, C; Stanley, C; Koch, T

    1995-09-01

    Using an oligohistidine peptide nucleic acids (oligohistidine-PNA) chimera, we have developed a rapid hybrid selection method that allows efficient, sequence-specific purification of a target nucleic acid. The method exploits two fundamental features of PNA. First, that PNA binds with high affinity and specificity to its complementary nucleic acid. Second, that amino acids are easily attached to the PNA oligomer during synthesis. We show that a (His)6-PNA chimera exhibits strong binding to chelated Ni2+ ions without compromising its native PNA hybridization properties. We further show that these characteristics allow the (His)6-PNA/DNA complex to be purified by the well-established method of metal ion affinity chromatography using a Ni(2+)-NTA (nitrilotriactic acid) resin. Specificity and efficiency are the touchstones of any nucleic acid purification scheme. We show that the specificity of the (His)6-PNA selection approach is such that oligonucleotides differing by only a single nucleotide can be selectively purified. We also show that large RNAs (2224 nucleotides) can be captured with high efficiency by using multiple (His)6-PNA probes. PNA can hybridize to nucleic acids in low-salt concentrations that destabilize native nucleic acid structures. We demonstrate that this property of PNA can be utilized to purify an oligonucleotide in which the target sequence forms part of an intramolecular stem/loop structure. PMID:7495562

  14. Canine Polydactyl Mutations With Heterogeneous Origin in the Conserved Intronic Sequence of LMBR1

    PubMed Central

    Park, Kiyun; Kang, Joohyun; Subedi, Krishna Pd.; Ha, Ji-Hong; Park, Chankyu

    2008-01-01

    Canine preaxial polydactyly (PPD) in the hind limb is a developmental trait that restores the first digit lost during canine evolution. Using a linkage analysis, we previously demonstrated that the affected gene in a Korean breed is located on canine chromosome 16. The candidate locus was further limited to a linkage disequilibrium (LD) block of <213 kb composing the single gene, LMBR1, by LD mapping with single nucleotide polymorphisms (SNPs) for affected individuals from both Korean and Western breeds. The ZPA regulatory sequence (ZRS) in intron 5 of LMBR1 was implicated in mammalian polydactyly. An analysis of the LD haplotypes around the ZRS for various dog breeds revealed that only a subset is assigned to Western breeds. Furthermore, two distinct affected haplotypes for Asian and Western breeds were found, each containing different single-base changes in the upstream sequence (pZRS) of the ZRS. Unlike the previously characterized cases of PPD identified in the mouse and human ZRS regions, the canine mutations in pZRS lacked the ectopic expression of sonic hedgehog in the anterior limb bud, distinguishing its role in limb development from that of the ZRS. PMID:18689889

  15. Extremely Acidophilic Protists from Acid Mine Drainage Host Rickettsiales-Lineage Endosymbionts That Have Intervening Sequences in Their 16S rRNA Genes

    PubMed Central

    Baker, Brett J.; Hugenholtz, Philip; Dawson, Scott C.; Banfield, Jillian F.

    2003-01-01

    During a molecular phylogenetic survey of extremely acidic (pH < 1), metal-rich acid mine drainage habitats in the Richmond Mine at Iron Mountain, Calif., we detected 16S rRNA gene sequences of a novel bacterial group belonging to the order Rickettsiales in the Alphaproteobacteria. The closest known relatives of this group (92% 16S rRNA gene sequence identity) are endosymbionts of the protist Acanthamoeba. Oligonucleotide 16S rRNA probes were designed and used to observe members of this group within acidophilic protists. To improve visualization of eukaryotic populations in the acid mine drainage samples, broad-specificity probes for eukaryotes were redesigned and combined to highlight this component of the acid mine drainage community. Approximately 4% of protists in the acid mine drainage samples contained endosymbionts. Measurements of internal pH of the protists showed that their cytosol is close to neutral, indicating that the endosymbionts may be neutrophilic. The endosymbionts had a conserved 273-nucleotide intervening sequence (IVS) in variable region V1 of their 16S rRNA genes. The IVS does not match any sequence in current databases, but the predicted secondary structure forms well-defined stem loops. IVSs are uncommon in rRNA genes and appear to be confined to bacteria living in close association with eukaryotes. Based on the phylogenetic novelty of the endosymbiont sequences and initial culture-independent characterization, we propose the name “Candidatus Captivus acidiprotistae.” To our knowledge, this is the first report of an endosymbiotic relationship in an extremely acidic habitat. PMID:12957940

  16. Conservative forgetful scholars: How people learn causal structure through sequences of interventions.

    PubMed

    Bramley, Neil R; Lagnado, David A; Speekenbrink, Maarten

    2015-05-01

    Interacting with a system is key to uncovering its causal structure. A computational framework for interventional causal learning has been developed over the last decade, but how real causal learners might achieve or approximate the computations entailed by this framework is still poorly understood. Here we describe an interactive computer task in which participants were incentivized to learn the structure of probabilistic causal systems through free selection of multiple interventions. We develop models of participants' intervention choices and online structure judgments, using expected utility gain, probability gain, and information gain and introducing plausible memory and processing constraints. We find that successful participants are best described by a model that acts to maximize information (rather than expected score or probability of being correct); that forgets much of the evidence received in earlier trials; but that mitigates this by being conservative, preferring structures consistent with earlier stated beliefs. We explore 2 heuristics that partly explain how participants might be approximating these models without explicitly representing or updating a hypothesis space. PMID:25329086

  17. Amino acid sequence of a vitamin K-dependent Ca2+-binding peptide from bovine prothrombin.

    PubMed

    Howard, J B; Fausch, M D

    1975-08-10

    The amino acid sequence of a 31-residue peptide from bovine prothrombin has been determined. This peptide has been shown to contain the vitamin K-dependent modification required for Ca2+ binding (Nelsestuen, G. L., and Suttie, J. W. (1973) Proc. Natl. Acad. Sci. U. S. A. 70, 3366-3370) and the modified amino acid, gamma-carboxyglutamic acid (Nelsestuen, G. L., Zytkovicz, T., and Howard, J. B. (1974) J. Biol. Chem. 249, 6347-6350). The peptide was shown to correspond to residues 12 to 42 of prothrombin. PMID:807581

  18. Amino acid sequences around the cysteine residues of rabbit muscle triose phosphate isomerase

    PubMed Central

    Miller, Janet C.; Waley, S. G.

    1971-01-01

    1. The nature of the subunits in rabbit muscle triose phosphate isomerase has been investigated. 2. Amino acid analyses show that there are five cysteine residues and two methionine residues/subunit. 3. The amino acid sequences around the cysteine residues have been determined; these account for about 75 residues. 4. Cleavage at the methionine residues with cyanogen bromide gave three fragments. 5. These results show that the subunits correspond to polypeptide chains, containing about 230 amino acid residues. The chains in triose phosphate isomerase seem to be shorter than those of other glycolytic enzymes. PMID:5165707

  19. Complete amino acid sequence of the Mu heavy chain of a human IgM immunoglobulin.

    PubMed

    Putnam, F W; Florent, G; Paul, C; Shinoda, T; Shimizu, A

    1973-10-19

    The amino acid sequence of the micro, chain of a human IgM immunoglobulin, including the location of all disulfide bridges and oligosaccharides, has been determined. The homology of the constant regions of immunoglobulin micro, gamma, alpha, and epsilon heavy chains reveals evolutionary relationships and suggests that two genes code for each heavy chain. PMID:4742735

  20. Draft Genome Sequence of the Butyric Acid Producer Clostridium tyrobutyricum Strain CIP I-776 (IFP923)

    PubMed Central

    Clément, Benjamin; Lopes Ferreira, Nicolas

    2016-01-01

    Here, we report the draft genome sequence of Clostridium tyrobutyricum CIP I-776 (IFP923), an efficient producer of butyric acid. The genome consists of a single chromosome of 3.19 Mb and provides useful data concerning the metabolic capacities of the strain. PMID:26941139

  1. Draft Genome Sequence of Perfluorooctane Acid-Degrading Bacterium Pseudomonas parafulva YAB-1

    PubMed Central

    Tang, Chongjian; Peng, Qingjing; Peng, Qingzhong

    2015-01-01

    Pseudomonas parafulva YAB-1, isolated from perfluorinated compound-contaminated soil, has the ability to degrade perfluorooctane acid (PFOA) compound. Here, we report the draft genome sequence and annotation of the PFOA-degrading bacterium P. parafulva YAB-1. The data provide the basis to investigate the molecular mechanism of PFOA metabolism. PMID:26337877

  2. Conserved sequences in the carboxyl terminus of integrase that are essential for human immunodeficiency virus type 1 replication.

    PubMed

    Cannon, P M; Byles, E D; Kingsman, S M; Kingsman, A J

    1996-01-01

    We have previously identified a residue in the carboxyl terminus of human immunodeficiency virus type 1 integrase (HIV-1 IN), W-235, the requirement for which is only revealed in viral assays for integrase function (P. M. Cannon, W. Wilson, E. Byles, S. M. Kingsman, and A. J. Kingsman, J. Virol. 68:4768-4775, 1994). Our further analysis of this region of retroviral IN has now identified several sequence motifs which are conserved in all the retroviruses we examined, apart from human spumaretrovirus. We have made mutations within these motifs in HIV-1 IN and examined their phenotypes when reintroduced into an infectious proviral clone. The deleterious effects of several of these mutations demonstrate the importance of these regions for IN function in vivo. We observed a further discrepancy, at a motif that is only conserved in the lentiviruses, in the ability of mutants to function in in vitro and in vivo assays. Substitutions both in this region and at W-235 abolish HIV-1 infectivity but do not affect particle production, morphology, reverse transcription, or nuclear import in T-cell lines. Taken together with the in vitro data suggesting that neither of these residues is directly involved in the catalytic reactions of IN, it seems likely that we have identified regions of IN that are essential for interactions with other components of the integration machinery. PMID:8523588

  3. Structural and sequence similarities of hydra xeroderma pigmentosum A protein to human homolog suggest early evolution and conservation.

    PubMed

    Barve, Apurva; Ghaskadbi, Saroj; Ghaskadbi, Surendra

    2013-01-01

    Xeroderma pigmentosum group A (XPA) is a protein that binds to damaged DNA, verifies presence of a lesion, and recruits other proteins of the nucleotide excision repair (NER) pathway to the site. Though its homologs from yeast, Drosophila, humans, and so forth are well studied, XPA has not so far been reported from protozoa and lower animal phyla. Hydra is a fresh-water cnidarian with a remarkable capacity for regeneration and apparent lack of organismal ageing. Cnidarians are among the first metazoa with a defined body axis, tissue grade organisation, and nervous system. We report here for the first time presence of XPA gene in hydra. Putative protein sequence of hydra XPA contains nuclear localization signal and bears the zinc-finger motif. It contains two conserved Pfam domains and various characterized features of XPA proteins like regions for binding to excision repair cross-complementing protein-1 (ERCC1) and replication protein A 70 kDa subunit (RPA70) proteins. Hydra XPA shows a high degree of similarity with vertebrate homologs and clusters with deuterostomes in phylogenetic analysis. Homology modelling corroborates the very close similarity between hydra and human XPA. The protein thus most likely functions in hydra in the same manner as in other animals, indicating that it arose early in evolution and has been conserved across animal phyla. PMID:24083246

  4. Structural and Sequence Similarities of Hydra Xeroderma Pigmentosum A Protein to Human Homolog Suggest Early Evolution and Conservation

    PubMed Central

    Ghaskadbi, Saroj

    2013-01-01

    Xeroderma pigmentosum group A (XPA) is a protein that binds to damaged DNA, verifies presence of a lesion, and recruits other proteins of the nucleotide excision repair (NER) pathway to the site. Though its homologs from yeast, Drosophila, humans, and so forth are well studied, XPA has not so far been reported from protozoa and lower animal phyla. Hydra is a fresh-water cnidarian with a remarkable capacity for regeneration and apparent lack of organismal ageing. Cnidarians are among the first metazoa with a defined body axis, tissue grade organisation, and nervous system. We report here for the first time presence of XPA gene in hydra. Putative protein sequence of hydra XPA contains nuclear localization signal and bears the zinc-finger motif. It contains two conserved Pfam domains and various characterized features of XPA proteins like regions for binding to excision repair cross-complementing protein-1 (ERCC1) and replication protein A 70 kDa subunit (RPA70) proteins. Hydra XPA shows a high degree of similarity with vertebrate homologs and clusters with deuterostomes in phylogenetic analysis. Homology modelling corroborates the very close similarity between hydra and human XPA. The protein thus most likely functions in hydra in the same manner as in other animals, indicating that it arose early in evolution and has been conserved across animal phyla. PMID:24083246

  5. High-Throughput Sequencing Reveals Diverse Sets of Conserved, Nonconserved, and Species-Specific miRNAs in Jute

    PubMed Central

    Islam, Md. Tariqul; Ferdous, Ahlan Sabah; Najnin, Rifat Ara; Sarker, Suprovath Kumar; Khan, Haseena

    2015-01-01

    MicroRNAs play a pivotal role in regulating a broad range of biological processes, acting by cleaving mRNAs or by translational repression. A group of plant microRNAs are evolutionarily conserved; however, others are expressed in a species-specific manner. Jute is an agroeconomically important fibre crop; nonetheless, no practical information is available for microRNAs in jute to date. In this study, Illumina sequencing revealed a total of 227 known microRNAs and 17 potential novel microRNA candidates in jute, of which 164 belong to 23 conserved families and the remaining 63 belong to 58 nonconserved families. Among a total of 81 identified microRNA families, 116 potential target genes were predicted for 39 families and 11 targets were predicted for 4 among the 17 identified novel microRNAs. For understanding better the functions of microRNAs, target genes were analyzed by Gene Ontology and their pathways illustrated by KEGG pathway analyses. The presence of microRNAs identified in jute was validated by stem-loop RT-PCR followed by end point PCR and qPCR for randomly selected 20 known and novel microRNAs. This study exhaustively identifies microRNAs and their target genes in jute which will ultimately pave the way for understanding their role in this crop and other crops. PMID:25861616

  6. Complete mitochondrial DNA sequence of the endangered giant sable antelope (Hippotragus niger variani): insights into conservation and taxonomy.

    PubMed

    Espregueira Themudo, Gonçalo; Rufino, Ana C; Campos, Paula F

    2015-02-01

    The giant sable antelope is one of the most endangered African bovids. Populations of this iconic animal, the national symbol of Angola, were recently rediscovered, after many decades of presumed extinction. Even so, their numbers are scarce and hence conservation plans are essential. However, fundamental information such as its taxonomic position, time of divergence and degree of genetic variation are still lacking. Here, we used a museum preserved horn as a source of DNA to describe, for the first time, the complete mitochondrial genome of the giant sable antelope, and provide insights into its evolutionary history. Reads generated by shotgun sequencing were mapped against the mitochondrial genome of common sable antelope and the nuclear genomes of cow and sheep. Phylogenetic reconstruction and divergence time estimate give support to the monophyly of the giant sable and a maximum divergence time of 170 thousand years to the closest subspecies. About 7% of the nuclear genome was mapped against the reference. The genetic resources reported here are now available for future work in the field of conservation genetics and phylogeny, in this and related species. PMID:25527983

  7. High-Throughput Sequencing Reveals Diverse Sets of Conserved, Nonconserved, and Species-Specific miRNAs in Jute.

    PubMed

    Islam, Md Tariqul; Ferdous, Ahlan Sabah; Najnin, Rifat Ara; Sarker, Suprovath Kumar; Khan, Haseena

    2015-01-01

    MicroRNAs play a pivotal role in regulating a broad range of biological processes, acting by cleaving mRNAs or by translational repression. A group of plant microRNAs are evolutionarily conserved; however, others are expressed in a species-specific manner. Jute is an agroeconomically important fibre crop; nonetheless, no practical information is available for microRNAs in jute to date. In this study, Illumina sequencing revealed a total of 227 known microRNAs and 17 potential novel microRNA candidates in jute, of which 164 belong to 23 conserved families and the remaining 63 belong to 58 nonconserved families. Among a total of 81 identified microRNA families, 116 potential target genes were predicted for 39 families and 11 targets were predicted for 4 among the 17 identified novel microRNAs. For understanding better the functions of microRNAs, target genes were analyzed by Gene Ontology and their pathways illustrated by KEGG pathway analyses. The presence of microRNAs identified in jute was validated by stem-loop RT-PCR followed by end point PCR and qPCR for randomly selected 20 known and novel microRNAs. This study exhaustively identifies microRNAs and their target genes in jute which will ultimately pave the way for understanding their role in this crop and other crops. PMID:25861616

  8. The amino acid sequence of cytochrome c-555 from the methane-oxidizing bacterium Methylococcus capsulatus.

    PubMed Central

    Ambler, R P; Dalton, H; Meyer, T E; Bartsch, R G; Kamen, M D

    1986-01-01

    The amino acid sequence of the cytochrome c-555 from the obligate methanotroph Methylococcus capsulatus strain Bath (N.C.I.B. 11132) was determined. It is a single polypeptide chain of 96 residues, binding a haem group through the cysteine residues at positions 19 and 22, and the only methionine residue is a position 59. The sequence does not closely resemble that of any other cytochrome c that has yet been characterized. Detailed evidence for the amino acid sequence of the protein has been deposited as Supplementary Publication SUP 50131 (12 pages) at the British Library Lending Division, Boston Spa, West Yorkshire LS23 7BQ, U.K., from whom copies are available on prepayment. PMID:3006666

  9. Conserved biosynthetic pathways for phosalacine, bialaphos and newly discovered phosphonic acid natural products.

    PubMed

    Blodgett, Joshua A V; Zhang, Jun Kai; Yu, Xiaomin; Metcalf, William W

    2016-01-01

    Natural products containing phosphonic or phosphinic acid functionalities often display potent biological activities with applications in medicine and agriculture. The herbicide phosphinothricin-tripeptide (PTT) was the first phosphinate natural product discovered, yet despite numerous studies, questions remain surrounding key transformations required for its biosynthesis. In particular, the enzymology required to convert phosphonoformate to carboxyphosphonoenolpyruvate and the mechanisms underlying phosphorus methylation remain poorly understood. In addition, the model for non-ribosomal peptide synthetase assembly of the intact tripeptide product has undergone numerous revisions that have yet to be experimentally tested. To further investigate the biosynthesis of this unusual natural product, we completely sequenced the PTT biosynthetic locus from Streptomyces hygroscopicus and compared it with the orthologous cluster from Streptomyces viridochromogenes. We also sequenced and analyzed the closely related phosalacine (PAL) biosynthetic locus from Kitasatospora phosalacinea. Using data drawn from the comparative analysis of the PTT and PAL pathways, we also evaluate three related recently discovered phosphonate biosynthetic loci from Streptomyces sviceus, Streptomyces sp. WM6386 and Frankia alni. Our observations address long-standing biosynthetic questions related to PTT and PAL production and suggest that additional members of this pharmacologically important class await discovery. PMID:26328935

  10. Conserved biosynthetic pathways for phosalacine, bialaphos and newly discovered phosphonic acid natural products

    PubMed Central

    Blodgett, Joshua A. V; Zhang, Jun Kai; Yu, Xiaomin; Metcalf, William W.

    2015-01-01

    Natural products containing phosphonic or phosphinic acid functionalities often display potent biological activities with applications in medicine and agriculture. The herbicide phosphinothricin-tripeptide (PTT) was the first phosphinate natural product discovered, yet despite numerous studies, questions remain surrounding key transformations required for its biosynthesis. In particular, the enzymology required to convert phosphonoformate to carboxyphosphonoenolpyruvate and the mechanisms underlying phosphorus-methylation remain poorly understood. In addition, the model for NRPS assembly of the intact tripeptide product has undergone numerous revisions that have yet to be experimentally tested. To further investigate the biosynthesis of this unusual natural product, we completely sequenced the PTT biosynthetic locus from Streptomyces hygroscopicus and compared it to the orthologous cluster from Streptomyces viridochromogenes. We also sequenced and analysed the closely related phosalacine (PAL) biosynthetic locus from Kitasatospora phosalacinea. Using data drawn from the comparative analysis of the PTT and PAL pathways, we also evaluate three related recently discovered phosphonate biosynthetic loci from Streptomyces sviceus, Streptomyces sp. WM6386 and Frankia alni. Our observations address long-standing biosynthetic questions related to PTT and PAL production and suggest that additional members of this pharmacologically important class await discovery. PMID:26328935

  11. Unique genomic sequences in human chromosome 16p are conserved in the great apes.

    PubMed

    Tarzami, S T; Kringstein, A M; Conte, R A; Verma, R S

    1997-01-27

    In humans, acute myelomonocytic leukemia (AMML) with abnormal bone marrow eosinophilia is diagnosed by the presence of a pericentric inversion in chromosome 16, involving breakpoints p13;q23 [i.e., inv(16)(p13;q23)]. A pericentric inversion involves breaks that have occurred on the p and q arms and the segment in between is rotated 180 degrees and reattaches. The recent development of a "human micro-coatasome" painting probe for 16p contains unique DNA sequences that fluorescently label only the short arm of chromosome 16, which facilitates the identification of such inversions and represents an ideal tool for analyzing the "divergence/convergence" of the equivalent human chromosome 16 (PTR 18, GGO 17 and PPY 19) in the great apes, chimpanzee, gorilla and orangutan. When the probe is used on the type of pericentric inversion characteristic of AMML, signals are observed on the proximal portions (the regions closest to the centromere) of the long and short arms of chromosome 16. The probe hybridized to only the short arm of all three ape chromosomes and signals were not observed on the long arms, suggesting that a pericentric inversion similar to that seen in AMML has not occurred in any of these great apes. PMID:9037113

  12. Allelic polymorphism in arabian camel ribonuclease and the amino acid sequence of bactrian camel ribonuclease.

    PubMed

    Welling, G W; Mulder, H; Beintema, J J

    1976-04-01

    Pancreatic ribonucleases from several species (whitetail deer, roe deer, guinea pig, and arabian camel) exhibit more than one amino acid at particular positions in their amino acid sequences. Since these enzymes were isolated from pooled pancreas, the origin of this heterogeneity is not clear. The pancreatic ribonucleases from 11 individual arabian camels (Camelus dromedarius) have been investigated with respect to the lysine-glutamine heterogeneity at position 103 (Welling et al., 1975). Six ribonucleases showed only one basic band and five showed two bands after polyacrylamide gel electrophoresis, suggesting a gene frequency of about 0.75 for the Lys gene and about 0.25 for the Gln gene. The amino acid sequence of bactrian camel (Camelus bactrianus) ribonuclease isolated from individual pancreatic tissue was determined and compared with that of arabian camel ribonuclease. The only difference was observed at position 103. In the ribonucleases from two unrelated bactrian camels, only glutamine was observed at that position. PMID:962846

  13. Conservation of the function counts: homologous neurons express sequence-related neuropeptides that originate from different genes.

    PubMed

    Neupert, Susanne; Huetteroth, Wolf; Schachtner, Joachim; Predel, Reinhard

    2009-11-01

    By means of single-cell matrix assisted laser desorption/ionization time-of-flight mass spectrometry, we analysed neuropeptide expression in all FXPRLamide/pheromone biosynthesis activating neuropeptide synthesizing neurons of the adult tobacco hawk moth, Manduca sexta. Mass spectra clearly suggest a completely identical processing of the pheromone biosynthesis activating neuropeptide-precursor in the mandibular, maxillary and labial neuromeres of the subesophageal ganglion. Only in the pban-neurons of the labial neuromere, products of two neuropeptide genes, namely the pban-gene and the capa-gene, were detected. Both of these genes expressed, amongst others, sequence-related neuropeptides (extended WFGPRLamides). We speculate that the expression of the two neuropeptide genes is a plesiomorph character typical of moths. A detailed examination of the neuroanatomy and the peptidome of the (two) pban-neurons in the labial neuromere of moths with homologous neurons of different insects indicates a strong conservation of the function of this neuroendocrine system. In other insects, however, the labial neurons either express products of the fxprl-gene or products of the capa-gene. The processing of the respective genes is reduced to extended WFGPRLamides in each case and yields a unique peptidome in the labial cells. Thus, sequence-related messenger molecules are always produced in these cells and it seems that the respective neurons recruited different neuropeptide genes for this motif. PMID:19712058

  14. The BEN domain is a novel sequence-specific DNA-binding domain conserved in neural transcriptional repressors

    PubMed Central

    Dai, Qi; Ren, Aiming; Westholm, Jakub O.; Serganov, Artem A.; Patel, Dinshaw J.; Lai, Eric C.

    2013-01-01

    We recently reported that Drosophila Insensitive (Insv) promotes sensory organ development and has activity as a nuclear corepressor for the Notch transcription factor Suppressor of Hairless [Su(H)]. Insv lacks domains of known biochemical function but contains a single BEN domain (i.e., a “BEN-solo” protein). Our chromatin immunoprecipitation (ChIP) sequencing (ChIP-seq) analysis confirmed binding of Insensitive to Su(H) target genes in the Enhancer of split gene complex [E(spl)-C]; however, de novo motif analysis revealed a novel site strongly enriched in Insv peaks (TCYAATHRGAA). We validate binding of endogenous Insv to genomic regions bearing such sites, whose associated genes are enriched for neural functions and are functionally repressed by Insv. Unexpectedly, we found that the Insv BEN domain binds specifically to this sequence motif and that Insv directly regulates transcription via this motif. We determined the crystal structure of the BEN–DNA target complex, revealing homodimeric binding of the BEN domain and extensive nucleotide contacts via α helices and a C-terminal loop. Point mutations in key DNA-contacting residues severely impair DNA binding in vitro and capacity for transcriptional regulation in vivo. We further demonstrate DNA-binding and repression activities by the mammalian neural BEN-solo protein BEND5. Altogether, we define novel DNA-binding activity in a conserved family of transcriptional repressors, opening a molecular window on this extensive gene family. PMID:23468431

  15. The highly conserved codon following the slippery sequence supports -1 frameshift efficiency at the HIV-1 frameshift site.

    PubMed

    Mathew, Suneeth F; Crowe-McAuliffe, Caillan; Graves, Ryan; Cardno, Tony S; McKinney, Cushla; Poole, Elizabeth S; Tate, Warren P

    2015-01-01

    HIV-1 utilises -1 programmed ribosomal frameshifting to translate structural and enzymatic domains in a defined proportion required for replication. A slippery sequence, U UUU UUA, and a stem-loop are well-defined RNA features modulating -1 frameshifting in HIV-1. The GGG glycine codon immediately following the slippery sequence (the 'intercodon') contributes structurally to the start of the stem-loop but has no defined role in current models of the frameshift mechanism, as slippage is inferred to occur before the intercodon has reached the ribosomal decoding site. This GGG codon is highly conserved in natural isolates of HIV. When the natural intercodon was replaced with a stop codon two different decoding molecules-eRF1 protein or a cognate suppressor tRNA-were able to access and decode the intercodon prior to -1 frameshifting. This implies significant slippage occurs when the intercodon is in the (perhaps distorted) ribosomal A site. We accommodate the influence of the intercodon in a model of frame maintenance versus frameshifting in HIV-1. PMID:25807539

  16. Plastid genome sequences of Gymnochlora stellata, Lotharella vacuolata, and Partenskyella glossopodia reveal remarkable structural conservation among chlorarachniophyte species.

    PubMed

    Suzuki, Shigekatsu; Hirakawa, Yoshihisa; Kofuji, Rumiko; Sugita, Mamoru; Ishida, Ken-Ichiro

    2016-07-01

    Chlorarachniophyte algae have complex plastids acquired by the uptake of a green algal endosymbiont, and this event is called secondary endosymbiosis. Interestingly, the plastids possess a relict endosymbiont nucleus, referred to as the nucleomorph, in the intermembrane space, and the nucleomorphs contain an extremely reduced and compacted genome in comparison with green algal nuclear genomes. Therefore, chlorarachniophyte plastids consist of two endosymbiotically derived genomes, i.e., the plastid and nucleomorph genomes. To date, complete nucleomorph genomes have been sequenced in four different species, whereas plastid genomes have been reported in only two species in chlorarachniophytes. To gain further insight into the evolution of endosymbiotic genomes in chlorarachniophytes, we newly sequenced the plastid genomes of three species, Gymnochlora stellata, Lotharella vacuolata, and Partenskyella glossopodia. Our findings reveal that chlorarachniophyte plastid genomes are highly conserved in size, gene content, and gene order among species, but their nucleomorph genomes are divergent in such features. Accordingly, the current architecture of the plastid genomes of chlorarachniophytes evolved in a common ancestor, and changed very little during their subsequent diversification. Furthermore, our phylogenetic analyses using multiple plastid genes suggest that chlorarachniophyte plastids are derived from a green algal lineage that is closely related to Bryopsidales in the Ulvophyceae group. PMID:26920842

  17. Use of a structural alphabet to find compatible folds for amino acid sequences

    PubMed Central

    Mahajan, Swapnil; de Brevern, Alexandre G; Sanejouand, Yves-Henri; Srinivasan, Narayanaswamy; Offmann, Bernard

    2015-01-01

    The structural annotation of proteins with no detectable homologs of known 3D structure identified using sequence-search methods is a major challenge today. We propose an original method that computes the conditional probabilities for the amino-acid sequence of a protein to fit to known protein 3D structures using a structural alphabet, known as “Protein Blocks” (PBs). PBs constitute a library of 16 local structural prototypes that approximate every part of protein backbone structures. It is used to encode 3D protein structures into 1D PB sequences and to capture sequence to structure relationships. Our method relies on amino acid occurrence matrices, one for each PB, to score global and local threading of query amino acid sequences to protein folds encoded into PB sequences. It does not use any information from residue contacts or sequence-search methods or explicit incorporation of hydrophobic effect. The performance of the method was assessed with independent test datasets derived from SCOP 1.75A. With a Z-score cutoff that achieved 95% specificity (i.e., less than 5% false positives), global and local threading showed sensitivity of 64.1% and 34.2%, respectively. We further tested its performance on 57 difficult CASP10 targets that had no known homologs in PDB: 38 compatible templates were identified by our approach and 66% of these hits yielded correctly predicted structures. This method scales-up well and offers promising perspectives for structural annotations at genomic level. It has been implemented in the form of a web-server that is freely available at http://www.bo-protscience.fr/forsa. PMID:25297700

  18. Use of a structural alphabet to find compatible folds for amino acid sequences.

    PubMed

    Mahajan, Swapnil; de Brevern, Alexandre G; Sanejouand, Yves-Henri; Srinivasan, Narayanaswamy; Offmann, Bernard

    2015-01-01

    The structural annotation of proteins with no detectable homologs of known 3D structure identified using sequence-search methods is a major challenge today. We propose an original method that computes the conditional probabilities for the amino-acid sequence of a protein to fit to known protein 3D structures using a structural alphabet, known as "Protein Blocks" (PBs). PBs constitute a library of 16 local structural prototypes that approximate every part of protein backbone structures. It is used to encode 3D protein structures into 1D PB sequences and to capture sequence to structure relationships. Our method relies on amino acid occurrence matrices, one for each PB, to score global and local threading of query amino acid sequences to protein folds encoded into PB sequences. It does not use any information from residue contacts or sequence-search methods or explicit incorporation of hydrophobic effect. The performance of the method was assessed with independent test datasets derived from SCOP 1.75A. With a Z-score cutoff that achieved 95% specificity (i.e., less than 5% false positives), global and local threading showed sensitivity of 64.1% and 34.2%, respectively. We further tested its performance on 57 difficult CASP10 targets that had no known homologs in PDB: 38 compatible templates were identified by our approach and 66% of these hits yielded correctly predicted structures. This method scales-up well and offers promising perspectives for structural annotations at genomic level. It has been implemented in the form of a web-server that is freely available at http://www.bo-protscience.fr/forsa. PMID:25297700

  19. Polypurine (A)-rich sequences promote cross-kingdom conservation of internal ribosome entry

    PubMed Central

    Dorokhov, Yuri L.; Skulachev, Maxim V.; Ivanov, Peter A.; Zvereva, Svetlana D.; Tjulkina, Lydia G.; Merits, Andres; Gleba, Yuri Y.; Hohn, Thomas; Atabekov, Joseph G.

    2002-01-01

    The internal ribosome entry sites (IRES), IRES\\documentclass[10pt]{article} \\usepackage{amsmath} \\usepackage{wasysym} \\usepackage{amsfonts} \\usepackage{amssymb} \\usepackage{amsbsy} \\usepackage{mathrsfs} \\usepackage{pmc} \\usepackage[Euler]{upgreek} \\pagestyle{empty} \\oddsidemargin -1.0in \\begin{document} \\begin{equation*}{\\mathrm{_{CP,148}^{CR}}}\\end{equation*}\\end{document} and IRES\\documentclass[10pt]{article} \\usepackage{amsmath} \\usepackage{wasysym} \\usepackage{amsfonts} \\usepackage{amssymb} \\usepackage{amsbsy} \\usepackage{mathrsfs} \\usepackage{pmc} \\usepackage[Euler]{upgreek} \\pagestyle{empty} \\oddsidemargin -1.0in \\begin{document} \\begin{equation*}{\\mathrm{_{MP,75}^{CR}}}\\end{equation*}\\end{document}, precede the coat protein (CP) and movement protein (MP) genes of crucifer-infecting tobamovirus (crTMV), respectively. In the present work, we analyzed the activity of these elements in transgenic plants and other organisms. Comparison of the relative activities of the crTMV IRES elements and the IRES from an animal virus—encephalomyocarditis virus—in plant, yeast, and HeLa cells identified the 148-nt IRES\\documentclass[10pt]{article} \\usepackage{amsmath} \\usepackage{wasysym} \\usepackage{amsfonts} \\usepackage{amssymb} \\usepackage{amsbsy} \\usepackage{mathrsfs} \\usepackage{pmc} \\usepackage[Euler]{upgreek} \\pagestyle{empty} \\oddsidemargin -1.0in \\begin{document} \\begin{equation*}{\\mathrm{_{CP,148}^{CR}}}\\end{equation*}\\end{document} as the strongest element that also displayed IRES activity across all kingdoms. Deletion analysis suggested that the polypurine (A)-rich sequences (PARSs) contained in IRES\\documentclass[10pt]{article} \\usepackage{amsmath} \\usepackage{wasysym} \\usepackage{amsfonts} \\usepackage{amssymb} \\usepackage{amsbsy} \\usepackage{mathrsfs} \\usepackage{pmc} \\usepackage[Euler]{upgreek} \\pagestyle{empty} \\oddsidemargin -1.0in \\begin{document} \\begin{equation*}{\\mathrm{_{CP,148}^{CR

  20. Software scripts for quality checking of high-throughput nucleic acid sequencers.

    PubMed

    Lazo, G R; Tong, J; Miller, R; Hsia, C; Rausch, C; Kang, Y; Anderson, O D

    2001-06-01

    We have developed a graphical interface to allow the researcher to view and assess the quality of sequencing results using a series of program scripts developed to process data generated by automated sequencers. The scripts are written in Perl programming language and are executable under the cgibin directory of a Web server environment. The scripts direct nucleic acid sequencing trace file data output from automated sequencers to be analyzed by the phred molecular biology program and are displayed as graphical hypertext mark-up language (HTML) pages. The scripts are mainly designed to handle 96-well microtiter dish samples, but the scripts are also able to read data from 384-well microtiter dishes 96 samples at a time. The scripts may be customized for different laboratory environments and computer configurations. Web links to the sources and discussion page are provided. PMID:11414222

  1. Identification of conserved hepatic transcriptomic responses to 17β-estradiol using high-throughput sequencing in brown trout

    PubMed Central

    Uren Webster, Tamsyn M.; Shears, Janice A.; Moore, Karen

    2015-01-01

    Estrogenic chemicals are major contaminants of surface waters and can threaten the sustainability of natural fish populations. Characterization of the global molecular mechanisms of toxicity of environmental contaminants has been conducted primarily in model species rather than species with limited existing transcriptomic or genomic sequence information. We aimed to investigate the global mechanisms of toxicity of an endocrine disrupting chemical of environmental concern [17β-estradiol (E2)] using high-throughput RNA sequencing (RNA-Seq) in an environmentally relevant species, brown trout (Salmo trutta). We exposed mature males to measured concentrations of 1.94, 18.06, and 34.38 ng E2/l for 4 days and sequenced three individual liver samples per treatment using an Illumina HiSeq 2500 platform. Exposure to 34.4 ng E2/L resulted in 2,113 differentially regulated transcripts (FDR < 0.05). Functional analysis revealed upregulation of processes associated with vitellogenesis, including lipid metabolism, cellular proliferation, and ribosome biogenesis, together with a downregulation of carbohydrate metabolism. Using real-time quantitative PCR, we validated the expression of eight target genes and identified significant differences in the regulation of several known estrogen-responsive transcripts in fish exposed to the lower treatment concentrations (including esr1 and zp2.5). We successfully used RNA-Seq to identify highly conserved responses to estrogen and also identified some estrogen-responsive transcripts that have been less well characterized, including nots and tgm2l. These results demonstrate the potential application of RNA-Seq as a valuable tool for assessing mechanistic effects of pollutants in ecologically relevant species for which little genomic information is available. PMID:26082144

  2. Identification of conserved hepatic transcriptomic responses to 17β-estradiol using high-throughput sequencing in brown trout.

    PubMed

    Uren Webster, Tamsyn M; Shears, Janice A; Moore, Karen; Santos, Eduarda M

    2015-09-01

    Estrogenic chemicals are major contaminants of surface waters and can threaten the sustainability of natural fish populations. Characterization of the global molecular mechanisms of toxicity of environmental contaminants has been conducted primarily in model species rather than species with limited existing transcriptomic or genomic sequence information. We aimed to investigate the global mechanisms of toxicity of an endocrine disrupting chemical of environmental concern [17β-estradiol (E2)] using high-throughput RNA sequencing (RNA-Seq) in an environmentally relevant species, brown trout (Salmo trutta). We exposed mature males to measured concentrations of 1.94, 18.06, and 34.38 ng E2/l for 4 days and sequenced three individual liver samples per treatment using an Illumina HiSeq 2500 platform. Exposure to 34.4 ng E2/L resulted in 2,113 differentially regulated transcripts (FDR < 0.05). Functional analysis revealed upregulation of processes associated with vitellogenesis, including lipid metabolism, cellular proliferation, and ribosome biogenesis, together with a downregulation of carbohydrate metabolism. Using real-time quantitative PCR, we validated the expression of eight target genes and identified significant differences in the regulation of several known estrogen-responsive transcripts in fish exposed to the lower treatment concentrations (including esr1 and zp2.5). We successfully used RNA-Seq to identify highly conserved responses to estrogen and also identified some estrogen-responsive transcripts that have been less well characterized, including nots and tgm2l. These results demonstrate the potential application of RNA-Seq as a valuable tool for assessing mechanistic effects of pollutants in ecologically relevant species for which little genomic information is available. PMID:26082144

  3. Codon Usage Patterns in Corynebacterium glutamicum: Mutational Bias, Natural Selection and Amino Acid Conservation.

    PubMed

    Liu, Guiming; Wu, Jinyu; Yang, Huanming; Bao, Qiyu

    2010-01-01

    The alternative synonymous codons in Corynebacterium glutamicum, a well-known bacterium used in industry for the production of amino acid, have been investigated by multivariate analysis. As C. glutamicum is a GC-rich organism, G and C are expected to predominate at the third position of codons. Indeed, overall codon usage analyses have indicated that C and/or G ending codons are predominant in this organism. Through multivariate statistical analysis, apart from mutational selection, we identified three other trends of codon usage variation among the genes. Firstly, the majority of highly expressed genes are scattered towards the positive end of the first axis, whereas the majority of lowly expressed genes are clustered towards the other end of the first axis. Furthermore, the distinct difference in the two sets of genes was that the C ending codons are predominate in putatively highly expressed genes, suggesting that the C ending codons are translationally optimal in this organism. Secondly, the majority of the putatively highly expressed genes have a tendency to locate on the leading strand, which indicates that replicational and transciptional selection might be invoked. Thirdly, highly expressed genes are more conserved than lowly expressed genes by synonymous and nonsynonymous substitutions among orthologous genes fromthe genomes of C. glutamicum and C. diphtheriae. We also analyzed other factors such as the length of genes and hydrophobicity that might influence codon usage and found their contributions to be weak. PMID:20445740

  4. A comparative study of 2',3'-cyclic-nucleotide 3'-phosphodiesterase in vertebrates: cDNA cloning and amino acid sequences for chicken and bullfrog enzymes.

    PubMed

    Kasama-Yoshida, H; Tohyama, Y; Kurihara, T; Sakuma, M; Kojima, H; Tamai, Y

    1997-10-01

    In mammalian brain, two 2',3'-cyclic-nucleotide 3'-phosphodiesterase (EC 3.1.4.37) isoforms, CNP1 and CNP2, are translated, respectively, from the two mRNAs, which have been transcribed and processed by alternative use of the two transcription start points and by differential splicing. In the present study, the cDNAs encoding chicken CNP2 and bullfrog CNP1, respectively, were isolated, and the amino acid sequences of chicken CNP2 and bullfrog CNP1 were deduced. Western blot analysis showed that chicken brain contains a major CNP2-type protein together with a minor unidentified isoform, and bullfrog brain contains only a CNP1-type protein. All available amino acid sequences of vertebrate 2',3'-cyclic-nucleotide 3'-phosphodiesterases were aligned and compared. Three conserved motif sequences were noted: (a) an ATP-binding site near the amino terminus, (b) an isoprenylation site at the carboxyl terminus, and (c) a probable catalytic site resembling the active site of beta-ketoacyl synthase (EC 2.3.1.41). The second and the third motifs are conserved also in goldfish RICH (regeneration-induced 2',3'-cyclic-nucleotide 3'-phosphodiesterase homologue), which has been shown recently to have 2',3'-cyclic-nucleotide 3'-phosphodiesterase activity. The third motif (probably catalytic site) was assigned for the first time in the present report. PMID:9326261

  5. Efficient Nucleic Acid Extraction and 16S rRNA Gene Sequencing for Bacterial Community Characterization.

    PubMed

    Anahtar, Melis N; Bowman, Brittany A; Kwon, Douglas S

    2016-01-01

    There is a growing appreciation for the role of microbial communities as critical modulators of human health and disease. High throughput sequencing technologies have allowed for the rapid and efficient characterization of bacterial communities using 16S rRNA gene sequencing from a variety of sources. Although readily available tools for 16S rRNA sequence analysis have standardized computational workflows, sample processing for DNA extraction remains a continued source of variability across studies. Here we describe an efficient, robust, and cost effective method for extracting nucleic acid from swabs. We also delineate downstream methods for 16S rRNA gene sequencing, including generation of sequencing libraries, data quality control, and sequence analysis. The workflow can accommodate multiple samples types, including stool and swabs collected from a variety of anatomical locations and host species. Additionally, recovered DNA and RNA can be separated and used for other applications, including whole genome sequencing or RNA-seq. The method described allows for a common processing approach for multiple sample types and accommodates downstream analysis of genomic, metagenomic and transcriptional information. PMID:27168460

  6. Efficient Nucleic Acid Extraction and 16S rRNA Gene Sequencing for Bacterial Community Characterization

    PubMed Central

    Anahtar, Melis N.; Bowman, Brittany A.; Kwon, Douglas S.

    2016-01-01

    There is a growing appreciation for the role of microbial communities as critical modulators of human health and disease. High throughput sequencing technologies have allowed for the rapid and efficient characterization of bacterial communities using 16S rRNA gene sequencing from a variety of sources. Although readily available tools for 16S rRNA sequence analysis have standardized computational workflows, sample processing for DNA extraction remains a continued source of variability across studies. Here we describe an efficient, robust, and cost effective method for extracting nucleic acid from swabs. We also delineate downstream methods for 16S rRNA gene sequencing, including generation of sequencing libraries, data quality control, and sequence analysis. The workflow can accommodate multiple samples types, including stool and swabs collected from a variety of anatomical locations and host species. Additionally, recovered DNA and RNA can be separated and used for other applications, including whole genome sequencing or RNA-seq. The method described allows for a common processing approach for multiple sample types and accommodates downstream analysis of genomic, metagenomic and transcriptional information. PMID:27168460

  7. Preparation of Nucleic Acid Libraries for Personalized Sequencing Systems Using an Integrated Microfluidic Hub Technology (Seventh Annual Sequencing, Finishing, Analysis in the Future (SFAF) Meeting 2012)

    ScienceCinema

    Patel, Kamlesh D [Ken]; SNL,

    2013-01-25

    Kamlesh (Ken) Patel from Sandia National Laboratories (Livermore, California) presents "Preparation of Nucleic Acid Libraries for Personalized Sequencing Systems Using an Integrated Microfluidic Hub Technology " at the 7th Annual Sequencing, Finishing, Analysis in the Future (SFAF) Meeting held in June, 2012 in Santa Fe, NM.

  8. The amino acid sequence of ribonuclease U2 from Ustilago sphaerogena.

    PubMed Central

    Sato, S; Uchida, T

    1975-01-01

    1. RNAase (ribonuclease) U2, a purine-specific RNAase, was reduced, aminoethylated and hydrolysed with trypsin, chymotrypsin and thermolysin. On the basis of the analyses of the resulting peptides, the complete amino acid sequence of RNAase U2 was determined, 2. When the sequence was compared with the amino acid sequence of RNAase T1 (EC 3.1.4.8), the following regions were found to be similar in the two enzymes; Tyr-Pro-His-Gln-Tyr (38-42) in RNAase U2 and Tyr-Pro-His-Lys-Tyr (38-42) in RNAase T1, Glu-Phe-Pro-Leu-Val (61-65) in RNAase U2 and Glu-Trp-Pro-Ile-Leu (58-62) in RNAase T1, Asp-Arg-Val-Ile-Tyr-Gln (83-88) in RNAase U2 and Asp-Arg-Val-Phe-Asn (76-81) in RNAase T1 and Val-Thr-His-Thr-Gly-Ala (98-103) in RNAase U2 and Ile-Thr-His-Thr-Gly-Ala (90-95) in RNAase T1. All of the amino acid residues, histidine-40, glutamate-58, arginine-77 and histidine-92, which were found to play a crucial role in the biological activity of RNAase T1, were included in the regions cited here. 3. Detailed evidence for the amino acid sequence of the sequence of the proteins has been deposited as Supplementary Publication SUP 50041 (33 PAGES) AT THE British Library (Lending Division)(formerly the National Lending Library for Science and Technology), Boston Spa, Yorks. LS23 7BQ, U.K., from whom copies can be obtained on the terms indicated in Biochem. J. (1975), 145, 5. PMID:1156364

  9. Deduced amino acid sequence of human pulmonary surfactant proteolipid: SPL(pVal)

    SciTech Connect

    Whitsett, J.A.; Glasser, S.W.; Korfhagen, T.R.; Weaver, T.E.; Clark, J.; Pilot-Matias, T.; Meuth, J.; Fox, J.L.

    1987-05-01

    Hydrophobic, proteolipid-like protein of Mr 6500 was isolated from ether/ethanol extracts of human, canine and bovine pulmonary surfactant. Amino acid composition of the protein demonstrated a remarkable abundance of hydrophobic residues, particularly valine and leucine. The N-terminal amino acid sequence of the human protein was determined: N-Leu-Ile-Pro-Cys-Cys-Pro-Val-Asn-Leu-Lys-Arg-Leu-Leu-Ile-Val4... An oligonucleotide probe was used to screen an adult human lung cDNA library and resulted in detection of cDNA clones with predicted amino acid sequence with close identity to the N-terminal amino acid sequence of the human peptide. SPL(pVal) was found within the reading frame of a larger peptide. SPL(pVal) results from proteolytic processing of a larger preprotein. Northern blot analysis detected in a single 1.0 kilobase SPL(pVal) RNA which was less abundant in fetal than in adult lung. Mixtures of purified canine and bovine SPL(pVal) and synthetic phospholipids display properties of rapid adsorption and surface tension lowering activity characteristic of surfactant. Human SPL(pVal) is a pulmonary surfactant proteolipid which may therefore be useful in combination with phospholipids and/or other surfactant proteins for the treatment of surfactant deficiency such as hyaline membrane disease in newborn infants.

  10. HIGHLY CONSERVED N-TERMINAL SEQUENCE FOR TELEOST VITELLOGENIN WITH POTENTIAL VALUE TO THE BIOCHEMISTRY, MOLECULAR BIOLOGY AND PATHOLOGY OF VITELLOGENESIS

    EPA Science Inventory

    N-terminal amino acid sequences for vitellogenin (Vtg) from six species of teleost fish: striped bass, Morone saxatillus; mummichog, Fundulus heteroclitus; pinfish, Lagodon rhomboides; brown bullhead, Ameiurus nebulosus; medaka, Oryzias latipes; yellow perch, Percaflavescens and ...

  11. Complete nucleic acid sequence of Penaeus stylirostris densovirus (PstDNV) from India.

    PubMed

    Rai, Praveen; Safeena, Muhammed P; Karunasagar, Iddya; Karunasagar, Indrani

    2011-06-01

    Infectious hypodermal and hematopoietic necrosis virus (IHHNV) of shrimp, recently been classified as Penaeus stylirostris densovirus (PstDNV). The complete nucleic acid sequence of PstDNV from India was obtained by cloning and sequencing of different DNA fragment of the virus. The genome organisation of PstDNV revealed that there were three major coding domains: a left ORF (NS1) of 2001 bp, a mid ORF (NS2) of 1092 bp and a right ORF (VP) of 990 bp. The complete genome and amino acid sequences of three proteins viz., NS1, NS2 and VP were compared with the genomes of the virus reported from Hawaii, China and Mexico and with partial sequence available from isolates from different regions. The phylogenetic analysis of shrimp, insect and vertebrate parvovirus sequences showed that the Indian PstDNV isolate is phylogenetically more closely related to one of the three isolates from Taiwan (AY355307), and two isolates (AY362547 and AY102034) from Thailand. PMID:21402111

  12. Human liver type pyruvate kinase: complete amino acid sequence and the expression in mammalian cells.

    PubMed Central

    Tani, K; Fujii, H; Nagata, S; Miwa, S

    1988-01-01

    Pyruvate kinase (PK) has four isozymes (L, R, M1, M2) that are encoded by two different genes. Among these isozymes, abnormalities of liver (L)-type PK is considered to be associated with hereditary nonspherocytic hemolytic anemia in humans. We isolated and determined the full-length sequence of human L-type PK cDNA. The cDNA contains 1629 base pairs encoding 543 amino acids, 68 base pairs of 5'-noncoding sequence, and 734 base pairs of 3'-noncoding sequence. The similarity between human and rat L-type PK was 86.9% at the nucleotide sequence level and 92.4% at the amino acid sequence level. The full-length L-type PK cDNA was placed under the promoter of simian virus 40 and introduced into monkey COS cells. Human L-type PK activity was detected in the extract of COS cells by the classical PK electrophoresis method. Images PMID:3126495

  13. Human liver type pyruvate kinase: Complete amino acid sequence and the expression in mammalian cells

    SciTech Connect

    Tani, Kenzaburo; Nagata, Shigekazu ); Fujii, Hisaichi ); Miwa, Shiro )

    1988-03-01

    Pyruvate kinase (PK) has four isozymes (L, R, M{sub 1}, M{sub 2}) that are encoded by two different genes. Among these isozymes, abnormalities of liver (L)-type PK is considered to be associated with hereditary nonspherocytic hemolytic anemia in humans. The authors isolated and determined the full-length sequence of human L-type PK cDNA. The cDNA contains 1,629 base pairs encoding 543 amino acids, 68 base pairs of 5{prime}-noncoding sequence, and 734 base pairs of 3{prime}-noncoding sequence. The similarity between human and rat L-type PK was 86.9% at the nucleotide sequence level and 92.4% at the amino acid sequence level. The full-length L-type PK cDNA was placed under the promoter of simian virus 40 and introduced into monkey COS cells. Human L-type PK activity was detected in the extract of COS cells by the classical PK electrophoresis method.

  14. Site-directed mutagenesis of conserved amino acids in the alpha subunit of toluene dioxygenase: potential mononuclear non-heme iron coordination sites.

    PubMed Central

    Jiang, H; Parales, R E; Lynch, N A; Gibson, D T

    1996-01-01

    The terminal oxygenase component of toluene dioxygenase from Pseudomonas putida F1 is an iron-sulfur protein (ISP(TOL)) that requires mononuclear iron for enzyme activity. Alignment of all available predicted amino acid sequences for the large (alpha) subunits of terminal oxygenases showed a conserved cluster of potential mononuclear iron-binding residues. These were between amino acids 210 and 230 in the alpha subunit (TodC1) of ISP(TOL). The conserved amino acids, Glu-214, Asp-219, Tyr-221, His-222, and His-228, were each independently replaced with an alanine residue by site-directed mutagenesis. Tyr-266 in TodC1, which has been suggested as an iron ligand, was treated in an identical manner. To assay toluene dioxygenase activity in the presence of TodC1 and its mutant forms, conditions for the reconstitution of wild-type ISP(TOL) activity from TodC1 and purified TodC2 (beta subunit) were developed and optimized. A mutation at Glu-214, Asp-219, His-222, or His-228 completely abolished toluene dioxygenase activity. TodC1 with an alanine substitution at either Tyr-221 or Tyr-266 retained partial enzyme activity (42 and 12%, respectively). In experiments with [14C]toluene, the two Tyr-->Ala mutations caused a reduction in the amount of Cis-[14C]-toluene dihydrodiol formed, whereas a mutation at Glu-214, Asp-219, His-222, or His-228 eliminated cis-toluene dihydrodiol formation. The expression level of all of the mutated TWO proteins was equivalent to that of wild-type TodC1 as judged by sodium dodecyl sulfate-polyacrylamide gel electrophoresis and Western blot (immunoblot) analyses. These results, in conjunction with the predicted amino acid sequences of 22 oxygenase components, suggest that the conserved motif Glu-X3-4,-Asp-X2-His-X4-5-His is critical for catalytic function and the glutamate, aspartate, and histidine residues may act as mononuclear iron ligands at the site of oxygen activation. PMID:8655491

  15. HMMerThread: Detecting Remote, Functional Conserved Domains in Entire Genomes by Combining Relaxed Sequence-Database Searches with Fold Recognition

    PubMed Central

    Bradshaw, Charles Richard; Surendranath, Vineeth; Henschel, Robert; Mueller, Matthias Stefan; Habermann, Bianca Hermine

    2011-01-01

    Conserved domains in proteins are one of the major sources of functional information for experimental design and genome-level annotation. Though search tools for conserved domain databases such as Hidden Markov Models (HMMs) are sensitive in detecting conserved domains in proteins when they share sufficient sequence similarity, they tend to miss more divergent family members, as they lack a reliable statistical framework for the detection of low sequence similarity. We have developed a greatly improved HMMerThread algorithm that can detect remotely conserved domains in highly divergent sequences. HMMerThread combines relaxed conserved domain searches with fold recognition to eliminate false positive, sequence-based identifications. With an accuracy of 90%, our software is able to automatically predict highly divergent members of conserved domain families with an associated 3-dimensional structure. We give additional confidence to our predictions by validation across species. We have run HMMerThread searches on eight proteomes including human and present a rich resource of remotely conserved domains, which adds significantly to the functional annotation of entire proteomes. We find ∼4500 cross-species validated, remotely conserved domain predictions in the human proteome alone. As an example, we find a DNA-binding domain in the C-terminal part of the A-kinase anchor protein 10 (AKAP10), a PKA adaptor that has been implicated in cardiac arrhythmias and premature cardiac death, which upon stress likely translocates from mitochondria to the nucleus/nucleolus. Based on our prediction, we propose that with this HLH-domain, AKAP10 is involved in the transcriptional control of stress response. Further remotely conserved domains we discuss are examples from areas such as sporulation, chromosome segregation and signalling during immune response. The HMMerThread algorithm is able to automatically detect the presence of remotely conserved domains in proteins based on weak

  16. Molecular cytogenetics by polymerase catalyzed amplification or in situ labelling of specific nucleic acid sequences

    SciTech Connect

    Bolund, L.; Brandt, C.; Hindkjaer, J.; Koch, J.; Koelvraa, S.; Pedersen, S. )

    1993-01-01

    The Polymerase Chain Reaction (PCR) can be performed on isolated cells or chromosomes and the product can be analyzed by DNA technology or by FISH to test metaphases. The authors have good experiences analyzing aberrant chromosomes by FACS sorting, PCR with degenerated primers and painting of test metaphases with the PCR product. They also utilize polymerases for PRimed IN Situ labelling (PRINS) of specific nucleic acid sequences. In PRINS oligonucleotides are hybridized to their target sequences and labeled nucleotides are incorporated at the site of hybridization with the oligonucleotide as primer. PRINS may eventually allow the study of individual genes, gene expression and even somatic mutations (in mRNA) in single cells.

  17. DNA Cloning of Plasmodium falciparum Circumsporozoite Gene: Amino Acid Sequence of Repetitive Epitope

    NASA Astrophysics Data System (ADS)

    Enea, Vincenzo; Ellis, Joan; Zavala, Fidel; Arnot, David E.; Asavanich, Achara; Masuda, Aoi; Quakyi, Isabella; Nussenzweig, Ruth S.

    1984-08-01

    A clone of complementary DNA encoding the circumsporozoite (CS) protein of the human malaria parasite Plasmodium falciparum has been isolated by screening an Escherichia coli complementary DNA library with a monoclonal antibody to the CS protein. The DNA sequence of the complementary DNA insert encodes a four-amino acid sequence: proline-asparagine-alanine-asparagine, tandemly repeated 23 times. The CS β -lactamase fusion protein specifically binds monoclonal antibodies to the CS protein and inhibits the binding of these antibodies to native Plasmodium falciparum CS protein. These findings provide a basis for the development of a vaccine against Plasmodium falciparum malaria.

  18. Method for high-volume sequencing of nucleic acids: random and directed priming with libraries of oligonucleotides

    DOEpatents

    Studier, F.W.

    1995-04-18

    Random and directed priming methods for determining nucleotide sequences by enzymatic sequencing techniques, using libraries of primers of lengths 8, 9 or 10 bases, are disclosed. These methods permit direct sequencing of nucleic acids as large as 45,000 base pairs or larger without the necessity for subcloning. Individual primers are used repeatedly to prime sequence reactions in many different nucleic acid molecules. Libraries containing as few as 10,000 octamers, 14,200 nonamers, or 44,000 decamers would have the capacity to determine the sequence of almost any cosmid DNA. Random priming with a fixed set of primers from a smaller library can also be used to initiate the sequencing of individual nucleic acid molecules, with the sequence being completed by directed priming with primers from the library. In contrast to random cloning techniques, a combined random and directed priming strategy is far more efficient. 2 figs.

  19. Method for high-volume sequencing of nucleic acids: random and directed priming with libraries of oligonucleotides

    DOEpatents

    Studier, F. William

    1995-04-18

    Random and directed priming methods for determining nucleotide sequences by enzymatic sequencing techniques, using libraries of primers of lengths 8, 9 or 10 bases, are disclosed. These methods permit direct sequencing of nucleic acids as large as 45,000 base pairs or larger without the necessity for subcloning. Individual primers are used repeatedly to prime sequence reactions in many different nucleic acid molecules. Libraries containing as few as 10,000 octamers, 14,200 nonamers, or 44,000 decamers would have the capacity to determine the sequence of almost any cosmid DNA. Random priming with a fixed set of primers from a smaller library can also be used to initiate the sequencing of individual nucleic acid molecules, with the sequence being completed by directed priming with primers from the library. In contrast to random cloning techniques, a combined random and directed priming strategy is far more efficient.

  20. PTS-Mediated Regulation of the Transcription Activator MtlR from Different Species: Surprising Differences despite Strong Sequence Conservation.

    PubMed

    Joyet, Philippe; Derkaoui, Meriem; Bouraoui, Houda; Deutscher, Josef

    2015-01-01

    The hexitol D-mannitol is transported by many bacteria via a phosphoenolpyruvate (PEP):carbohydrate phosphotransferase system (PTS). In most Firmicutes, the transcription activator MtlR controls the expression of the genes encoding the D-mannitol-specific PTS components and D-mannitol-1-P dehydrogenase. MtlR contains an N-terminal helix-turn-helix motif followed by an Mga-like domain, two PTS regulation domains (PRDs), an EIIB(Gat)- and an EIIA(Mtl)-like domain. The four regulatory domains are the target of phosphorylation by PTS components. Despite strong sequence conservation, the mechanisms controlling the activity of MtlR from Lactobacillus casei, Bacillus subtilis and Geobacillus stearothermophilus are quite different. Owing to the presence of a tyrosine in place of the second conserved histidine (His) in PRD2, L. casei MtlR is not phosphorylated by Enzyme I (EI) and HPr. When the corresponding His in PRD2 of MtlR from B. subtilis and G. stearothermophilus was replaced with alanine, the transcription regulator was no longer phosphorylated and remained inactive. Surprisingly, L. casei MtlR functions without phosphorylation in PRD2 because in a ptsI (EI) mutant MtlR is constitutively active. EI inactivation prevents not only phosphorylation of HPr, but also of the PTS(Mtl) components, which inactivate MtlR by phosphorylating its EIIB(Gat)- or EIIA(Mtl)-like domain. This explains the constitutive phenotype of the ptsI mutant. The absence of EIIB(Mtl)-mediated phosphorylation leads to induction of the L. caseimtl operon. This mechanism resembles mtlARFD induction in G. stearothermophilus, but differs from EIIA(Mtl)-mediated induction in B. subtilis. In contrast to B. subtilis MtlR, L. casei MtlR activation does not require sequestration to the membrane via the unphosphorylated EIIB(Mtl) domain. PMID:26159071

  1. Cross-species conservation of complementary amino acid-ribonucleobase interactions and their potential for ribosome-free encoding

    PubMed Central

    Cannon, John G. D.; Sherman, Rachel M.; Wang, Victoria M. Y.; Newman, Grace A.

    2015-01-01

    The role of amino acid-RNA nucleobase interactions in the evolution of RNA translation and protein-mRNA autoregulation remains an open area of research. We describe the inference of pairwise amino acid-RNA nucleobase interaction preferences using structural data from known RNA-protein complexes. We observed significant matching between an amino acid’s nucleobase affinity and corresponding codon content in both the standard genetic code and mitochondrial variants. Furthermore, we showed that knowledge of nucleobase preferences allows statistically significant prediction of protein primary sequence from mRNA using purely physiochemical information. Interestingly, ribosomal primary sequences were more accurately predicted than non-ribosomal sequences, suggesting a potential role for direct amino acid-nucleobase interactions in the genesis of amino acid-based ribosomal components. Finally, we observed matching between amino acid-nucleobase affinities and corresponding mRNA sequences in 35 evolutionarily diverse proteomes. We believe these results have important implications for the study of the evolutionary origins of the genetic code and protein-mRNA cross-regulation. PMID:26656258

  2. Partial amino acid sequence of apolipoprotein(a) shows that it is homologous to plasminogen

    SciTech Connect

    Eaton, D.L.; Fless, G.M.; Kohr, W.J.; McLean, J.W.; Xu, Q.T.; Miller, C.G.; Lawn, R.M.; Scanu, A.M.

    1987-05-01

    Apolipoprotein(a) (apo(a)) is a glycoprotein with M/sub r/ approx. 280,000 that is disulfide linked to apolipoprotein B in lipoprotein(a) particles. Elevated plasma levels of lipoprotein(a) are correlated with atherosclerosis. Partial amino acid sequence of apo(a) shows that it has striking homology to plasminogen. Plasminogen is a plasma serine protease zymogen that consists of five homologous and tandemly repeated domains called kringles and a trypsin-like protease domain. The amino-terminal sequence obtained for apo(a) is homologous to the beginning of kringle 4 but not the amino terminus of plasminogen. Apo(a) was subjected to limited proteolysis by trypsin or V8 protease, and fragments generated were isolated and sequenced. Sequences obtained from several of these fragments are highly (77-100%) homologous to plasminogen residues 391-421, which reside within kringle 4. Analysis of these internal apo(a) sequences revealed that apo(a) may contain at least two kringle 4-like domains. A sequence obtained from another tryptic fragment also shows homology to the end of kringle 4 and the beginning of kringle 5. Sequence data obtained from the two tryptic fragments shows homology with the protease domain of plasminogen. One of these sequences is homologous to the sequences surrounding the activation site of plasminogen. Plasminogen is activated by the cleavage of a specific arginine residue by urokinase and tissue plasminogen activator; however, the corresponding site in apo(a) is a serine that would not be cleaved by tissue plasminogen activator or urokinase. Using a plasmin-specific assay, no proteolytic activity could be demonstrated for lipoprotein(a) particles. These results suggest that apo(a) contains kringle-like domains and an inactive protease domain.

  3. Marker genes that are less conserved in their sequences are useful for predicting genome-wide similarity levels between closely related prokaryotic strains

    DOE PAGESBeta

    Lan, Yemin; Rosen, Gail; Hershberg, Ruth

    2016-05-03

    The 16s rRNA gene is so far the most widely used marker for taxonomical classification and separation of prokaryotes. Since it is universally conserved among prokaryotes, it is possible to use this gene to classify a broad range of prokaryotic organisms. At the same time, it has often been noted that the 16s rRNA gene is too conserved to separate between prokaryotes at finer taxonomic levels. In this paper, we examine how well levels of similarity of 16s rRNA and 73 additional universal or nearly universal marker genes correlate with genome-wide levels of gene sequence similarity. We demonstrate that themore » percent identity of 16s rRNA predicts genome-wide levels of similarity very well for distantly related prokaryotes, but not for closely related ones. In closely related prokaryotes, we find that there are many other marker genes for which levels of similarity are much more predictive of genome-wide levels of gene sequence similarity. Finally, we show that the identities of the markers that are most useful for predicting genome-wide levels of similarity within closely related prokaryotic lineages vary greatly between lineages. However, the most useful markers are always those that are least conserved in their sequences within each lineage. In conclusion, our results show that by choosing markers that are less conserved in their sequences within a lineage of interest, it is possible to better predict genome-wide gene sequence similarity between closely related prokaryotes than is possible using the 16s rRNA gene. We point readers towards a database we have created (POGO-DB) that can be used to easily establish which markers show lowest levels of sequence conservation within different prokaryotic lineages.« less

  4. The Complete Genome Sequence of the Lactic Acid Bacterium Lactococcus lactis ssp. lactis IL1403

    PubMed Central

    Bolotin, Alexander; Wincker, Patrick; Mauger, Stéphane; Jaillon, Olivier; Malarme, Karine; Weissenbach, Jean; Ehrlich, S. Dusko; Sorokin, Alexei

    2001-01-01

    Lactococcus lactis is a nonpathogenic AT-rich gram-positive bacterium closely related to the genus Streptococcus and is the most commonly used cheese starter. It is also the best-characterized lactic acid bacterium. We sequenced the genome of the laboratory strain IL1403, using a novel two-step strategy that comprises diagnostic sequencing of the entire genome and a shotgun polishing step. The genome contains 2,365,589 base pairs and encodes 2310 proteins, including 293 protein-coding genes belonging to six prophages and 43 insertion sequence (IS) elements. Nonrandom distribution of IS elements indicates that the chromosome of the sequenced strain may be a product of recent recombination between two closely related genomes. A complete set of late competence genes is present, indicating the ability of L. lactis to undergo DNA transformation. Genomic sequence revealed new possibilities for fermentation pathways and for aerobic respiration. It also indicated a horizontal transfer of genetic information from Lactococcus to gram-negative enteric bacteria of Salmonella-Escherichia group. [The sequence data described in this paper has been submitted to the GenBank data library under accession no. AE005176.] PMID:11337471

  5. On human disease-causing amino acid variants: statistical study of sequence and structural patterns

    PubMed Central

    Alexov, Emil

    2015-01-01

    Statistical analysis was carried out on large set of naturally occurring human amino acid variations and it was demonstrated that there is a preference for some amino acid substitutions to be associated with diseases. At an amino acid sequence level, it was shown that the disease-causing variants frequently involve drastic changes of amino acid physico-chemical properties of proteins such as charge, hydrophobicity and geometry. Structural analysis of variants involved in diseases and being frequently observed in human population showed similar trends: disease-causing variants tend to cause more changes of hydrogen bond network and salt bridges as compared with harmless amino acid mutations. Analysis of thermodynamics data reported in literature, both experimental and computational, indicated that disease-causing variants tend to destabilize proteins and their interactions, which prompted us to investigate the effects of amino acid mutations on large databases of experimentally measured energy changes in unrelated proteins. Although the experimental datasets were linked neither to diseases nor exclusory to human proteins, the observed trends were the same: amino acid mutations tend to destabilize proteins and their interactions. Having in mind that structural and thermodynamics properties are interrelated, it is pointed out that any large change of any of them is anticipated to cause a disease. PMID:25689729

  6. Self-sequencing of amino acids and origins of polyfunctional protocells.

    PubMed

    Fox, S W

    1984-01-01

    The primal role of the origins of proteins in molecular evolution is discussed. On the basis of this premise, the significance of the experimentally established self-sequencing of amino acids under simulated geological conditions is explained as due to the fact that the products are highly nonrandom and accordingly contain many kinds of information. When such thermal proteins are aggregated into laboratory protocells, an action that occurs readily, the resultant protocells also contain many kinds of information. Residue-by-residue order, enzymic activities, and lipid quality accordingly occur within each preparation of proteinoid (thermal protein). In this paper are reviewed briefly the phenomenon of self-sequencing of amino acids, its relationship to evolutionary processes, other significance of such self-ordering, and the experimental evidence for original polyfunctional protocells. PMID:6462684

  7. Self-Sequencing of Amino Acids and Origins of Polyfunctional Protocells

    NASA Astrophysics Data System (ADS)

    Fox, Sidney W.

    1984-12-01

    The primal role of the origins of proteins in molecular evolution is discussed. On the basis of this premise, the significance of the experimentally established self-sequencing of amino acids under simulated geological conditions is explained as due to the fact that the products are highly nonrandom and accordingly contain many kinds of information. When such thermal proteins are aggregated into laboratory protocells, an action that occurs readily, the resultant protocells also contain many kinds of information. Residue-by-residue order, enzymic activities, and lipid quality accordingly occur within each preparation of proteinoid (thermal protein). In this paper are reviewed briefly the phenomenon of self-sequencing of amino acids, its relationship to evolutionary processes, other significance of such self-ordering, and the experimental evidence for original polyfunctional protocells.

  8. A conserved predicted pseudoknot in the NS2A-encoding sequence of West Nile and Japanese encephalitis flaviviruses suggests NS1' may derive from ribosomal frameshifting

    PubMed Central

    Firth, Andrew E; Atkins, John F

    2009-01-01

    Japanese encephalitis, West Nile, Usutu and Murray Valley encephalitis viruses form a tight subgroup within the larger Flavivirus genus. These viruses utilize a single-polyprotein expression strategy, resulting in ~10 mature proteins. Plotting the conservation at synonymous sites along the polyprotein coding sequence reveals strong conservation peaks at the very 5' end of the coding sequence, and also at the 5' end of the sequence encoding the NS2A protein. Such peaks are generally indicative of functionally important non-coding sequence elements. The second peak corresponds to a predicted stable pseudoknot structure whose biological importance is supported by compensatory mutations that preserve the structure. The pseudoknot is preceded by a conserved slippery heptanucleotide (Y CCU UUU), thus forming a classical stimulatory motif for -1 ribosomal frameshifting. We hypothesize, therefore, that the functional importance of the pseudoknot is to stimulate a portion of ribosomes to shift -1 nt into a short (45 codon), conserved, overlapping open reading frame, termed foo. Since cleavage at the NS1-NS2A boundary is known to require synthesis of NS2A in cis, the resulting transframe fusion protein is predicted to be NS1-NS2AN-term-FOO. We hypothesize that this may explain the origin of the previously identified NS1 'extension' protein in JEV-group flaviviruses, known as NS1'. PMID:19196463

  9. Sequence of morphological transitions in two-dimensional pattern growth from aqueous ascorbic Acid solutions.

    PubMed

    Paranjpe, A S

    2002-08-12

    A sequence of morphological transitions in two-dimensional dehydration patterns of aqueous solutions of ascorbic acid is observed with humidity as a control parameter. Change in morphology occurs due to humidity induced variation in the concentration of the metastable supersaturated solution phase formed after initial solvent evaporation. As percent humidity is varied from 40 to 80, patterns change from compact circular --> radial --> density modulated radial (a new morphology) --> density modulated circular --> density modulated dendritic (a new morphology) --> dense branching. PMID:12190528

  10. Self-sequencing of amino acids and origins of polyfunctional protocells

    NASA Technical Reports Server (NTRS)

    Fox, S. W.

    1984-01-01

    The role of proteins in the origin of living things is discussed. It has been experimentally established that amino acids can sequence themselves under simulated geological conditions with highly nonrandom products which accordingly contain diverse information. Multiple copies of each type of macromolecule are formed, resulting in greater power for any protoenzymic molecule than would accrue from a single copy of each type. Thermal proteins are readily incorporated into laboratory protocells. The experimental evidence for original polyfunctional protocells is discussed.

  11. Snake venom. The amino acid sequence of protein A from Dendroaspis polylepis polylepis (black mamba) venom.

    PubMed

    Joubert, F J; Strydom, D J

    1980-12-01

    Protein A from Dendroaspis polylepis polylepis venom comprises 81 amino acids, including ten half-cystine residues. The complete primary structures of protein A and its variant A' were elucidated. The sequences of proteins A and A', which differ in a single position, show no homology with various neurotoxins and non-neurotoxic proteins and represent a new type of elapid venom protein. PMID:7461607

  12. Characterization of the microbial acid mine drainage microbial community using culturing and direct sequencing techniques.

    PubMed

    Auld, Ryan R; Myre, Maxine; Mykytczuk, Nadia C S; Leduc, Leo G; Merritt, Thomas J S

    2013-05-01

    We characterized the bacterial community from an AMD tailings pond using both classical culturing and modern direct sequencing techniques and compared the two methods. Acid mine drainage (AMD) is produced by the environmental and microbial oxidation of minerals dissolved from mining waste. Surprisingly, we know little about the microbial communities associated with AMD, despite the fundamental ecological roles of these organisms and large-scale economic impact of these waste sites. AMD microbial communities have classically been characterized by laboratory culturing-based techniques and more recently by direct sequencing of marker gene sequences, primarily the 16S rRNA gene. In our comparison of the techniques, we find that their results are complementary, overall indicating very similar community structure with similar dominant species, but with each method identifying some species that were missed by the other. We were able to culture the majority of species that our direct sequencing results indicated were present, primarily species within the Acidithiobacillus and Acidiphilium genera, although estimates of relative species abundance were only obtained from direct sequencing. Interestingly, our culture-based methods recovered four species that had been overlooked from our sequencing results because of the rarity of the marker gene sequences, likely members of the rare biosphere. Further, direct sequencing indicated that a single genus, completely missed in our culture-based study, Legionella, was a dominant member of the microbial community. Our results suggest that while either method does a reasonable job of identifying the dominant members of the AMD microbial community, together the methods combine to give a more complete picture of the true diversity of this environment. PMID:23485423

  13. 37 CFR 1.822 - Symbols and format to be used for nucleotide and/or amino acid sequence data.

    Code of Federal Regulations, 2011 CFR

    2011-07-01

    ... approved by the Director of the Federal Register in accordance with 5 U.S.C. 552(a) and 1 CFR part 51... base or modified or unusual amino acid may be presented in a given sequence as the corresponding unmodified base or amino acid if the modified base or modified or unusual amino acid is one of those...

  14. 37 CFR 1.822 - Symbols and format to be used for nucleotide and/or amino acid sequence data.

    Code of Federal Regulations, 2010 CFR

    2010-07-01

    ... approved by the Director of the Federal Register in accordance with 5 U.S.C. 552(a) and 1 CFR part 51... base or modified or unusual amino acid may be presented in a given sequence as the corresponding unmodified base or amino acid if the modified base or modified or unusual amino acid is one of those...

  15. Two Cytoplasmic Acylation Sites and an Adjacent Hydrophobic Residue, but No Other Conserved Amino Acids in the Cytoplasmic Tail of HA from Influenza A Virus Are Crucial for Virus Replication.

    PubMed

    Siche, Stefanie; Brett, Katharina; Möller, Lars; Kordyukova, Larisa V; Mintaev, Ramil R; Alexeevski, Andrei V; Veit, Michael

    2015-12-01

    Recruitment of the matrix protein M1 to the assembly site of the influenza virus is thought to be mediated by interactions with the cytoplasmic tail of hemagglutinin (HA). Based on a comprehensive sequence comparison of all sequences present in the database, we analyzed the effect of mutating conserved residues in the cytosol-facing part of the transmembrane region and cytoplasmic tail of HA (A/WSN/33 (H1N1) strain) on virus replication and morphology of virions. Removal of the two cytoplasmic acylation sites and substitution of a neighboring isoleucine by glutamine prevented rescue of infectious virions. In contrast, a conservative exchange of the same isoleucine, non-conservative exchanges of glycine and glutamine, deletion of the acylation site at the end of the transmembrane region and shifting it into the tail did not affect virus morphology and had only subtle effects on virus growth and on the incorporation of M1 and Ribo-Nucleoprotein Particles (RNPs). Thus, assuming that essential amino acids are conserved between HA subtypes we suggest that, besides the two cytoplasmic acylation sites (including adjacent hydrophobic residues), no other amino acids in the cytoplasmic tail of HA are indispensable for virus assembly and budding. PMID:26670246

  16. Two Cytoplasmic Acylation Sites and an Adjacent Hydrophobic Residue, but No Other Conserved Amino Acids in the Cytoplasmic Tail of HA from Influenza A Virus Are Crucial for Virus Replication

    PubMed Central

    Siche, Stefanie; Brett, Katharina; Möller, Lars; Kordyukova, Larisa V.; Mintaev, Ramil R.; Alexeevski, Andrei V.; Veit, Michael

    2015-01-01

    Recruitment of the matrix protein M1 to the assembly site of the influenza virus is thought to be mediated by interactions with the cytoplasmic tail of hemagglutinin (HA). Based on a comprehensive sequence comparison of all sequences present in the database, we analyzed the effect of mutating conserved residues in the cytosol-facing part of the transmembrane region and cytoplasmic tail of HA (A/WSN/33 (H1N1) strain) on virus replication and morphology of virions. Removal of the two cytoplasmic acylation sites and substitution of a neighboring isoleucine by glutamine prevented rescue of infectious virions. In contrast, a conservative exchange of the same isoleucine, non-conservative exchanges of glycine and glutamine, deletion of the acylation site at the end of the transmembrane region and shifting it into the tail did not affect virus morphology and had only subtle effects on virus growth and on the incorporation of M1 and Ribo-Nucleoprotein Particles (RNPs). Thus, assuming that essential amino acids are conserved between HA subtypes we suggest that, besides the two cytoplasmic acylation sites (including adjacent hydrophobic residues), no other amino acids in the cytoplasmic tail of HA are indispensable for virus assembly and budding. PMID:26670246

  17. Design and analysis of structure-activity relationship of novel antimicrobial peptides derived from the conserved sequence of cecropin.

    PubMed

    Hao, Gang; Shi, Yong-Hui; Han, Jing-Hui; Li, Qi-Hui; Tang, Ya-Li; Le, Guo-Wei

    2008-03-01

    We have de novo designed four antimicrobial peptides AMP-A/B/C/D, the 51-residues peptides, which are based on the conserved sequence of cecropin. In the present study, the four peptides were chemically synthesized and their activities assayed. Their secondary structure, amphipathic property, electric field distribution and transmembrane domain were subsequently predicted by bioinformatics tools. Finally, the structure-activity relationship was analyzed from the results of activity experiments and prediction. The results of activity experiments indicated that AMP-B/C/D clearly possessed excellent broad-spectrum activity against bacteria, whereas AMP-A was almost inactive against most of the bacterial strains tested. AMP-B/C/D showed more potent activity against Gram-positive bacteria than against Gram-negative bacteria. By utilizing bioinformatics analysis tools, we found that the secondary structure of the four cation peptides was mainly alpha-helix, and the result of CD spectrum also displayed that all the peptides had considerable alpha-helix in the presence of either 50% TFE or SDS micelles. AMP-C showed much better activity than other peptides against most of the bacteria tested, owing to its remarkable cation property and the amphipathic character of its N-terminal. The study of structure-activity relationship of the designed peptides confirmed that amphipathic structure and high net positive charge were prerequisites for maintaining their activities. PMID:17929330

  18. Dose-sensitivity, conserved non-coding sequences, and duplicate gene retention through multiple tetraploidies in the grasses.

    PubMed

    Schnable, James C; Pedersen, Brent S; Subramaniam, Sabarinath; Freeling, Michael

    2011-01-01

    Whole genome duplications, or tetraploidies, are an important source of increased gene content. Following whole genome duplication, duplicate copies of many genes are lost from the genome. This loss of genes is biased both in the classes of genes deleted and the subgenome from which they are lost. Many or all classes are genes preferentially retained as duplicate copies are engaged in dose sensitive protein-protein interactions, such that deletion of any one duplicate upsets the status quo of subunit concentrations, and presumably lowers fitness as a result. Transcription factors are also preferentially retained following every whole genome duplications studied. This has been explained as a consequence of protein-protein interactions, just as for other highly retained classes of genes. We show that the quantity of conserved noncoding sequences (CNSs) associated with genes predicts the likelihood of their retention as duplicate pairs following whole genome duplication. As many CNSs likely represent binding sites for transcriptional regulators, we propose that the likelihood of gene retention following tetraploidy may also be influenced by dose-sensitive protein-DNA interactions between the regulatory regions of CNS-rich genes - nicknamed bigfoot genes - and the proteins that bind to them. Using grass genomes, we show that differential loss of CNSs from one member of a pair following the pre-grass tetraploidy reduces its chance of retention in the subsequent maize lineage tetraploidy. PMID:22645525

  19. Dose–Sensitivity, Conserved Non-Coding Sequences, and Duplicate Gene Retention Through Multiple Tetraploidies in the Grasses

    PubMed Central

    Schnable, James C.; Pedersen, Brent S.; Subramaniam, Sabarinath; Freeling, Michael

    2011-01-01

    Whole genome duplications, or tetraploidies, are an important source of increased gene content. Following whole genome duplication, duplicate copies of many genes are lost from the genome. This loss of genes is biased both in the classes of genes deleted and the subgenome from which they are lost. Many or all classes are genes preferentially retained as duplicate copies are engaged in dose sensitive protein–protein interactions, such that deletion of any one duplicate upsets the status quo of subunit concentrations, and presumably lowers fitness as a result. Transcription factors are also preferentially retained following every whole genome duplications studied. This has been explained as a consequence of protein–protein interactions, just as for other highly retained classes of genes. We show that the quantity of conserved noncoding sequences (CNSs) associated with genes predicts the likelihood of their retention as duplicate pairs following whole genome duplication. As many CNSs likely represent binding sites for transcriptional regulators, we propose that the likelihood of gene retention following tetraploidy may also be influenced by dose–sensitive protein–DNA interactions between the regulatory regions of CNS-rich genes – nicknamed bigfoot genes – and the proteins that bind to them. Using grass genomes, we show that differential loss of CNSs from one member of a pair following the pre-grass tetraploidy reduces its chance of retention in the subsequent maize lineage tetraploidy. PMID:22645525

  20. Nanopore Analysis of Nucleic Acids: Single-Molecule Studies of Molecular Dynamics, Structure, and Base Sequence

    NASA Astrophysics Data System (ADS)

    Olasagasti, Felix; Deamer, David W.

    Nucleic acids are linear polynucleotides in which each base is covalently linked to a pentose sugar and a phosphate group carrying a negative charge. If a pore having roughly the crosssectional diameter of a single-stranded nucleic acid is embedded in a thin membrane and a voltage of 100 mV or more is applied, individual nucleic acids in solution can be captured by the electrical field in the pore and translocated through by single-molecule electrophoresis. The dimensions of the pore cannot accommodate anything larger than a single strand, so each base in the molecule passes through the pore in strict linear sequence. The nucleic acid strand occupies a large fraction of the pore's volume during translocation and therefore produces a transient blockade of the ionic current created by the applied voltage. If it could be demonstrated that each nucleotide in the polymer produced a characteristic modulation of the ionic current during its passage through the nanopore, the sequence of current modulations would reflect the sequence of bases in the polymer. According to this basic concept, nanopores are analogous to a Coulter counter that detects nanoscopic molecules rather than microscopic [1,2]. However, the advantage of nanopores is that individual macromolecules can be characterized because different chemical and physical properties affect their passage through the pore. Because macromolecules can be captured in the pore as well as translocated, the nanopore can be used to detect individual functional complexes that form between a nucleic acid and an enzyme. No other technique has this capability.

  1. Deep parallel sequencing reveals conserved and novel miRNAs in gill and hepatopancreas of giant freshwater prawn.

    PubMed

    Tan, Tian Tian; Chen, Maoshan; Harikrishna, Jennifer Ann; Khairuddin, Norliana; Mohd Shamsudin, Maizatul Izzah; Zhang, Guojie; Bhassu, Subha

    2013-10-01

    MicroRNAs (miRNAs) are ~20-22 nucleotides, non protein-coding RNA regulatory genes that post-transcriptionally regulate many protein-coding genes, influencing critical biological and metabolic processes. While the number of known microRNA is increasing, there is currently no published data for miRNA from giant freshwater prawns, Macrobrachium rosenbergii (M. rosenbergii), a commercially cultured and economically important food species. In this study, we identified novel miRNAs in the gill and hepatopancreas of M. rosenbergii. Through a deep parallel sequencing analysis and an in silico data analysis approach, 327 miRNA families were identified from small RNA libraries with reference to both the de novo transcriptome of M. rosenbergii obtained from RNA-Seq and to miRBase (Release 18.0, November 2012). Based on the identified mature miRNA and recovered precursor sequences that form appropriate hairpin structures, three conserved miRNA (miR125, miR750, miR993) and 27 novel miRNA candidates encoding messenger-like non-coding RNA were identified. miR-125, miR-750, G-m0002/H-m0009, G-m0005, G-m0008/H-m0016, G-m0011/H-m0027 and G-m0015 were selected for experimental validation with stem-loop quantitative RT-PCR and were found to be coherent with the expression profile of deep sequencing data as evaluated with Pearson's correlation coefficient (r = 0.835178 for miRNA in gill, r = 0.724131 for miRNA in hepatopancreas). Using a combinatorial approach of pathway enrichment analysis and inverse expression relationship of miRNA and mRNA, four co-expressed novel miRNA candidates (G-m0005, G-m0008/H-m0016, G-m0011/H-m0027, and G-m0015) were found to be associated with energy metabolism. In addition, the expression of the three novel miRNA candidates (G-m0005, G-m0008/H-m0016, and G-m0011/H-m0027) were also found to be significantly reduced at 9 and 24 h post infection in M. rosenbergii challenged with infectious hypodermal and hematopoietic necrosis virus, suggesting a functional

  2. Highly conserved influenza A virus epitope sequences as candidates of H3N2 flu vaccine targets.

    PubMed

    Wu, Ko-Wen; Chien, Chih-Yi; Li, Shiao-Wen; King, Chwan-Chuen; Chang, Chuan-Hsiung

    2012-08-01

    This study focused on identifying the conserved epitopes in a single subtype A (H3N2)-as candidates for vaccine targets. We identified a total of 32 conserved epitopes in four viral proteins [22 HA, 4PB1, 3 NA, 3 NP]. Evaluation of conserved epitopes in coverage during 1968-2010 revealed that (1) 12 HA conserved epitopes were highly present in the circulating viruses; (2) the remaining 10 HA conserved epitopes appeared with lower percentage but a significantly increasing trend after 1989 [p<0.001]; and (3) the conserved epitopes in NA, NP and PB1 are also highly frequent in wild-type viruses. These conserved epitopes also covered an extremely high percentage of the 16 vaccine strains during the 42 year period. The identification of highly conserved epitopes using our approach can also be applied to develop broad-spectrum vaccines. PMID:22698979

  3. Complete amino acid sequence of a histidine-rich proteolytic fragment of human ceruloplasmin.

    PubMed

    Kingston, I B; Kingston, B L; Putnam, F W

    1979-04-01

    The complete amino acid sequence has been determined for a fragment of human ceruloplasmin [ferroxidase; iron(II):oxygen oxidoreductase, EC 1.16.3.1]. The fragment (designated Cp F5) contains 159 amino acid residues and has a molecular weight of 18,650; it lacks carbohydrate, is rich in histidine, and contains one free cysteine that may be part of a copper-binding site. This fragment is present in most commercial preparations of ceruloplasmin, probably owing to proteolytic degradation, but can also be obtained by limited cleavage of single-chain ceruloplasmin with plasmin. Cp F5 probably is an intact domain attached to the COOH-terminal end of single-chain ceruloplasmin via a labile interdomain peptide bond. A model of the secondary structure predicted by empirical methods suggests that almost one-third of the amino acid residues are distributed in alpha helices, about a third in beta-sheet structure, and the remainder in beta turns and unidentified structures. Computer analysis of the amino acid sequence has not demonstrated a statistically significant relationship between this ceruloplasmin fragment and any other protein, but there is some evidence for an internal duplication. PMID:287005

  4. The amino acid sequence of Lady Amherst's pheasant (Chrysolophus amherstiae) and golden pheasant (Chrysolophus pictus) egg-white lysozymes.

    PubMed

    Araki, T; Kuramoto, M; Torikata, T

    1990-09-01

    The amino acids of Lady Amherst's pheasant and golden pheasant egg-white lysozymes have been sequenced. The carboxymethylated lysozymes were digested with trypsin followed by sequencing of the tryptic peptides. Lady Amherst's pheasant lysozyme proved to consist of 129 amino acid residues, and a relative molecular mass of 14,423 Da was calculated. This lysozyme had 6 amino acids substitutions when compared with hen egg-white lysozyme: Phe3 to Tyr, His15 to Leu, Gln41 to His, Asn77 to His, Gln 121 to Asn, and a newly found substitution of Ile124 to Thr. The amino acid sequence of golden pheasant lysozyme was identical to that of Lady Amherst's phesant lysozyme. The phylogenetic tree constructured by the comparison of amino acid sequences of phasianoid birds lysozymes revealed a minimum genetic distance between these pheasants and the turkey-peafowl group. PMID:1368578

  5. 5'-terminal nucleotide sequences of mammalian type C helper viruses are conserved in the genomes of replication-defective mammalian transforming viruses.

    PubMed Central

    Tronick, S R; Cabradilla, C D; Aaronson, S A; Haseltine, W A

    1978-01-01

    The RNAs of replication-defective murine and primate type C transforming viruses were analyzed for the presence of nucleotide sequences homologous to the genomes of their respective helper type C viruses by using DNAs complementary (cDNA) to either the 5'-terminal (cDNA5') or total (cDNAtotal) nucleotide sequences of the helper virus RNA. The defective viruses examined have previously been shown to vary in their ability to express helper viral gag gene proteins. With cDNAtotal as a probe, these transforming viruses were shown to vary in their representation of helper sequences (15 to 60% hybridization of cDNAtotal). In striking contrast, 5'-terminal-specific sequences of the helper virus were conserved in the RNAs of every transforming virus tested (is greater than 80% hybridization of cDNA5'). These findings suggest a critical role for these sequences in the life cycle of the defective transforming virus. PMID:209210

  6. Complete Genome Sequence of a thermotolerant sporogenic lactic acid bacterium, Bacillus coagulans strain 36D1

    PubMed Central

    Rhee, Mun Su; Moritz, Brélan E.; Xie, Gary; Glavina del Rio, T.; Dalin, E.; Tice, H.; Bruce, D.; Goodwin, L.; Chertkov, O.; Brettin, T.; Han, C.; Detter, C.; Pitluck, S.; Land, Miriam L.; Patel, Milind; Ou, Mark; Harbrucker, Roberta; Ingram, Lonnie O.; Shanmugam, K. T.

    2011-01-01

    Bacillus coagulans is a ubiquitous soil bacterium that grows at 50-55 °C and pH 5.0 and ferments various sugars that constitute plant biomass to L (+)-lactic acid. The ability of this sporogenic lactic acid bacterium to grow at 50-55 °C and pH 5.0 makes this organism an attractive microbial biocatalyst for production of optically pure lactic acid at industrial scale not only from glucose derived from cellulose but also from xylose, a major constituent of hemicellulose. This bacterium is also considered as a potential probiotic. Complete genome sequence of a representative strain, B. coagulans strain 36D1, is presented and discussed. PMID:22675583

  7. Complete amino acid sequence of globin chains and biological activity of fragmented crocodile hemoglobin (Crocodylus siamensis).

    PubMed

    Srihongthong, Saowaluck; Pakdeesuwan, Anawat; Daduang, Sakda; Araki, Tomohiro; Dhiravisit, Apisak; Thammasirirak, Sompong

    2012-08-01

    Hemoglobin, α-chain, β-chain and fragmented hemoglobin of Crocodylus siamensis demonstrated both antibacterial and antioxidant activities. Antibacterial and antioxidant properties of the hemoglobin did not depend on the heme structure but could result from the compositions of amino acid residues and structures present in their primary structure. Furthermore, thirteen purified active peptides were obtained by RP-HPLC analyses, corresponding to fragments in the α-globin chain and the β-globin chain which are mostly located at the N-terminal and C-terminal parts. These active peptides operate on the bacterial cell membrane. The globin chains of Crocodylus siamensis showed similar amino acids to the sequences of Crocodylus niloticus. The novel amino acid substitutions of α-chain and β-chain are not associated with the heme binding site or the bicarbonate ion binding site, but could be important through their interactions with membranes of bacteria. PMID:22648692

  8. Comparative characterization of random-sequence proteins consisting of 5, 12, and 20 kinds of amino acids.

    PubMed

    Tanaka, Junko; Doi, Nobuhide; Takashima, Hideaki; Yanagawa, Hiroshi

    2010-04-01

    Screening of functional proteins from a random-sequence library has been used to evolve novel proteins in the field of evolutionary protein engineering. However, random-sequence proteins consisting of the 20 natural amino acids tend to aggregate, and the occurrence rate of functional proteins in a random-sequence library is low. From the viewpoint of the origin of life, it has been proposed that primordial proteins consisted of a limited set of amino acids that could have been abundantly formed early during chemical evolution. We have previously found that members of a random-sequence protein library constructed with five primitive amino acids show high solubility (Doi et al., Protein Eng Des Sel 2005;18:279-284). Although such a library is expected to be appropriate for finding functional proteins, the functionality may be limited, because they have no positively charged amino acid. Here, we constructed three libraries of 120-amino acid, random-sequence proteins using alphabets of 5, 12, and 20 amino acids by preselection using mRNA display (to eliminate sequences containing stop codons and frameshifts) and characterized and compared the structural properties of random-sequence proteins arbitrarily chosen from these libraries. We found that random-sequence proteins constructed with the 12-member alphabet (including five primitive amino acids and positively charged amino acids) have higher solubility than those constructed with the 20-member alphabet, though other biophysical properties are very similar in the two libraries. Thus, a library of moderate complexity constructed from 12 amino acids may be a more appropriate resource for functional screening than one constructed from 20 amino acids. PMID:20162614

  9. N-Terminal Amino Acid Sequence Determination of Proteins by N-Terminal Dimethyl Labeling: Pitfalls and Advantages When Compared with Edman Degradation Sequence Analysis.

    PubMed

    Chang, Elizabeth; Pourmal, Sergei; Zhou, Chun; Kumar, Rupesh; Teplova, Marianna; Pavletich, Nikola P; Marians, Kenneth J; Erdjument-Bromage, Hediye

    2016-07-01

    In recent history, alternative approaches to Edman sequencing have been investigated, and to this end, the Association of Biomolecular Resource Facilities (ABRF) Protein Sequencing Research Group (PSRG) initiated studies in 2014 and 2015, looking into bottom-up and top-down N-terminal (Nt) dimethyl derivatization of standard quantities of intact proteins with the aim to determine Nt sequence information. We have expanded this initiative and used low picomole amounts of myoglobin to determine the efficiency of Nt-dimethylation. Application of this approach on protein domains, generated by limited proteolysis of overexpressed proteins, confirms that it is a universal labeling technique and is very sensitive when compared with Edman sequencing. Finally, we compared Edman sequencing and Nt-dimethylation of the same polypeptide fragments; results confirm that there is agreement in the identity of the Nt amino acid sequence between these 2 methods. PMID:27006647

  10. N-Terminal Amino Acid Sequence Determination of Proteins by N-Terminal Dimethyl Labeling: Pitfalls and Advantages When Compared with Edman Degradation Sequence Analysis

    PubMed Central

    Chang, Elizabeth; Pourmal, Sergei; Zhou, Chun; Kumar, Rupesh; Teplova, Marianna; Pavletich, Nikola P.; Marians, Kenneth J.

    2016-01-01

    In recent history, alternative approaches to Edman sequencing have been investigated, and to this end, the Association of Biomolecular Resource Facilities (ABRF) Protein Sequencing Research Group (PSRG) initiated studies in 2014 and 2015, looking into bottom-up and top-down N-terminal (Nt) dimethyl derivatization of standard quantities of intact proteins with the aim to determine Nt sequence information. We have expanded this initiative and used low picomole amounts of myoglobin to determine the efficiency of Nt-dimethylation. Application of this approach on protein domains, generated by limited proteolysis of overexpressed proteins, confirms that it is a universal labeling technique and is very sensitive when compared with Edman sequencing. Finally, we compared Edman sequencing and Nt-dimethylation of the same polypeptide fragments; results confirm that there is agreement in the identity of the Nt amino acid sequence between these 2 methods. PMID:27006647

  11. Partial amino acid sequence of fructose-1,6-bisphosphatase from the blue-green algae Synechococcus leopoliensis.

    PubMed

    Marcus, F; Latshaw, S P; Steup, M; Gerbling, K P

    1989-08-01

    Purified fructose-1,6-bisphosphatase from the cyanobacterium Synechococcus leopoliensis was S-carboxymethylated and cleaved with trypsin. The resulting peptides were purified by reversed-phase high performance liquid chromatography and the amino acid sequence of six of the purified peptides was determined by gas-phase microsequencing. The results revealed sequence homology with other fructose-1,6-bisphosphatases. The obtained sequence data provides information required for the design of oligonucleotide hybridization probes to screen existing libraries of cyanobacterial DNA. The determination of the amino acid sequence of cyanobacterial proteins may yield important information with respect to the endosymbiotic theory of evolution. PMID:2550924

  12. Protein sequence analysis by incorporating modified chaos game and physicochemical properties into Chou's general pseudo amino acid composition.

    PubMed

    Xu, Chunrui; Sun, Dandan; Liu, Shenghui; Zhang, Yusen

    2016-10-01

    In this contribution we introduced a novel graphical method to compare protein sequences. By mapping a protein sequence into 3D space based on codons and physicochemical properties of 20 amino acids, we are able to get a unique P-vector from the 3D curve. This approach is consistent with wobble theory of amino acids. We compute the distance between sequences by their P-vectors to measure similarities/dissimilarities among protein sequences. Finally, we use our method to analyze four datasets and get better results compared with previous approaches. PMID:27375218

  13. A single amino acid change, Q114R, in the cleavage-site sequence of Newcastle disease virus fusion protein attenuates viral replication and pathogenicity.

    PubMed

    Samal, Sweety; Kumar, Sachin; Khattar, Sunil K; Samal, Siba K

    2011-10-01

    A key determinant of Newcastle disease virus (NDV) virulence is the amino acid sequence at the fusion (F) protein cleavage site. The NDV F protein is synthesized as an inactive precursor, F(0), and is activated by proteolytic cleavage between amino acids 116 and 117 to produce two disulfide-linked subunits, F(1) and F(2). The consensus sequence of the F protein cleavage site of virulent [(112)(R/K)-R-Q-(R/K)-R↓F-I(118)] and avirulent [(112)(G/E)-(K/R)-Q-(G/E)-R↓L-I(118)] strains contains a conserved glutamine residue at position 114. Recently, some NDV strains from Africa and Madagascar were isolated from healthy birds and have been reported to contain five basic residues (R-R-R-K-R↓F-I/V or R-R-R-R-R↓F-I/V) at the F protein cleavage site. In this study, we have evaluated the role of this conserved glutamine residue in the replication and pathogenicity of NDV by using the moderately pathogenic Beaudette C strain and by making Q114R, K115R and I118V mutants of the F protein in this strain. Our results showed that changing the glutamine to a basic arginine residue reduced viral replication and attenuated the pathogenicity of the virus in chickens. The pathogenicity was further reduced when the isoleucine at position 118 was substituted for valine. PMID:21677091

  14. An intronic peroxisome proliferator-activated receptor-binding sequence mediates fatty acid induction of the human carnitine palmitoyltransferase 1A.

    PubMed

    Napal, Laura; Marrero, Pedro F; Haro, Diego

    2005-12-01

    The liver plays a central role in the response to fasting. The hormonal profile in this condition, low insulin, and high concentrations of glucagon in plasma, induce the release of large amounts of fatty acids from adipose tissue. Prolonged starvation can therefore induce a dramatic change in the fatty acid oxidative capacity of liver metabolism. Modulation of gene expression by PPARalpha plays a crucial role in this response. While a major role for PPARalpha in the liver is to produce ketone bodies as fuel through beta-oxidation for peripheral tissues during fast, its participation in the control of CPT1A, the rate-limiting step of the pathway, remains controversial. Using Web-based software (VISTA) combining transcription factor binding site database searches with comparative sequence analyses, we have localized a conserved functional PPAR responsive element downstream of the transcriptional start site of the human CPT1A gene. We have shown that this sequence is fundamental for fatty acids or PGC1-induced transcriptional activation of the CPT1A gene. These results corroborate the hypothesis that PPARalpha regulates the limiting step in the oxidation of fatty acids in liver mitochondria. PMID:16271724

  15. Comparative Sequence and Structure Analysis Reveals the Conservation and Diversity of Nucleotide Positions and Their Associated Tertiary Interactions in the Riboswitches

    PubMed Central

    Appasamy, Sri D.; Ramlan, Effirul Ikhwan; Firdaus-Raih, Mohd

    2013-01-01

    The tertiary motifs in complex RNA molecules play vital roles to either stabilize the formation of RNA 3D structure or to provide important biological functionality to the molecule. In order to better understand the roles of these tertiary motifs in riboswitches, we examined 11 representative riboswitch PDB structures for potential agreement of both motif occurrences and conservations. A total of 61 unique tertiary interactions were found in the reference structures. In addition to the expected common A-minor motifs and base-triples mainly involved in linking distant regions the riboswitch structures three highly conserved variants of A-minor interactions called G-minors were found in the SAM-I and FMN riboswitches where they appear to be involved in the recognition of the respective ligand’s functional groups. From our structural survey as well as corresponding structure and sequence alignments, the agreement between motif occurrences and conservations are very prominent across the representative riboswitches. Our analysis provide evidence that some of these tertiary interactions are essential components to form the structure where their sequence positions are conserved despite a high degree of diversity in other parts of the respective riboswitches sequences. This is indicative of a vital role for these tertiary interactions in determining the specific biological function of riboswitch. PMID:24040136

  16. Human ERCC5 cDNA-cosmid complementation for excision repair and bipartite amino acid domains conserved with RAD proteins of saccharomyces cerevisiae and schizosaccharomyces pombe

    SciTech Connect

    MacInnes, M.A.; Dickson, J.A.; Hernandez, R.R.; Lin, G.Y.; Park, M.S.; Schauer, S.; Reynolds, R.J.; Strniste, G.F. ); Learmonth, D. ); Mudgett, J.S. ); Yu, J.Y. )

    1993-10-01

    Several human genes related to DNA excision repair (ER) have been isolated via ER cross-species complementation (ERCC) of UV-sensitive CHO cells. The authors have now isolated and characterized cDNAs for the human ERCC5 gene that complement CHO UV135 cells. The ERCC5 mRNA size is about 4.6 kb. Their available cDNA clones are partial length, and no single clone was active for UV135 complementation. When cDNAs were mixed pairwise with a cosmid clone containing an overlapping 5[prime]-end segment of the ERCC5 gene, DNA transfer produced UV-resistant colonies with 60 to 95% correction of UV resistance relative to either a genomic ERCC5 DNA transformant or the CHO AA8 progenitor cells. cDNA-cosmid transformants regained intermediate levels (20 to 45%) of ER-dependent reactivation of a UV-damaged pSVCATgpt reporter plasmid. Their evidence strongly implicates an in situ recombination mechanism in cDNA-cosmid complementation for ER. The complete deduced amino acid sequence of ERCC5 was reconstructed for several cDNA clones encoding a predicted protein of 1,186 amino acids. The ERCC5 protein has extensive sequence similarities, in bipartite domains A and B, to products of RAD repair genes of two yeast, Saccharomyces cerevisiae RAD2 and Schizosaccharomyces pombe rad13. Sequence, structural, and functional data taken together indicate that ERCC5 and its relatives are probable functional homologs. A second locus represented by S. cerevisiae YKL510 and S. pombe rad2 genes is structurally distinct from the ERCC5 locus but retains vestigial A and B domain similarities. Their analyses suggest that ERCC5 is a nuclear-localized protein with one or more highly conserved helix-loop-helix segments within domains A and B. 69 refs., 6 figs., 1 tab.

  17. Nucleotide sequence of the phosphoglycerate kinase gene from the extreme thermophile Thermus thermophilus. Comparison of the deduced amino acid sequence with that of the mesophilic yeast phosphoglycerate kinase.

    PubMed Central

    Bowen, D; Littlechild, J A; Fothergill, J E; Watson, H C; Hall, L

    1988-01-01

    Using oligonucleotide probes derived from amino acid sequencing information, the structural gene for phosphoglycerate kinase from the extreme thermophile, Thermus thermophilus, was cloned in Escherichia coli and its complete nucleotide sequence determined. The gene consists of an open reading frame corresponding to a protein of 390 amino acid residues (calculated Mr 41,791) with an extreme bias for G or C (93.1%) in the codon third base position. Comparison of the deduced amino acid sequence with that of the corresponding mesophilic yeast enzyme indicated a number of significant differences. These are discussed in terms of the unusual codon bias and their possible role in enhanced protein thermal stability. Images Fig. 1. PMID:3052437

  18. Bacteria obtained from a sequencing batch reactor that are capable of growth on dehydroabietic acid.

    PubMed Central

    Mohn, W W

    1995-01-01

    Eleven isolates capable of growth on the resin acid dehydroabietic acid (DhA) were obtained from a sequencing batch reactor designed to treat a high-strength process stream from a paper mill. The isolates belonged to two groups, represented by strains DhA-33 and DhA-35, which were characterized. In the bioreactor, bacteria like DhA-35 were more abundant than those like DhA-33. The population in the bioreactor of organisms capable of growth on DhA was estimated to be 1.1 x 10(6) propagules per ml, based on a most-probable-number determination. Analysis of small-subunit rRNA partial sequences indicated that DhA-33 was most closely related to Sphingomonas yanoikuyae (Sab = 0.875) and that DhA-35 was most closely related to Zoogloea ramigera (Sab = 0.849). Both isolates additionally grew on other abietanes, i.e., abietic and palustric acids, but not on the pimaranes, pimaric and isopimaric acids. For DhA-33 and DhA-35 with DhA as the sole organic substrate, doubling times were 2.7 and 2.2 h, respectively, and growth yields were 0.30 and 0.25 g of protein per g of DhA, respectively. Glucose as a cosubstrate stimulated growth of DhA-33 on DhA and stimulated DhA degradation by the culture. Pyruvate as a cosubstrate did not stimulate growth of DhA-35 on DhA and reduced the specific rate of DhA degradation of the culture. DhA induced DhA and abietic acid degradation activities in both strains, and these activities were heat labile. Cell suspensions of both strains consumed DhA at a rate of 6 mumol mg of protein-1 h-1.(ABSTRACT TRUNCATED AT 250 WORDS) PMID:7793937

  19. Sequence Conservation and Sexually Dimorphic Expression of the Ftz-F1 Gene in the Crustacean Daphnia magna

    PubMed Central

    Mohamad Ishak, Nur Syafiqah; Kato, Yasuhiko; Matsuura, Tomoaki; Watanabe, Hajime

    2016-01-01

    Identifying the genes required for environmental sex determination is important for understanding the evolution of diverse sex determination mechanisms in animals. Orthologs of Drosophila orphan receptor Fushi tarazu factor-1 (Ftz-F1) are known to function in genetic sex determination. In contrast, their roles in environmental sex determination remain unknown. In this study, we have cloned and characterized the Ftz-F1 ortholog in the branchiopod crustacean Daphnia magna, which produces males in response to environmental stimuli. Similar to that observed in Drosophila, D. magna Ftz-F1 (DapmaFtz-F1) produces two splicing variants, αFtz-F1 and βFtz-F1, which encode 699 and 777 amino acids, respectively. Both isoforms share a DNA-binding domain, a ligand-binding domain, and an AF-2 activation domain and differ only at the A/B domain. The phylogenetic position and genomic structure of DapmaFtz-F1 suggested that this gene has diverged from an ancestral gene common to branchiopod crustacean and insect Ftz-F1 genes. qRT-PCR showed that at the one cell and gastrulation stages, both DapmaFtz-F1 isoforms are two-fold more abundant in males than in females. In addition, in later stages, their sexual dimorphic expressions were maintained in spite of reduced expression. Time-lapse imaging of DapmaFtz-F1 RNAi embryos was performed in H2B-GFP expressing transgenic Daphnia, demonstrating that development of the RNAi embryos slowed down after the gastrulation stage and stopped at 30–48 h after ovulation. DapmaFtz-F1 shows high homology to insect Ftz-F1 orthologs based on its amino acid sequence and exon-intron organization. The sexually dimorphic expression of DapmaFtz-F1 suggests that it plays a role in environmental sex determination of D. magna. PMID:27138373

  20. Sequence Conservation and Sexually Dimorphic Expression of the Ftz-F1 Gene in the Crustacean Daphnia magna.

    PubMed

    Mohamad Ishak, Nur Syafiqah; Kato, Yasuhiko; Matsuura, Tomoaki; Watanabe, Hajime

    2016-01-01

    Identifying the genes required for environmental sex determination is important for understanding the evolution of diverse sex determination mechanisms in animals. Orthologs of Drosophila orphan receptor Fushi tarazu factor-1 (Ftz-F1) are known to function in genetic sex determination. In contrast, their roles in environmental sex determination remain unknown. In this study, we have cloned and characterized the Ftz-F1 ortholog in the branchiopod crustacean Daphnia magna, which produces males in response to environmental stimuli. Similar to that observed in Drosophila, D. magna Ftz-F1 (DapmaFtz-F1) produces two splicing variants, αFtz-F1 and βFtz-F1, which encode 699 and 777 amino acids, respectively. Both isoforms share a DNA-binding domain, a ligand-binding domain, and an AF-2 activation domain and differ only at the A/B domain. The phylogenetic position and genomic structure of DapmaFtz-F1 suggested that this gene has diverged from an ancestral gene common to branchiopod crustacean and insect Ftz-F1 genes. qRT-PCR showed that at the one cell and gastrulation stages, both DapmaFtz-F1 isoforms are two-fold more abundant in males than in females. In addition, in later stages, their sexual dimorphic expressions were maintained in spite of reduced expression. Time-lapse imaging of DapmaFtz-F1 RNAi embryos was performed in H2B-GFP expressing transgenic Daphnia, demonstrating that development of the RNAi embryos slowed down after the gastrulation stage and stopped at 30-48 h after ovulation. DapmaFtz-F1 shows high homology to insect Ftz-F1 orthologs based on its amino acid sequence and exon-intron organization. The sexually dimorphic expression of DapmaFtz-F1 suggests that it plays a role in environmental sex determination of D. magna. PMID:27138373

  1. Nucleic and amino acid sequences relating to a novel transketolase, and methods for the expression thereof

    DOEpatents

    Croteau, Rodney Bruce; Wildung, Mark Raymond; Lange, Bernd Markus; McCaskill, David G.

    2001-01-01

    cDNAs encoding 1-deoxyxylulose-5-phosphate synthase from peppermint (Mentha piperita) have been isolated and sequenced, and the corresponding amino acid sequences have been determined. Accordingly, isolated DNA sequences (SEQ ID NO:3, SEQ ID NO:5, SEQ ID NO:7) are provided which code for the expression of 1-deoxyxylulose-5-phosphate synthase from plants. In another aspect the present invention provides for isolated, recombinant DXPS proteins, such as the proteins having the sequences set forth in SEQ ID NO:4, SEQ ID NO:6 and SEQ ID NO:8. In other aspects, replicable recombinant cloning vehicles are provided which code for plant 1-deoxyxylulose-5-phosphate synthases, or for a base sequence sufficiently complementary to at least a portion of 1-deoxyxylulose-5-phosphate synthase DNA or RNA to enable hybridization therewith. In yet other aspects, modified host cells are provided that have been transformed, transfected, infected and/or injected with a recombinant cloning vehicle and/or DNA sequence encoding a plant 1-deoxyxylulose-5-phosphate synthase. Thus, systems and methods are provided for the recombinant expression of the aforementioned recombinant 1-deoxyxylulose-5-phosphate synthase that may be used to facilitate its production, isolation and purification in significant amounts. Recombinant 1-deoxyxylulose-5-phosphate synthase may be used to obtain expression or enhanced expression of 1-deoxyxylulose-5-phosphate synthase in plants in order to enhance the production of 1-deoxyxylulose-5-phosphate, or its derivatives such as isopentenyl diphosphate (BP), or may be otherwise employed for the regulation or expression of 1-deoxyxylulose-5-phosphate synthase, or the production of its products.

  2. Novel method for PIK3CA mutation analysis: locked nucleic acid--PCR sequencing.

    PubMed

    Ang, Daphne; O'Gara, Rebecca; Schilling, Amy; Beadling, Carol; Warrick, Andrea; Troxell, Megan L; Corless, Christopher L

    2013-05-01

    Somatic mutations in PIK3CA are commonly seen in invasive breast cancer and several other carcinomas, occurring in three hotspots: codons 542 and 545 of exon 9 and in codon 1047 of exon 20. We designed a locked nucleic acid (LNA)-PCR sequencing assay to detect low levels of mutant PIK3CA DNA with attention to avoiding amplification of a pseudogene on chromosome 22 that has >95% homology to exon 9 of PIK3CA. We tested 60 FFPE breast DNA samples with known PIK3CA mutation status (48 cases had one or more PIK3CA mutations, and 12 were wild type) as identified by PCR-mass spectrometry. PIK3CA exons 9 and 20 were amplified in the presence or absence of LNA-oligonucleotides designed to bind to the wild-type sequences for codons 542, 545, and 1047, and partially suppress their amplification. LNA-PCR sequencing confirmed all 51 PIK3CA mutations; however, the mutation detection rate by standard Sanger sequencing was only 69% (35 of 51). Of the 12 PIK3CA wild-type cases, LNA-PCR sequencing detected three additional H1047R mutations in "normal" breast tissue and one E545K in usual ductal hyperplasia. Histopathological review of these three normal breast specimens showed columnar cell change in two (both with known H1047R mutations) and apocrine metaplasia in one. The novel LNA-PCR shows higher sensitivity than standard Sanger sequencing and did not amplify the known pseudogene. PMID:23541593

  3. Genome Sequence Analysis of the Naphthenic Acid Degrading and Metal Resistant Bacterium Cupriavidus gilardii CR3.

    PubMed

    Wang, Xiaoyu; Chen, Meili; Xiao, Jingfa; Hao, Lirui; Crowley, David E; Zhang, Zhewen; Yu, Jun; Huang, Ning; Huo, Mingxin; Wu, Jiayan

    2015-01-01

    Cupriavidus sp. are generally heavy metal tolerant bacteria with the ability to degrade a variety of aromatic hydrocarbon compounds, although the degradation pathways and substrate versatilities remain largely unknown. Here we studied the bacterium Cupriavidus gilardii strain CR3, which was isolated from a natural asphalt deposit, and which was shown to utilize naphthenic acids as a sole carbon source. Genome sequencing of C. gilardii CR3 was carried out to elucidate possible mechanisms for the naphthenic acid biodegradation. The genome of C. gilardii CR3 was composed of two circular chromosomes chr1 and chr2 of respectively 3,539,530 bp and 2,039,213 bp in size. The genome for strain CR3 encoded 4,502 putative protein-coding genes, 59 tRNA genes, and many other non-coding genes. Many genes were associated with xenobiotic biodegradation and metal resistance functions. Pathway prediction for degradation of cyclohexanecarboxylic acid, a representative naphthenic acid, suggested that naphthenic acid undergoes initial ring-cleavage, after which the ring fission products can be degraded via several plausible degradation pathways including a mechanism similar to that used for fatty acid oxidation. The final metabolic products of these pathways are unstable or volatile compounds that were not toxic to CR3. Strain CR3 was also shown to have tolerance to at least 10 heavy metals, which was mainly achieved by self-detoxification through ion efflux, metal-complexation and metal-reduction, and a powerful DNA self-repair mechanism. Our genomic analysis suggests that CR3 is well adapted to survive the harsh environment in natural asphalts containing naphthenic acids and high concentrations of heavy metals. PMID:26301592

  4. Genome Sequence Analysis of the Naphthenic Acid Degrading and Metal Resistant Bacterium Cupriavidus gilardii CR3

    PubMed Central

    Xiao, Jingfa; Hao, Lirui; Crowley, David E.; Zhang, Zhewen; Yu, Jun; Huang, Ning; Huo, Mingxin; Wu, Jiayan

    2015-01-01

    Cupriavidus sp. are generally heavy metal tolerant bacteria with the ability to degrade a variety of aromatic hydrocarbon compounds, although the degradation pathways and substrate versatilities remain largely unknown. Here we studied the bacterium Cupriavidus gilardii strain CR3, which was isolated from a natural asphalt deposit, and which was shown to utilize naphthenic acids as a sole carbon source. Genome sequencing of C. gilardii CR3 was carried out to elucidate possible mechanisms for the naphthenic acid biodegradation. The genome of C. gilardii CR3 was composed of two circular chromosomes chr1 and chr2 of respectively 3,539,530 bp and 2,039,213 bp in size. The genome for strain CR3 encoded 4,502 putative protein-coding genes, 59 tRNA genes, and many other non-coding genes. Many genes were associated with xenobiotic biodegradation and metal resistance functions. Pathway prediction for degradation of cyclohexanecarboxylic acid, a representative naphthenic acid, suggested that naphthenic acid undergoes initial ring-cleavage, after which the ring fission products can be degraded via several plausible degradation pathways including a mechanism similar to that used for fatty acid oxidation. The final metabolic products of these pathways are unstable or volatile compounds that were not toxic to CR3. Strain CR3 was also shown to have tolerance to at least 10 heavy metals, which was mainly achieved by self-detoxification through ion efflux, metal-complexation and metal-reduction, and a powerful DNA self-repair mechanism. Our genomic analysis suggests that CR3 is well adapted to survive the harsh environment in natural asphalts containing naphthenic acids and high concentrations of heavy metals. PMID:26301592

  5. Bile acid sulfotransferase I from rat liver sulfates bile acids and 3-hydroxy steroids: purification, N-terminal amino acid sequence, and kinetic properties.

    PubMed

    Barnes, S; Buchina, E S; King, R J; McBurnett, T; Taylor, K B

    1989-04-01

    A bile acid:3'phosphoadenosine-5'phosphosulfate:sulfotransferase (BAST I) from adult female rat liver cytosol has been purified 157-fold by a two-step isolation procedure. The N-terminal amino acid sequence of the 30,000 subunit has been determined for the first 35 residues. The Vmax of purified BAST I is 18.7 nmol/min per mg protein with N-(3-hydroxy-5 beta-cholanoyl)glycine (glycolithocholic acid) as substrate, comparable to that of the corresponding purified human BAST (Chen, L-J., and I. H. Segel, 1985. Arch. Biochem. Biophys. 241: 371-379). BAST I activity has a broad pH optimum from 5.5-7.5. Although maximum activity occurs with 5 mM MgCl2, Mg2+ is not essential for BAST I activity. The greatest sulfotransferase activity and the highest substrate affinity is observed with bile acids or steroids that have a steroid nucleus containing a 3 beta-hydroxy group and a 5-6 double bond or a trans A-B ring junction. These substrates have normal hyperbolic initial velocity curves with substrate inhibition occurring above 5 microM. Of the saturated 5 beta-bile acids, those with a single 3-hydroxy group are the most active. The addition of a second hydroxy group at the 6- or 7-position eliminates more than 99% of the activity. In contrast, 3 alpha,12 alpha-dihydroxy-5 beta-cholan-24-oic acid (deoxycholic acid) is an excellent substrate. The initial velocity curves for glycolithocholic and deoxycholic acid conjugates are sigmoidal rather than hyperbolic, suggestive of an allosteric effect. Maximum activity is observed at 80 microM for glycolithocholic acid. All substrates, bile acids and steroids, are inhibited by the 5 beta-bile acid, 3-keto-5 beta-cholanoic acid. The data suggest that BAST I is the same protein as hydrosteroid sulfotransferase 2 (Marcus, C. J., et al. 1980. Anal. Biochem. 107: 296-304). PMID:2754334

  6. Sequence-defined bioactive macrocycles via an acid-catalysed cascade reaction

    NASA Astrophysics Data System (ADS)

    Porel, Mintu; Thornlow, Dana N.; Phan, Ngoc N.; Alabi, Christopher A.

    2016-06-01

    Synthetic macrocycles derived from sequence-defined oligomers are a unique structural class whose ring size, sequence and structure can be tuned via precise organization of the primary sequence. Similar to peptides and other peptidomimetics, these well-defined synthetic macromolecules become pharmacologically relevant when bioactive side chains are incorporated into their primary sequence. In this article, we report the synthesis of oligothioetheramide (oligoTEA) macrocycles via a one-pot acid-catalysed cascade reaction. The versatility of the cyclization chemistry and modularity of the assembly process was demonstrated via the synthesis of >20 diverse oligoTEA macrocycles. Structural characterization via NMR spectroscopy revealed the presence of conformational isomers, which enabled the determination of local chain dynamics within the macromolecular structure. Finally, we demonstrate the biological activity of oligoTEA macrocycles designed to mimic facially amphiphilic antimicrobial peptides. The preliminary results indicate that macrocyclic oligoTEAs with just two-to-three cationic charge centres can elicit potent antibacterial activity against Gram-positive and Gram-negative bacteria.

  7. Unconventional amino acid sequence of the sun anemone (Stoichactis helianthus) polypeptide neurotoxin

    SciTech Connect

    Kem, W.; Dunn, B.; Parten, B.; Pennington, M.; Price, D.

    1986-05-01

    A 5000 dalton polypeptide neurotoxin (Sh-NI) purified by G50 Sephadex, P-cellulose, and SP-Sephadex chromatography was homogeneous by isoelectric focusing. Sh-NI was highly toxic to crayfish (LD/sub 50/ 0.6 ..mu..g/kg) but without effect upon mice at 15,000 ..mu..g/kg (i.p. injection). The reduced, /sup 3/H-carboxymethylated toxin and its fragments were subjected to automatic Edman degradation and the resulting PTH-amino acids were identified by HPLC, back hydrolysis, and scintillation counting. Peptides resulting from proteolytic (clostripain, staphylococcal protease) and chemical (tryptophan) cleavage were sequenced. The sequence is: AACKCDDEGPDIRTAPLTGTVDLGSCNAGWEKCASYYTIIADCCRKKK. This sequence differs considerably from the homologous Anemonia and Anthopleura toxins; many of the identical residues (6 half-cystines, G9, P10, R13, G19, G29, W30) are probably critical for folding rather than receptor recognition. However, the Sh-NI sequence closely resembles Radioanthus macrodactylus neurotoxin III and r. paumotensis II. The authors propose that Sh-NI and related Radioanthus toxins act upon a different site on the sodium channel.

  8. Repeat sequence chromosome specific nucleic acid probes and methods of preparing and using

    DOEpatents

    Weier, H.U.G.; Gray, J.W.

    1995-06-27

    A primer directed DNA amplification method to isolate efficiently chromosome-specific repeated DNA wherein degenerate oligonucleotide primers are used is disclosed. The probes produced are a heterogeneous mixture that can be used with blocking DNA as a chromosome-specific staining reagent, and/or the elements of the mixture can be screened for high specificity, size and/or high degree of repetition among other parameters. The degenerate primers are sets of primers that vary in sequence but are substantially complementary to highly repeated nucleic acid sequences, preferably clustered within the template DNA, for example, pericentromeric alpha satellite repeat sequences. The template DNA is preferably chromosome-specific. Exemplary primers and probes are disclosed. The probes of this invention can be used to determine the number of chromosomes of a specific type in metaphase spreads, in germ line and/or somatic cell interphase nuclei, micronuclei and/or in tissue sections. Also provided is a method to select arbitrarily repeat sequence probes that can be screened for chromosome-specificity. 18 figs.

  9. Repeat sequence chromosome specific nucleic acid probes and methods of preparing and using

    DOEpatents

    Weier, Heinz-Ulrich G.; Gray, Joe W.

    1995-01-01

    A primer directed DNA amplification method to isolate efficiently chromosome-specific repeated DNA wherein degenerate oligonucleotide primers are used is disclosed. The probes produced are a heterogeneous mixture that can be used with blocking DNA as a chromosome-specific staining reagent, and/or the elements of the mixture can be screened for high specificity, size and/or high degree of repetition among other parameters. The degenerate primers are sets of primers that vary in sequence but are substantially complementary to highly repeated nucleic acid sequences, preferably clustered within the template DNA, for example, pericentromeric alpha satellite repeat sequences. The template DNA is preferably chromosome-specific. Exemplary primers ard probes are disclosed. The probes of this invention can be used to determine the number of chromosomes of a specific type in metaphase spreads, in germ line and/or somatic cell interphase nuclei, micronuclei and/or in tissue sections. Also provided is a method to select arbitrarily repeat sequence probes that can be screened for chromosome-specificity.

  10. Identification of highly conserved residues involved in inhibition of HIV-1 RNase H function by Diketo acid derivatives.

    PubMed

    Corona, Angela; Di Leva, Francesco Saverio; Thierry, Sylvain; Pescatori, Luca; Cuzzucoli Crucitti, Giuliana; Subra, Frederic; Delelis, Olivier; Esposito, Francesca; Rigogliuso, Giuseppe; Costi, Roberta; Cosconati, Sandro; Novellino, Ettore; Di Santo, Roberto; Tramontano, Enzo

    2014-10-01

    HIV-1 reverse transcriptase (RT)-associated RNase H activity is an essential function in viral genome retrotranscription. RNase H is a promising drug target for which no inhibitor is available for therapy. Diketo acid (DKA) derivatives are active site Mg(2+)-binding inhibitors of both HIV-1 RNase H and integrase (IN) activities. To investigate the DKA binding site of RNase H and the mechanism of action, six couples of ester and acid DKAs, derived from 6-[1-(4-fluorophenyl)methyl-1H-pyrrol-2-yl)]-2,4-dioxo-5-hexenoic acid ethyl ester (RDS1643), were synthesized and tested on both RNase H and IN functions. Most of the ester derivatives showed selectivity for HIV-1 RNase H versus IN, while acids inhibited both functions. Molecular modeling and site-directed mutagenesis studies on the RNase H domain demonstrated different binding poses for ester and acid DKAs and proved that DKAs interact with residues (R448, N474, Q475, Y501, and R557) involved not in the catalytic motif but in highly conserved portions of the RNase H primer grip motif. The ester derivative RDS1759 selectively inhibited RNase H activity and viral replication in the low micromolar range, making contacts with residues Q475, N474, and Y501. Quantitative PCR studies and fluorescence-activated cell sorting (FACS) analyses showed that RDS1759 selectively inhibited reverse transcription in cell-based assays. Overall, we provide the first demonstration that RNase H inhibition by DKAs is due not only to their chelating properties but also to specific interactions with highly conserved amino acid residues in the RNase H domain, leading to effective targeting of HIV retrotranscription in cells and hence offering important insights for the rational design of RNase H inhibitors. PMID:25092689

  11. Detection of Nucleic Acids with Graphene Nanopores: Ab Initio Characterization of a Novel Sequencing Device

    NASA Astrophysics Data System (ADS)

    Nelson, Tammie; Zhang, Bo; Prezhdo, Oleg

    2010-03-01

    We report an ab initio study of the interaction of two nucleobases, cytosine and adenine, with a novel graphene nanopore device for detecting the base sequence of a single-stranded nucleic acid (ssDNA or RNA). The nucleobases were inserted into a pore in a graphene nanoribbon, and the electrical current and conductance spectra were calculated as functions of voltage applied across the nanoribbon. The conductance spectra and charge densities were analyzed in the presence of each nucleobase in the graphene nanopore. The results indicate that, due to significant differences in the conductance spectra, the proposed device has adequate sensitivity to discriminate between different nucleotides. Moreover, we show that the nucleotide conductance spectra is not affected by its orientation inside the graphene nanopore. The proposed technique may be extremely useful for real applications in developing ultrafast, low cost DNA sequencing methods.

  12. Identification of G and P genotype-specific motifs in the predicted VP7 and VP4 amino acid sequences.

    PubMed

    Ma, Yongping

    2015-12-01

    Equine rotavirus (ERV) strain L338 (G13P[18]) has a unique G and P genotype. However, the evolutionary relationship of L338 with other ERVs is still unknown. Here whole genome analysis of the L338 ERV strain was independently performed. Its genotype constellations were determined as G13-P[18]-I6-R9-C9-M6-A6-N9-T12-E14-H11, confirming previous genotype assignments. The L338 strain only shared the P[18] and I6 genotypes with other ERVs. The nucleotide sequences of the other 9 RNA segments were different from those of cogent genes of all other group A rotavirus (RVA) strains including ERVs and formed unique phylogenetic lineages. The L338 evolutionary footprints were tentatively identified in both VP7 and VP4 amino acid sequences: two regions were found in VP7 and twelve in VP4. The conserved regions shared between L338 and other group A rotavirus strains (RVAs) indicated that L338 was more closely related genomically to animal and human RVAs other than ERVs, suggesting that L338 may not be an endogenous equine RV but have emerged as an interspecies reassortant with other RVA strains. Furthermore, genotype-specific motifs of all 27 G and 37 P types were identified in regions 7-1a (aa 91-100) of VP7 and regions 8-1 (aa146-151) and 8-3 (aa113-118 and 125-135) of VP4 (VP8*). PMID:26321159

  13. The sexually dimorphic on the Y-chromosome gene (sdY) is a conserved male-specific Y-chromosome sequence in many salmonids

    PubMed Central

    Yano, Ayaka; Nicol, Barbara; Jouanno, Elodie; Quillet, Edwige; Fostier, Alexis; Guyomard, René; Guiguen, Yann

    2013-01-01

    All salmonid species investigated to date have been characterized with a male heterogametic sex-determination system. However, as these species do not share any Y-chromosome conserved synteny, there remains a debate on whether they share a common master sex-determining gene. In this study, we investigated the extent of conservation and evolution of the rainbow trout (Oncorhynchus mykiss) master sex-determining gene, sdY (sexually dimorphic on the Y-chromosome), in 15 different species of salmonids. We found that the sdY sequence is highly conserved in all salmonids and that sdY is a male-specific Y-chromosome gene in the majority of these species. These findings demonstrate that most salmonids share a conserved sex-determining locus and also strongly suggest that sdY may be this conserved master sex-determining gene. However, in two whitefish species (subfamily Coregoninae), sdY was found both in males and females, suggesting that alternative sex-determination systems may have also evolved in this family. Based on the wide conservation of sdY as a male-specific Y-chromosome gene, efficient and easy molecular sexing techniques can now be developed that will be of great interest for studying these economically and environmentally important species. PMID:23745140

  14. Isolation of a new marker and conserved sequences close to the DiGeorge syndrome marker HP500 (D22S134).

    PubMed Central

    Wadey, R; Daw, S; Wickremasinghe, A; Roberts, C; Wilson, D; Goodship, J; Burn, J; Halford, S; Scambler, P J

    1993-01-01

    End fragment cloning from a YAC at the D22S134 locus allowed the isolation of a new probe HD7k. This marker detects hemizygosity in two patients previously shown to be dizygous for D22S134. This positions the distal deletion breakpoint in these patients to the sequences within the YAC, and confirms that HD7k is proximal to D22S134. In a search for coding sequences within the region commonly deleted in DGS we have identified a conserved sequence at D22S134. Although no cDNAs have yet been isolated, genomic sequencing shows a short open reading frame with weak similarity to collagen proteins. Images PMID:8230156

  15. Fast computational methods for predicting protein structure from primary amino acid sequence

    DOEpatents

    Agarwal, Pratul Kumar

    2011-07-19

    The present invention provides a method utilizing primary amino acid sequence of a protein, energy minimization, molecular dynamics and protein vibrational modes to predict three-dimensional structure of a protein. The present invention also determines possible intermediates in the protein folding pathway. The present invention has important applications to the design of novel drugs as well as protein engineering. The present invention predicts the three-dimensional structure of a protein independent of size of the protein, overcoming a significant limitation in the prior art.

  16. Amino-terminal amino acid sequence of the major structural polypeptides of avian retroviruses: sequence homology between reticuloendotheliosis virus p30 and p30s of mammalian retroviruses.

    PubMed Central

    Hunter, E; Bhown, A S; Bennett, J C

    1978-01-01

    The major structural polypeptides, p30 of reticuloendotheliosis virus (REV) (strain T) and p27 of avian sarcoma virus B77, have been compared with regard to amino acid composition. NH2-terminal amino acid sequence, and immunological crossreactions. The amino acid composition of the two polypeptides is distinct, and a comparison of the first 30 NH2-terminal amino acids of REV p30 with that for the first 25 of B77 p27 yields only three homologous residues. In competition radioimmunoassays the polypeptides show no crossreactivity. A comparison of the amino acid composition and NH2-terminal amino acid sequence of REV p30 with those reported for several mammalian retrovirus p30s shows remarkable similarities. Both REV and mammalian p30s contain a large number of polar residues in their amino acid composition and show approximately 40% homology in the first 30 NH2-terminal amino acids. No crossreactivity could be observed, however, in competition radioimmunoassays between Rauscher murine leukemia virus p30 and that of REV. The observations reported here suggest a close evolutionary relationship between REV and the mammalian retroviruses. Images PMID:208072

  17. Purification and amino acid sequence of aminopeptidase P from pig kidney.

    PubMed

    Vergas Romero, C; Neudorfer, I; Mann, K; Schäfer, W

    1995-04-01

    Aminopeptidase P from kidney cortex was purified in high yield (recovery greater than or equal to 20%) by a series of column chromatographic steps after solubilization of the membrane-bound glycoprotein with n-butanol. A coupled enzymic assay, using Gly-Pro-Pro-NH-Nap as substrate and dipeptidyl-peptidase IV as auxilliary enzyme, was used to monitor the purification. The purification procedure yielded two forms of aminopeptidase P differing in their carbohydrate composition (glycoforms). Both enzyme preparations were homogeneous as assessed by SDS/PAGE silver staining, and isoelectric focusing. Both forms possessed the same substrate specificity, catalysed the same reaction, and consisted of identical protein chains. The amino acid sequence determined by Edman degradation and mass spectrometry consisted of 623 amino acids. Six N-glycosylation sites, all contained in the N-terminal half of the protein, were characterized. PMID:7744038

  18. Draft Genome Sequence of Cupriavidus sp. Strain SK-3, a 4-Chlorobiphenyl- and 4-Clorobenzoic Acid-Degrading Bacterium

    PubMed Central

    Vilo, Claudia; Benedik, Michael J.; Ilori, Matthew

    2014-01-01

    We report the draft genome sequence of Cupriavidus sp. strain SK-3, which can use 4-chlorobiphenyl and 4-clorobenzoic acid as the sole carbon source for growth. The draft genome sequence allowed the study of the polychlorinated biphenyl degradation mechanism and the recharacterization of the strain SK-3 as a Cupriavidus species. PMID:24994805

  19. Draft Genome Sequence of Bacillus subtilis subsp. natto Strain CGMCC 2108, a High Producer of Poly-γ-Glutamic Acid

    PubMed Central

    Tan, Siyuan; Su, Anping; Zhang, Chen; Ren, Yuanyuan

    2016-01-01

    Here, we report the 4.1-Mb draft genome sequence of Bacillus subtilis subsp. natto strain CGMCC 2108, a high producer of poly-γ-glutamic acid (γ-PGA). This sequence will provide further help for the biosynthesis of γ-PGA and will greatly facilitate research efforts in metabolic engineering of B. subtilis subsp. natto strain CGMCC 2108. PMID:27231363

  20. New monoclonal antibodies to the Ebola virus glycoprotein: Identification and analysis of the amino acid sequence of the variable domains.

    PubMed

    Panina, A A; Aliev, T K; Shemchukova, O B; Dement'yeva, I G; Varlamov, N E; Pozdnyakova, L P; Bokov, M N; Dolgikh, D A; Sveshnikov, P G; Kirpichnikov, M P

    2016-03-01

    We determined the nucleotide and amino acid sequences of variable domains of three new monoclonal antibodies to the glycoprotein of Ebola virus capsid. The framework and hypervariable regions of immunoglobulin heavy and light chains were identified. The primary structures were confirmed using massspectrometry analysis. Immunoglobulin database search showed the uniqueness of the sequences obtained. PMID:27193713

  1. Genome Sequence of the Lactic Acid Bacterium Lactococcus lactis subsp. lactis TOMSC161, Isolated from a Nonscalded Curd Pressed Cheese

    PubMed Central

    Velly, H.; Abraham, A.-L.; Loux, V.; Delacroix-Buchet, A.; Fonseca, F.; Bouix, M.

    2014-01-01

    Lactococcus lactis is a lactic acid bacterium used in the production of many fermented foods, such as dairy products. Here, we report the genome sequence of L. lactis subsp. lactis TOMSC161, isolated from nonscalded curd pressed cheese. This genome sequence provides information in relation to dairy environment adaptation. PMID:25377704

  2. Draft Genome Sequence of Bacillus subtilis subsp. natto Strain CGMCC 2108, a High Producer of Poly-γ-Glutamic Acid.

    PubMed

    Tan, Siyuan; Meng, Yonghong; Su, Anping; Zhang, Chen; Ren, Yuanyuan

    2016-01-01

    Here, we report the 4.1-Mb draft genome sequence of Bacillus subtilis subsp. natto strain CGMCC 2108, a high producer of poly-γ-glutamic acid (γ-PGA). This sequence will provide further help for the biosynthesis of γ-PGA and will greatly facilitate research efforts in metabolic engineering of B. subtilis subsp. natto strain CGMCC 2108. PMID:27231363

  3. Functional characterization of the conserved amino acids in Pop1p, the largest common protein subunit of yeast RNases P and MRP

    PubMed Central

    Xiao, Shaohua; Hsieh, John; Nugent, Rebecca L.; Coughlin, Daniel J.; Fierke, Carol A.; Engelke, David R.

    2006-01-01

    RNase P and RNase MRP are ribonucleoprotein enzymes required for 5′-end maturation of precursor tRNAs (pre-tRNAs) and processing of precursor ribosomal RNAs, respectively. In yeast, RNase P and MRP holoenzymes have eight protein subunits in common, with Pop1p being the largest at >100 kDa. Little is known about the functions of Pop1p, beyond the fact that it binds specifically to the RNase P RNA subunit, RPR1 RNA. In this study, we refined the previous Pop1 phylogenetic sequence alignment and found four conserved regions. Highly conserved amino acids in yeast Pop1p were mutagenized by randomization and conditionally defective mutations were obtained. Effects of the Pop1p mutations on pre-tRNA processing, pre-rRNA processing, and stability of the RNA subunits of RNase P and MRP were examined. In most cases, functional defects in RNase P and RNase MRP in vivo were consistent with assembly defects of the holoenzymes, although moderate kinetic defects in RNase P were also observed. Most mutations affected both pre-tRNA and pre-rRNA processing, but a few mutations preferentially interfered with only RNase P or only RNase MRP. In addition, one temperature-sensitive mutation had no effect on either tRNA or rRNA processing, consistent with an additional role for RNase P, RNase MRP, or Pop1p in some other form. This study shows that the Pop1p subunit plays multiple roles in the assembly and function of of RNases P and MRP, and that the functions can be differentiated through the mutations in conserved residues. PMID:16618965

  4. Sequence Evaluation of FGF and FGFR Gene Conserved Non-Coding Elements in Non-Syndromic Cleft Lip and Palate Cases

    PubMed Central

    Riley, Bridget M.; Murray, Jeffrey C.

    2009-01-01

    Non-syndromic cleft lip and palate (NS CLP) is a complex birth defect resulting from multiple genetic and environmental factors. We have previously reported the sequencing of the coding region of genes in the fibroblast growth factor (FGF) signaling pathway, in which missense and non-sense mutations contribute to approximately 5%–6% NS CLP cases. In this article we report the sequencing of conserved non-coding elements (CNEs) in and around 11 of the FGF and FGFR genes, which identified 55 novel variants. Seven of variants are highly conserved among ≥8 species and 31 variants alter transcription factor binding sites, 8 of which are important for craniofacial development. Additionally, 15 NS CLP patients had a combination of coding mutations and CNE variants, suggesting that an accumulation of variants in the FGF signaling pathway may contribute to clefting. PMID:17963255

  5. Formation Sequences of Iron Minerals in the Acidic Alteration Products and Variation of Hydrothermal Fluid Conditions

    NASA Astrophysics Data System (ADS)

    Isobe, H.; Yoshizawa, M.

    2008-12-01

    Iron minerals have important role in environmental issues not only on the Earth but also other terrestrial planets. Iron mineral species related to alteration products of primary minerals with surface or subsurface fluids are characterized by temperature, acidity and redox conditions of the fluids. We can see various iron- bearing alteration products in alteration products around fumaroles in geothermal/volcanic areas. In this study, zonal structures of iron minerals in alteration products of the geothermal area are observed to elucidate temporal and spatial variation of hydrothermal fluids. Alteration of the pyroxene-amphibole andesite of Garan-dake volcano, Oita, Japan occurs by the acidic hydrothermal fluid to form cristobalite leaching out elements other than Si. Hand specimens with unaltered or weakly altered core and cristobalite crust show various sequences of layers. XRD analysis revealed that the alteration degree is represented by abundance of cristobalite. Intermediately altered layers are characterized by occurrence including alunite, pyrite, kaolinite, goethite and hematite. A specimen with reddish brown core surrounded by cristobalite-rich white crust has brown colored layers at the boundary of core and the crust. Reddish core is characterized by occurrence of crystalline hematite by XRD. Another hand specimen has light gray core, which represents reduced conditions, and white cristobalite crust with light brown and reddish brown layers of ferric iron minerals between the core and the crust. On the other hand, hornblende crystals, typical ferrous iron-bearing mineral of the host rock, are well preserved in some samples with strongly decolorized cristobalite-rich groundmass. Hydrothermal alteration experiments of iron-rich basaltic material shows iron mineral species depend on acidity and temperature of the fluid. Oxidation states of the iron-bearing mineral species are strongly influenced by the acidity and redox conditions. Variations of alteration

  6. Amino acid sequence and carbohydrate-binding analysis of the N-acetyl-D-galactosamine-specific C-type lectin, CEL-I, from the Holothuroidea, Cucumaria echinata.

    PubMed

    Hatakeyama, Tomomitsu; Matsuo, Noriaki; Shiba, Kouhei; Nishinohara, Shoichi; Yamasaki, Nobuyuki; Sugawara, Hajime; Aoyagi, Haruhiko

    2002-01-01

    CEL-I is one of the Ca2+-dependent lectins that has been isolated from the sea cucumber, Cucumaria echinata. This protein is composed of two identical subunits held by a single disulfide bond. The complete amino acid sequence of CEL-I was determined by sequencing the peptides produced by proteolytic fragmentation of S-pyridylethylated CEL-I. A subunit of CEL-I is composed of 140 amino acid residues. Two intrachain (Cys3-Cys14 and Cys31-Cys135) and one interchain (Cys36) disulfide bonds were also identified from an analysis of the cystine-containing peptides obtained from the intact protein. The similarity between the sequence of CEL-I and that of other C-type lectins was low, while the C-terminal region, including the putative Ca2+ and carbohydrate-binding sites, was relatively well conserved. When the carbohydrate-binding activity was examined by a solid-phase microplate assay, CEL-I showed much higher affinity for N-acetyl-D-galactosamine than for other galactose-related carbohydrates. The association constant of CEL-I for p-nitrophenyl N-acetyl-beta-D-galactosaminide (NP-GalNAc) was determined to be 2.3 x 10(4) M(-1), and the maximum number of bound NP-GalNAc was estimated to be 1.6 by an equilibrium dialysis experiment. PMID:11866098

  7. The complete mitochondrial genome sequence of the hornwort Phaeoceros laevis: retention of many ancient pseudogenes and conservative evolution of mitochondrial genomes in hornworts.

    PubMed

    Xue, Jia-Yu; Liu, Yang; Li, Libo; Wang, Bin; Qiu, Yin-Long

    2010-02-01

    Plants have large and complex mitochondrial genomes in comparison to other eukaryotes. In bryophytes, the mitochondrial genomes exhibit a mixed mode of conservative and dynamic evolution. Here, we sequenced the complete mitochondrial genome from hornwort Phaeoceros laevis, to investigate the level of conservation in mitochondrial genome evolution within hornworts. The circular molecule consists of 209,482 base pairs and represents the largest known mitochondrial genome of bryophytes. It contains 30 protein genes, 3 rRNA genes, and 21 tRNA genes, with 34 cis-spliced group II introns disrupting 16 protein genes. There are 11 pseudogenes in this genome, and nine of them are shared with the other fully sequenced hornwort chondriome from Megaceros aenigmaticus, a distant relative of P. laevis. These pseudogenes were likely formed during an early stage of hornwort evolution. The two hornwort chondriomes differ by four inversions and translocations, seven genes, and four introns in the genome structure and organization. At the sequence level, they are very similar, with the identity values ranging mostly from 80 to 95% in intergenic spacers, introns, and exons. These data indicate that mitochondrial genome evolution in hornworts is less conservative than in liverworts, but has not reached the dynamic level as seen in seed plants. PMID:19998039

  8. Cytoplasmic protein binding to highly conserved sequences in the 3' untranslated region of mouse protamine 2 mRNA, a translationally regulated transcript of male germ cells.

    PubMed

    Kwon, Y K; Hecht, N B

    1991-05-01

    The expression of the protamines, the predominant nuclear proteins of mammalian spermatozoa, is regulated translationally during male germ-cell development. The 3' untranslated region (UTR) of protamine 1 mRNA has been reported to control its time of translation. To understand the mechanisms controlling translation of the protamine mRNAs, we have sought to identify cis elements of the 3' UTR of protamine 2 mRNA that are recognized by cytoplasmic factors. From gel retardation assays, two sequence elements are shown to form specific RNA-protein complexes. Protein binding sites of the two complexes were determined by RNase T1 mapping, by blocking the putative binding sites with antisense oligonucleotides, and by competition assays. The sequences of these elements, located between nucleotides + 537 and + 572 in protamine 2 mRNA, are highly conserved among postmeiotic translationally regulated nuclear proteins of the mammalian testis. Two closely linked protein binding sites were detected. UV-crosslinking studies revealed that a protein of about 18 kDa binds to one of the conserved sequences. These data demonstrate specific protein binding to a highly conserved 3' UTR of translationally regulated testicular mRNA. PMID:2023906

  9. Clofibrate-induced cytochrome P450-lauric acid omega hydroxylase(P450LA omega):purification, cDNA cloning, sequence and regulation

    SciTech Connect

    Hardwick, J.P.; Song, B.J.; Gonzalez, F.J.

    1986-05-01

    A cytochrome P450 that hydroxylates lauric acid at the 12 position (P450LA omega) was isolated from liver microsomes of clofibrate treated rats. P450LA omega was immunologically distinct from P450s a,b,c,d,e,f,g,h,j,PB1, and PCN1. Polyclonal antibody against P450LA omega was utilized to screen a gt11 cDNA library. A clone (pP450LA omega), was isolated and its sequence determined. The P450LA omega mRNA is a minimum 2387 nts in length and codes for a P450 of Mr.58,222 daltons. This protein shares less than 35% amino acid similarity with P450s b,c,d,e,f,PB1, and PCN1; however, it does contain a hydrophobic amino terminal peptide and a conserved sequence surrounding the Cys residue at position 456, which is similar to other microsomal P450s. P450LA omega is present at high levels in untreated rat kidney and is induced by clofibrate in both kidney and liver. This induction is the result of an accumulation of mRNA through a rapid transcriptional activation of the P450LA gene. Southern blotting data suggest the presence of 2 or 3 genes in the P450LA omega family. This P450 gene family may be associated with arachidonic acid and prostraglandin metabolism in kidney and other tissues.

  10. Draft Genome Sequences of Gluconobacter cerinus CECT 9110 and Gluconobacter japonicus CECT 8443, Acetic Acid Bacteria Isolated from Grape Must

    PubMed Central

    Sainz, Florencia

    2016-01-01

    We report here the draft genome sequences of Gluconobacter cerinus strain CECT9110 and Gluconobacter japonicus CECT8443, acetic acid bacteria isolated from grape must. Gluconobacter species are well known for their ability to oxidize sugar alcohols into the corresponding acids. Our objective was to select strains to oxidize effectively d-glucose. PMID:27365351

  11. Members of a unique histidine acid phosphatase family are conserved amongst a group of primitive eukaryotic human pathogens.

    PubMed

    Shakarian, Alison M; Joshi, Manju B; Yamage, Mat; Ellis, Stephanie L; Debrabant, Alain; Dwyer, Dennis M

    2003-03-01

    Recently, we identified and characterized the genes encoding several distinct members of the histidine-acid phosphatase enzyme family from Leishmania donovani, a primitive protozoan pathogen of humans. These included genes encoding the heavily phosphorylated/glycosylated, tartrate-sensitive, secretory acid phosphatases (Ld SAcP-1 and Ld SAcP-2) and the unique, tartrate-resistant, externally-oriented, surface membrane-bound acid phosphatase (Ld MAcP) of this parasite. It had been previously suggested that these enzymes may play essential roles in the growth, development and survival of this organism. In this report, to further examine this hypothesis, we assessed whether members of the L. donovani histidine-acid phosphatase enzyme family were conserved amongst other pathogenic Leishmania and related trypanosomatid parasites. Such phylogenetic conservation would clearly indicate an evolutionary selection for this family of enzymes and strongly suggest and support an important functional role for acid phosphatases to the survival of these parasites. Results of pulsed field gel electrophoresis and Southern blotting showed that homologs of both the Ld SAcPs and Ld MAcP were present in each of the visceral and cutaneous Leishmania species examined (i.e. isolates of L. donovani, L. infantum, L. tropica, L. major and L. mexicana, respectively). Further, results of enzyme assays showed that all of these organisms expressed both tartrate-sensitive and tartrate-resistant acid phosphatase activities. In addition, homologs of both the Ld SAcPs and Ld MAcP genes and their corresponding enzyme activities were also identified in two Crithidia species (C. fasciculata and C. luciliae) and in Leptomonas seymouri. In contrast, Trypanosoma brucei, Trypanosoma cruzi and Phytomonas serpens had only very-low levels of such enzyme activities. Cumulatively, results of this study showed that homologs of the Ld SAcPs and Ld MAcP are conserved amongst all pathogenic Leishmania sps. suggesting

  12. Swfoldrate: predicting protein folding rates from amino acid sequence with sliding window method.

    PubMed

    Cheng, Xiang; Xiao, Xuan; Wu, Zhi-cheng; Wang, Pu; Lin, Wei-zhong

    2013-01-01

    Protein folding is the process by which a protein processes from its denatured state to its specific biologically active conformation. Understanding the relationship between sequences and the folding rates of proteins remains an important challenge. Most previous methods of predicting protein folding rate require the tertiary structure of a protein as an input. In this study, the long-range and short-range contact in protein were used to derive extended version of the pseudo amino acid composition based on sliding window method. This method is capable of predicting the protein folding rates just from the amino acid sequence without the aid of any structural class information. We systematically studied the contributions of individual features to folding rate prediction. The optimal feature selection procedures are adopted by means of combining the forward feature selection and sequential backward selection method. Using the jackknife cross validation test, the method was demonstrated on the large dataset. The predictor was achieved on the basis of multitudinous physicochemical features and statistical features from protein using nonlinear support vector machine (SVM) regression model, the method obtained an excellent agreement between predicted and experimentally observed folding rates of proteins. The correlation coefficient is 0.9313 and the standard error is 2.2692. The prediction server is freely available at http://www.jci-bioinfo.cn/swfrate/input.jsp. PMID:22933332

  13. From amino acid sequence to bioactivity: The biomedical potential of antitumor peptides.

    PubMed

    Blanco-Míguez, Aitor; Gutiérrez-Jácome, Alberto; Pérez-Pérez, Martín; Pérez-Rodríguez, Gael; Catalán-García, Sandra; Fdez-Riverola, Florentino; Lourenço, Anália; Sánchez, Borja

    2016-06-01

    Chemoprevention is the use of natural and/or synthetic substances to block, reverse, or retard the process of carcinogenesis. In this field, the use of antitumor peptides is of interest as, (i) these molecules are small in size, (ii) they show good cell diffusion and permeability, (iii) they affect one or more specific molecular pathways involved in carcinogenesis, and (iv) they are not usually genotoxic. We have checked the Web of Science Database (23/11/2015) in order to collect papers reporting on bioactive peptide (1691 registers), which was further filtered searching terms such as "antiproliferative," "antitumoral," or "apoptosis" among others. Works reporting the amino acid sequence of an antiproliferative peptide were kept (60 registers), and this was complemented with the peptides included in CancerPPD, an extensive resource for antiproliferative peptides and proteins. Peptides were grouped according to one of the following mechanism of action: inhibition of cell migration, inhibition of tumor angiogenesis, antioxidative mechanisms, inhibition of gene transcription/cell proliferation, induction of apoptosis, disorganization of tubulin structure, cytotoxicity, or unknown mechanisms. The main mechanisms of action of those antiproliferative peptides with known amino acid sequences are presented and finally, their potential clinical usefulness and future challenges on their application is discussed. PMID:27010507

  14. The amino acid sequences and activities of synergistic hemolysins from Staphylococcus cohnii.

    PubMed

    Mak, Pawel; Maszewska, Agnieszka; Rozalska, Malgorzata

    2008-10-01

    Staphylococcus cohnii ssp. cohnii and S. cohnii ssp. urealyticus are a coagulase-negative staphylococci considered for a long time as unable to cause infections. This situation changed recently and pathogenic strains of these bacteria were isolated from hospital environments, patients and medical staff. Most of the isolated strains were resistant to many antibiotics. The present work describes isolation and characterization of several synergistic peptide hemolysins produced by these bacteria and acting as virulence factors responsible for hemolytic and cytotoxic activities. Amino acid sequences of respective hemolysins from S. cohnii ssp. cohnii (named as H1C, H2C and H3C) and S. cohnii ssp. urealyticus (H1U, H2U and H3U) were identical. Peptides H1 and H3 possessed significant amino acid homology to three synergistic hemolysins secreted by Staphylococcus lugdunensis and to putative antibacterial peptide produced by Staphylococcus saprophyticus ssp. saprophyticus. On the other hand, hemolysin H2 had a unique sequence. All isolated peptides lysed red cells from different mammalian species and exerted a cytotoxic effect on human fibroblasts. PMID:18752624

  15. Complete amino acid sequence of the myoglobin from the Pacific spotted dolphin, Stenella attenuata graffmani.

    PubMed

    Jones, B N; Wang, C C; Dwulet, F E; Lehman, L D; Meuth, J L; Bogardt, R A; Gurd, F R

    1979-04-25

    The complete amino acid sequence of the major component myoglobin from the Pacific spotted dolphin, Stenella attenuata graffmani, was determined by the automated Edman degradation of several large peptides obtained by specific cleavage of the protein. The acetimidated apomyoglobin was selectively cleaved at its two methionyl residues with cyanogen bromide and at its three arginyl residues by trypsin. By subjecting four of these peptides and the apomyoglobin to automated Edman degradation, over 80% of the primary structure of the protein was obtained. The remainder of the covalent structure was determined by the sequence analysis of peptides that resulted from further digestion of the central cyanogen bromide fragment. This fragment was cleaved at its glutamyl residues with staphylococcal protease and its lysyl residues with trypsin. The action of trypsin was restricted to the lysyl residues by chemical modification of the single arginyl residue of the fragment with 1,2-cyclohexanedione. The primary structure of this myoglobin proved to be identical with that from the Atlantic bottlenosed dolphin and Pacific common dolphin but differs from the myoglobins of the killer whale and pilot whale at two positions. The above sequence identities and differences reflect the close taxonomic relationship of these five species of Cetacea. PMID:454657

  16. Isolation and amino acid sequences of squirrel monkey (Saimiri sciurea) insulin and glucagon.

    PubMed Central

    Yu, J H; Eng, J; Yalow, R S

    1990-01-01

    It was reported two decades ago that insulin was not detectable in the glucose-stimulated state in Saimiri sciurea, the New World squirrel monkey, by a radioimmunoassay system developed with guinea pig anti-pork insulin antibody and labeled pork insulin. With the same system, reasonable levels were observed in rhesus monkeys and chimpanzees. This suggested that New World monkeys, like the New World hystricomorph rodents such as the guinea pig and the coypu, might have insulins whose sequences differ markedly from those of Old World mammals. In this report we describe the purification and amino acid sequences of squirrel monkey insulin and glucagon. We demonstrate that the substitutions at B29, B27, A2, A4, and A17 of squirrel monkey insulin are identical with those previously found in another New World primate, the owl monkey (Aotus trivirgatus). The immunologic cross-reactivity of this insulin in our immunoassay system is only a few percent of that of human insulin. Squirrel monkey glucagon is identical with the usual glucagon found in Old World mammals, which predicts that the glucagons of other New World monkeys would not differ from the usual Old World mammalian glucagon. It appears that the peptides of the New World monkeys have diverged less from those of the Old World mammals than have those of the New World hystricomorph rodents. The striking improvements in peptide purification and sequencing have the potential for adding new information concerning the evolutionary divergence of species. PMID:2263627

  17. Isolation and amino acid sequences of squirrel monkey (Saimiri sciurea) insulin and glucagon

    SciTech Connect

    Yu, Jinghua ); Eng, J.; Yalow, R.S. City Univ. of New York, NY )

    1990-12-01

    It was reported two decades ago that insulin was not detectable in the glucose-stimulated state in Saimiri sciurea, the New World squirrel monkey, by a radioimmunoassay system developed with guinea pig anti-pork insulin antibody and labeled park insulin. With the same system, reasonable levels were observed in rhesus monkeys and chimpanzees. This suggested that New World monkeys, like the New World hystricomorph rodents such as the guinea pig and the coypu, might have insulins whose sequences differ markedly from those of Old World mammals. In this report the authors describe the purification and amino acid sequences of squirrel monkey insulin and glucagon. They demonstrate that the substitutions at B29, B27, A2, A4, and A17 of squirrel monkey insulin are identical with those previously found in another New World primate, the owl monkey (Aotus trivirgatus). The immunologic cross-reactivity of this insulin in their immunoassay system is only a few percent of that of human insulin. It appears that the peptides of the New World monkeys have diverged less from those of the Old World mammals than have those of the New World hystricomorph rodents. The striking improvements in peptide purification and sequencing have the potential for adding new information concerning the evolutionary divergence of species.

  18. Binding site discovery from nucleic acid sequences by discriminative learning of hidden Markov models

    PubMed Central

    Maaskola, Jonas; Rajewsky, Nikolaus

    2014-01-01

    We present a discriminative learning method for pattern discovery of binding sites in nucleic acid sequences based on hidden Markov models. Sets of positive and negative example sequences are mined for sequence motifs whose occurrence frequency varies between the sets. The method offers several objective functions, but we concentrate on mutual information of condition and motif occurrence. We perform a systematic comparison of our method and numerous published motif-finding tools. Our method achieves the highest motif discovery performance, while being faster than most published methods. We present case studies of data from various technologies, including ChIP-Seq, RIP-Chip and PAR-CLIP, of embryonic stem cell transcription factors and of RNA-binding proteins, demonstrating practicality and utility of the method. For the alternative splicing factor RBM10, our analysis finds motifs known to be splicing-relevant. The motif discovery method is implemented in the free software package Discrover. It is applicable to genome- and transcriptome-scale data, makes use of available repeat experiments and aside from binary contrasts also more complex data configurations can be utilized. PMID:25389269

  19. Nucleotide and derived amino acid sequences of the major porin of Comamonas acidovorans and comparison of porin primary structures.

    PubMed Central

    Gerbl-Rieger, S; Peters, J; Kellermann, J; Lottspeich, F; Baumeister, W

    1991-01-01

    The DNA sequence of the gene which codes for the major outer membrane porin (Omp32) of Comamonas acidovorans has been determined. The structural gene encodes a precursor consisting of 351 amino acid residues with a signal peptide of 19 amino acid residues. Comparisons with amino acid sequences of outer membrane proteins and porins from several other members of the class Proteobacteria and of the Chlamydia trachomatis porin and the Neurospora crassa mitochondrial porin revealed a motif of eight regions of local homology. The results of this analysis are discussed with regard to common structural features of porins. PMID:1848840

  20. Evolution of conserved non-coding sequences within the vertebrate Hox clusters through the two-round whole genome duplications revealed by phylogenetic footprinting analysis.

    PubMed

    Matsunami, Masatoshi; Sumiyama, Kenta; Saitou, Naruya

    2010-12-01

    As a result of two-round whole genome duplications, four or more paralogous Hox clusters exist in vertebrate genomes. The paralogous genes in the Hox clusters show similar expression patterns, implying shared regulatory mechanisms for expression of these genes. Previous studies partly revealed the expression mechanisms of Hox genes. However, cis-regulatory elements that control these paralogous gene expression are still poorly understood. Toward solving this problem, the authors searched conserved non-coding sequences (CNSs), which are candidates of cis-regulatory elements. When comparing orthologous Hox clusters of 19 vertebrate species, 208 intergenic conserved regions were found. The authors then searched for CNSs that were conserved not only between orthologous clusters but also among the four paralogous Hox clusters. The authors found three regions that are conserved among all the four clusters and eight regions that are conserved between intergenic regions of two paralogous Hox clusters. In total, 28 CNSs were identified in the paralogous Hox clusters, and nine of them were newly found in this study. One of these novel regions bears a RARE motif. These CNSs are candidates for gene expression regulatory regions among paralogous Hox clusters. The authors also compared vertebrate CNSs with amphioxus CNSs within the Hox cluster, and found that two CNSs in the HoxA and HoxB clusters retain homology with amphioxus CNSs through the two-round whole genome duplications. PMID:20981416

  1. Phylogenetic comparison of the pre-mRNA adenosine deaminase ADAR2 genes and transcripts: conservation and diversity in editing site sequence and alternative splicing patterns.

    PubMed

    Slavov, D; Gardiner, K

    2002-10-16

    Adenosine deaminase that acts on RNA -2 (ADAR2) is a member of a family of vertebrate genes that encode adenosine (A)-to-inosine (I) RNA deaminases, enzymes that deaminate specific A residues in specific pre-mRNAs to produce I. Known substrates of ADAR2 include sites within the coding regions of pre-mRNAs of the ionotropic glutamate receptors, GluR2-6, and the serotonin receptor, 5HT2C. Mammalian ADAR2 expression is itself regulated by A-to-I editing and by several alternative splicing events. Because the biological consequences of ADAR2 function are significant, we have undertaken a phylogenetic comparison of these features. Here we report a comparison of cDNA sequences, genomic organization, editing site sequences and patterns of alternative splicing of ADAR2 genes from human, mouse, chicken, pufferfish and zebrafish. Coding sequences and intron/exon organization are highly conserved. All ADAR2 genes show evidence of transcript editing with required sequences and predicted secondary structures very highly conserved. Patterns and levels of editing and alternative splicing vary among organisms, and include novel N-terminal exons and splicing events. PMID:12459255

  2. The glove-like structure of the conserved membrane protein TatC provides insight into signal sequence recognition in twin-arginine translocation.

    PubMed

    Ramasamy, Sureshkumar; Abrol, Ravinder; Suloway, Christian J M; Clemons, William M

    2013-05-01

    In bacteria, two signal-sequence-dependent secretion pathways translocate proteins across the cytoplasmic membrane. Although the mechanism of the ubiquitous general secretory pathway is becoming well understood, that of the twin-arginine translocation pathway, responsible for translocation of folded proteins across the bilayer, is more mysterious. TatC, the largest and most conserved of three integral membrane components, provides the initial binding site of the signal sequence prior to pore assembly. Here, we present two crystal structures of TatC from the thermophilic bacteria Aquifex aeolicus at 4.0 Å and 6.8 Å resolution. The membrane architecture of TatC includes a glove-shaped structure with a lipid-exposed pocket predicted by molecular dynamics to distort the membrane. Correlating the biochemical literature to these results suggests that the signal sequence binds in this pocket, leading to structural changes that facilitate higher order assemblies. PMID:23583035

  3. The glove-like structure of the conserved membrane protein TatC provides insight into signal sequence recognition in twin-arginine translocation

    PubMed Central

    Ramasamy, Sureshkumar; Abrol, Ravinder; Suloway, Christian J.M.; Clemons, William M.

    2013-01-01

    SUMMARY In bacteria, two signal sequence dependent secretion pathways translocate proteins across the cytoplasmic membrane. While the mechanism of the ubiquitous general secretory pathway (SEC) is becoming well understood, that of the twin-arginine translocation pathway (TAT), responsible for translocation of folded proteins across the bilayer, is more mysterious. TatC, the largest and most conserved of three integral membrane components, provides the initial binding site of the signal sequence prior to pore assembly. Here, we present two crystal structures of TatC from the thermophilic bacteria Aquifex aeolicus at 4.0Å and 6.8Å resolution. The novel membrane architecture of TatC includes a glove-shaped structure with a lipid-exposed pocket predicted by molecular dynamics to distort the membrane. Correlating the biochemical literature to these results suggests that the signal sequence binds in this pocket leading to structural changes that facilitate higher order assemblies. PMID:23583035

  4. Conservation Weighting Functions Enable Covariance Analyses to Detect Functionally Important Amino Acids

    PubMed Central

    Colwell, Lucy J.; Brenner, Michael P.; Murray, Andrew W.

    2014-01-01

    The explosive growth in the number of protein sequences gives rise to the possibility of using the natural variation in sequences of homologous proteins to find residues that control different protein phenotypes. Because in many cases different phenotypes are each controlled by a group of residues, the mutations that separate one version of a phenotype from another will be correlated. Here we incorporate biological knowledge about protein phenotypes and their variability in the sequence alignment of interest into algorithms that detect correlated mutations, improving their ability to detect the residues that control those phenotypes. We demonstrate the power of this approach using simulations and recent experimental data. Applying these principles to the protein families encoded by Dscam and Protocadherin allows us to make testable predictions about the residues that dictate the specificity of molecular interactions. PMID:25379728

  5. Conservation of sequence in the internal transcribed spacers and 5.8S ribosomal RNA among geographically separated isolates of parasitic scuticociliates (Ciliophora, Orchitophryidae).

    PubMed

    Goggin, C L; Murphy, N E

    2000-02-24

    Nucleotide sequence from the internal transcribed spacers (ITS1 and ITS2) and the 5.8S gene from the ribosomal RNA gene cluster of isolates of the scuticociliate Orchitophrya stellarum from 4 asteroid hosts were compared. Surprisingly, these data (495 bp) were identical for O. stellarum isolated from the testes of Asterias amurensis from Japan; Pisaster ochraceus from British Columbia, Canada; Asterias rubens from The Netherlands; and Asterias vulgaris from Prince Edward Island, Canada. These sequence data were compared to those from 3 scuticociliates which parasitise crustaceans: Mesanophrys pugettensis, M. chesapeakensis and Anophryoides haemophila. No difference was found in this region between the nucleotide sequence of M. pugettensis and M. chesapeakensis. The sequence of Mesanophrys spp. differed by 9.2% in the ITS1 and 4.7% in the ITS2 from that of O. stellarum. The sequence from the ITS1 (135 bp) and ITS2 (233 bp) of A. haemophila differed by 42.6 and 20.5% respectively from those of O. stellarum. Therefore, nucleotide sequence of the ITS regions in these scuticociliates is highly conserved. PMID:10785865

  6. Comparative genomic analysis of a neurotoxigenic Clostridium species using partial genome sequence: Phylogenetic analysis of a few conserved proteins involved in cellular processes and metabolism.

    PubMed

    Alam, Syed Imteyaz; Dixit, Aparna; Tomar, Arvind; Singh, Lokendra

    2010-04-01

    Clostridial organisms produce neurotoxins, which are generally regarded as the most potent toxic substances of biological origin and potential biological warfare agents. Clostridium tetani produces tetanus neurotoxin and is responsible for the fatal tetanus disease. In spite of the extensive immunization regimen, the disease is an important cause of death especially among neonates. Strains of C. tetani have not been genetically characterized except the complete genome sequencing of strain E88. The present study reports the genetic makeup and phylogenetic affiliations of an environmental strain of this bacterium with respect to C. tetani E88 and other clostridia. A shot gun library was constructed from the genomic DNA of C. tetani drde, isolated from decaying fish sample. Unique clones were sequenced and sequences compared with its closest relative C. tetani E88. A total of 275 clones were obtained and 32,457 bases of non-redundant sequence were generated. A total of 150 base changes were observed over the entire length of sequence obtained, including, additions, deletions and base substitutions. Of the total 120 ORFs detected, 48 exhibited closest similarity to E88 proteins of which three are hypothetical proteins. Eight of the ORFs exhibited similarity with hypothetical proteins from other organisms and 10 aligned with other proteins from unrelated organisms. There is an overall conservation of protein sequences among the two strains of C. tetani and. Selected ORFs involved in cellular processes and metabolism were subjected to phylogenetic analysis. PMID:19527791

  7. Rdh10a Provides a Conserved Critical Step in the Synthesis of Retinoic Acid during Zebrafish Embryogenesis

    PubMed Central

    D’Aniello, Enrico; Ravisankar, Padmapriyadarshini; Waxman, Joshua S.

    2015-01-01

    The first step in the conversion of vitamin A into retinoic acid (RA) in embryos requires retinol dehydrogenases (RDHs). Recent studies have demonstrated that RDH10 is a critical core component of the machinery that produces RA in mouse and Xenopus embryos. If the conservation of Rdh10 function in the production of RA extends to teleost embryos has not been investigated. Here, we report that zebrafish Rdh10a deficient embryos have defects consistent with loss of RA signaling, including anteriorization of the nervous system and enlarged hearts with increased cardiomyocyte number. While knockdown of Rdh10a alone produces relatively mild RA deficient phenotypes, Rdh10a can sensitize embryos to RA deficiency and enhance phenotypes observed when Aldh1a2 function is perturbed. Moreover, excess Rdh10a enhances embryonic sensitivity to retinol, which has relatively mild teratogenic effects compared to retinal and RA treatment. Performing Rdh10a regulatory expression analysis, we also demonstrate that a conserved teleost rdh10a enhancer requires Pax2 sites to drive expression in the eyes of transgenic embryos. Altogether, our results demonstrate that Rdh10a has a conserved requirement in the first step of RA production within vertebrate embryos. PMID:26394147

  8. Plant fatty acid hydroxylases

    DOEpatents

    Somerville, Chris; Broun, Pierre; van de Loo, Frank

    2001-01-01

    This invention relates to plant fatty acyl hydroxylases. Methods to use conserved amino acid or nucleotide sequences to obtain plant fatty acyl hydroxylases are described. Also described is the use of cDNA clones encoding a plant hydroxylase to produce a family of hydroxylated fatty acids in transgenic plants. In addition, the use of genes encoding fatty acid hydroxylases or desaturases to alter the level of lipid fatty acid unsaturation in transgenic plants is described.

  9. Large-scale nucleotide sequence alignment and sequence variability assessment to identify the evolutionarily highly conserved regions for universal screening PCR assay design: an example of influenza A virus.

    PubMed

    Nagy, Alexander; Jiřinec, Tomáš; Černíková, Lenka; Jiřincová, Helena; Havlíčková, Martina

    2015-01-01

    The development of a diagnostic polymerase chain reaction (PCR) or quantitative PCR (qPCR) assay for universal detection of highly variable viral genomes is always a difficult task. The purpose of this chapter is to provide a guideline on how to align, process, and evaluate a huge set of homologous nucleotide sequences in order to reveal the evolutionarily most conserved positions suitable for universal qPCR primer and hybridization probe design. Attention is paid to the quantification and clear graphical visualization of the sequence variability at each position of the alignment. In addition, specific problems related to the processing of the extremely large sequence pool are highlighted. All of these steps are performed using an ordinary desktop computer without the need for extensive mathematical or computational skills. PMID:25697651

  10. Full Genome Virus Detection in Fecal Samples Using Sensitive Nucleic Acid Preparation, Deep Sequencing, and a Novel Iterative Sequence Classification Algorithm

    PubMed Central

    Cotten, Matthew; Oude Munnink, Bas; Canuti, Marta; Deijs, Martin; Watson, Simon J.; Kellam, Paul; van der Hoek, Lia

    2014-01-01

    We have developed a full genome virus detection process that combines sensitive nucleic acid preparation optimised for virus identification in fecal material with Illumina MiSeq sequencing and a novel post-sequencing virus identification algorithm. Enriched viral nucleic acid was converted to double-stranded DNA and subjected to Illumina MiSeq sequencing. The resulting short reads were processed with a novel iterative Python algorithm SLIM for the identification of sequences with homology to known viruses. De novo assembly was then used to generate full viral genomes. The sensitivity of this process was demonstrated with a set of fecal samples from HIV-1 infected patients. A quantitative assessment of the mammalian, plant, and bacterial virus content of this compartment was generated and the deep sequencing data were sufficient to assembly 12 complete viral genomes from 6 virus families. The method detected high levels of enteropathic viruses that are normally controlled in healthy adults, but may be involved in the pathogenesis of HIV-1 infection and will provide a powerful tool for virus detection and for analyzing changes in the fecal virome associated with HIV-1 progression and pathogenesis. PMID:24695106

  11. A highly conserved DNA replication module from Streptococcus thermophilus phages is similar in sequence and topology to a module from Lactococcus lactis phages.

    PubMed

    Desiere, F; Lucchini, S; Bruttin, A; Zwahlen, M C; Brüssow, H

    1997-08-01

    A highly conserved DNA region extending over 5 kb was observed in Streptococcus thermophilus bacteriophages. Comparative sequencing of one temperate and 26 virulent phages demonstrated in the most extreme case an 18% aa difference for a predicted protein, while the majority of the phages showed fewer, if any aa changes. The relative degree of aa conservation was not homogeneous over the DNA segment investigated. Sequence analysis of the conserved segment revealed genes possibly involved in DNA transactions. Three predicted proteins (orf 233, 443, and 382 gene product (gp)) showed nucleoside triphosphate binding motifs. Orf 443 gp showed in addition a DEAH box motif, characteristically found in a subgroup of helicases, and a variant zinc finger motif known from a phage T7 helicase/primase. Tree analysis classified orf 443 gp as a distant member of the helicase superfamily. Orf 382 gp showed similarity to putative plasmid DNA primases. Downstream of orf 382 a noncoding repeat region was identified that showed similarity to a putative minus origin from a cryptic S. thermophilus plasmid. Four predicted proteins showed not only high degrees of aa identity (34 to 63%) with proteins from Lactococcus lactis phages, but their genes showed a similar topological organization. We interpret this as evidence for a horizontal gene transfer event between phages of the two bacterial genera in the distant past. PMID:9268169

  12. Comparative sequence analysis of Solanum and Arabidopsis in a hot spot for pathogen resistance on potato chromosome V reveals a patchwork of conserved and rapidly evolving genome segments

    PubMed Central

    2007-01-01

    Background Quantitative phenotypic variation of agronomic characters in crop plants is controlled by environmental and genetic factors (quantitative trait loci = QTL). To understand the molecular basis of such QTL, the identification of the underlying genes is of primary interest and DNA sequence analysis of the genomic regions harboring QTL is a prerequisite for that. QTL mapping in potato (Solanum tuberosum) has identified a region on chromosome V tagged by DNA markers GP21 and GP179, which contains a number of important QTL, among others QTL for resistance to late blight caused by the oomycete Phytophthora infestans and to root cyst nematodes. Results To obtain genomic sequence for the targeted region on chromosome V, two local BAC (bacterial artificial chromosome) contigs were constructed and sequenced, which corresponded to parts of the homologous chromosomes of the diploid, heterozygous genotype P6/210. Two contiguous sequences of 417,445 and 202,781 base pairs were assembled and annotated. Gene-by-gene co-linearity was disrupted by non-allelic insertions of retrotransposon elements, stretches of diverged intergenic sequences, differences in gene content and gene order. The latter was caused by inversion of a 70 kbp genomic fragment. These features were also found in comparison to orthologous sequence contigs from three homeologous chromosomes of Solanum demissum, a wild tuber bearing species. Functional annotation of the sequence identified 48 putative open reading frames (ORF) in one contig and 22 in the other, with an average of one ORF every 9 kbp. Ten ORFs were classified as resistance-gene-like, 11 as F-box-containing genes, 13 as transposable elements and three as transcription factors. Comparing potato to Arabidopsis thaliana annotated proteins revealed five micro-syntenic blocks of three to seven ORFs with A. thaliana chromosomes 1, 3 and 5. Conclusion Comparative sequence analysis revealed highly conserved collinear regions that flank regions

  13. Evolutionary connections of biological kingdoms based on protein and nucleic acid sequence evidence

    NASA Technical Reports Server (NTRS)

    Dayhoff, M. O.

    1983-01-01

    Prokaryotic and eukaryotic evolutionary trees are developed from protein and nucleic-acid sequences by the methods of numerical taxonomy. Trees are presented for bacterial ferredoxins, 5S ribosomal RNA, c-type cytochromes , cytochromes c2 and c', and 5.8S ribosomal RNA; the implications for early evolution are discussed; and a composite tree showing the branching of the anaerobes, aerobes, archaebacteria, and eukaryotes is shown. Single lines are found for all oxygen-evolving photosynthetic forms and for the salt-loving and high-temperature forms of archaebacteria. It is argued that the eukaryote mitochondria, chloroplasts, and cytoplasmic host material are descended from free-living prokaryotes that formed symbiotic associations, with more than one symbiotic event involved in the evolution of each organelle.

  14. The amino acid alphabet and the architecture of the protein sequence-structure map. I. Binary alphabets.

    PubMed

    Ferrada, Evandro

    2014-12-01

    The correspondence between protein sequences and structures, or sequence-structure map, relates to fundamental aspects of structural, evolutionary and synthetic biology. The specifics of the mapping, such as the fraction of accessible sequences and structures, or the sequences' ability to fold fast, are dictated by the type of interactions between the monomers that compose the sequences. The set of possible interactions between monomers is encapsulated by the potential energy function. In this study, I explore the impact of the relative forces of the potential on the architecture of the sequence-structure map. My observations rely on simple exact models of proteins and random samples of the space of potential energy functions of binary alphabets. I adopt a graph perspective and study the distribution of viable sequences and the structures they produce, as networks of sequences connected by point mutations. I observe that the relative proportion of attractive, neutral and repulsive forces defines types of potentials, that induce sequence-structure maps of vastly different architectures. I characterize the properties underlying these differences and relate them to the structure of the potential. Among these properties are the expected number and relative distribution of sequences associated to specific structures and the diversity of structures as a function of sequence divergence. I study the types of binary potentials observed in natural amino acids and show that there is a strong bias towards only some types of potentials, a bias that seems to characterize the folding code of natural proteins. I discuss implications of these observations for the architecture of the sequence-structure map of natural proteins, the construction of random libraries of peptides, and the early evolution of the natural amino acid alphabet. PMID:25473967

  15. The Amino Acid Alphabet and the Architecture of the Protein Sequence-Structure Map. I. Binary Alphabets

    PubMed Central

    Ferrada, Evandro

    2014-01-01

    The correspondence between protein sequences and structures, or sequence-structure map, relates to fundamental aspects of structural, evolutionary and synthetic biology. The specifics of the mapping, such as the fraction of accessible sequences and structures, or the sequences' ability to fold fast, are dictated by the type of interactions between the monomers that compose the sequences. The set of possible interactions between monomers is encapsulated by the potential energy function. In this study, I explore the impact of the relative forces of the potential on the architecture of the sequence-structure map. My observations rely on simple exact models of proteins and random samples of the space of potential energy functions of binary alphabets. I adopt a graph perspective and study the distribution of viable sequences and the structures they produce, as networks of sequences connected by point mutations. I observe that the relative proportion of attractive, neutral and repulsive forces defines types of potentials, that induce sequence-structure maps of vastly different architectures. I characterize the properties underlying these differences and relate them to the structure of the potential. Among these properties are the expected number and relative distribution of sequences associated to specific structures and the diversity of structures as a function of sequence divergence. I study the types of binary potentials observed in natural amino acids and show that there is a strong bias towards only some types of potentials, a bias that seems to characterize the folding code of natural proteins. I discuss implications of these observations for the architecture of the sequence-structure map of natural proteins, the construction of random libraries of peptides, and the early evolution of the natural amino acid alphabet. PMID:25473967

  16. Trypsin inhibitors from ridged gourd (Luffa acutangula Linn.) seeds: purification, properties, and amino acid sequences.

    PubMed

    Haldar, U C; Saha, S K; Beavis, R C; Sinha, N K

    1996-02-01

    Two trypsin inhibitors, LA-1 and LA-2, have been isolated from ridged gourd (Luffa acutangula Linn.) seeds and purified to homogeneity by gel filtration followed by ion-exchange chromatography. The isoelectric point is at pH 4.55 for LA-1 and at pH 5.85 for LA-2. The Stokes radius of each inhibitor is 11.4 A. The fluorescence emission spectrum of each inhibitor is similar to that of the free tyrosine. The biomolecular rate constant of acrylamide quenching is 1.0 x 10(9) M-1 sec-1 for LA-1 and 0.8 x 10(9) M-1 sec-1 for LA-2 and that of K2HPO4 quenching is 1.6 x 10(11) M-1 sec-1 for LA-1 and 1.2 x 10(11) M-1 sec-1 for LA-2. Analysis of the circular dichroic spectra yields 40% alpha-helix and 60% beta-turn for La-1 and 45% alpha-helix and 55% beta-turn for LA-2. Inhibitors LA-1 and LA-2 consist of 28 and 29 amino acid residues, respectively. They lack threonine, alanine, valine, and tryptophan. Both inhibitors strongly inhibit trypsin by forming enzyme-inhibitor complexes at a molar ratio of unity. A chemical modification study suggests the involvement of arginine of LA-1 and lysine of LA-2 in their reactive sites. The inhibitors are very similar in their amino acid sequences, and show sequence homology with other squash family inhibitors. PMID:8924202

  17. Microfluidic platform for isolating nucleic acid targets using sequence specific hybridization

    PubMed Central

    Wang, Jingjing; Morabito, Kenneth; Tang, Jay X.; Tripathi, Anubhav

    2013-01-01

    The separation of target nucleic acid sequences from biological samples has emerged as a significant process in today's diagnostics and detection strategies. In addition to the possible clinical applications, the fundamental understanding of target and sequence specific hybridization on surface modified magnetic beads is of high value. In this paper, we describe a novel microfluidic platform that utilizes a mobile magnetic field in static microfluidic channels, where single stranded DNA (ssDNA) molecules are isolated via nucleic acid hybridization. We first established efficient isolation of biotinylated capture probe (BP) using streptavidin-coated magnetic beads. Subsequently, we investigated the hybridization of target ssDNA with BP bound to beads and explained these hybridization kinetics using a dual-species kinetic model. The number of hybridized target ssDNA molecules was determined to be about 6.5 times less than that of BP on the bead surface, due to steric hindrance effects. The hybridization of target ssDNA with non-complementary BP bound to bead was also examined, and non-specific hybridization was found to be insignificant. Finally, we demonstrated highly efficient capture and isolation of target ssDNA in the presence of non-target ssDNA, where as low as 1% target ssDNA can be detected from mixture. The microfluidic method described in this paper is significantly relevant and is broadly applicable, especially towards point-of-care biological diagnostic platforms that require binding and separation of known target biomolecules, such as RNA, ssDNA, or protein. PMID:24404041

  18. Molecular characterization of the body site-specific human epidermal cytokeratin 9: cDNA cloning, amino acid sequence, and tissue specificity of gene expression.

    PubMed

    Langbein, L; Heid, H W; Moll, I; Franke, W W

    1993-12-01

    Differentiation of human plantar and palmar epidermis is characterized by the suprabasal synthesis of a major special intermediate-sized filament (IF) protein, the type I (acidic) cytokeratin 9 (CK 9). Using partial amino acid (aa) sequence information obtained by direct Edman sequencing of peptides resulting from proteolytic digestion of purified CK 9, we synthesized several redundant primers by 'back-translation'. Amplification by polymerase chain reaction (PCR) of cDNAs obtained by reverse transcription of mRNAs from human foot sole epidermis, including 5'-primer extension, resulted in multiple overlapping cDNA clones, from which the complete cDNA (2353 bp) could be constructed. This cDNA encoded the CK 9 polypeptide with a calculated molecular weight of 61,987 and an isoelectric point at about pH 5.0. The aa sequence deduced from cDNA was verified in several parts by comparison with the peptide sequences and showed the typical structure of type I CKs, with a head (153 aa), and alpha-helical coiled-coil-forming rod (306 aa), and a tail (163 aa) domain. The protein displayed the highest homology to human CK 10, not only in the highly conserved rod domain but also in large parts of the head and the tail domains. On the other hand, the aa sequence revealed some remarkable differences from CK 10 and other CKs, even in the most conserved segments of the rod domain. The nuclease digestion pattern seen on Southern blot analysis of human genomic DNA indicated the existence of a unique CK 9 gene. Using CK 9-specific riboprobes for hybridization on Northern blots of RNAs from various epithelia, a mRNA of about 2.4 kb in length could be identified only in foot sole epidermis, and a weaker cross-hybridization signal was seen in RNA from bovine heel pad epidermis at about 2.0 kb. A large number of tissues and cell cultures were examined by PCR of mRNA-derived cDNAs, using CK 9-specific primers. But even with this very sensitive signal amplification, only palmar

  19. Low-pass shotgun sequencing of the barley genome facilitates rapid identification of genes, conserved non-coding sequences and novel repeats

    Technology Transfer Automated Retrieval System (TEKTRAN)

    Background: Barley has one of the largest and most complex genomes of all economically important food crops. The rise of new short read sequencing technologies such as Illumina/Solexa permits such large genomes to be effectively sampled at relatively low costs. An MDR (Mathematically Defined Repeat)...

  20. Characterization of N-glycosylation and amino acid sequence features of immunoglobulins from swine.

    PubMed

    Lopez, Paul G; Girard, Lauren; Buist, Marjorie; de Oliveira, Andrey Giovanni Gomes; Bodnar, Edward; Salama, Apolline; Soulillou, Jean-Paul; Perreault, Hélène

    2016-02-01

    The primary goal of this study was to develop a method to study the N-glycosylation of IgG from swine in order to detect epitopes containing N-glycolylneuraminic acid (Neu5Gc) and/or terminal galactose residues linked in α1-3 susceptible to cause xenograft-related problems. Samples of immunoglobulin were isolated from porcine serum using protein-A affinity chromatography. The eluate was then separated on electrophoretic gel, and bands corresponding to the N-glycosylated heavy chains were cut off the gel and subjected to tryptic digestion. Peptides and glycopeptides were separated by reversed phase liquid chromatography and fractions were collected for matrix-assisted laser desorption/ionization time-of-flight mass spectrometric (MALDI-TOF-MS) analysis. Overall no α1-3 galactose was detected, as demonstrated by complete susceptibility of terminal galactose residues to β-galactosidase digestion. Neu5Gc was detected on singly sialylated structures. Two major N-glycopeptides were found, EEQFNSTYR and EAQFNSTYR as determined by tandem MS (MS/MS), as previously reported by Butler et al. (Immunogenetics, 61, 2009, 209-230), who found 11 subclasses for porcine IgG. Out of the 11, ten include the sequence corresponding to EEQFNSTYR, and only one codes for EAQFNSTYR. In this study, glycosylation patterns associated with both chains were slightly different, in that EEQFNSTYR had a higher content of galactose. The last step of this study consisted of peptide-mapping the 11 reported porcine IgG sequences. Although there was considerable overlap, at least one unique tryptic peptide was found per IgG sequence. The workflow presented in this manuscript constitutes the first study to use MALDI-TOF-MS in the investigation of porcine IgG structural features. PMID:26586247

  1. Human Retroviruses and AIDS. A compilation and analysis of nucleic acid and amino acid sequences: I--II; III--V

    SciTech Connect

    Myers, G.; Korber, B.; Wain-Hobson, S.; Smith, R.F.; Pavlakis, G.N.

    1993-12-31

    This compendium and the accompanying floppy diskettes are the result of an effort to compile and rapidly publish all relevant molecular data concerning the human immunodeficiency viruses (HIV) and related retroviruses. The scope of the compendium and database is best summarized by the five parts that it comprises: (I) HIV and SIV Nucleotide Sequences; (II) Amino Acid Sequences; (III) Analyses; (IV) Related Sequences; and (V) Database Communications. Information within all the parts is updated at least twice in each year, which accounts for the modes of binding and pagination in the compendium.

  2. Chicken interferon consensus sequence-binding protein (ICSBP) and interferon regulatory factor (IRF) 1 genes reveal evolutionary conservation in the IRF gene family.

    PubMed Central

    Jungwirth, C; Rebbert, M; Ozato, K; Degen, H J; Schultz, U; Dawid, I B

    1995-01-01

    Members of the IRF family mediate transcriptional responses to interferons (IFNs) and to virus infection. So far, proteins of this family have been studied only among mammalian species. Here we report the isolation of cDNA clones encoding two members of this family from chicken, interferon consensus sequence-binding protein (ICSBP) and IRF-1. The predicted chicken ICSBP and IRF-1 proteins show high levels of sequence similarity to their corresponding human and mouse counterparts. Sequence identities in the putative DNA-binding domains of chicken and human ICSBP and IRF-1 were 97% and 89%, respectively, whereas the C-terminal regions showed identities of 64% and 51%; sequence relationships with mouse ICSBP and IRF-1 are very similar. Chicken ICSBP was found to be expressed in several embryonic tissues, and both chicken IRF-1 and ICSBP were strongly induced in chicken fibroblasts by IFN treatment, supporting the involvement of these factors in IFN-regulated gene expression. The presence of proteins homologous to mammalian IRF family members, together with earlier observations on the occurrence of functionally homologous IFN-responsive elements in chicken and mammalian genes, highlights the conservation of transcriptional mechanisms in the IFN system, a finding that contrasts with the extensive sequence and functional divergence of the IFNs. Images Fig. 3 Fig. 4 Fig. 5 PMID:7536924

  3. Whole-Exome Sequencing in a South American Cohort Links ALDH1A3, FOXN1 and Retinoic Acid Regulation Pathways to Autism Spectrum Disorders

    PubMed Central

    Moreno-Ramos, Oscar A.; Olivares, Ana María; Haider, Neena B.; de Autismo, Liga Colombiana; Lattig, María Claudia

    2015-01-01

    Autism spectrum disorders (ASDs) are a range of complex neurodevelopmental conditions principally characterized by dysfunctions linked to mental development. Previous studies have shown that there are more than 1000 genes likely involved in ASD, expressed mainly in brain and highly interconnected among them. We applied whole exome sequencing in Colombian—South American trios. Two missense novel SNVs were found in the same child: ALDH1A3 (RefSeq NM_000693: c.1514T>C (p.I505T)) and FOXN1 (RefSeq NM_003593: c.146C>T (p.S49L)). Gene expression studies reveal that Aldh1a3 and Foxn1 are expressed in ~E13.5 mouse embryonic brain, as well as in adult piriform cortex (PC; ~P30). Conserved Retinoic Acid Response Elements (RAREs) upstream of human ALDH1A3 and FOXN1 and in mouse Aldh1a3 and Foxn1 genes were revealed using bioinformatic approximation. Chromatin immunoprecipitation (ChIP) assay using Retinoid Acid Receptor B (Rarb) as the immunoprecipitation target suggests RA regulation of Aldh1a3 and Foxn1 in mice. Our results frame a possible link of RA regulation in brain to ASD etiology, and a feasible non-additive effect of two apparently unrelated variants in ALDH1A3 and FOXN1 recognizing that every result given by next generation sequencing should be cautiously analyzed, as it might be an incidental finding. PMID:26352270

  4. Whole-Exome Sequencing in a South American Cohort Links ALDH1A3, FOXN1 and Retinoic Acid Regulation Pathways to Autism Spectrum Disorders.

    PubMed

    Moreno-Ramos, Oscar A; Olivares, Ana María; Haider, Neena B; de Autismo, Liga Colombiana; Lattig, María Claudia

    2015-01-01

    Autism spectrum disorders (ASDs) are a range of complex neurodevelopmental conditions principally characterized by dysfunctions linked to mental development. Previous studies have shown that there are more than 1000 genes likely involved in ASD, expressed mainly in brain and highly interconnected among them. We applied whole exome sequencing in Colombian-South American trios. Two missense novel SNVs were found in the same child: ALDH1A3 (RefSeq NM_000693: c.1514T>C (p.I505T)) and FOXN1 (RefSeq NM_003593: c.146C>T (p.S49L)). Gene expression studies reveal that Aldh1a3 and Foxn1 are expressed in ~E13.5 mouse embryonic brain, as well as in adult piriform cortex (PC; ~P30). Conserved Retinoic Acid Response Elements (RAREs) upstream of human ALDH1A3 and FOXN1 and in mouse Aldh1a3 and Foxn1 genes were revealed using bioinformatic approximation. Chromatin immunoprecipitation (ChIP) assay using Retinoid Acid Receptor B (Rarb) as the immunoprecipitation target suggests RA regulation of Aldh1a3 and Foxn1 in mice. Our results frame a possible link of RA regulation in brain to ASD etiology, and a feasible non-additive effect of two apparently unrelated variants in ALDH1A3 and FOXN1 recognizing that every result given by next generation sequencing should be cautiously analyzed, as it might be an incidental finding. PMID:26352270

  5. Microcollinearity in an ethylene receptor coding gene region of the Coffea canephora genome is extensively conserved with Vitis vinifera and other distant dicotyledonous sequenced genomes

    PubMed Central

    Guyot, Romain; de la Mare, Marion; Viader, Véronique; Hamon, Perla; Coriton, Olivier; Bustamante-Porras, José; Poncet, Valérie; Campa, Claudine; Hamon, Serge; de Kochko, Alexandre

    2009-01-01

    Background Coffea canephora, also called Robusta, belongs to the Rubiaceae, the fourth largest angiosperm family. This diploid species (2x = 2n = 22) has a fairly small genome size of ≈ 690 Mb and despite its extreme economic importance, particularly for developing countries, knowledge on the genome composition, structure and evolution remain very limited. Here, we report the 160 kb of the first C. canephora Bacterial Artificial Chromosome (BAC) clone ever sequenced and its fine analysis. Results This clone contains the CcEIN4 gene, encoding an ethylene receptor, and twenty other predicted genes showing a high gene density of one gene per 7.8 kb. Most of them display perfect matches with C. canephora expressed sequence tags or show transcriptional activities through PCR amplifications on cDNA libraries. Twenty-three transposable elements, mainly Class II transposon derivatives, were identified at this locus. Most of these Class II elements are Miniature Inverted-repeat Transposable Elements (MITE) known to be closely associated with plant genes. This BAC composition gives a pattern similar to those found in gene rich regions of Solanum lycopersicum and Medicago truncatula genomes indicating that the CcEIN4 regions may belong to a gene rich region in the C. canephora genome. Comparative sequence analysis indicated an extensive conservation between C. canephora and most of the reference dicotyledonous genomes studied in this work, such as tomato (S. lycopersicum), grapevine (V. vinifera), barrel medic M. truncatula, black cottonwood (Populus trichocarpa) and Arabidopsis thaliana. The higher degree of microcollinearity was found between C. canephora and V. vinifera, which belong respectively to the Asterids and Rosids, two clades that diverged more than 114 million years ago. Conclusion This study provides a first glimpse of C. canephora genome composition and evolution. Our data revealed a remarkable conservation of the microcollinearity between C. canephora and V

  6. Lactic acid production from potato peel waste by anaerobic sequencing batch fermentation using undefined mixed culture.

    PubMed

    Liang, Shaobo; McDonald, Armando G; Coats, Erik R

    2015-11-01

    Lactic acid (LA) is a necessary industrial feedstock for producing the bioplastic, polylactic acid (PLA), which is currently produced by pure culture fermentation of food carbohydrates. This work presents an alternative to produce LA from potato peel waste (PPW) by anaerobic fermentation in a sequencing batch reactor (SBR) inoculated with undefined mixed culture from a municipal wastewater treatment plant. A statistical design of experiments approach was employed using set of 0.8L SBRs using gelatinized PPW at a solids content range from 30 to 50 g L(-1), solids retention time of 2-4 days for yield and productivity optimization. The maximum LA production yield of 0.25 g g(-1) PPW and highest productivity of 125 mg g(-1) d(-1) were achieved. A scale-up SBR trial using neat gelatinized PPW (at 80 g L(-1) solids content) at the 3 L scale was employed and the highest LA yield of 0.14 g g(-1) PPW and a productivity of 138 mg g(-1) d(-1) were achieved with a 1 d SRT. PMID:25708409

  7. Bacterial community compositions in sediment polluted by perfluoroalkyl acids (PFAAs) using Illumina high-throughput sequencing.

    PubMed

    Sun, Yajun; Wang, Tieyu; Peng, Xiawei; Wang, Pei; Lu, Yonglong

    2016-06-01

    The characterization of bacterial community compositions and the change in perfluoroalkyl acids (PFAAs) along a natural river distribution system were explored in the present study. Illumina high-throughput sequencing was used to explore bacterial community diversity and structure in sediment polluted by PFAAs from the Xiaoqing River, the area with concentrated fluorochemical facilities in China. The concentration of PFAAs was in the range of 8.44-465.60 ng/g dry weight (dw) in sediment. Perfluorooctanoic acid (PFOA) was the dominant PFAA in all samples, which accounted for 94.2 % of total PFAAs. High-level PFOA could lead to an obvious increase in relative abundance of Proteobacteria, ε-Proteobacteria, Thiobacillus, and Sulfurimonas and the decrease in relative abundance of other bacteria. Redundancy analysis revealed that PFOA played an important role in the formation of bacterial community, and PFOA at higher concentration could reduce the diversity of bacterial community. When the concentration of PFOA was below 100 ng/g dw in sediment, no significant effect on microbial community structure was observed. Thiobacillus and Sulfurimonas were positively correlated with the concentration of PFOA, suggesting that both genera were resistant to PFOA contamination. PMID:26780047

  8. Mass spectrometric detection of the amino acid sequence polymorphism of the hepatitis C virus antigen.

    PubMed

    Kaysheva, A L; Ivanov, Yu D; Frantsuzov, P A; Krohin, N V; Pavlova, T I; Uchaikin, V F; Konev, V А; Kovalev, O B; Ziborov, V S; Archakov, A I

    2016-03-01

    A method for detection and identification of the hepatitis C virus antigen (HCVcoreAg) in human serum with consideration for possible amino acid substitutions is proposed. The method is based on a combination of biospecific capturing and concentrating of the target protein on the surface of the chip for atomic force microscope (AFM chip) with subsequent protein identification by tandem mass spectrometric (MS/MS) analysis. Biospecific AFM-capturing of viral particles containing HCVcoreAg from serum samples was performed by use of AFM chips with monoclonal antibodies (anti-HCVcore) covalently immobilized on the surface. Biospecific complexes were registered and counted by AFM. Further MS/MS analysis allowed to reliably identify the HCVcoreAg in the complexes formed on the AFM chip surface. Analysis of MS/MS spectra, with the account taken of the possible polymorphisms in the amino acid sequence of the HCVcoreAg, enabled us to increase the number of identified peptides. PMID:26773170

  9. The Eps1p Protein Disulfide Isomerase Conserves Classic Thioredoxin Superfamily Amino Acid Motifs but Not Their Functional Geometries

    PubMed Central

    Biran, Shai; Gat, Yair; Fass, Deborah

    2014-01-01

    The widespread thioredoxin superfamily enzymes typically share the following features: a characteristic α-β fold, the presence of a Cys-X-X-Cys (or Cys-X-X-Ser) redox-active motif, and a proline in the cis configuration abutting the redox-active site in the tertiary structure. The Cys-X-X-Cys motif is at the solvent-exposed amino terminus of an α-helix, allowing the first cysteine to engage in nucleophilic attack on substrates, or substrates to attack the Cys-X-X-Cys disulfide, depending on whether the enzyme functions to reduce, isomerize, or oxidize its targets. We report here the X-ray crystal structure of an enzyme that breaks many of our assumptions regarding the sequence-structure relationship of thioredoxin superfamily proteins. The yeast Protein Disulfide Isomerase family member Eps1p has Cys-X-X-Cys motifs and proline residues at the appropriate primary structural positions in its first two predicted thioredoxin-fold domains. However, crystal structures show that the Cys-X-X-Cys of the second domain is buried and that the adjacent proline is in the trans, rather than the cis isomer. In these configurations, neither the “active-site” disulfide nor the backbone carbonyl preceding the proline is available to interact with substrate. The Eps1p structures thus expand the documented diversity of the PDI oxidoreductase family and demonstrate that conserved sequence motifs in common folds do not guarantee structural or functional conservation. PMID:25437863

  10. Peptide sequencing by using a combination of partial acid hydrolysis and fast-atom-bombardment mass spectrometry.

    PubMed Central

    De Angelis, F; Botta, M; Ceccarelli, S; Nicoletti, R

    1986-01-01

    To overcome the limit of the intensity of ions carrying sequence information in structural determinations of peptides by fast-atom-bombardment m.s., we have developed a method that consists in taking spectra of the peptide acid hydrolysates at different hydrolysis times. Peaks correspond to the oligomers arising from the peptide partial hydrolysis. The sequence can then be identified from the structurally overlapping fragments. PMID:2428356

  11. G glycoprotein amino acid residues required for human monoclonal antibody RAB1 neutralization are conserved in rabies virus street isolates.

    PubMed

    Wang, Yang; Rowley, Kirk J; Booth, Brian J; Sloan, Susan E; Ambrosino, Donna M; Babcock, Gregory J

    2011-08-01

    Replacement of polyclonal anti-rabies immunoglobulin (RIG) used in rabies post-exposure prophylaxis (PEP) with a monoclonal antibody will eliminate cost and availability constraints that currently exist using RIG in the developing world. The human monoclonal antibody RAB1 has been shown to neutralize all rabies street isolates tested; however for the laboratory-adapted fixed strain, CVS-11, mutation in the G glycoprotein of amino acid 336 from asparagine (N) to aspartic acid (D) resulted in resistance to neutralization. Interestingly, this same mutation in the G glycoprotein of a second laboratory-adapted fixed strain (ERA) did not confer resistance to RAB1 neutralization. Using cell surface staining and lentivirus pseudotyped with rabies virus G glycoprotein (RABVpp), we identified an amino acid alteration in CVS-11 (K346), not present in ERA (R346), which was required in combination with D336 to confer resistance to RAB1. A complete analysis of G glycoprotein sequences from GenBank demonstrated that no identified rabies isolates contain the necessary combination of G glycoprotein mutations for resistance to RAB1 neutralization, consistent with the broad neutralization of RAB1 observed in direct viral neutralization experiments with street isolates. All combinations of amino acids 336 and 346 reported in the sequence database were engineered into the ERA G glycoprotein and RAB1 was able to neutralize RABVpp bearing ERA G glycoprotein containing all known combinations at these critical residues. These data demonstrate that RAB1 has the capacity to neutralize all identified rabies isolates and a minimum of two distinct mutations in the G glycoprotein are required for abrogation of RAB1 neutralization. PMID:21693135

  12. Canine preprorelaxin: nucleic acid sequence and localization within the canine placenta.

    PubMed

    Klonisch, T; Hombach-Klonisch, S; Froehlich, C; Kauffold, J; Steger, K; Steinetz, B G; Fischer, B

    1999-03-01

    Employing uteroplacental tissue at Day 35 of gestation, we determined the nucleic acid sequence of canine preprorelaxin using reverse transcription- and rapid amplification of cDNA ends-polymerase chain reaction. Canine preprorelaxin cDNA consisted of 534 base pairs encoding a protein of 177 amino acids with a signal peptide of 25 amino acids (aa), a B domain of 35 aa, a C domain of 93 aa, and an A domain of 24 aa. The putative receptor binding region in the N'-terminal part of the canine relaxin B domain GRDYVR contained two substitutions from the classical motif (E-->D and L-->Y). Canine preprorelaxin shared highest homology with porcine and equine preprorelaxin. Northern analysis revealed a 1-kilobase transcript present in total RNA of canine uteroplacental tissue but not of kidney tissue. Uteroplacental tissue from two bitches each at Days 30 and 35 of gestation were studied by in situ hybridization to localize relaxin mRNA. Immunohistochemistry for relaxin, cytokeratin, vimentin, and von Willebrand factor was performed on uteroplacental tissue at Day 30 of gestation. The basal cell layer at the core of the chorionic villi was devoid of relaxin mRNA and immunoreactive relaxin or vimentin but was immunopositive for cytokeratin and identified as cytotrophoblast cells. The cell layer surrounding the chorionic villi displayed specific hybridization signals for relaxin mRNA and immunoreactivity for relaxin and cytokeratin but not for vimentin, and was identified as syncytiotrophoblast. Those areas of the chorioallantoic tissue with most intense relaxin immunoreactivity were highly vascularized as demonstrated by immunoreactive von Willebrand factor expressed on vascular endothelium. The uterine glands and nonplacental uterine areas of the canine zonary girdle placenta were devoid of relaxin mRNA and relaxin. We conclude that the syncytiotrophoblast is the source of relaxin in the canine placenta. PMID:10026098

  13. Purification and partial amino acid sequence of the chloroplast cytochrome b-559.

    PubMed

    Widger, W R; Cramer, W A; Hermodson, M; Meyer, D; Gullifor, M

    1984-03-25

    The hydrophobic cytochrome b-559, purified from unstacked, ethanol-washed spinach thylakoid membranes, using extraction with 2% Triton X-100 in 4 M urea and three chromatographic steps in the presence of protease inhibitors, has a dominant band on sodium dodecyl sulfate-urea gels corresponding to Mr = 10,000. The yield of this preparation is 30-50% (5-10 mg) starting with 600 mg of chlorophyll. The heme content yields a calculated molecular weight of no more than 17,500/heme, and perhaps somewhat smaller after correction for impurities. The Mr = 10,000 band is stained by the tetramethylbenzidine-H2O2 heme reagent on lithium dodecyl sulfate gels run at 0 degrees C. The Mr = 10,000 protein, further separated by high performance liquid chromatography, contains a unique NH2 terminus that is not blocked, and the amino acid sequence for the first 27 residues is NH2-Ser-Gly-Ser-Thr-Gly-Glu-Arg-Ser-Phe-Ala-Asp-Ile-Ile-Thr-Ser-Ile-Arg-Tyr-Trp -Val-Ile-X-Ser-Ile-Thr-Ile-Pro. . . COOH. Approximately 55% of the amino acids are hydrophobic, based on amino acid analysis of the Mr = 10,000 peptide, which also indicated the presence of at least one histidine. Only one cytochrome b-559 component could be identified, whose yield indicated that it arises from a single b-559 protein in chloroplasts corresponding to the in situ high potential cytochrome of the chloroplast photosystem II. PMID:6706983

  14. Sequence-Specific Electrical Purification of Nucleic Acids with Nanoporous Gold Electrodes.

    PubMed

    Daggumati, Pallavi; Appelt, Sandra; Matharu, Zimple; Marco, Maria L; Seker, Erkin

    2016-06-22

    Nucleic-acid-based biosensors have enabled rapid and sensitive detection of pathogenic targets; however, these devices often require purified nucleic acids for analysis since the constituents of complex biological fluids adversely affect sensor performance. This purification step is typically performed outside the device, thereby increasing sample-to-answer time and introducing contaminants. We report a novel approach using a multifunctional matrix, nanoporous gold (np-Au), which enables both detection of specific target sequences in a complex biological sample and their subsequent purification. The np-Au electrodes modified with 26-mer DNA probes (via thiol-gold chemistry) enabled sensitive detection and capture of complementary DNA targets in the presence of complex media (fetal bovine serum) and other interfering DNA fragments in the range of 50-1500 base pairs. Upon capture, the noncomplementary DNA fragments and serum constituents of varying sizes were washed away. Finally, the surface-bound DNA-DNA hybrids were released by electrochemically cleaving the thiol-gold linkage, and the hybrids were iontophoretically eluted from the nanoporous matrix. The optical and electrophoretic characterization of the analytes before and after the detection-purification process revealed that low target DNA concentrations (80 pg/μL) can be successfully detected in complex biological fluids and subsequently released to yield pure hybrids free of polydisperse digested DNA fragments and serum biomolecules. Taken together, this multifunctional platform is expected to enable seamless integration of detection and purification of nucleic acid biomarkers of pathogens and diseases in miniaturized diagnostic devices. PMID:27244455

  15. Recognition of Conserved Amino Acid Motifs of Common Viruses and Its Role in Autoimmunity

    PubMed Central

    2005-01-01

    The triggers of autoimmune diseases such as multiple sclerosis (MS) remain elusive. Epidemiological studies suggest that common pathogens can exacerbate and also induce MS, but it has been difficult to pinpoint individual organisms. Here we demonstrate that in vivo clonally expanded CD4+ T cells isolated from the cerebrospinal fluid of a MS patient during disease exacerbation respond to a poly-arginine motif of the nonpathogenic and ubiquitous Torque Teno virus. These T cell clones also can be stimulated by arginine-enriched protein domains from other common viruses and recognize multiple autoantigens. Our data suggest that repeated infections with common pathogenic and even nonpathogenic viruses could expand T cells specific for conserved protein domains that are able to cross-react with tissue-derived and ubiquitous autoantigens. PMID:16362076

  16. Negative Ion In-Source Decay Matrix-Assisted Laser Desorption/Ionization Mass Spectrometry for Sequencing Acidic Peptides

    NASA Astrophysics Data System (ADS)

    McMillen, Chelsea L.; Wright, Patience M.; Cassady, Carolyn J.

    2016-05-01

    Matrix-assisted laser desorption/ionization (MALDI) in-source decay was studied in the negative ion mode on deprotonated peptides to determine its usefulness for obtaining extensive sequence information for acidic peptides. Eight biological acidic peptides, ranging in size from 11 to 33 residues, were studied by negative ion mode ISD (nISD). The matrices 2,5-dihydroxybenzoic acid, 2-aminobenzoic acid, 2-aminobenzamide, 1,5-diaminonaphthalene, 5-amino-1-naphthol, 3-aminoquinoline, and 9-aminoacridine were used with each peptide. Optimal fragmentation was produced with 1,5-diaminonphthalene (DAN), and extensive sequence informative fragmentation was observed for every peptide except hirudin(54-65). Cleavage at the N-Cα bond of the peptide backbone, producing c' and z' ions, was dominant for all peptides. Cleavage of the N-Cα bond N-terminal to proline residues was not observed. The formation of c and z ions is also found in electron transfer dissociation (ETD), electron capture dissociation (ECD), and positive ion mode ISD, which are considered to be radical-driven techniques. Oxidized insulin chain A, which has four highly acidic oxidized cysteine residues, had less extensive fragmentation. This peptide also exhibited the only charged localized fragmentation, with more pronounced product ion formation adjacent to the highly acidic residues. In addition, spectra were obtained by positive ion mode ISD for each protonated peptide; more sequence informative fragmentation was observed via nISD for all peptides. Three of the peptides studied had no product ion formation in ISD, but extensive sequence informative fragmentation was found in their nISD spectra. The results of this study indicate that nISD can be used to readily obtain sequence information for acidic peptides.

  17. Maize peroxidase Px5 has a highly conserved sequence in inbreds resistant to mycotoxin producing fungi which enhances fungal and insect resistance.

    PubMed

    Dowd, Patrick F; Johnson, Eric T

    2016-01-01

    Mycotoxin presence in maize causes health and economic issues for humans and animals. Although many studies have investigated expression differences of genes putatively governing resistance to producing fungi, few have confirmed a resistance role, or examined putative resistance gene structure in more than a couple of inbreds. The pericarp expression of maize Px5 has previously been associated with resistance to Aspergillus flavus growth and insects in a set of inbreds. Genes from 14 different inbreds that included ones with resistance and susceptibility to A. flavus, Fusarium proliferatum, F. verticillioides and F. graminearum and/or mycotoxin production were cloned using high fidelity enzymes, and sequenced. The sequence of Px5 from all resistant inbreds was identical, except for a single base change in two inbreds, only one of which affected the amino acid sequence. Conversely, the Px5 sequence from several susceptible inbreds had several base variations, some of which affected amino acid sequence that would potentially alter secondary structure, and thus enzyme function. The sequence of the maize peroxidase Px5 common to inbreds resistant to mycotoxigenic fungi was overexpressed in maize callus. Callus transformants overexpressing the gene caused significant reductions in growth for fall armyworms, corn earworms, and F. graminearum compared to transformant callus with a β-glucuronidase gene. This study demonstrates rarer transcripts of potential resistance genes overlooked by expression screens can be identified by sequence comparisons. A role in pest resistance can be verified by callus expression of the candidate genes, which can thereby justify larger scale transformation and regeneration of transgenic plants expressing the resistance gene for further evaluation. PMID:26659597

  18. A family of conserved bacterial effectors inhibits salicylic acid-mediated basal immunity and promotes disease necrosis in plants.

    PubMed

    DebRoy, Sruti; Thilmony, Roger; Kwack, Yong-Bum; Nomura, Kinya; He, Sheng Yang

    2004-06-29

    Salicylic acid (SA)-mediated host immunity plays a central role in combating microbial pathogens in plants. Inactivation of SA-mediated immunity, therefore, would be a critical step in the evolution of a successful plant pathogen. It is known that mutations in conserved effector loci (CEL) in the plant pathogens Pseudomonas syringae (the Delta CEL mutation), Erwinia amylovora (the dspA/E mutation), and Pantoea stewartii subsp. stewartii (the wtsE mutation) exert particularly strong negative effects on bacterial virulence in their host plants by unknown mechanisms. We found that the loss of virulence in Delta CEL and dspA/E mutants was linked to their inability to suppress cell wall-based defenses and to cause normal disease necrosis in Arabidopsis and apple host plants. The Delta CEL mutant activated SA-dependent callose deposition in wild-type Arabidopsis but failed to elicit high levels of callose-associated defense in Arabidopsis plants blocked in SA accumulation or synthesis. This mutant also multiplied more aggressively in SA-deficient plants than in wild-type plants. The hopPtoM and avrE genes in the CEL of P. syringae were found to encode suppressors of this SA-dependent basal defense. The widespread conservation of the HopPtoM and AvrE families of effectors in various bacteria suggests that suppression of SA-dependent basal immunity and promotion of host cell death are important virulence strategies for bacterial infection of plants. PMID:15210989

  19. Conservation of inner nuclear membrane targeting sequences in mammalian Pom121 and yeast Heh2 membrane proteins.

    PubMed

    Kralt, Annemarie; Jagalur, Noorjahan B; van den Boom, Vincent; Lokareddy, Ravi K; Steen, Anton; Cingolani, Gino; Fornerod, Maarten; Veenhoff, Liesbeth M

    2015-09-15

    Endoplasmic reticulum-synthesized membrane proteins traffic through the nuclear pore complex (NPC) en route to the inner nuclear membrane (INM). Although many membrane proteins pass the NPC by simple diffusion, two yeast proteins, ScSrc1/ScHeh1 and ScHeh2, are actively imported. In these proteins, a nuclear localization signal (NLS) and an intrinsically disordered linker encode the sorting signal for recruiting the transport factors for FG-Nup and RanGTP-dependent transport through the NPC. Here we address whether a similar import mechanism applies in metazoans. We show that the (putative) NLSs of metazoan HsSun2, MmLem2, HsLBR, and HsLap2β are not sufficient to drive nuclear accumulation of a membrane protein in yeast, but the NLS from RnPom121 is. This NLS of Pom121 adapts a similar fold as the NLS of Heh2 when transport factor bound and rescues the subcellular localization and synthetic sickness of Heh2ΔNLS mutants. Consistent with the conservation of these NLSs, the NLS and linker of Heh2 support INM localization in HEK293T cells. The conserved features of the NLSs of ScHeh1, ScHeh2, and RnPom121 and the effective sorting of Heh2-derived reporters in human cells suggest that active import is conserved but confined to a small subset of INM proteins. PMID:26179916

  20. Conservation of inner nuclear membrane targeting sequences in mammalian Pom121 and yeast Heh2 membrane proteins

    PubMed Central

    Kralt, Annemarie; Jagalur, Noorjahan B.; van den Boom, Vincent; Lokareddy, Ravi K.; Steen, Anton; Cingolani, Gino; Fornerod, Maarten; Veenhoff, Liesbeth M.

    2015-01-01

    Endoplasmic reticulum–synthesized membrane proteins traffic through the nuclear pore complex (NPC) en route to the inner nuclear membrane (INM). Although many membrane proteins pass the NPC by simple diffusion, two yeast proteins, ScSrc1/ScHeh1 and ScHeh2, are actively imported. In these proteins, a nuclear localization signal (NLS) and an intrinsically disordered linker encode the sorting signal for recruiting the transport factors for FG-Nup and RanGTP-dependent transport through the NPC. Here we address whether a similar import mechanism applies in metazoans. We show that the (putative) NLSs of metazoan HsSun2, MmLem2, HsLBR, and HsLap2β are not sufficient to drive nuclear accumulation of a membrane protein in yeast, but the NLS from RnPom121 is. This NLS of Pom121 adapts a similar fold as the NLS of Heh2 when transport factor bound and rescues the subcellular localization and synthetic sickness of Heh2ΔNLS mutants. Consistent with the conservation of these NLSs, the NLS and linker of Heh2 support INM localization in HEK293T cells. The conserved features of the NLSs of ScHeh1, ScHeh2, and RnPom121 and the effective sorting of Heh2-derived reporters in human cells suggest that active import is conserved but confined to a small subset of INM proteins. PMID:26179916

  1. Complete amino acid sequence of the medium-chain S-acyl fatty acid synthetase thio ester hydrolase from rat mammary gland

    SciTech Connect

    Randhawa, Z.I.; Smith, S.

    1987-03-10

    The complete amino acid sequence of the medium-chain S-acyl fatty acid synthetase thio ester hydrolase (thioesterase II) from rat mammary gland is presented. Most of the sequence was derived by analysis of (/sup 14/C)-labelled peptide fragments produced by cleavage at methionyl, glutamyl, lysyl, arginyl, and tryptophanyl residues. A small section of the sequence was deduced from a previously analyzed cDNA clone. The protein consists of 260 residues and has a blocked amino-terminal methionine and calculated M/sub r/ of 29,212. The carboxy-terminal sequence, verified by Edman degradation of the carboxy-terminal cyanogen bromide fragment and carboxypeptidase Y digestion of the intact thioesterase II, terminates with a serine residue and lacks three additional residues predicted by the cDNA sequence. The native enzyme contains three cysteine residues but no disulfide bridges. The active site serine residue is located at position 101. The rat mammary gland thioesterase II exhibits approximately 40% homology with a thioesterase from mallard uropygial gland, the sequence of which was recently determined by cDNA analysis. Thus the two enzymes may share similar structural features and a common evolutionary origin. The location of the active site in these thioesterases differs from that of other serine active site esterases; indeed, the enzymes do not exhibit any significant homology with other serine esterases, suggesting that they may constitute a separate new family of serine active site enzymes.

  2. The complete amino acid sequence of the A-chain of human plasma alpha 2HS-glycoprotein.

    PubMed

    Yoshioka, Y; Gejyo, F; Marti, T; Rickli, E E; Bürgi, W; Offner, G D; Troxler, R F; Schmid, K

    1986-02-01

    Normal human plasma alpha 2HS-glycoprotein has earlier been shown to be comprised of two polypeptide chains. Recently, the amino acid and carbohydrate sequences of the short chain were elucidated (Gejyo, F., Chang, J.-L., Bürgi, W., Schmid, K., Offner, G. D., Troxler, R.F., van Halbeck, H., Dorland, L., Gerwig, G. J., and Vliegenthart, J.F.G. (1983) J. Biol. Chem. 258, 4966-4971). In the present study, the amino acid sequence of the long chain of this protein, designated A-chain, was determined and found to consist of 282 amino acid residues. Twenty-four amino acid doublets were found; the most abundant of these are Pro-Pro and Ala-Ala which each occur five times. Of particular interest is the presence of three Gly-X-Pro and one Gly-Pro-X sequences that are characteristic of the repeating sequences of collagens. Chou-Fasman evaluation of the secondary structure suggested that the A-chain contains 29% alpha-helix, 24% beta-pleated sheet, and 26% reverse turns and, thus, approximately 80% of the polypeptide chain may display ordered structure. Four glycosylation sites were identified. The two N-glycosidic oligosaccharides were found in the center region (residues 138 and 158), whereas the two O-glycosidic heterosaccharides, both linked to threonine (residues 238 and 252), occur within the carboxyl-terminal region. The N-glycans are linked to Asn residues in beta-turns, while the O-glycans are located in short random segments. Comparison of the sequence of the amino- and carboxyl-terminal 30 residues with protein sequences in a data bank demonstrated that the A-chain is not significantly related to any known proteins. However, the proline-rich carboxyl-terminal region of the A-chain displays some sequence similarity to collagens and the collagen-like domains of complement subcomponent C1q. PMID:3944104

  3. Comparative application of direct sequencing, PCR-RFLP, and cytogenetic markers in the genetic characterization of Pimelodus (Siluriformes: Pimelodidae) species: possible implications for fish conservation.

    PubMed

    Ferreira, M; Bressane, K C O; Moresco, A R C; Moreira-Filho, O; Almeida-Toledo, L F; Garcia, C

    2014-01-01

    Pimelodus (Pimelodidae) is a genus comprising a group of South American species with complex taxonomic relationships. Cytogenetics, polymerase chain reaction restriction fragment length polymorphism (PCR-RFLP), and sequencing data of mitochondrial genes were analyzed to characterize 4 Pimelodus species: P. fur, P. heraldoi, P. maculatus, and Pimelodus sp. All populations presented 2n=56 chromosomes and distinct karyotypic formulae. The heterochromatin distribution pattern and the number and location of 5S and 18S rDNA sites are discussed. The application of PCR-RFLP markers and sequencing of mitochondrial DNA genes provided species-specific haplotypes, which allowed us to differentiate the species studied. The mitochondrial gene sequences presented nucleotide mutations in the restriction sites and throughout the sequences, and they were mostly related to synonymous substitutions in the coded proteins; however, they did not affect the protein and its function. Comparing the data obtained using these 3 methodologies, the existence of a species complex in P. maculatus along the basins studied might be inferred, showing that cytogenetics is an important tool in studies focusing on the conservation or management of both natural and captive populations of these fishes. PMID:25036358

  4. Replication origins and a sequence involved in coordinate induction of the immediate-early gene family are conserved in an intergenic region of herpes simplex virus.

    PubMed Central

    Whitton, J L; Clements, J B

    1984-01-01

    We have determined the structure of the 5' portion of herpes simplex virus type 2 (HSF-2) immediate-early (IE) mRNA-3 and have obtained the DNA sequence specifying the N terminus of its encoded polypeptide, Vmw182, its untranslated leader and the intergenic region between IEmRNAs-3 & 4/5. Comparison of the HSV-2 intergenic sequences with the HSV-1 equivalent region identifies several conserved regions: (1) an AT-rich element with core consensus TAATGARAT which is likely to be the 'activator' sequence through which coordinate induction of the IE gene family is mediated. (2) GC-rich and GA-rich tracts, found in a wide variety of eukaryotic promoters, which vary in position and orientation between HSV-2 and HSV-1 and which represent modulators of transcription. (3) TATA homologies present 15-25 base pairs (bp) upstream of mRNA 5' termini. (4) a 137bp direct repeat in HSV-2 which contains sequence almost identical to the HSV-1 replication origin. Images PMID:6322134

  5. Branched-chain amino acid catabolism is a conserved regulator of physiological ageing.

    PubMed

    Mansfeld, Johannes; Urban, Nadine; Priebe, Steffen; Groth, Marco; Frahm, Christiane; Hartmann, Nils; Gebauer, Juliane; Ravichandran, Meenakshi; Dommaschk, Anne; Schmeisser, Sebastian; Kuhlow, Doreen; Monajembashi, Shamci; Bremer-Streck, Sibylle; Hemmerich, Peter; Kiehntopf, Michael; Zamboni, Nicola; Englert, Christoph; Guthke, Reinhard; Kaleta, Christoph; Platzer, Matthias; Sühnel, Jürgen; Witte, Otto W; Zarse, Kim; Ristow, Michael

    2015-01-01

    Ageing has been defined as a global decline in physiological function depending on both environmental and genetic factors. Here we identify gene transcripts that are similarly regulated during physiological ageing in nematodes, zebrafish and mice. We observe the strongest extension of lifespan when impairing expression of the branched-chain amino acid transferase-1 (bcat-1) gene in C. elegans, which leads to excessive levels of branched-chain amino acids (BCAAs). We further show that BCAAs reduce a LET-363/mTOR-dependent neuro-endocrine signal, which we identify as DAF-7/TGFβ, and that impacts lifespan depending on its related receptors, DAF-1 and DAF-4, as well as ultimately on DAF-16/FoxO and HSF-1 in a cell-non-autonomous manner. The transcription factor HLH-15 controls and epistatically synergizes with BCAT-1 to modulate physiological ageing. Lastly and consistent with previous findings in rodents, nutritional supplementation of BCAAs extends nematodal lifespan. Taken together, BCAAs act as periphery-derived metabokines that induce a central neuro-endocrine response, culminating in extended healthspan. PMID:26620638

  6. Branched-chain amino acid catabolism is a conserved regulator of physiological ageing

    PubMed Central

    Mansfeld, Johannes; Urban, Nadine; Priebe, Steffen; Groth, Marco; Frahm, Christiane; Hartmann, Nils; Gebauer, Juliane; Ravichandran, Meenakshi; Dommaschk, Anne; Schmeisser, Sebastian; Kuhlow, Doreen; Monajembashi, Shamci; Bremer-Streck, Sibylle; Hemmerich, Peter; Kiehntopf, Michael; Zamboni, Nicola; Englert, Christoph; Guthke, Reinhard; Kaleta, Christoph; Platzer, Matthias; Sühnel, Jürgen; Witte, Otto W.; Zarse, Kim; Ristow, Michael

    2015-01-01

    Ageing has been defined as a global decline in physiological function depending on both environmental and genetic factors. Here we identify gene transcripts that are similarly regulated during physiological ageing in nematodes, zebrafish and mice. We observe the strongest extension of lifespan when impairing expression of the branched-chain amino acid transferase-1 (bcat-1) gene in C. elegans, which leads to excessive levels of branched-chain amino acids (BCAAs). We further show that BCAAs reduce a LET-363/mTOR-dependent neuro-endocrine signal, which we identify as DAF-7/TGFβ, and that impacts lifespan depending on its related receptors, DAF-1 and DAF-4, as well as ultimately on DAF-16/FoxO and HSF-1 in a cell-non-autonomous manner. The transcription factor HLH-15 controls and epistatically synergizes with BCAT-1 to modulate physiological ageing. Lastly and consistent with previous findings in rodents, nutritional supplementation of BCAAs extends nematodal lifespan. Taken together, BCAAs act as periphery-derived metabokines that induce a central neuro-endocrine response, culminating in extended healthspan. PMID:26620638

  7. Analysis of the functional domains of biosynthetic threonine deaminase by comparison of the amino acid sequences of three wild-type alleles to the amino acid sequence of biodegradative threonine deaminase.

    PubMed

    Taillon, B E; Little, R; Lawther, R P

    1988-03-31

    The nucleotide sequence of the gene, ilvA, for biosynthetic threonine deaminase (Tda) from Salmonella typhimurium was determined. The deduced amino acid sequence was compared with the deduced amino acid sequences of the biosynthetic Tda from Escherichia coli K-12 (ilvA) and Saccharomyces cerevisiae (ILV1) and the biodegradative Tda from E. coli K-12 (tdc). The comparison indicated the presence of two types of blocks of homologous amino acids. The first type of homology is in the N-terminal portion of all four isozymes of Tda and probably indicates amino acids involved in catalysis. The second type of homology is found in the C-terminal portion of the three biosynthetic isozymes and presumably is involved in either (i) the binding or interaction of the allosteric effector isoleucine with the enzyme, or (ii) subunit interactions. The sites of amino acid changes of two E. coli K-12 ilvA alleles with altered response to isoleucine are consistent with the conclusion that the C-terminal portion of biosynthetic Tda is involved in allosteric regulation. PMID:3290055

  8. The developmental transcriptome landscape of bovine skeletal muscle defined by Ribo-Zero ribonucleic acid sequencing.

    PubMed

    Sun, X; Li, M; Sun, Y; Cai, H; Li, R; Wei, X; Lan, X; Huang, Y; Lei, C; Chen, H

    2015-12-01

    Ribonucleic acid sequencing (RNA-Seq) libraries are normally prepared with oligo(dT) selection of poly(A)+ mRNA, but it depends on intact total RNA samples. Recent studies have described Ribo-Zero technology, a novel method that can capture both poly(A)+ and poly(A)- transcripts from intact or fragmented RNA samples. We report here the first application of Ribo-Zero RNA-Seq for the analysis of the bovine embryonic, neonatal, and adult skeletal muscle whole transcriptome at an unprecedented depth. Overall, 19,893 genes were found to be expressed, with a high correlation of expression levels between the calf and the adult. Hundreds of genes were found to be highly expressed in the embryo and decreased at least 10-fold after birth, indicating their potential roles in embryonic muscle development. In addition, we present for the first time the analysis of global transcript isoform discovery in bovine skeletal muscle and identified 36,694 transcript isoforms. Transcriptomic data were also analyzed to unravel sequence variations; 185,036 putative SNP and 12,428 putative short insertions-deletions (InDel) were detected. Specifically, many stop-gain, stop-loss, and frameshift mutations were identified that probably change the relative protein production and sequentially affect the gene function. Notably, the numbers of stage-specific transcripts, alternative splicing events, SNP, and InDel were greater in the embryo than in the calf and the adult, suggesting that gene expression is most active in the embryo. The resulting view of the transcriptome at a single-base resolution greatly enhances the comprehensive transcript catalog and uncovers the global trends in gene expression during bovine skeletal muscle development. PMID:26641174

  9. Sequence repeats and protein structure

    NASA Astrophysics Data System (ADS)

    Hoang, Trinh X.; Trovato, Antonio; Seno, Flavio; Banavar, Jayanth R.; Maritan, Amos

    2012-11-01

    Repeats are frequently found in known protein sequences. The level of sequence conservation in tandem repeats correlates with their propensities to be intrinsically disordered. We employ a coarse-grained model of a protein with a two-letter amino acid alphabet, hydrophobic (H) and polar (P), to examine the sequence-structure relationship in the realm of repeated sequences. A fraction of repeated sequences comprises a distinct class of bad folders, whose folding temperatures are much lower than those of random sequences. Imperfection in sequence repetition improves the folding properties of the bad folders while deteriorating those of the good folders. Our results may explain why nature has utilized repeated sequences for their versatility and especially to design functional proteins that are intrinsically unstructured at physiological temperatures.

  10. Method for the detection of specific nucleic acid sequences by polymerase nucleotide incorporation

    DOEpatents

    Castro, Alonso

    2004-06-01

    A method for rapid and efficient detection of a target DNA or RNA sequence is provided. A primer having a 3'-hydroxyl group at one end and having a sequence of nucleotides sufficiently homologous with an identifying sequence of nucleotides in the target DNA is selected. The primer is hybridized to the identifying sequence of nucleotides on the DNA or RNA sequence and a reporter molecule is synthesized on the target sequence by progressively binding complementary nucleotides to the primer, where the complementary nucleotides include nucleotides labeled with a fluorophore. Fluorescence emitted by fluorophores on single reporter molecules is detected to identify the target DNA or RNA sequence.

  11. A Conserved Acidic Residue in Phenylalanine Hydroxylase Contributes to Cofactor Affinity and Catalysis

    PubMed Central

    2015-01-01

    The catalytic domains of aromatic amino acid hydroxylases (AAAHs) contain a non-heme iron coordinated to a 2-His-1-carboxylate facial triad and two water molecules. Asp139 from Chromobacterium violaceum PAH (cPAH) resides within the second coordination sphere and contributes key hydrogen bonds with three active site waters that mediate its interaction with an oxidized form of the cofactor, 7,8-dihydro-l-biopterin, in crystal structures. To determine the catalytic role of this residue, various point mutants were prepared and characterized. Our isothermal titration calorimetry (ITC) analysis of iron binding implies that polarity at position 139 is not the sole criterion for metal affinity, as binding studies with D139E suggest that the size of the amino acid side chain also appears to be important. High-resolution crystal structures of the mutants reveal that Asp139 may not be essential for holding the bridging water molecules together, because many of these waters are retained even in the Ala mutant. However, interactions via the bridging waters contribute to cofactor binding at the active site, interactions for which charge of the residue is important, as the D139N mutant shows a 5-fold decrease in its affinity for pterin as revealed by ITC (compared to a 16-fold loss of affinity in the case of the Ala mutant). The Asn and Ala mutants show a much more pronounced defect in their kcat values, with nearly 16- and 100-fold changes relative to that of the wild type, respectively, indicating a substantial role of this residue in stabilization of the transition state by aligning the cofactor in a productive orientation, most likely through direct binding with the cofactor, supported by data from molecular dynamics simulations of the complexes. Our results indicate that the intervening water structure between the cofactor and the acidic residue masks direct interaction between the two, possibly to prevent uncoupled hydroxylation of the cofactor before the arrival of

  12. CONSERVED REGULATOR ELEMENTS IDENTIFIED FROM A COMPARATIVE PUROINDOLINE GENE SEQUENCE SURVEY OF TRITICUM AND AEGILOPS DIPLOID TAXA

    Technology Transfer Automated Retrieval System (TEKTRAN)

    Kernel texture (“hardness”) is an important trait that determines end-use quality of wheat (Triticum aestivum L. and T. turgidum ssp. durum [Desf.] Husn.). Variation in texture is associated with the presence/absence or sequence polymorphism of two proteins, puroindoline a and puroindoline b. This...

  13. OVINE HERPESVIRUS-2 GLYCOPROTEIN B SEQUENCES FROM TISSUES OF RUMINANT MALIGNANT CATARRHAL FEVER AND HEALTHY SHEEP ARE HIGHLY CONSERVED.

    Technology Transfer Automated Retrieval System (TEKTRAN)

    Ovine herpesvirus-2 (OHV-2) infection has been associated with malignant catarrhal fever (MCF) in susceptible ruminants. In order to further investigate whether OHV-2 is an aetiological agent for sheep-associated (SA) MCF in cattle and bison, the entire sequences of OHV-2 glycoprotein B (gB) from di...

  14. Characterization and cDNA sequence of Bothriechis schlegeliil-amino acid oxidase with antibacterial activity.

    PubMed

    Vargas Muñoz, Leidy Johana; Estrada-Gomez, Sebastian; Núñez, Vitelbina; Sanz, Libia; Calvete, Juan J

    2014-08-01

    Snake venoms are complex mixtures of proteins including l-amino acid oxidase (lAAO). A lAAO (named BslAAO) with a mass of 56kDa and a theoretical Ip of 5.79, was purified from Bothriechis schlegelii venom through size-exclusion, ion exchange and affinity chromatography. The entire protein sequence of 498 amino acids, was determined from cDNA using reverse-transcribed mRNA isolated from venom gland. The enzyme showed dose-dependent inhibition of bacterial growth. BslAAO showed inhibitory effect against S. aureus with a MIC of 4μg/mL and a MBC of 8μg/mL. Against Acinetobacter baumannii, showed a MIC of 2μg/mL and MBC of 4μg/mL, No effect was observed in Escherichia coli. This antibacterial activity was inhibited by catalase, indicating that antimicrobial activity was due to H2O2 production. BslAAO did not show any cytotoxic activity toward mouse myoblast cell line C2C12 or peripheral blood mononuclear cells. The enzyme oxidated l-Leu, with a Km of 16.37μM and a Vmax of 0.39μM/min. Snake venoms lAAOs, are potential frames of different therapeutics molecules since these enzymes exhibit low MICs and MBCs and show to be harmless to human cells due to microorganisms being generally several fold more sensitive to reactive oxygen species than human tissues. PMID:24875315

  15. Hygienisation and nutrient conservation of sewage sludge or cattle manure by lactic acid fermentation.

    PubMed

    Scheinemann, Hendrik A; Dittmar, Katja; Stöckel, Frank S; Müller, Hermann; Krüger, Monika E

    2015-01-01

    Manure from animal farms and sewage sludge contain pathogens and opportunistic organisms in various concentrations depending on the health of the herds and human sources. Other than for the presence of pathogens, these waste substances are excellent nutrient sources and constitute a preferred organic fertilizer. However, because of the pathogens, the risks of infection of animals or humans increase with the indiscriminate use of manure, especially liquid manure or sludge, for agriculture. This potential problem can increase with the global connectedness of animal herds fed imported feed grown on fields fertilized with local manures. This paper describes a simple, easy-to-use, low-tech hygienization method which conserves nutrients and does not require large investments in infrastructure. The proposed method uses the microbiotic shift during mesophilic fermentation of cow manure or sewage sludge during which gram-negative bacteria, enterococci and yeasts were inactivated below the detection limit of 3 log10 cfu/g while lactobacilli increased up to a thousand fold. Pathogens like Salmonella, Listeria monocytogenes, Staphylococcus aureus, E. coli EHEC O:157 and vegetative Clostridium perfringens were inactivated within 3 days of fermentation. In addition, ECBO-viruses and eggs of Ascaris suum were inactivated within 7 and 56 days, respectively. Compared to the mass lost through composting (15-57%), the loss of mass during fermentation (< 2.45%) is very low and provides strong economic and ecological benefits for this process. This method might be an acceptable hygienization method for developed as well as undeveloped countries, and could play a key role in public and animal health while safely closing the nutrient cycle by reducing the necessity of using energy-inefficient inorganic fertilizer for crop production. PMID:25786255

  16. Hygienisation and Nutrient Conservation of Sewage Sludge or Cattle Manure by Lactic Acid Fermentation

    PubMed Central

    Scheinemann, Hendrik A.; Dittmar, Katja; Stöckel, Frank S.; Müller, Hermann; Krüger, Monika E.

    2015-01-01

    Manure from animal farms and sewage sludge contain pathogens and opportunistic organisms in various concentrations depending on the health of the herds and human sources. Other than for the presence of pathogens, these waste substances are excellent nutrient sources and constitute a preferred organic fertilizer. However, because of the pathogens, the risks of infection of animals or humans increase with the indiscriminate use of manure, especially liquid manure or sludge, for agriculture. This potential problem can increase with the global connectedness of animal herds fed imported feed grown on fields fertilized with local manures. This paper describes a simple, easy-to-use, low-tech hygienization method which conserves nutrients and does not require large investments in infrastructure. The proposed method uses the microbiotic shift during mesophilic fermentation of cow manure or sewage sludge during which gram-negative bacteria, enterococci and yeasts were inactivated below the detection limit of 3 log10 cfu/g while lactobacilli increased up to a thousand fold. Pathogens like Salmonella, Listeria monocytogenes, Staphylococcus aureus, E. coli EHEC O:157 and vegetative Clostridium perfringens were inactivated within 3 days of fermentation. In addition, ECBO-viruses and eggs of Ascaris suum were inactivated within 7 and 56 days, respectively. Compared to the mass lost through composting (15–57%), the loss of mass during fermentation (< 2.45%) is very low and provides strong economic and ecological benefits for this process. This method might be an acceptable hygienization method for developed as well as undeveloped countries, and could play a key role in public and animal health while safely closing the nutrient cycle by reducing the necessity of using energy-inefficient inorganic fertilizer for crop production. PMID:25786255

  17. Genome Sequence of a Candidate World Health Organization Reference Strain of Zika Virus for Nucleic Acid Testing

    PubMed Central

    Trösemeier, Jan-Hendrik; Musso, Didier; Blümel, Johannes; Thézé, Julien; Pybus, Oliver G.

    2016-01-01

    We report here the sequence of a candidate reference strain of Zika virus (ZIKV) developed on behalf of the World Health Organization (WHO). The ZIKV reference strain is intended for use in nucleic acid amplification (NAT)-based assays for the detection and quantification of ZIKV RNA. PMID:27587826

  18. Genome Sequence of Schizochytrium sp. CCTCC M209059, an Effective Producer of Docosahexaenoic Acid-Rich Lipids

    PubMed Central

    Ji, Xiao-Jun; Mo, Kai-Qiang; Ren, Lu-Jing; Li, Gan-Lu; Huang, Jian-Zhong

    2015-01-01

    Schizochytrium is an effective species for producing omega-3 docosahexaenoic acid (DHA). Here, we report a genome sequence of Schizochytrium sp. CCTCC M209059, which has a genome size of 39.09 Mb. It will provide the genomic basis for further insights into the metabolic and regulatory mechanisms underlying the DHA formation. PMID:26251485

  19. Evolutionary Distance of Amino Acid Sequence Orthologs across Macaque Subspecies: Identifying Candidate Genes for SIV Resistance in Chinese Rhesus Macaques

    PubMed Central

    Ross, Cody T.; Roodgar, Morteza; Smith, David Glenn

    2015-01-01

    We use the Reciprocal Smallest Distance (RSD) algorithm to identify amino acid sequence orthologs in the Chinese and Indian rhesus macaque draft sequences and estimate the evolutionary distance between such orthologs. We then use GOanna to map gene function annotations and human gene identifiers to the rhesus macaque amino acid sequences. We conclude methodologically by cross-tabulating a list of amino acid orthologs with large divergence scores with a list of genes known to be involved in SIV or HIV pathogenesis. We find that many of the amino acid sequences with large evolutionary divergence scores, as calculated by the RSD algorithm, have been shown to be related to HIV pathogenesis in previous laboratory studies. Four of the strongest candidate genes for SIVmac resistance in Chinese rhesus macaques identified in this study are CDK9, CXCL12, TRIM21, and TRIM32. Additionally, ANKRD30A, CTSZ, GORASP2, GTF2H1, IL13RA1, MUC16, NMDAR1, Notch1, NT5M, PDCD5, RAD50, and TM9SF2 were identified as possible candidates, among others. We failed to find many laboratory experiments contrasting the effects of Indian and Chinese orthologs at these sites on SIVmac pathogenesis, but future comparative studies might hold fertile ground for research into the biological mechanisms underlying innate resistance to SIVmac in Chinese rhesus macaques. PMID:25884674

  20. Evolutionary distance of amino acid sequence orthologs across macaque subspecies: identifying candidate genes for SIV resistance in Chinese rhesus macaques.

    PubMed

    Ross, Cody T; Roodgar, Morteza; Smith, David Glenn

    2015-01-01

    We use the Reciprocal Smallest Distance (RSD) algorithm to identify amino acid sequence orthologs in the Chinese and Indian rhesus macaque draft sequences and estimate the evolutionary distance between such orthologs. We then use GOanna to map gene function annotations and human gene identifiers to the rhesus macaque amino acid sequences. We conclude methodologically by cross-tabulating a list of amino acid orthologs with large divergence scores with a list of genes known to be involved in SIV or HIV pathogenesis. We find that many of the amino acid sequences with large evolutionary divergence scores, as calculated by the RSD algorithm, have been shown to be related to HIV pathogenesis in previous laboratory studies. Four of the strongest candidate genes for SIVmac resistance in Chinese rhesus macaques identified in this study are CDK9, CXCL12, TRIM21, and TRIM32. Additionally, ANKRD30A, CTSZ, GORASP2, GTF2H1, IL13RA1, MUC16, NMDAR1, Notch1, NT5M, PDCD5, RAD50, and TM9SF2 were identified as possible candidates, among others. We failed to find many laboratory experiments contrasting the effects of Indian and Chinese orthologs at these sites on SIVmac pathogenesis, but future comparative studies might hold fertile ground for research into the biological mechanisms underlying innate resistance to SIVmac in Chinese rhesus macaques. PMID:25884674

  1. Draft Genome Sequence of Lactobacillus delbrueckii subsp. bulgaricus CFL1, a Lactic Acid Bacterium Isolated from French Handcrafted Fermented Milk.

    PubMed

    Meneghel, Julie; Dugat-Bony, Eric; Irlinger, Françoise; Loux, Valentin; Vidal, Marie; Passot, Stéphanie; Béal, Catherine; Layec, Séverine; Fonseca, Fernanda

    2016-01-01

    Lactobacillus delbrueckii subsp. bulgaricus (L. bulgaricus) is a lactic acid bacterium widely used for the production of yogurt and cheeses. Here, we report the genome sequence of L. bulgaricus CFL1 to improve our knowledge on its stress-induced damages following production and end-use processes. PMID:26941141

  2. Draft Genome Sequence of Cutaneotrichosporon curvatus DSM 101032 (Formerly Cryptococcus curvatus), an Oleaginous Yeast Producing Polyunsaturated Fatty Acids.

    PubMed

    Hofmeyer, Thomas; Hackenschmidt, Silke; Nadler, Florian; Thürmer, Andrea; Daniel, Rolf; Kabisch, Johannes

    2016-01-01

    Cutaneotrichosporon curvatus DSM 101032 is an oleaginous yeast that can be isolated from various habitats and is capable of producing substantial amounts of polyunsaturated fatty acids. Here, we present the first draft genome sequence of any C. curvatus species. PMID:27174275

  3. Complete genome sequence of Lactobacillus plantarum ZS2058, a probiotic strain with high conjugated linoleic acid production ability.

    PubMed

    Yang, Bo; Chen, Haiqin