Transcription factor clusters regulate genes in eukaryotic cells
Hedlund, Erik G; Friemann, Rosmarie; Hohmann, Stefan
2017-01-01
Transcription is regulated through binding factors to gene promoters to activate or repress expression, however, the mechanisms by which factors find targets remain unclear. Using single-molecule fluorescence microscopy, we determined in vivo stoichiometry and spatiotemporal dynamics of a GFP tagged repressor, Mig1, from a paradigm signaling pathway of Saccharomyces cerevisiae. We find the repressor operates in clusters, which upon extracellular signal detection, translocate from the cytoplasm, bind to nuclear targets and turnover. Simulations of Mig1 configuration within a 3D yeast genome model combined with a promoter-specific, fluorescent translation reporter confirmed clusters are the functional unit of gene regulation. In vitro and structural analysis on reconstituted Mig1 suggests that clusters are stabilized by depletion forces between intrinsically disordered sequences. We observed similar clusters of a co-regulatory activator from a different pathway, supporting a generalized cluster model for transcription factors that reduces promoter search times through intersegment transfer while stabilizing gene expression. PMID:28841133
Yan, Bin; Yang, Xinping; Lee, Tin-Lap; Friedman, Jay; Tang, Jun; Van Waes, Carter; Chen, Zhong
2007-01-01
Background Differentially expressed gene profiles have previously been observed among pathologically defined cancers by microarray technologies, including head and neck squamous cell carcinomas (HNSCCs). However, the molecular expression signatures and transcriptional regulatory controls that underlie the heterogeneity in HNSCCs are not well defined. Results Genome-wide cDNA microarray profiling of ten HNSCC cell lines revealed novel gene expression signatures that distinguished cancer cell subsets associated with p53 status. Three major clusters of over-expressed genes (A to C) were defined through hierarchical clustering, Gene Ontology, and statistical modeling. The promoters of genes in these clusters exhibited different patterns and prevalence of transcription factor binding sites for p53, nuclear factor-κB (NF-κB), activator protein (AP)-1, signal transducer and activator of transcription (STAT)3 and early growth response (EGR)1, as compared with the frequency in vertebrate promoters. Cluster A genes involved in chromatin structure and function exhibited enrichment for p53 and decreased AP-1 binding sites, whereas clusters B and C, containing cytokine and antiapoptotic genes, exhibited a significant increase in prevalence of NF-κB binding sites. An increase in STAT3 and EGR1 binding sites was distributed among the over-expressed clusters. Novel regulatory modules containing p53 or NF-κB concomitant with other transcription factor binding motifs were identified, and experimental data supported the predicted transcriptional regulation and binding activity. Conclusion The transcription factors p53, NF-κB, and AP-1 may be important determinants of the heterogeneous pattern of gene expression, whereas STAT3 and EGR1 may broadly enhance gene expression in HNSCCs. Defining these novel gene signatures and regulatory mechanisms will be important for establishing new molecular classifications and subtyping, which in turn will promote development of targeted therapeutics for HNSCC. PMID:17498291
Dover, Nir; Barash, Jason R.; Burke, Julianne N.; ...
2014-05-22
Botulinum neurotoxin (BoNT) is the most poisonous substances known and its eight toxin types (A to H) are distinguished by the inability of polyclonal antibodies that neutralize one toxin type to neutralize any of the other seven toxin types. Infant botulism, an intestinal toxemia orphan disease, is the most common form of human botulism in the United States. It results from swallowed spores of Clostridium botulinum (or rarely, neurotoxigenic Clostridium butyricum or Clostridium baratii) that germinate and temporarily colonize the lumen of the large intestine, where, as vegetative cells, they produce botulinum toxin. Botulinum neurotoxin is encoded by the bontmore » gene that is part of a toxin gene cluster that includes several accessory genes. In this paper, we sequenced for the first time the complete botulinum neurotoxin gene cluster of nonproteolytic C. baratii type F7. Like the type E and the nonproteolytic type F6 botulinum toxin gene clusters, the C. baratii type F7 had an orfX toxin gene cluster that lacked the regulatory botR gene which is found in proteolytic C. botulinum strains and codes for an alternative σ factor. In the absence of botR, we identified a putative alternative regulatory gene located upstream of the C. baratii type F7 toxin gene cluster. This putative regulatory gene codes for a predicted σ factor that contains DNA-binding-domain homologues to the DNA-binding domains both of BotR and of other members of the TcdR-related group 5 of the σ 70 family that are involved in the regulation of toxin gene expression in clostridia. We showed that this TcdR-related protein in association with RNA polymerase core enzyme specifically binds to the C. baratii type F7 botulinum toxin gene cluster promoters. Finally, this TcdR-related protein may therefore be involved in regulating the expression of the genes of the botulinum toxin gene cluster in neurotoxigenic C. baratii.« less
Fragmentation of an aflatoxin-like gene cluster in a forest pathogen
USDA-ARS?s Scientific Manuscript database
Secondary metabolic pathway genes are typically clustered in fungi. An exception to this paradigm is seen for genes required for the production of dothistromin, an aflatoxin-like virulence factor produced by the pine needle pathogen Dothistroma septosporum. In contrast to the tight clustering of gen...
Kettle, Andrew J; Carere, Jason; Batley, Jacqueline; Manners, John M; Kazan, Kemal; Gardiner, Donald M
2016-03-01
A number of cereals produce the benzoxazolinone class of phytoalexins. Fusarium species pathogenic towards these hosts can typically degrade these compounds via an aminophenol intermediate, and the ability to do so is encoded by a group of genes found in the Fusarium Detoxification of Benzoxazolinone (FDB) cluster. A zinc finger transcription factor encoded by one of the FDB cluster genes (FDB3) has been proposed to regulate the expression of other genes in the cluster and hence is potentially involved in benzoxazolinone degradation. Herein we show that Fdb3 is essential for the ability of Fusarium pseudograminearum to efficiently detoxify the predominant wheat benzoxazolinone, 6-methoxy-benzoxazolin-2-one (MBOA), but not benzoxazoline-2-one (BOA). Furthermore, additional genes thought to be part of the FDB gene cluster, based upon transcriptional response to benzoxazolinones, are regulated by Fdb3. However, deletion mutants for these latter genes remain capable of benzoxazolinone degradation, suggesting that they are not essential for this process. Crown Copyright © 2016. Published by Elsevier Inc. All rights reserved.
Challis, Rachel C; Araujo, Geisilaine S R; Wong, Edwin K S; Anderson, Holly E; Awan, Atif; Dorman, Anthony M; Waldron, Mary; Wilson, Valerie; Brocklebank, Vicky; Strain, Lisa; Morgan, B Paul; Harris, Claire L; Marchbank, Kevin J; Goodship, Timothy H J; Kavanagh, David
2016-06-01
The regulators of complement activation cluster at chromosome 1q32 contains the complement factor H (CFH) and five complement factor H-related (CFHR) genes. This area of the genome arose from several large genomic duplications, and these low-copy repeats can cause genome instability in this region. Genomic disorders affecting these genes have been described in atypical hemolytic uremic syndrome, arising commonly through nonallelic homologous recombination. We describe a novel CFH/CFHR3 hybrid gene secondary to a de novo 6.3-kb deletion that arose through microhomology-mediated end joining rather than nonallelic homologous recombination. We confirmed a transcript from this hybrid gene and showed a secreted protein product that lacks the recognition domain of factor H and exhibits impaired cell surface complement regulation. The fact that the formation of this hybrid gene arose as a de novo event suggests that this cluster is a dynamic area of the genome in which additional genomic disorders may arise. Copyright © 2016 by the American Society of Nephrology.
Esplin, M Sean; Manuck, Tracy A.; Varner, Michael W.; Christensen, Bryce; Biggio, Joseph; Bukowski, Radek; Parry, Samuel; Zhang, Heping; Huang, Hao; Andrews, William; Saade, George; Sadovsky, Yoel; Reddy, Uma M.; Ilekis, John
2015-01-01
Objective We sought to employ an innovative tool based on common biological pathways to identify specific phenotypes among women with spontaneous preterm birth (SPTB), in order to enhance investigators' ability to identify to highlight common mechanisms and underlying genetic factors responsible for SPTB. Study Design A secondary analysis of a prospective case-control multicenter study of SPTB. All cases delivered a preterm singleton at SPTB ≤34.0 weeks gestation. Each woman was assessed for the presence of underlying SPTB etiologies. A hierarchical cluster analysis was used to identify groups of women with homogeneous phenotypic profiles. One of the phenotypic clusters was selected for candidate gene association analysis using VEGAS software. Results 1028 women with SPTB were assigned phenotypes. Hierarchical clustering of the phenotypes revealed five major clusters. Cluster 1 (N=445) was characterized by maternal stress, cluster 2 (N=294) by premature membrane rupture, cluster 3 (N=120) by familial factors, and cluster 4 (N=63) by maternal comorbidities. Cluster 5 (N=106) was multifactorial, characterized by infection (INF), decidual hemorrhage (DH) and placental dysfunction (PD). These three phenotypes were highly correlated by Chi-square analysis [PD and DH (p<2.2e-6); PD and INF (p=6.2e-10); INF and DH (p=0.0036)]. Gene-based testing identified the INS (insulin) gene as significantly associated with cluster 3 of SPTB. Conclusion We identified 5 major clusters of SPTB based on a phenotype tool and hierarchal clustering. There was significant correlation between several of the phenotypes. The INS gene was associated with familial factors underlying SPTB. PMID:26070700
DOE Office of Scientific and Technical Information (OSTI.GOV)
Berman, Benjamin P.; Pfeiffer, Barret D.; Laverty, Todd R.
2004-08-06
The identification of sequences that control transcription in metazoans is a major goal of genome analysis. In a previous study, we demonstrated that searching for clusters of predicted transcription factor binding sites could discover active regulatory sequences, and identified 37 regions of the Drosophila melanogaster genome with high densities of predicted binding sites for five transcription factors involved in anterior-posterior embryonic patterning. Nine of these clusters overlapped known enhancers. Here, we report the results of in vivo functional analysis of 27 remaining clusters. We generated transgenic flies carrying each cluster attached to a basal promoter and reporter gene, and assayedmore » embryos for reporter gene expression. Six clusters are enhancers of adjacent genes: giant, fushi tarazu, odd-skipped, nubbin, squeeze and pdm2; three drive expression in patterns unrelated to those of neighboring genes; the remaining 18 do not appear to have enhancer activity. We used the Drosophila pseudoobscura genome to compare patterns of evolution in and around the 15 positive and 18 false-positive predictions. Although conservation of primary sequence cannot distinguish true from false positives, conservation of binding-site clustering accurately discriminates functional binding-site clusters from those with no function. We incorporated conservation of binding-site clustering into a new genome-wide enhancer screen, and predict several hundred new regulatory sequences, including 85 adjacent to genes with embryonic patterns. Measuring conservation of sequence features closely linked to function--such as binding-site clustering--makes better use of comparative sequence data than commonly used methods that examine only sequence identity.« less
Horizontal transfer of a large and highly toxic secondary metabolic gene cluster between fungi.
Slot, Jason C; Rokas, Antonis
2011-01-25
Genes involved in intermediary and secondary metabolism in fungi are frequently physically linked or clustered. For example, in Aspergillus nidulans the entire pathway for the production of sterigmatocystin (ST), a highly toxic secondary metabolite and a precursor to the aflatoxins (AF), is located in a ∼54 kb, 23 gene cluster. We discovered that a complete ST gene cluster in Podospora anserina was horizontally transferred from Aspergillus. Phylogenetic analysis shows that most Podospora cluster genes are adjacent to or nested within Aspergillus cluster genes, although the two genera belong to different taxonomic classes. Furthermore, the Podospora cluster is highly conserved in content, sequence, and microsynteny with the Aspergillus ST/AF clusters and its intergenic regions contain 14 putative binding sites for AflR, the transcription factor required for activation of the ST/AF biosynthetic genes. Examination of ∼52,000 Podospora expressed sequence tags identified transcripts for 14 genes in the cluster, with several expressed at multiple life cycle stages. The presence of putative AflR-binding sites and the expression evidence for several cluster genes, coupled with the recent independent discovery of ST production in Podospora [1], suggest that this HGT event probably resulted in a functional cluster. Given the abundance of metabolic gene clusters in fungi, our finding that one of the largest known metabolic gene clusters moved intact between species suggests that such transfers might have significantly contributed to fungal metabolic diversity. PAPERFLICK: Copyright © 2011 Elsevier Ltd. All rights reserved.
Stevens, David Cole; Conway, Kyle R.; Pearce, Nelson; Villegas-Peñaranda, Luis Roberto; Garza, Anthony G.; Boddy, Christopher N.
2013-01-01
Background Heterologous expression of bacterial biosynthetic gene clusters is currently an indispensable tool for characterizing biosynthetic pathways. Development of an effective, general heterologous expression system that can be applied to bioprospecting from metagenomic DNA will enable the discovery of a wealth of new natural products. Methodology We have developed a new Escherichia coli-based heterologous expression system for polyketide biosynthetic gene clusters. We have demonstrated the over-expression of the alternative sigma factor σ54 directly and positively regulates heterologous expression of the oxytetracycline biosynthetic gene cluster in E. coli. Bioinformatics analysis indicates that σ54 promoters are present in nearly 70% of polyketide and non-ribosomal peptide biosynthetic pathways. Conclusions We have demonstrated a new mechanism for heterologous expression of the oxytetracycline polyketide biosynthetic pathway, where high-level pleiotropic sigma factors from the heterologous host directly and positively regulate transcription of the non-native biosynthetic gene cluster. Our bioinformatics analysis is consistent with the hypothesis that heterologous expression mediated by the alternative sigma factor σ54 may be a viable method for the production of additional polyketide products. PMID:23724102
Esplin, M Sean; Manuck, Tracy A; Varner, Michael W; Christensen, Bryce; Biggio, Joseph; Bukowski, Radek; Parry, Samuel; Zhang, Heping; Huang, Hao; Andrews, William; Saade, George; Sadovsky, Yoel; Reddy, Uma M; Ilekis, John
2015-09-01
We sought to use an innovative tool that is based on common biologic pathways to identify specific phenotypes among women with spontaneous preterm birth (SPTB) to enhance investigators' ability to identify and to highlight common mechanisms and underlying genetic factors that are responsible for SPTB. We performed a secondary analysis of a prospective case-control multicenter study of SPTB. All cases delivered a preterm singleton at SPTB ≤34.0 weeks' gestation. Each woman was assessed for the presence of underlying SPTB causes. A hierarchic cluster analysis was used to identify groups of women with homogeneous phenotypic profiles. One of the phenotypic clusters was selected for candidate gene association analysis with the use of VEGAS software. One thousand twenty-eight women with SPTB were assigned phenotypes. Hierarchic clustering of the phenotypes revealed 5 major clusters. Cluster 1 (n = 445) was characterized by maternal stress; cluster 2 (n = 294) was characterized by premature membrane rupture; cluster 3 (n = 120) was characterized by familial factors, and cluster 4 (n = 63) was characterized by maternal comorbidities. Cluster 5 (n = 106) was multifactorial and characterized by infection (INF), decidual hemorrhage (DH), and placental dysfunction (PD). These 3 phenotypes were correlated highly by χ(2) analysis (PD and DH, P < 2.2e-6; PD and INF, P = 6.2e-10; INF and DH, (P = .0036). Gene-based testing identified the INS (insulin) gene as significantly associated with cluster 3 of SPTB. We identified 5 major clusters of SPTB based on a phenotype tool and hierarch clustering. There was significant correlation between several of the phenotypes. The INS gene was associated with familial factors that were underlying SPTB. Copyright © 2015 Elsevier Inc. All rights reserved.
Clustering approaches to identifying gene expression patterns from DNA microarray data.
Do, Jin Hwan; Choi, Dong-Kug
2008-04-30
The analysis of microarray data is essential for large amounts of gene expression data. In this review we focus on clustering techniques. The biological rationale for this approach is the fact that many co-expressed genes are co-regulated, and identifying co-expressed genes could aid in functional annotation of novel genes, de novo identification of transcription factor binding sites and elucidation of complex biological pathways. Co-expressed genes are usually identified in microarray experiments by clustering techniques. There are many such methods, and the results obtained even for the same datasets may vary considerably depending on the algorithms and metrics for dissimilarity measures used, as well as on user-selectable parameters such as desired number of clusters and initial values. Therefore, biologists who want to interpret microarray data should be aware of the weakness and strengths of the clustering methods used. In this review, we survey the basic principles of clustering of DNA microarray data from crisp clustering algorithms such as hierarchical clustering, K-means and self-organizing maps, to complex clustering algorithms like fuzzy clustering.
DOE Office of Scientific and Technical Information (OSTI.GOV)
Berman, Benjamin P.; Pfeiffer, Barret D.; Laverty, Todd R.
2004-08-06
Background The identification of sequences that control transcription in metazoans is a major goal of genome analysis. In a previous study, we demonstrated that searching for clusters of predicted transcription factor binding sites could discover active regulatory sequences, and identified 37 regions of the Drosophila melanogaster genome with high densities of predicted binding sites for five transcription factors involved in anterior-posterior embryonic patterning. Nine of these clusters overlapped known enhancers. Here, we report the results of in vivo functional analysis of 27 remaining clusters. Results We generated transgenic flies carrying each cluster attached to a basal promoter and reporter gene,more » and assayed embryos for reporter gene expression. Six clusters are enhancers of adjacent genes: giant, fushi tarazu, odd-skipped, nubbin, squeeze and pdm2; three drive expression in patterns unrelated to those of neighboring genes; the remaining 18 do not appear to have enhancer activity. We used the Drosophila pseudoobscura genome to compare patterns of evolution in and around the 15 positive and 18 false-positive predictions. Although conservation of primary sequence cannot distinguish true from false positives, conservation of binding-site clustering accurately discriminates functional binding-site clusters from those with no function. We incorporated conservation of binding-site clustering into a new genome-wide enhancer screen, and predict several hundred new regulatory sequences, including 85 adjacent to genes with embryonic patterns. Conclusions Measuring conservation of sequence features closely linked to function - such as binding-site clustering - makes better use of comparative sequence data than commonly used methods that examine only sequence identity.« less
USDA-ARS?s Scientific Manuscript database
An intronless cluster of three class I small heat shock protein (sHSP) chaperone genes, Sl17.6, Sl20.0 and Sl20.1, resident on the short arm of chromosome 6 in tomato, was previously characterized (Goyal et al., 2012). This shsp chaperone gene cluster was found decorated with cis sequences known to ...
Derntl, Christian; Rassinger, Alice; Srebotnik, Ewald; Mach, Robert L.
2016-01-01
ABSTRACT The industrially used ascomycete Trichoderma reesei secretes a typical yellow pigment during cultivation, while other Trichoderma species do not. A comparative genomic analysis suggested that a putative secondary metabolism cluster, containing two polyketide-synthase encoding genes, is responsible for the yellow pigment synthesis. This cluster is conserved in a set of rather distantly related fungi, including Acremonium chrysogenum and Penicillium chrysogenum. In an attempt to silence the cluster in T. reesei, two genes of the cluster encoding transcription factors were individually deleted. For a complete genetic proof-of-function, the genes were reinserted into the genomes of the respective deletion strains. The deletion of the first transcription factor (termed yellow pigment regulator 1 [Ypr1]) resulted in the full abolishment of the yellow pigment formation and the expression of most genes of this cluster. A comparative high-pressure liquid chromatography (HPLC) analysis of supernatants of the ypr1 deletion and its parent strain suggested the presence of several yellow compounds in T. reesei that are all derived from the same cluster. A subsequent gas chromatography/mass spectrometry analysis strongly indicated the presence of sorbicillin in the major HPLC peak. The presence of the second transcription factor, termed yellow pigment regulator 2 (Ypr2), reduces the yellow pigment formation and the expression of most cluster genes, including the gene encoding the activator Ypr1. IMPORTANCE Trichoderma reesei is used for industry-scale production of carbohydrate-active enzymes. During growth, it secretes a typical yellow pigment. This is not favorable for industrial enzyme production because it makes the downstream process more complicated and thus increases operating costs. In this study, we demonstrate which regulators influence the synthesis of the yellow pigment. Based on these data, we also provide indication as to which genes are under the control of these regulators and are finally responsible for the biosynthesis of the yellow pigment. These genes are organized in a cluster that is also found in other industrially relevant fungi, such as the two antibiotic producers Penicillium chrysogenum and Acremonium chrysogenum. The targeted manipulation of a secondary metabolism cluster is an important option for any biotechnologically applied microorganism. PMID:27520818
Clustering change patterns using Fourier transformation with time-course gene expression data.
Kim, Jaehee
2011-01-01
To understand the behavior of genes, it is important to explore how the patterns of gene expression change over a period of time because biologically related gene groups can share the same change patterns. In this study, the problem of finding similar change patterns is induced to clustering with the derivative Fourier coefficients. This work is aimed at discovering gene groups with similar change patterns which share similar biological properties. We developed a statistical model using derivative Fourier coefficients to identify similar change patterns of gene expression. We used a model-based method to cluster the Fourier series estimation of derivatives. We applied our model to cluster change patterns of yeast cell cycle microarray expression data with alpha-factor synchronization. It showed that, as the method clusters with the probability-neighboring data, the model-based clustering with our proposed model yielded biologically interpretable results. We expect that our proposed Fourier analysis with suitably chosen smoothing parameters could serve as a useful tool in classifying genes and interpreting possible biological change patterns.
Surles-Zeigler, Monique C; Li, Yonggang; Distel, Timothy J; Omotayo, Hakeem; Ge, Shaokui; Ford, Byron D
2018-01-01
Ischemic stroke is a major cause of mortality in the United States. We previously showed that neuregulin-1 (NRG1) was neuroprotective in rat models of ischemic stroke. We used gene expression profiling to understand the early cellular and molecular mechanisms of NRG1's effects after the induction of ischemia. Ischemic stroke was induced by middle cerebral artery occlusion (MCAO). Rats were allocated to 3 groups: (1) control, (2) MCAO and (3) MCAO + NRG1. Cortical brain tissues were collected three hours following MCAO and NRG1 treatment and subjected to microarray analysis. Data and statistical analyses were performed using R/Bioconductor platform alongside Genesis, Ingenuity Pathway Analysis and Enrichr software packages. There were 2693 genes differentially regulated following ischemia and NRG1 treatment. These genes were organized by expression patterns into clusters using a K-means clustering algorithm. We further analyzed genes in clusters where ischemia altered gene expression, which was reversed by NRG1 (clusters 4 and 10). NRG1, IRS1, OPA3, and POU6F1 were central linking (node) genes in cluster 4. Conserved Transcription Factor Binding Site Finder (CONFAC) identified ETS-1 as a potential transcriptional regulator of NRG1 suppressed genes following ischemia. A transcription factor activity array showed that ETS-1 activity was increased 2-fold, 3 hours following ischemia and this activity was attenuated by NRG1. These findings reveal key early transcriptional mechanisms associated with neuroprotection by NRG1 in the ischemic penumbra.
Integrating Data Clustering and Visualization for the Analysis of 3D Gene Expression Data
DOE Office of Scientific and Technical Information (OSTI.GOV)
Data Analysis and Visualization; nternational Research Training Group ``Visualization of Large and Unstructured Data Sets,'' University of Kaiserslautern, Germany; Computational Research Division, Lawrence Berkeley National Laboratory, One Cyclotron Road, Berkeley, CA 94720, USA
2008-05-12
The recent development of methods for extracting precise measurements of spatial gene expression patterns from three-dimensional (3D) image data opens the way for new analyses of the complex gene regulatory networks controlling animal development. We present an integrated visualization and analysis framework that supports user-guided data clustering to aid exploration of these new complex datasets. The interplay of data visualization and clustering-based data classification leads to improved visualization and enables a more detailed analysis than previously possible. We discuss (i) integration of data clustering and visualization into one framework; (ii) application of data clustering to 3D gene expression data; (iii)more » evaluation of the number of clusters k in the context of 3D gene expression clustering; and (iv) improvement of overall analysis quality via dedicated post-processing of clustering results based on visualization. We discuss the use of this framework to objectively define spatial pattern boundaries and temporal profiles of genes and to analyze how mRNA patterns are controlled by their regulatory transcription factors.« less
Randise-Hinchliff, Carlo; Coukos, Robert; Sood, Varun; Sumner, Michael Chas; Zdraljevic, Stefan; Meldi Sholl, Lauren; Garvey Brickner, Donna; Ahmed, Sara; Watchmaker, Lauren; Brickner, Jason H
2016-03-14
In budding yeast, targeting of active genes to the nuclear pore complex (NPC) and interchromosomal clustering is mediated by transcription factor (TF) binding sites in the gene promoters. For example, the binding sites for the TFs Put3, Ste12, and Gcn4 are necessary and sufficient to promote positioning at the nuclear periphery and interchromosomal clustering. However, in all three cases, gene positioning and interchromosomal clustering are regulated. Under uninducing conditions, local recruitment of the Rpd3(L) histone deacetylase by transcriptional repressors blocks Put3 DNA binding. This is a general function of yeast repressors: 16 of 21 repressors blocked Put3-mediated subnuclear positioning; 11 of these required Rpd3. In contrast, Ste12-mediated gene positioning is regulated independently of DNA binding by mitogen-activated protein kinase phosphorylation of the Dig2 inhibitor, and Gcn4-dependent targeting is up-regulated by increasing Gcn4 protein levels. These different regulatory strategies provide either qualitative switch-like control or quantitative control of gene positioning over different time scales. © 2016 Randise-Hinchliff et al.
Blanco-Rojo, Ruth; Toxqui, Laura; López-Parra, Ana M; Baeza-Richer, Carlos; Pérez-Granados, Ana M; Arroyo-Pardo, Eduardo; Vaquero, M Pilar
2014-03-06
The aim of this study was to investigate the combined influence of diet, menstruation and genetic factors on iron status in Spanish menstruating women (n = 142). Dietary intake was assessed by a 72-h detailed dietary report and menstrual blood loss by a questionnaire, to determine a Menstrual Blood Loss Coefficient (MBLC). Five selected SNPs were genotyped: rs3811647, rs1799852 (Tf gene); rs1375515 (CACNA2D3 gene); and rs1800562 and rs1799945 (HFE gene, mutations C282Y and H63D, respectively). Iron biomarkers were determined and cluster analysis was performed. Differences among clusters in dietary intake, menstrual blood loss parameters and genotype frequencies distribution were studied. A categorical regression was performed to identify factors associated with cluster belonging. Three clusters were identified: women with poor iron status close to developing iron deficiency anemia (Cluster 1, n = 26); women with mild iron deficiency (Cluster 2, n = 59) and women with normal iron status (Cluster 3, n = 57). Three independent factors, red meat consumption, MBLC and mutation C282Y, were included in the model that better explained cluster belonging (R2 = 0.142, p < 0.001). In conclusion, the combination of high red meat consumption, low menstrual blood loss and the HFE C282Y mutation may protect from iron deficiency in women of childbearing age. These findings could be useful to implement adequate strategies to prevent iron deficiency anemia.
Ryge, Jesper; Winther, Ole; Wienecke, Jacob; Sandelin, Albin; Westerdahl, Ann-Charlotte; Hultborn, Hans; Kiehn, Ole
2010-06-09
Spinal cord injury leads to neurological dysfunctions affecting the motor, sensory as well as the autonomic systems. Increased excitability of motor neurons has been implicated in injury-induced spasticity, where the reappearance of self-sustained plateau potentials in the absence of modulatory inputs from the brain correlates with the development of spasticity. Here we examine the dynamic transcriptional response of motor neurons to spinal cord injury as it evolves over time to unravel common gene expression patterns and their underlying regulatory mechanisms. For this we use a rat-tail-model with complete spinal cord transection causing injury-induced spasticity, where gene expression profiles are obtained from labeled motor neurons extracted with laser microdissection 0, 2, 7, 21 and 60 days post injury. Consensus clustering identifies 12 gene clusters with distinct time expression profiles. Analysis of these gene clusters identifies early immunological/inflammatory and late developmental responses as well as a regulation of genes relating to neuron excitability that support the development of motor neuron hyper-excitability and the reappearance of plateau potentials in the late phase of the injury response. Transcription factor motif analysis identifies differentially expressed transcription factors involved in the regulation of each gene cluster, shaping the expression of the identified biological processes and their associated genes underlying the changes in motor neuron excitability. This analysis provides important clues to the underlying mechanisms of transcriptional regulation responsible for the increased excitability observed in motor neurons in the late chronic phase of spinal cord injury suggesting alternative targets for treatment of spinal cord injury. Several transcription factors were identified as potential regulators of gene clusters containing elements related to motor neuron hyper-excitability, the manipulation of which potentially could be used to alter the transcriptional response to prevent the motor neurons from entering a state of hyper-excitability.
Transcriptional regulation of gene expression clusters in motor neurons following spinal cord injury
2010-01-01
Background Spinal cord injury leads to neurological dysfunctions affecting the motor, sensory as well as the autonomic systems. Increased excitability of motor neurons has been implicated in injury-induced spasticity, where the reappearance of self-sustained plateau potentials in the absence of modulatory inputs from the brain correlates with the development of spasticity. Results Here we examine the dynamic transcriptional response of motor neurons to spinal cord injury as it evolves over time to unravel common gene expression patterns and their underlying regulatory mechanisms. For this we use a rat-tail-model with complete spinal cord transection causing injury-induced spasticity, where gene expression profiles are obtained from labeled motor neurons extracted with laser microdissection 0, 2, 7, 21 and 60 days post injury. Consensus clustering identifies 12 gene clusters with distinct time expression profiles. Analysis of these gene clusters identifies early immunological/inflammatory and late developmental responses as well as a regulation of genes relating to neuron excitability that support the development of motor neuron hyper-excitability and the reappearance of plateau potentials in the late phase of the injury response. Transcription factor motif analysis identifies differentially expressed transcription factors involved in the regulation of each gene cluster, shaping the expression of the identified biological processes and their associated genes underlying the changes in motor neuron excitability. Conclusions This analysis provides important clues to the underlying mechanisms of transcriptional regulation responsible for the increased excitability observed in motor neurons in the late chronic phase of spinal cord injury suggesting alternative targets for treatment of spinal cord injury. Several transcription factors were identified as potential regulators of gene clusters containing elements related to motor neuron hyper-excitability, the manipulation of which potentially could be used to alter the transcriptional response to prevent the motor neurons from entering a state of hyper-excitability. PMID:20534130
Janevska, Slavica; Arndt, Birgit; Baumann, Leonie; Apken, Lisa Helene; Mauriz Marques, Lucas Maciel; Humpf, Hans-Ulrich; Tudzynski, Bettina
2017-01-01
The PKS-NRPS-derived tetramic acid equisetin and its N-desmethyl derivative trichosetin exhibit remarkable biological activities against a variety of organisms, including plants and bacteria, e.g., Staphylococcus aureus. The equisetin biosynthetic gene cluster was first described in Fusarium heterosporum, a species distantly related to the notorious rice pathogen Fusarium fujikuroi. Here we present the activation and characterization of a homologous, but silent, gene cluster in F. fujikuroi. Bioinformatic analysis revealed that this cluster does not contain the equisetin N-methyltransferase gene eqxD and consequently, trichosetin was isolated as final product. The adaption of the inducible, tetracycline-dependent Tet-on promoter system from Aspergillus niger achieved a controlled overproduction of this toxic metabolite and a functional characterization of each cluster gene in F. fujikuroi. Overexpression of one of the two cluster-specific transcription factor (TF) genes, TF22, led to an activation of the three biosynthetic cluster genes, including the PKS-NRPS key gene. In contrast, overexpression of TF23, encoding a second Zn(II)2Cys6 TF, did not activate adjacent cluster genes. Instead, TF23 was induced by the final product trichosetin and was required for expression of the transporter-encoding gene MFS-T. TF23 and MFS-T likely act in consort and contribute to detoxification of trichosetin and therefore, self-protection of the producing fungus. PMID:28379186
Janevska, Slavica; Arndt, Birgit; Baumann, Leonie; Apken, Lisa Helene; Mauriz Marques, Lucas Maciel; Humpf, Hans-Ulrich; Tudzynski, Bettina
2017-04-05
The PKS-NRPS-derived tetramic acid equisetin and its N -desmethyl derivative trichosetin exhibit remarkable biological activities against a variety of organisms, including plants and bacteria, e.g., Staphylococcus aureus . The equisetin biosynthetic gene cluster was first described in Fusarium heterosporum , a species distantly related to the notorious rice pathogen Fusarium fujikuroi . Here we present the activation and characterization of a homologous, but silent, gene cluster in F. fujikuroi . Bioinformatic analysis revealed that this cluster does not contain the equisetin N -methyltransferase gene eqxD and consequently, trichosetin was isolated as final product. The adaption of the inducible, tetracycline-dependent Tet-on promoter system from Aspergillus niger achieved a controlled overproduction of this toxic metabolite and a functional characterization of each cluster gene in F. fujikuroi . Overexpression of one of the two cluster-specific transcription factor (TF) genes, TF22 , led to an activation of the three biosynthetic cluster genes, including the PKS-NRPS key gene. In contrast, overexpression of TF23 , encoding a second Zn(II)₂Cys₆ TF, did not activate adjacent cluster genes. Instead, TF23 was induced by the final product trichosetin and was required for expression of the transporter-encoding gene MFS-T . TF23 and MFS-T likely act in consort and contribute to detoxification of trichosetin and therefore, self-protection of the producing fungus.
Li, Liangtao; Miao, Ren; Bertram, Sophie; Jia, Xuan; Ward, Diane M.; Kaplan, Jerry
2012-01-01
Yeast respond to increased cytosolic iron by activating the transcription factor Yap5 increasing transcription of CCC1, which encodes a vacuolar iron importer. Using a genetic screen to identify genes involved in Yap5 iron sensing, we discovered that a mutation in SSQ1, which encodes a mitochondrial chaperone involved in iron-sulfur cluster synthesis, prevented expression of Yap5 target genes. We demonstrated that mutation or reduced expression of other genes involved in mitochondrial iron-sulfur cluster synthesis (YFH1, ISU1) prevented induction of the Yap5 response. We took advantage of the iron-dependent catalytic activity of Pseudaminobacter salicylatoxidans gentisate 1,2-dioxygenase expressed in yeast to measure changes in cytosolic iron. We determined that reductions in iron-sulfur cluster synthesis did not affect the activity of cytosolic gentisate 1,2-dioxygenase. We show that loss of activity of the cytosolic iron-sulfur cluster assembly complex proteins or deletion of cytosolic glutaredoxins did not reduce expression of Yap5 target genes. These results suggest that the high iron transcriptional response, as well as the low iron transcriptional response, senses iron-sulfur clusters. PMID:22915593
Bhattacharya, Anindya; De, Rajat K
2010-08-01
Distance based clustering algorithms can group genes that show similar expression values under multiple experimental conditions. They are unable to identify a group of genes that have similar pattern of variation in their expression values. Previously we developed an algorithm called divisive correlation clustering algorithm (DCCA) to tackle this situation, which is based on the concept of correlation clustering. But this algorithm may also fail for certain cases. In order to overcome these situations, we propose a new clustering algorithm, called average correlation clustering algorithm (ACCA), which is able to produce better clustering solution than that produced by some others. ACCA is able to find groups of genes having more common transcription factors and similar pattern of variation in their expression values. Moreover, ACCA is more efficient than DCCA with respect to the time of execution. Like DCCA, we use the concept of correlation clustering concept introduced by Bansal et al. ACCA uses the correlation matrix in such a way that all genes in a cluster have the highest average correlation values with the genes in that cluster. We have applied ACCA and some well-known conventional methods including DCCA to two artificial and nine gene expression datasets, and compared the performance of the algorithms. The clustering results of ACCA are found to be more significantly relevant to the biological annotations than those of the other methods. Analysis of the results show the superiority of ACCA over some others in determining a group of genes having more common transcription factors and with similar pattern of variation in their expression profiles. Availability of the software: The software has been developed using C and Visual Basic languages, and can be executed on the Microsoft Windows platforms. The software may be downloaded as a zip file from http://www.isical.ac.in/~rajat. Then it needs to be installed. Two word files (included in the zip file) need to be consulted before installation and execution of the software. Copyright 2010 Elsevier Inc. All rights reserved.
Co-clustering phenome–genome for phenotype classification and disease gene discovery
Hwang, TaeHyun; Atluri, Gowtham; Xie, MaoQiang; Dey, Sanjoy; Hong, Changjin; Kumar, Vipin; Kuang, Rui
2012-01-01
Understanding the categorization of human diseases is critical for reliably identifying disease causal genes. Recently, genome-wide studies of abnormal chromosomal locations related to diseases have mapped >2000 phenotype–gene relations, which provide valuable information for classifying diseases and identifying candidate genes as drug targets. In this article, a regularized non-negative matrix tri-factorization (R-NMTF) algorithm is introduced to co-cluster phenotypes and genes, and simultaneously detect associations between the detected phenotype clusters and gene clusters. The R-NMTF algorithm factorizes the phenotype–gene association matrix under the prior knowledge from phenotype similarity network and protein–protein interaction network, supervised by the label information from known disease classes and biological pathways. In the experiments on disease phenotype–gene associations in OMIM and KEGG disease pathways, R-NMTF significantly improved the classification of disease phenotypes and disease pathway genes compared with support vector machines and Label Propagation in cross-validation on the annotated phenotypes and genes. The newly predicted phenotypes in each disease class are highly consistent with human phenotype ontology annotations. The roles of the new member genes in the disease pathways are examined and validated in the protein–protein interaction subnetworks. Extensive literature review also confirmed many new members of the disease classes and pathways as well as the predicted associations between disease phenotype classes and pathways. PMID:22735708
Conditions for the Evolution of Gene Clusters in Bacterial Genomes
Ballouz, Sara; Francis, Andrew R.; Lan, Ruiting; Tanaka, Mark M.
2010-01-01
Genes encoding proteins in a common pathway are often found near each other along bacterial chromosomes. Several explanations have been proposed to account for the evolution of these structures. For instance, natural selection may directly favour gene clusters through a variety of mechanisms, such as increased efficiency of coregulation. An alternative and controversial hypothesis is the selfish operon model, which asserts that clustered arrangements of genes are more easily transferred to other species, thus improving the prospects for survival of the cluster. According to another hypothesis (the persistence model), genes that are in close proximity are less likely to be disrupted by deletions. Here we develop computational models to study the conditions under which gene clusters can evolve and persist. First, we examine the selfish operon model by re-implementing the simulation and running it under a wide range of conditions. Second, we introduce and study a Moran process in which there is natural selection for gene clustering and rearrangement occurs by genome inversion events. Finally, we develop and study a model that includes selection and inversion, which tracks the occurrence and fixation of rearrangements. Surprisingly, gene clusters fail to evolve under a wide range of conditions. Factors that promote the evolution of gene clusters include a low number of genes in the pathway, a high population size, and in the case of the selfish operon model, a high horizontal transfer rate. The computational analysis here has shown that the evolution of gene clusters can occur under both direct and indirect selection as long as certain conditions hold. Under these conditions the selfish operon model is still viable as an explanation for the evolution of gene clusters. PMID:20168992
Furuya, Toshiki; Hirose, Satomi; Semba, Hisashi; Kino, Kuniki
2011-01-01
The mimABCD gene cluster encodes the binuclear iron monooxygenase that oxidizes propane and phenol in Mycobacterium smegmatis strain MC2 155 and Mycobacterium goodii strain 12523. Interestingly, expression of the mimABCD gene cluster is induced by acetone. In this study, we investigated the regulator gene responsible for this acetone-responsive expression. In the genome sequence of M. smegmatis strain MC2 155, the mimABCD gene cluster is preceded by a gene designated mimR, which is divergently transcribed. Sequence analysis revealed that MimR exhibits amino acid similarity with the NtrC family of transcriptional activators, including AcxR and AcoR, which are involved in acetone and acetoin metabolism, respectively. Unexpectedly, many homologs of the mimR gene were also found in the sequenced genomes of actinomycetes. A plasmid carrying a transcriptional fusion of the intergenic region between the mimR and mimA genes with a promoterless green fluorescent protein (GFP) gene was constructed and introduced into M. smegmatis strain MC2 155. Using a GFP reporter system, we confirmed by deletion and complementation analyses that the mimR gene product is the positive regulator of the mimABCD gene cluster expression that is responsive to acetone. M. goodii strain 12523 also utilized the same regulatory system as M. smegmatis strain MC2 155. Although transcriptional activators of the NtrC family generally control transcription using the σ54 factor, a gene encoding the σ54 factor was absent from the genome sequence of M. smegmatis strain MC2 155. These results suggest the presence of a novel regulatory system in actinomycetes, including mycobacteria. PMID:21856847
Suzuki, Kuta; Tanaka, Mizuki; Konno, Yui; Ichikawa, Takanori; Ichinose, Sakurako; Hasegawa-Shiro, Sachiko; Shintani, Takahiro; Gomi, Katsuya
2015-02-01
The production of amylolytic enzymes in Aspergillus oryzae is induced in the presence of starch or maltose, and two Zn2Cys6-type transcription factors, AmyR and MalR, are involved in this regulation. AmyR directly regulates the expression of amylase genes, and MalR controls the expression of maltose-utilizing (MAL) cluster genes. Deletion of malR gene resulted in poor growth on starch medium and reduction in α-amylase production level. To elucidate the activation mechanisms of these two transcription factors in amylase production, the expression profiles of amylases and MAL cluster genes under carbon catabolite derepression condition and subcellular localization of these transcription factors fused with a green fluorescent protein (GFP) were examined. Glucose, maltose, and isomaltose induced the expression of amylase genes, and GFP-AmyR was translocated from the cytoplasm to nucleus after the addition of these sugars. Rapid induction of amylase gene expression and nuclear localization of GFP-AmyR by isomaltose suggested that this sugar was the strongest inducer for AmyR activation. In contrast, GFP-MalR was constitutively localized in the nucleus and the expression of MAL cluster genes was induced by maltose, but not by glucose or isomaltose. In the presence of maltose, the expression of amylase genes was preceded by MAL cluster gene expression. Furthermore, deletion of the malR gene resulted in a significant decrease in the α-amylase activity induced by maltose, but had apparently no effect on the expression of α-amylase genes in the presence of isomaltose. These results suggested that activation of AmyR and MalR is regulated in a different manner, and the preceding activation of MalR is essential for the utilization of maltose as an inducer for AmyR activation.
1988-01-01
We report the organization of the human genes encoding the complement components C4-binding protein (C4BP), C3b/C4b receptor (CR1), decay accelerating factor (DAF), and C3dg receptor (CR2) within the regulator of complement activation (RCA) gene cluster. Using pulsed field gel electrophoresis analysis these genes have been physically linked and aligned as CR1-CR2-DAF-C4BP in an 800-kb DNA segment. The very tight linkage between the CR1 and the C4BP loci, contrasted with the relative long DNA distance between these genes, suggests the existence of mechanisms interfering with recombination within the RCA gene cluster. PMID:2450163
Zaag, Rim; Tamby, Jean Philippe; Guichard, Cécile; Tariq, Zakia; Rigaill, Guillem; Delannoy, Etienne; Renou, Jean-Pierre; Balzergue, Sandrine; Mary-Huard, Tristan; Aubourg, Sébastien; Martin-Magniette, Marie-Laure; Brunaud, Véronique
2015-01-01
CATdb (http://urgv.evry.inra.fr/CATdb) is a database providing a public access to a large collection of transcriptomic data, mainly for Arabidopsis but also for other plants. This resource has the rare advantage to contain several thousands of microarray experiments obtained with the same technical protocol and analyzed by the same statistical pipelines. In this paper, we present GEM2Net, a new module of CATdb that takes advantage of this homogeneous dataset to mine co-expression units and decipher Arabidopsis gene functions. GEM2Net explores 387 stress conditions organized into 18 biotic and abiotic stress categories. For each one, a model-based clustering is applied on expression differences to identify clusters of co-expressed genes. To characterize functions associated with these clusters, various resources are analyzed and integrated: Gene Ontology, subcellular localization of proteins, Hormone Families, Transcription Factor Families and a refined stress-related gene list associated to publications. Exploiting protein-protein interactions and transcription factors-targets interactions enables to display gene networks. GEM2Net presents the analysis of the 18 stress categories, in which 17,264 genes are involved and organized within 681 co-expression clusters. The meta-data analyses were stored and organized to compose a dynamic Web resource. © The Author(s) 2014. Published by Oxford University Press on behalf of Nucleic Acids Research.
Hox gene clusters in the Indonesian coelacanth, Latimeria menadoensis
Koh, Esther G. L.; Lam, Kevin; Christoffels, Alan; Erdmann, Mark V.; Brenner, Sydney; Venkatesh, Byrappa
2003-01-01
The Hox genes encode transcription factors that play a key role in specifying body plans of metazoans. They are organized into clusters that contain up to 13 paralogue group members. The complex morphology of vertebrates has been attributed to the duplication of Hox clusters during vertebrate evolution. In contrast to the single Hox cluster in the amphioxus (Branchiostoma floridae), an invertebrate-chordate, mammals have four clusters containing 39 Hox genes. Ray-finned fishes (Actinopterygii) such as zebrafish and fugu possess more than four Hox clusters. The coelacanth occupies a basal phylogenetic position among lobe-finned fishes (Sarcopterygii), which gave rise to the tetrapod lineage. The lobe fins of sarcopterygians are considered to be the evolutionary precursors of tetrapod limbs. Thus, the characterization of Hox genes in the coelacanth should provide insights into the origin of tetrapod limbs. We have cloned the complete second exon of 33 Hox genes from the Indonesian coelacanth, Latimeria menadoensis, by extensive PCR survey and genome walking. Phylogenetic analysis shows that 32 of these genes have orthologs in the four mammalian HOX clusters, including three genes (HoxA6, D1, and D8) that are absent in ray-finned fishes. The remaining coelacanth gene is an ortholog of hoxc1 found in zebrafish but absent in mammals. Our results suggest that coelacanths have four Hox clusters bearing a gene complement more similar to mammals than to ray-finned fishes, but with an additional gene, HoxC1, which has been lost during the evolution of mammals from lobe-finned fishes. PMID:12547909
Hox gene clusters in the Indonesian coelacanth, Latimeria menadoensis.
Koh, Esther G L; Lam, Kevin; Christoffels, Alan; Erdmann, Mark V; Brenner, Sydney; Venkatesh, Byrappa
2003-02-04
The Hox genes encode transcription factors that play a key role in specifying body plans of metazoans. They are organized into clusters that contain up to 13 paralogue group members. The complex morphology of vertebrates has been attributed to the duplication of Hox clusters during vertebrate evolution. In contrast to the single Hox cluster in the amphioxus (Branchiostoma floridae), an invertebrate-chordate, mammals have four clusters containing 39 Hox genes. Ray-finned fishes (Actinopterygii) such as zebrafish and fugu possess more than four Hox clusters. The coelacanth occupies a basal phylogenetic position among lobe-finned fishes (Sarcopterygii), which gave rise to the tetrapod lineage. The lobe fins of sarcopterygians are considered to be the evolutionary precursors of tetrapod limbs. Thus, the characterization of Hox genes in the coelacanth should provide insights into the origin of tetrapod limbs. We have cloned the complete second exon of 33 Hox genes from the Indonesian coelacanth, Latimeria menadoensis, by extensive PCR survey and genome walking. Phylogenetic analysis shows that 32 of these genes have orthologs in the four mammalian HOX clusters, including three genes (HoxA6, D1, and D8) that are absent in ray-finned fishes. The remaining coelacanth gene is an ortholog of hoxc1 found in zebrafish but absent in mammals. Our results suggest that coelacanths have four Hox clusters bearing a gene complement more similar to mammals than to ray-finned fishes, but with an additional gene, HoxC1, which has been lost during the evolution of mammals from lobe-finned fishes.
Holland, Peter W H
2013-01-01
Many homeobox genes encode transcription factors with regulatory roles in animal and plant development. Homeobox genes are found in almost all eukaryotes, and have diversified into 11 gene classes and over 100 gene families in animal evolution, and 10 to 14 gene classes in plants. The largest group in animals is the ANTP class which includes the well-known Hox genes, plus other genes implicated in development including ParaHox (Cdx, Xlox, Gsx), Evx, Dlx, En, NK4, NK3, Msx, and Nanog. Genomic data suggest that the ANTP class diversified by extensive tandem duplication to generate a large array of genes, including an NK gene cluster and a hypothetical ProtoHox gene cluster that duplicated to generate Hox and ParaHox genes. Expression and functional data suggest that NK, Hox, and ParaHox gene clusters acquired distinct roles in patterning the mesoderm, nervous system, and gut. The PRD class is also diverse and includes Pax2/5/8, Pax3/7, Pax4/6, Gsc, Hesx, Otx, Otp, and Pitx genes. PRD genes are not generally arranged in ancient genomic clusters, although the Dux, Obox, and Rhox gene clusters arose in mammalian evolution as did several non-clustered PRD genes. Tandem duplication and genome duplication expanded the number of homeobox genes, possibly contributing to the evolution of developmental complexity, but homeobox gene loss must not be ignored. Evolutionary changes to homeobox gene expression have also been documented, including Hox gene expression patterns shifting in concert with segmental diversification in vertebrates and crustaceans, and deletion of a Pitx1 gene enhancer in pelvic-reduced sticklebacks. WIREs Dev Biol 2013, 2:31-45. doi: 10.1002/wdev.78 For further resources related to this article, please visit the WIREs website. The author declares that he has no conflicts of interest. Copyright © 2012 Wiley Periodicals, Inc.
The human RHOX gene cluster: target genes and functional analysis of gene variants in infertile men.
Borgmann, Jennifer; Tüttelmann, Frank; Dworniczak, Bernd; Röpke, Albrecht; Song, Hye-Won; Kliesch, Sabine; Wilkinson, Miles F; Laurentino, Sandra; Gromoll, Jörg
2016-11-15
The X-linked reproductive homeobox (RHOX) gene cluster encodes transcription factors preferentially expressed in reproductive tissues. This gene cluster has important roles in male fertility based on phenotypic defects of Rhox-mutant mice and the finding that aberrant RHOX promoter methylation is strongly associated with abnormal human sperm parameters. However, little is known about the molecular mechanism of RHOX function in humans. Using gene expression profiling, we identified genes regulated by members of the human RHOX gene cluster. Some genes were uniquely regulated by RHOXF1 or RHOXF2/2B, while others were regulated by both of these transcription factors. Several of these regulated genes encode proteins involved in processes relevant to spermatogenesis; e.g. stress protection and cell survival. One of the target genes of RHOXF2/2B is RHOXF1, suggesting cross-regulation to enhance transcriptional responses. The potential role of RHOX in human infertility was addressed by sequencing all RHOX exons in a group of 250 patients with severe oligozoospermia. This revealed two mutations in RHOXF1 (c.515G > A and c.522C > T) and four in RHOXF2/2B (-73C > G, c.202G > A, c.411C > T and c.679G > A), of which only one (c.202G > A) was found in a control group of men with normal sperm concentration. Functional analysis demonstrated that c.202G > A and c.679G > A significantly impaired the ability of RHOXF2/2B to regulate downstream genes. Molecular modelling suggested that these mutations alter RHOXF2/F2B protein conformation. By combining clinical data with in vitro functional analysis, we demonstrate how the X-linked RHOX gene cluster may function in normal human spermatogenesis and we provide evidence that it is impaired in human male fertility.
Janevska, Slavica; Tudzynski, Bettina
2018-01-01
The fungus Fusarium fujikuroi causes bakanae disease of rice due to its ability to produce the plant hormones, the gibberellins. The fungus is also known for producing harmful mycotoxins (e.g., fusaric acid and fusarins) and pigments (e.g., bikaverin and fusarubins). However, for a long time, most of these well-known products could not be linked to biosynthetic gene clusters. Recent genome sequencing has revealed altogether 47 putative gene clusters. Most of them were orphan clusters for which the encoded natural product(s) were unknown. In this review, we describe the current status of our research on identification and functional characterizations of novel secondary metabolite gene clusters. We present several examples where linking known metabolites to the respective biosynthetic genes has been achieved and describe recent strategies and methods to access new natural products, e.g., by genetic manipulation of pathway-specific or global transcritption factors. In addition, we demonstrate that deletion and over-expression of histone-modifying genes is a powerful tool to activate silent gene clusters and to discover their products.
A mixture model-based approach to the clustering of microarray expression data.
McLachlan, G J; Bean, R W; Peel, D
2002-03-01
This paper introduces the software EMMIX-GENE that has been developed for the specific purpose of a model-based approach to the clustering of microarray expression data, in particular, of tissue samples on a very large number of genes. The latter is a nonstandard problem in parametric cluster analysis because the dimension of the feature space (the number of genes) is typically much greater than the number of tissues. A feasible approach is provided by first selecting a subset of the genes relevant for the clustering of the tissue samples by fitting mixtures of t distributions to rank the genes in order of increasing size of the likelihood ratio statistic for the test of one versus two components in the mixture model. The imposition of a threshold on the likelihood ratio statistic used in conjunction with a threshold on the size of a cluster allows the selection of a relevant set of genes. However, even this reduced set of genes will usually be too large for a normal mixture model to be fitted directly to the tissues, and so the use of mixtures of factor analyzers is exploited to reduce effectively the dimension of the feature space of genes. The usefulness of the EMMIX-GENE approach for the clustering of tissue samples is demonstrated on two well-known data sets on colon and leukaemia tissues. For both data sets, relevant subsets of the genes are able to be selected that reveal interesting clusterings of the tissues that are either consistent with the external classification of the tissues or with background and biological knowledge of these sets. EMMIX-GENE is available at http://www.maths.uq.edu.au/~gjm/emmix-gene/
Qurrat-ul-Ain; Seemab, Umair; Nawaz, Sulaman; Rashid, Sajid
2011-01-01
In human, WNT gene clusters are highly conserved at specie level and associated with carcinogenesis. Among them, WNT-10A and WNT-6 genes clustered in chromosome 2q35 are homologous to WNT-10B and WNT-1 located in chromosome 12q13, respectively. In an attempt to study co-regulation, the coordinated expression of these genes was monitored in human breast cancer tissues. As compared to normal tissue, both WNT-10A and WNT-10B genes exhibited lower expression while WNT-6 and WNT-1 showed increased expression in breast cancer tissues. The co-expression pattern was elaborated by detailed phylogenetic and syntenic analyses. Moreover, the intergenic and intragenic regions for these gene clusters were analyzed for studying the transcriptional regulation. In this context, adequate conserved binding sites for SOX and TCF family of transcriptional factors were observed. We propose that SOX9 and TCF4 may compete for binding at the promoters of WNT family genes thus regulating the disease phenotype. PMID:22355234
Marenholz, Ingo; Grosche, Sarah; Kalb, Birgit; Rüschendorf, Franz; Blümchen, Katharina; Schlags, Rupert; Harandi, Neda; Price, Mareike; Hansen, Gesine; Seidenberg, Jürgen; Röblitz, Holger; Yürek, Songül; Tschirner, Sebastian; Hong, Xiumei; Wang, Xiaobin; Homuth, Georg; Schmidt, Carsten O; Nöthen, Markus M; Hübner, Norbert; Niggemann, Bodo; Beyer, Kirsten; Lee, Young-Ae
2017-10-20
Genetic factors and mechanisms underlying food allergy are largely unknown. Due to heterogeneity of symptoms a reliable diagnosis is often difficult to make. Here, we report a genome-wide association study on food allergy diagnosed by oral food challenge in 497 cases and 2387 controls. We identify five loci at genome-wide significance, the clade B serpin (SERPINB) gene cluster at 18q21.3, the cytokine gene cluster at 5q31.1, the filaggrin gene, the C11orf30/LRRC32 locus, and the human leukocyte antigen (HLA) region. Stratifying the results for the causative food demonstrates that association of the HLA locus is peanut allergy-specific whereas the other four loci increase the risk for any food allergy. Variants in the SERPINB gene cluster are associated with SERPINB10 expression in leukocytes. Moreover, SERPINB genes are highly expressed in the esophagus. All identified loci are involved in immunological regulation or epithelial barrier function, emphasizing the role of both mechanisms in food allergy.
Neubauer, Lisa; Dopstadt, Julian; Humpf, Hans-Ulrich; Tudzynski, Paul
2016-01-01
Claviceps purpurea is a phytopathogenic fungus infecting a broad range of grasses including economically important cereal crop plants. The infection cycle ends with the formation of the typical purple-black pigmented sclerotia containing the toxic ergot alkaloids. Besides these ergot alkaloids little is known about the secondary metabolism of the fungus. Red anthraquinone derivatives and yellow xanthone dimers (ergochromes) have been isolated from sclerotia and described as ergot pigments, but the corresponding gene cluster has remained unknown. Fungal pigments gain increasing interest for example as environmentally friendly alternatives to existing dyes. Furthermore, several pigments show biological activities and may have some pharmaceutical value. This study identified the gene cluster responsible for the synthesis of the ergot pigments. Overexpression of the cluster-specific transcription factor led to activation of the gene cluster and to the production of several known ergot pigments. Knock out of the cluster key enzyme, a nonreducing polyketide synthase, clearly showed that this cluster is responsible for the production of red anthraquinones as well as yellow ergochromes. Furthermore, a tentative biosynthetic pathway for the ergot pigments is proposed. By changing the culture conditions, pigment production was activated in axenic culture so that high concentration of phosphate and low concentration of sucrose induced pigment syntheses. This is the first functional analysis of a secondary metabolite gene cluster in the ergot fungus besides that for the classical ergot alkaloids. We demonstrated that this gene cluster is responsible for the typical purple-black color of the ergot sclerotia and showed that the red and yellow ergot pigments are products of the same biosynthetic pathway. Activation of the gene cluster in axenic culture opened up new possibilities for biotechnological applications like the dye production or the development of new pharmaceuticals.
Boyce, Kylie J.; McLauchlan, Alisha; Schreider, Lena; Andrianopoulos, Alex
2015-01-01
During infection, pathogens must utilise the available nutrient sources in order to grow while simultaneously evading or tolerating the host’s defence systems. Amino acids are an important nutritional source for pathogenic fungi and can be assimilated from host proteins to provide both carbon and nitrogen. The hpdA gene of the dimorphic fungus Penicillium marneffei, which encodes an enzyme which catalyses the second step of tyrosine catabolism, was identified as up-regulated in pathogenic yeast cells. As well as enabling the fungus to acquire carbon and nitrogen, tyrosine is also a precursor in the formation of two types of protective melanin; DOPA melanin and pyomelanin. Chemical inhibition of HpdA in P. marneffei inhibits ex vivo yeast cell production suggesting that tyrosine is a key nutrient source during infectious growth. The genes required for tyrosine catabolism, including hpdA, are located in a gene cluster and the expression of these genes is induced in the presence of tyrosine. A gene (hmgR) encoding a Zn(II)2-Cys6 binuclear cluster transcription factor is present within the cluster and is required for tyrosine induced expression and repression in the presence of a preferred nitrogen source. AreA, the GATA-type transcription factor which regulates the global response to limiting nitrogen conditions negatively regulates expression of cluster genes in the absence of tyrosine and is required for nitrogen metabolite repression. Deletion of the tyrosine catabolic genes in the cluster affects growth on tyrosine as either a nitrogen or carbon source and affects pyomelanin, but not DOPA melanin, production. In contrast to other genes of the tyrosine catabolic cluster, deletion of hpdA results in no growth within macrophages. This suggests that the ability to catabolise tyrosine is not required for macrophage infection and that HpdA has an additional novel role to that of tyrosine catabolism and pyomelanin production during growth in host cells. PMID:25812137
Complete Genome Sequence and Comparative Analysis of the Fish Pathogen Lactococcus garvieae
Oshima, Kenshiro; Yoshizaki, Mariko; Kawanishi, Michiko; Nakaya, Kohei; Suzuki, Takehito; Miyauchi, Eiji; Ishii, Yasuo; Tanabe, Soichi; Murakami, Masaru; Hattori, Masahira
2011-01-01
Lactococcus garvieae causes fatal haemorrhagic septicaemia in fish such as yellowtail. The comparative analysis of genomes of a virulent strain Lg2 and a non-virulent strain ATCC 49156 of L. garvieae revealed that the two strains shared a high degree of sequence identity, but Lg2 had a 16.5-kb capsule gene cluster that is absent in ATCC 49156. The capsule gene cluster was composed of 15 genes, of which eight genes are highly conserved with those in exopolysaccharide biosynthesis gene cluster often found in Lactococcus lactis strains. Sequence analysis of the capsule gene cluster in the less virulent strain L. garvieae Lg2-S, Lg2-derived strain, showed that two conserved genes were disrupted by a single base pair deletion, respectively. These results strongly suggest that the capsule is crucial for virulence of Lg2. The capsule gene cluster of Lg2 may be a genomic island from several features such as the presence of insertion sequences flanked on both ends, different GC content from the chromosomal average, integration into the locus syntenic to other lactococcal genome sequences, and distribution in human gut microbiomes. The analysis also predicted other potential virulence factors such as haemolysin. The present study provides new insights into understanding of the virulence mechanisms of L. garvieae in fish. PMID:21829716
Heterologous Production of a Novel Cyclic Peptide Compound, KK-1, in Aspergillus oryzae.
Yoshimi, Akira; Yamaguchi, Sigenari; Fujioka, Tomonori; Kawai, Kiyoshi; Gomi, Katsuya; Machida, Masayuki; Abe, Keietsu
2018-01-01
A novel cyclic peptide compound, KK-1, was originally isolated from the plant-pathogenic fungus Curvularia clavata . It consists of 10 amino acid residues, including five N -methylated amino acid residues, and has potent antifungal activity. Recently, the genome-sequencing analysis of C. clavata was completed, and the biosynthetic genes involved in KK-1 production were predicted by using a novel gene cluster mining tool, MIDDAS-M. These genes form an approximately 75-kb cluster, which includes nine open reading frames, containing a non-ribosomal peptide synthetase (NRPS) gene. To determine whether the predicted genes were responsible for the biosynthesis of KK-1, we performed heterologous production of KK-1 in Aspergillus oryzae by introduction of the cluster genes into the genome of A. oryzae . The NRPS gene was split in two fragments and then reconstructed in the A. oryzae genome, because the gene was quite large (approximately 40 kb). The remaining seven genes in the cluster, excluding the regulatory gene kkR , were simultaneously introduced into the strain of A. oryzae in which NRPS had already been incorporated. To evaluate the heterologous production of KK-1 in A. oryzae , gene expression was analyzed by RT-PCR and KK-1 productivity was quantified by HPLC. KK-1 was produced in variable quantities by a number of transformed strains, along with expression of the cluster genes. The amount of KK-1 produced by the strain with the greatest expression of all genes was lower than that produced by the original producer, C. clavata . Therefore, expression of the cluster genes is necessary and sufficient for the heterologous production of KK-1 in A. oryzae , although there may be unknown factors limiting productivity in this species.
Heterologous Production of a Novel Cyclic Peptide Compound, KK-1, in Aspergillus oryzae
Yoshimi, Akira; Yamaguchi, Sigenari; Fujioka, Tomonori; Kawai, Kiyoshi; Gomi, Katsuya; Machida, Masayuki; Abe, Keietsu
2018-01-01
A novel cyclic peptide compound, KK-1, was originally isolated from the plant-pathogenic fungus Curvularia clavata. It consists of 10 amino acid residues, including five N-methylated amino acid residues, and has potent antifungal activity. Recently, the genome-sequencing analysis of C. clavata was completed, and the biosynthetic genes involved in KK-1 production were predicted by using a novel gene cluster mining tool, MIDDAS-M. These genes form an approximately 75-kb cluster, which includes nine open reading frames, containing a non-ribosomal peptide synthetase (NRPS) gene. To determine whether the predicted genes were responsible for the biosynthesis of KK-1, we performed heterologous production of KK-1 in Aspergillus oryzae by introduction of the cluster genes into the genome of A. oryzae. The NRPS gene was split in two fragments and then reconstructed in the A. oryzae genome, because the gene was quite large (approximately 40 kb). The remaining seven genes in the cluster, excluding the regulatory gene kkR, were simultaneously introduced into the strain of A. oryzae in which NRPS had already been incorporated. To evaluate the heterologous production of KK-1 in A. oryzae, gene expression was analyzed by RT-PCR and KK-1 productivity was quantified by HPLC. KK-1 was produced in variable quantities by a number of transformed strains, along with expression of the cluster genes. The amount of KK-1 produced by the strain with the greatest expression of all genes was lower than that produced by the original producer, C. clavata. Therefore, expression of the cluster genes is necessary and sufficient for the heterologous production of KK-1 in A. oryzae, although there may be unknown factors limiting productivity in this species. PMID:29686660
de Marcos, Alberto; Triviño, Magdalena; Pérez-Bueno, María Luisa; Ballesteros, Isabel; Barón, Matilde; Mena, Montaña; Fenoll, Carmen
2015-01-01
Loss of function of the positive stomata development regulators SPCH or MUTE in Arabidopsis thaliana renders stomataless plants; spch-3 and mute-3 mutants are extreme dwarfs, but produce cotyledons and tiny leaves, providing a system to interrogate plant life in the absence of stomata. To this end, we compared their cotyledon transcriptomes with that of wild-type plants. K-means clustering of differentially expressed genes generated four clusters: clusters 1 and 2 grouped genes commonly regulated in the mutants, while clusters 3 and 4 contained genes distinctively regulated in mute-3. Classification in functional categories and metabolic pathways of genes in clusters 1 and 2 suggested that both mutants had depressed secondary, nitrogen and sulfur metabolisms, while only a few photosynthesis-related genes were down-regulated. In situ quenching analysis of chlorophyll fluorescence revealed limited inhibition of photosynthesis. This and other fluorescence measurements matched the mutant transcriptomic features. Differential transcriptomes of both mutants were enriched in growth-related genes, including known stomata development regulators, which paralleled their epidermal phenotypes. Analysis of cluster 3 was not informative for developmental aspects of mute-3. Cluster 4 comprised genes differentially up−regulated in mute−3, 35% of which were direct targets for SPCH and may relate to the unique cell types of mute−3. A screen of T-DNA insertion lines in genes differentially expressed in the mutants identified a gene putatively involved in stomata development. A collection of lines for conditional overexpression of transcription factors differentially expressed in the mutants rendered distinct epidermal phenotypes, suggesting that these proteins may be novel stomatal development regulators. Thus, our transcriptome analysis represents a useful source of new genes for the study of stomata development and for characterizing physiology and growth in the absence of stomata. PMID:26157447
Iacob, Eli; Light, Alan R.; Donaldson, Gary W.; Okifuji, Akiko; Hughen, Ronald W.; White, Andrea T.; Light, Kathleen C.
2015-01-01
Objective To determine if independent candidate genes can be grouped into meaningful biological factors and if these factors are associated with the diagnosis of chronic fatigue syndrome (CFS) and fibromyalgia (FMS) while controlling for co-morbid depression, sex, and age. Methods We included leukocyte mRNA gene expression from a total of 261 individuals including healthy controls (n=61), patients with FMS only (n=15), CFS only (n=33), co-morbid CFS and FMS (n=79), and medication-resistant (n=42) or medication-responsive (n=31) depression. We used Exploratory Factor Analysis (EFA) on 34 candidate genes to determine factor scores and regression analysis to examine if these factors were associated with specific diagnoses. Results EFA resulted in four independent factors with minimal overlap of genes between factors explaining 51% of the variance. We labeled these factors by function as: 1) Purinergic and cellular modulators; 2) Neuronal growth and immune function; 3) Nociception and stress mediators; 4) Energy and mitochondrial function. Regression analysis predicting these biological factors using FMS, CFS, depression severity, age, and sex revealed that greater expression in Factors 1 and 3 was positively associated with CFS and negatively associated with depression severity (QIDS score), but not associated with FMS. Conclusion Expression of candidate genes can be grouped into meaningful clusters, and CFS and depression are associated with the same 2 clusters but in opposite directions when controlling for co-morbid FMS. Given high co-morbid disease and interrelationships between biomarkers, EFA may help determine patient subgroups in this population based on gene expression. PMID:26097208
Clustering gene expression data based on predicted differential effects of GV interaction.
Pan, Hai-Yan; Zhu, Jun; Han, Dan-Fu
2005-02-01
Microarray has become a popular biotechnology in biological and medical research. However, systematic and stochastic variabilities in microarray data are expected and unavoidable, resulting in the problem that the raw measurements have inherent "noise" within microarray experiments. Currently, logarithmic ratios are usually analyzed by various clustering methods directly, which may introduce bias interpretation in identifying groups of genes or samples. In this paper, a statistical method based on mixed model approaches was proposed for microarray data cluster analysis. The underlying rationale of this method is to partition the observed total gene expression level into various variations caused by different factors using an ANOVA model, and to predict the differential effects of GV (gene by variety) interaction using the adjusted unbiased prediction (AUP) method. The predicted GV interaction effects can then be used as the inputs of cluster analysis. We illustrated the application of our method with a gene expression dataset and elucidated the utility of our approach using an external validation.
Schoberle, Taylor J.; Nguyen-Coleman, C. Kim; Herold, Jennifer; Yang, Ally; Weirauch, Matt; Hughes, Timothy R.; McMurray, John S.; May, Gregory S.
2014-01-01
Secondary metabolites are produced by numerous organisms and can either be beneficial, benign, or harmful to humans. Genes involved in the synthesis and transport of these secondary metabolites are frequently found in gene clusters, which are often coordinately regulated, being almost exclusively dependent on transcription factors that are located within the clusters themselves. Gliotoxin, which is produced by a variety of Aspergillus species, Trichoderma species, and Penicillium species, exhibits immunosuppressive properties and has therefore been the subject of research for many laboratories. There have been a few proteins shown to regulate the gliotoxin cluster, most notably GliZ, a Zn2Cys6 binuclear finger transcription factor that lies within the cluster, and LaeA, a putative methyltransferase that globally regulates secondary metabolism clusters within numerous fungal species. Using a high-copy inducer screen in A. fumigatus, our lab has identified a novel C2H2 transcription factor, which plays an important role in regulating the gliotoxin biosynthetic cluster. This transcription factor, named GipA, induces gliotoxin production when present in extra copies. Furthermore, loss of gipA reduces gliotoxin production significantly. Through protein binding microarray and mutagenesis, we have identified a DNA binding site recognized by GipA that is in extremely close proximity to a potential GliZ DNA binding site in the 5′ untranslated region of gliA, which encodes an efflux pump within the gliotoxin cluster. Not surprisingly, GliZ and GipA appear to work in an interdependent fashion to positively control gliA expression. PMID:24784729
Lu, Hong; Patil, Prabhu; Van Sluys, Marie-Anne; White, Frank F; Ryan, Robert P; Dow, J Maxwell; Rabinowicz, Pablo; Salzberg, Steven L; Leach, Jan E; Sonti, Ramesh; Brendel, Volker; Bogdanove, Adam J
2008-01-01
Xanthomonas is a large genus of plant-associated and plant-pathogenic bacteria. Collectively, members cause diseases on over 392 plant species. Individually, they exhibit marked host- and tissue-specificity. The determinants of this specificity are unknown. To assess potential contributions to host- and tissue-specificity, pathogenesis-associated gene clusters were compared across genomes of eight Xanthomonas strains representing vascular or non-vascular pathogens of rice, brassicas, pepper and tomato, and citrus. The gum cluster for extracellular polysaccharide is conserved except for gumN and sequences downstream. The xcs and xps clusters for type II secretion are conserved, except in the rice pathogens, in which xcs is missing. In the otherwise conserved hrp cluster, sequences flanking the core genes for type III secretion vary with respect to insertion sequence element and putative effector gene content. Variation at the rpf (regulation of pathogenicity factors) cluster is more pronounced, though genes with established functional relevance are conserved. A cluster for synthesis of lipopolysaccharide varies highly, suggesting multiple horizontal gene transfers and reassortments, but this variation does not correlate with host- or tissue-specificity. Phylogenetic trees based on amino acid alignments of gum, xps, xcs, hrp, and rpf cluster products generally reflect strain phylogeny. However, amino acid residues at four positions correlate with tissue specificity, revealing hpaA and xpsD as candidate determinants. Examination of genome sequences of xanthomonads Xylella fastidiosa and Stenotrophomonas maltophilia revealed that the hrp, gum, and xcs clusters are recent acquisitions in the Xanthomonas lineage. Our results provide insight into the ancestral Xanthomonas genome and indicate that differentiation with respect to host- and tissue-specificity involved not major modifications or wholesale exchange of clusters, but subtle changes in a small number of genes or in non-coding sequences, and/or differences outside the clusters, potentially among regulatory targets or secretory substrates.
A cross-species bi-clustering approach to identifying conserved co-regulated genes.
Sun, Jiangwen; Jiang, Zongliang; Tian, Xiuchun; Bi, Jinbo
2016-06-15
A growing number of studies have explored the process of pre-implantation embryonic development of multiple mammalian species. However, the conservation and variation among different species in their developmental programming are poorly defined due to the lack of effective computational methods for detecting co-regularized genes that are conserved across species. The most sophisticated method to date for identifying conserved co-regulated genes is a two-step approach. This approach first identifies gene clusters for each species by a cluster analysis of gene expression data, and subsequently computes the overlaps of clusters identified from different species to reveal common subgroups. This approach is ineffective to deal with the noise in the expression data introduced by the complicated procedures in quantifying gene expression. Furthermore, due to the sequential nature of the approach, the gene clusters identified in the first step may have little overlap among different species in the second step, thus difficult to detect conserved co-regulated genes. We propose a cross-species bi-clustering approach which first denoises the gene expression data of each species into a data matrix. The rows of the data matrices of different species represent the same set of genes that are characterized by their expression patterns over the developmental stages of each species as columns. A novel bi-clustering method is then developed to cluster genes into subgroups by a joint sparse rank-one factorization of all the data matrices. This method decomposes a data matrix into a product of a column vector and a row vector where the column vector is a consistent indicator across the matrices (species) to identify the same gene cluster and the row vector specifies for each species the developmental stages that the clustered genes co-regulate. Efficient optimization algorithm has been developed with convergence analysis. This approach was first validated on synthetic data and compared to the two-step method and several recent joint clustering methods. We then applied this approach to two real world datasets of gene expression during the pre-implantation embryonic development of the human and mouse. Co-regulated genes consistent between the human and mouse were identified, offering insights into conserved functions, as well as similarities and differences in genome activation timing between the human and mouse embryos. The R package containing the implementation of the proposed method in C ++ is available at: https://github.com/JavonSun/mvbc.git and also at the R platform https://www.r-project.org/ jinbo@engr.uconn.edu. © The Author 2016. Published by Oxford University Press.
Rankinen, Tuomo; Sarzynski, Mark A.; Ghosh, Sujoy; Bouchard, Claude
2015-01-01
Clustering of obesity, coronary artery disease, and cardiovascular disease risk factors is observed in epidemiological studies and clinical settings. Twin and family studies have provided some supporting evidence for the clustering hypothesis. Loci nearest a lead single nucleotide polymorphism (SNP) showing genome-wide significant associations with coronary artery disease, body mass index, C-reactive protein, blood pressure, lipids, and type 2 diabetes mellitus were selected for pathway and network analyses. Eighty-seven autosomal regions (181 SNPs), mapping to 56 genes, were found to be pleiotropic. Most pleiotropic regions contained genes associated with coronary artery disease and plasma lipids, whereas some exhibited coaggregation between obesity and cardiovascular disease risk factors. We observed enrichment for liver X receptor (LXR)/retinoid X receptor (RXR) and farnesoid X receptor/RXR nuclear receptor signaling among pleiotropic genes and for signatures of coronary artery disease and hepatic steatosis. In the search for functionally interacting networks, we found that 43 pleiotropic genes were interacting in a network with an additional 24 linker genes. ENCODE (Encyclopedia of DNA Elements) data were queried for distribution of pleiotropic SNPs among regulatory elements and coding sequence variations. Of the 181 SNPs, 136 were annotated to ≥1 regulatory feature. An enrichment analysis found over-representation of enhancers and DNAse hypersensitive regions when compared against all SNPs of the 1000 Genomes pilot project. In summary, there are genomic regions exerting pleiotropic effects on cardiovascular disease risk factors, although only a few included obesity. Further studies are needed to resolve the clustering in terms of DNA variants, genes, pathways, and actionable targets. PMID:25722444
eMBI: Boosting Gene Expression-based Clustering for Cancer Subtypes.
Chang, Zheng; Wang, Zhenjia; Ashby, Cody; Zhou, Chuan; Li, Guojun; Zhang, Shuzhong; Huang, Xiuzhen
2014-01-01
Identifying clinically relevant subtypes of a cancer using gene expression data is a challenging and important problem in medicine, and is a necessary premise to provide specific and efficient treatments for patients of different subtypes. Matrix factorization provides a solution by finding checker-board patterns in the matrices of gene expression data. In the context of gene expression profiles of cancer patients, these checkerboard patterns correspond to genes that are up- or down-regulated in patients with particular cancer subtypes. Recently, a new matrix factorization framework for biclustering called Maximum Block Improvement (MBI) is proposed; however, it still suffers several problems when applied to cancer gene expression data analysis. In this study, we developed many effective strategies to improve MBI and designed a new program called enhanced MBI (eMBI), which is more effective and efficient to identify cancer subtypes. Our tests on several gene expression profiling datasets of cancer patients consistently indicate that eMBI achieves significant improvements in comparison with MBI, in terms of cancer subtype prediction accuracy, robustness, and running time. In addition, the performance of eMBI is much better than another widely used matrix factorization method called nonnegative matrix factorization (NMF) and the method of hierarchical clustering, which is often the first choice of clinical analysts in practice.
eMBI: Boosting Gene Expression-based Clustering for Cancer Subtypes
Chang, Zheng; Wang, Zhenjia; Ashby, Cody; Zhou, Chuan; Li, Guojun; Zhang, Shuzhong; Huang, Xiuzhen
2014-01-01
Identifying clinically relevant subtypes of a cancer using gene expression data is a challenging and important problem in medicine, and is a necessary premise to provide specific and efficient treatments for patients of different subtypes. Matrix factorization provides a solution by finding checker-board patterns in the matrices of gene expression data. In the context of gene expression profiles of cancer patients, these checkerboard patterns correspond to genes that are up- or down-regulated in patients with particular cancer subtypes. Recently, a new matrix factorization framework for biclustering called Maximum Block Improvement (MBI) is proposed; however, it still suffers several problems when applied to cancer gene expression data analysis. In this study, we developed many effective strategies to improve MBI and designed a new program called enhanced MBI (eMBI), which is more effective and efficient to identify cancer subtypes. Our tests on several gene expression profiling datasets of cancer patients consistently indicate that eMBI achieves significant improvements in comparison with MBI, in terms of cancer subtype prediction accuracy, robustness, and running time. In addition, the performance of eMBI is much better than another widely used matrix factorization method called nonnegative matrix factorization (NMF) and the method of hierarchical clustering, which is often the first choice of clinical analysts in practice. PMID:25374455
The WRKY Transcription Factor Genes in Lotus japonicus.
Song, Hui; Wang, Pengfei; Nan, Zhibiao; Wang, Xingjun
2014-01-01
WRKY transcription factor genes play critical roles in plant growth and development, as well as stress responses. WRKY genes have been examined in various higher plants, but they have not been characterized in Lotus japonicus. The recent release of the L. japonicus whole genome sequence provides an opportunity for a genome wide analysis of WRKY genes in this species. In this study, we identified 61 WRKY genes in the L. japonicus genome. Based on the WRKY protein structure, L. japonicus WRKY (LjWRKY) genes can be classified into three groups (I-III). Investigations of gene copy number and gene clusters indicate that only one gene duplication event occurred on chromosome 4 and no clustered genes were detected on chromosomes 3 or 6. Researchers previously believed that group II and III WRKY domains were derived from the C-terminal WRKY domain of group I. Our results suggest that some WRKY genes in group II originated from the N-terminal domain of group I WRKY genes. Additional evidence to support this hypothesis was obtained by Medicago truncatula WRKY (MtWRKY) protein motif analysis. We found that LjWRKY and MtWRKY group III genes are under purifying selection, suggesting that WRKY genes will become increasingly structured and functionally conserved.
Genome-wide network of regulatory genes for construction of a chordate embryo.
Shoguchi, Eiichi; Hamaguchi, Makoto; Satoh, Nori
2008-04-15
Animal development is controlled by gene regulation networks that are composed of sequence-specific transcription factors (TF) and cell signaling molecules (ST). Although housekeeping genes have been reported to show clustering in the animal genomes, whether the genes comprising a given regulatory network are physically clustered on a chromosome is uncertain. We examined this question in the present study. Ascidians are the closest living relatives of vertebrates, and their tadpole-type larva represents the basic body plan of chordates. The Ciona intestinalis genome contains 390 core TF genes and 119 major ST genes. Previous gene disruption assays led to the formulation of a basic chordate embryonic blueprint, based on over 3000 genetic interactions among 79 zygotic regulatory genes. Here, we mapped the regulatory genes, including all 79 regulatory genes, on the 14 pairs of Ciona chromosomes by fluorescent in situ hybridization (FISH). Chromosomal localization of upstream and downstream regulatory genes demonstrates that the components of coherent developmental gene networks are evenly distributed over the 14 chromosomes. Thus, this study provides the first comprehensive evidence that the physical clustering of regulatory genes, or their target genes, is not relevant for the genome-wide control of gene expression during development.
Dixit, Shalabh; Kumar Biswal, Akshaya; Min, Aye; Henry, Amelia; Oane, Rowena H.; Raorane, Manish L.; Longkumer, Toshisangba; Pabuayon, Isaiah M.; Mutte, Sumanth K.; Vardarajan, Adithi R.; Miro, Berta; Govindan, Ganesan; Albano-Enriquez, Blesilda; Pueffeld, Mandy; Sreenivasulu, Nese; Slamet-Loedin, Inez; Sundarvelpandian, Kalaipandian; Tsai, Yuan-Ching; Raghuvanshi, Saurabh; Hsing, Yue-Ie C.; Kumar, Arvind; Kohli, Ajay
2015-01-01
Sub-QTLs and multiple intra-QTL genes are hypothesized to underpin large-effect QTLs. Known QTLs over gene families, biosynthetic pathways or certain traits represent functional gene-clusters of genes of the same gene ontology (GO). Gene-clusters containing genes of different GO have not been elaborated, except in silico as coexpressed genes within QTLs. Here we demonstrate the requirement of multiple intra-QTL genes for the full impact of QTL qDTY12.1 on rice yield under drought. Multiple evidences are presented for the need of the transcription factor ‘no apical meristem’ (OsNAM12.1) and its co-localized target genes of separate GO categories for qDTY12.1 function, raising a regulon-like model of genetic architecture. The molecular underpinnings of qDTY12.1 support its effectiveness in further improving a drought tolerant genotype and for its validity in multiple genotypes/ecosystems/environments. Resolving the combinatorial value of OsNAM12.1 with individual intra-QTL genes notwithstanding, identification and analyses of qDTY12.1has fast-tracked rice improvement towards food security. PMID:26507552
Moorthy, Sakthi D.; Davidson, Scott; Shchuka, Virlana M.; Singh, Gurdeep; Malek-Gilani, Nakisa; Langroudi, Lida; Martchenko, Alexandre; So, Vincent; Macpherson, Neil N.; Mitchell, Jennifer A.
2017-01-01
Transcriptional enhancers are critical for maintaining cell-type–specific gene expression and driving cell fate changes during development. Highly transcribed genes are often associated with a cluster of individual enhancers such as those found in locus control regions. Recently, these have been termed stretch enhancers or super-enhancers, which have been predicted to regulate critical cell identity genes. We employed a CRISPR/Cas9-mediated deletion approach to study the function of several enhancer clusters (ECs) and isolated enhancers in mouse embryonic stem (ES) cells. Our results reveal that the effect of deleting ECs, also classified as ES cell super-enhancers, is highly variable, resulting in target gene expression reductions ranging from 12% to as much as 92%. Partial deletions of these ECs which removed only one enhancer or a subcluster of enhancers revealed partially redundant control of the regulated gene by multiple enhancers within the larger cluster. Many highly transcribed genes in ES cells are not associated with a super-enhancer; furthermore, super-enhancer predictions ignore 81% of the potentially active regulatory elements predicted by cobinding of five or more pluripotency-associated transcription factors. Deletion of these additional enhancer regions revealed their robust regulatory role in gene transcription. In addition, select super-enhancers and enhancers were identified that regulated clusters of paralogous genes. We conclude that, whereas robust transcriptional output can be achieved by an isolated enhancer, clusters of enhancers acting on a common target gene act in a partially redundant manner to fine tune transcriptional output of their target genes. PMID:27895109
2018-01-01
ABSTRACT Botrytis cinerea is a plant-pathogenic fungus producing apothecia as sexual fruiting bodies. To study the function of mating type (MAT) genes, single-gene deletion mutants were generated in both genes of the MAT1-1 locus and both genes of the MAT1-2 locus. Deletion mutants in two MAT genes were entirely sterile, while mutants in the other two MAT genes were able to develop stipes but never formed an apothecial disk. Little was known about the reprogramming of gene expression during apothecium development. We analyzed transcriptomes of sclerotia, three stages of apothecium development (primordia, stipes, and apothecial disks), and ascospores by RNA sequencing. Ten secondary metabolite gene clusters were upregulated at the onset of sexual development and downregulated in ascospores released from apothecia. Notably, more than 3,900 genes were differentially expressed in ascospores compared to mature apothecial disks. Among the genes that were upregulated in ascospores were numerous genes encoding virulence factors, which reveals that ascospores are transcriptionally primed for infection prior to their arrival on a host plant. Strikingly, the massive transcriptional changes at the initiation and completion of the sexual cycle often affected clusters of genes, rather than randomly dispersed genes. Thirty-five clusters of genes were jointly upregulated during the onset of sexual reproduction, while 99 clusters of genes (comprising >900 genes) were jointly downregulated in ascospores. These transcriptional changes coincided with changes in expression of genes encoding enzymes participating in chromatin organization, hinting at the occurrence of massive epigenetic regulation of gene expression during sexual reproduction. PMID:29440571
Roque, Dario R; Makowski, Liza; Chen, Ting-Huei; Rashid, Naim; Hayes, D Neil; Bae-Jump, Victoria
2016-08-01
The Cancer Genome Atlas (TCGA) identified four integrated clusters for endometrial cancer (EC): POLE, MSI, CNL and CNH. We evaluated differences in gene expression profiles of obese and non-obese women with EC and examined the association of body mass index (BMI) within the clusters identified in TCGA. TCGA RNAseq data was used to identify genes related to increasing BMI among ECs. The POLE, MSI and CNL clusters were composed mostly of endometrioid EC. Patient BMI was compared between these three clusters with one-way ANOVA. Association between gene expression and BMI was also assessed while adjusting for confounding effects of potential confounding factors. p-Values testing the association between gene expression and BMI were adjusted for multiple hypothesis testing over the 20,531 genes considered. Mean BMI was statistically different between the ECs in the CNL (35.8) versus POLE (29.8) cluster (p=0.006) and approached significance for the MSI (33.0) versus CNL (35.8) cluster (p=0.05). 181 genes were significantly up- or down-regulated with increasing BMI in endometrioid EC (q-value<0.01), including LPL, IRS-1, IGFBP4, IGFBP7 and the progesterone receptor. DAVID functional annotation analysis revealed significant enrichment in "cell cycle" (adjusted p-value=1.5E-5) and "DNA metabolic processes" (adjusted p-value=1E-3) for the identified genes. Obesity related genes were found to be upregulated with increasing BMI among endometrioid ECs. Patients with POLE tumors have the lowest median BMI when compared to MSI and CNL. Given the heterogeneity among endometrioid EC, consideration should be given to abandoning the Type I and II classification of EC tumors. Copyright © 2016 Elsevier Inc. All rights reserved.
Siegel, Nicol; Hoegg, Simone; Salzburger, Walter; Braasch, Ingo; Meyer, Axel
2007-01-01
Background The evolutionary lineage leading to the teleost fish underwent a whole genome duplication termed FSGD or 3R in addition to two prior genome duplications that took place earlier during vertebrate evolution (termed 1R and 2R). Resulting from the FSGD, additional copies of genes are present in fish, compared to tetrapods whose lineage did not experience the 3R genome duplication. Interestingly, we find that ParaHox genes do not differ in number in extant teleost fishes despite their additional genome duplication from the genomic situation in mammals, but they are distributed over twice as many paralogous regions in fish genomes. Results We determined the DNA sequence of the entire ParaHox C1 paralogon in the East African cichlid fish Astatotilapia burtoni, and compared it to orthologous regions in other vertebrate genomes as well as to the paralogous vertebrate ParaHox D paralogons. Evolutionary relationships among genes from these four chromosomal regions were studied with several phylogenetic algorithms. We provide evidence that the genes of the ParaHox C paralogous cluster are duplicated in teleosts, just as it had been shown previously for the D paralogon genes. Overall, however, synteny and cluster integrity seems to be less conserved in ParaHox gene clusters than in Hox gene clusters. Comparative analyses of non-coding sequences uncovered conserved, possibly co-regulatory elements, which are likely to contain promoter motives of the genes belonging to the ParaHox paralogons. Conclusion There seems to be strong stabilizing selection for gene order as well as gene orientation in the ParaHox C paralogon, since with a few exceptions, only the lengths of the introns and intergenic regions differ between the distantly related species examined. The high degree of evolutionary conservation of this gene cluster's architecture in particular – but possibly clusters of genes more generally – might be linked to the presence of promoter, enhancer or inhibitor motifs that serve to regulate more than just one gene. Therefore, deletions, inversions or relocations of individual genes could destroy the regulation of the clustered genes in this region. The existence of such a regulation network might explain the evolutionary conservation of gene order and orientation over the course of hundreds of millions of years of vertebrate evolution. Another possible explanation for the highly conserved gene order might be the existence of a regulator not located immediately next to its corresponding gene but further away since a relocation or inversion would possibly interrupt this interaction. Different ParaHox clusters were found to have experienced differential gene loss in teleosts. Yet the complete set of these homeobox genes was maintained, albeit distributed over almost twice the number of chromosomes. Selection due to dosage effects and/or stoichiometric disturbance might act more strongly to maintain a modal number of homeobox genes (and possibly transcription factors more generally) per genome, yet permit the accumulation of other (non regulatory) genes associated with these homeobox gene clusters. PMID:17822543
Factor IX gene haplotypes in Amerindians.
Franco, R F; Araújo, A G; Zago, M A; Guerreiro, J F; Figueiredo, M S
1997-02-01
We have determined the haplotypes of the factor IX gene for 95 Indians from 5 Brazilian Amazon tribes: Wayampí, Wayana-Apalaí, Kayapó, Arára, and Yanomámi. Eight polymorphisms linked to the factor IX gene were investigated: MseI (at 5', nt -698), BamHI (at 5', nt -561), DdeI (intron 1), BamHI (intron 2), XmnI (intron 3), TaqI (intron 4), MspI (intron 4), and HhaI (at 3', approximately 8 kb). The results of the haplotype distribution and the allele frequencies for each of the factor IX gene polymorphisms in Amerindians were similar to the results reported for Asian populations but differed from results for other ethnic groups. Only five haplotypes were identified within the entire Amerindian study population, and the haplotype distribution was significantly different among the five tribes, with one (Arára) to four (Wayampí) haplotypes being found per tribe. These findings indicate a significant heterogeneity among the Indian tribes and contrast with the homogeneous distribution of the beta-globin gene cluster haplotypes but agree with our recent findings on the distribution of alpha-globin gene cluster haplotypes and the allele frequencies for six VNTRs in the same Amerindian tribes. Our data represent the first study of factor IX-associated polymorphisms in Amerindian populations and emphasizes the applicability of these genetic markers for population and human evolution studies.
Global Identification of Genes Affecting Iron-Sulfur Cluster Biogenesis and Iron Homeostasis
Hidese, Ryota; Kurihara, Tatsuo; Esaki, Nobuyoshi
2014-01-01
Iron-sulfur (Fe-S) clusters are ubiquitous cofactors that are crucial for many physiological processes in all organisms. In Escherichia coli, assembly of Fe-S clusters depends on the activity of the iron-sulfur cluster (ISC) assembly and sulfur mobilization (SUF) apparatus. However, the underlying molecular mechanisms and the mechanisms that control Fe-S cluster biogenesis and iron homeostasis are still poorly defined. In this study, we performed a global screen to identify the factors affecting Fe-S cluster biogenesis and iron homeostasis using the Keio collection, which is a library of 3,815 single-gene E. coli knockout mutants. The approach was based on radiolabeling of the cells with [2-14C]dihydrouracil, which entirely depends on the activity of an Fe-S enzyme, dihydropyrimidine dehydrogenase. We identified 49 genes affecting Fe-S cluster biogenesis and/or iron homeostasis, including 23 genes important only under microaerobic/anaerobic conditions. This study defines key proteins associated with Fe-S cluster biogenesis and iron homeostasis, which will aid further understanding of the cellular mechanisms that coordinate the processes. In addition, we applied the [2-14C]dihydrouracil-labeling method to analyze the role of amino acid residues of an Fe-S cluster assembly scaffold (IscU) as a model of the Fe-S cluster assembly apparatus. The analysis showed that Cys37, Cys63, His105, and Cys106 are essential for the function of IscU in vivo, demonstrating the potential of the method to investigate in vivo function of proteins involved in Fe-S cluster assembly. PMID:24415728
Deciphering the Anti-Aflatoxinogenic Properties of Eugenol Using a Large-Scale q-PCR Approach
Caceres, Isaura; El Khoury, Rhoda; Medina, Ángel; Lippi, Yannick; Naylies, Claire; Atoui, Ali; El Khoury, André; Oswald, Isabelle P.; Bailly, Jean-Denis; Puel, Olivier
2016-01-01
Produced by several species of Aspergillus, Aflatoxin B1 (AFB1) is a carcinogenic mycotoxin contaminating many crops worldwide. The utilization of fungicides is currently one of the most common methods; nevertheless, their use is not environmentally or economically sound. Thus, the use of natural compounds able to block aflatoxinogenesis could represent an alternative strategy to limit food and feed contamination. For instance, eugenol, a 4-allyl-2-methoxyphenol present in many essential oils, has been identified as an anti-aflatoxin molecule. However, its precise mechanism of action has yet to be clarified. The production of AFB1 is associated with the expression of a 70 kB cluster, and not less than 21 enzymatic reactions are necessary for its production. Based on former empirical data, a molecular tool composed of 60 genes targeting 27 genes of aflatoxin B1 cluster and 33 genes encoding the main regulatory factors potentially involved in its production, was developed. We showed that AFB1 inhibition in Aspergillus flavus following eugenol addition at 0.5 mM in a Malt Extract Agar (MEA) medium resulted in a complete inhibition of the expression of all but one gene of the AFB1 biosynthesis cluster. This transcriptomic effect followed a down-regulation of the complex composed by the two internal regulatory factors, AflR and AflS. This phenomenon was also influenced by an over-expression of veA and mtfA, two genes that are directly linked to AFB1 cluster regulation. PMID:27128940
Clustering of change patterns using Fourier coefficients.
Kim, Jaehee; Kim, Haseong
2008-01-15
To understand the behavior of genes, it is important to explore how the patterns of gene expression change over a time period because biologically related gene groups can share the same change patterns. Many clustering algorithms have been proposed to group observation data. However, because of the complexity of the underlying functions there have not been many studies on grouping data based on change patterns. In this study, the problem of finding similar change patterns is induced to clustering with the derivative Fourier coefficients. The sample Fourier coefficients not only provide information about the underlying functions, but also reduce the dimension. In addition, as their limiting distribution is a multivariate normal, a model-based clustering method incorporating statistical properties would be appropriate. This work is aimed at discovering gene groups with similar change patterns that share similar biological properties. We developed a statistical model using derivative Fourier coefficients to identify similar change patterns of gene expression. We used a model-based method to cluster the Fourier series estimation of derivatives. The model-based method is advantageous over other methods in our proposed model because the sample Fourier coefficients asymptotically follow the multivariate normal distribution. Change patterns are automatically estimated with the Fourier representation in our model. Our model was tested in simulations and on real gene data sets. The simulation results showed that the model-based clustering method with the sample Fourier coefficients has a lower clustering error rate than K-means clustering. Even when the number of repeated time points was small, the same results were obtained. We also applied our model to cluster change patterns of yeast cell cycle microarray expression data with alpha-factor synchronization. It showed that, as the method clusters with the probability-neighboring data, the model-based clustering with our proposed model yielded biologically interpretable results. We expect that our proposed Fourier analysis with suitably chosen smoothing parameters could serve as a useful tool in classifying genes and interpreting possible biological change patterns. The R program is available upon the request.
Regulation of apoptosis by peroxisome proliferators.
Roberts, Ruth A; Michel, Cecile; Coyle, Beth; Freathy, Caroline; Cain, Kelvin; Boitier, Eric
2004-04-01
Peroxisome proliferators (PPs) constitute a large and chemically diverse family of non-genotoxic rodent hepatocarcinogens that activate the PP-activated receptor alpha (PPARalpha). In order to investigate the hypothesis that PPs elicit their carcinogenic effects through the suppression of apoptosis, we established an in vitro assay for apoptosis using both primary rat hepatocytes and the FaO rat hepatoma cell line. Apoptosis was induced by transforming growth factor beta1 (TGFbeta1), the physiological negative regulator of liver growth. In this system, PPs could suppress both spontaneous and TGFbeta1-induced apoptosis. In order to understand the mechanisms of this regulation of apoptosis, we conducted microarray analysis followed by pathway-specific gene clustering in TGFbeta1-treated cells. After treatment, 76 genes were up-regulated and 185 were down-regulated more than 1.5-fold. Cluster analysis of up-regulated genes revealed three clusters, A-C. Cluster A (4h) was associated with 12% apoptosis and consisted of genes mainly of the cytoskeleton and extracellular matrix such as troponin and the proteoglycan SDC4. In cluster B (8h; 25% apoptosis), there were many pro- and anti-apoptotic genes such as XIAP, BAK1 and BAD, whereas at 16h (40% apoptosis) the regulated genes were mainly those of the cellular stress pathways such as the genes implicated in the activation of the transcription factor NFkappab. Genes found down-regulated in response to TGFbeta1 were mainly those associated with oxidative stress and several genes implicated in glutathione production and maintenance. Thus, TGFbeta1 may induce apoptosis via a down regulation of oxidant defence leading to the generation of reactive oxygen species. The ability of PPs to impact on these apoptosis pathways remains to be determined. To approach this question, we have developed a technique using laser capture microdissection of livers treated with the PP, clofibric acid coupled with gene expression array analysis. Results show that some of the key steps of the LCM process had an impact on the gene profiles generated. However, this did not preclude accurate determination of a PP-specific molecular signature. Thus, the choice of appropriate controls will ensure that meaningful gene expression analyses can be performed on tissue microdissected from the foci generated in clofibric acid treated livers. These data will allow the identification of specific genes that are regulated by PPs leading to changes in apoptosis and ultimately to tumours.
Identifying and Assessing Interesting Subgroups in a Heterogeneous Population.
Lee, Woojoo; Alexeyenko, Andrey; Pernemalm, Maria; Guegan, Justine; Dessen, Philippe; Lazar, Vladimir; Lehtiö, Janne; Pawitan, Yudi
2015-01-01
Biological heterogeneity is common in many diseases and it is often the reason for therapeutic failures. Thus, there is great interest in classifying a disease into subtypes that have clinical significance in terms of prognosis or therapy response. One of the most popular methods to uncover unrecognized subtypes is cluster analysis. However, classical clustering methods such as k-means clustering or hierarchical clustering are not guaranteed to produce clinically interesting subtypes. This could be because the main statistical variability--the basis of cluster generation--is dominated by genes not associated with the clinical phenotype of interest. Furthermore, a strong prognostic factor might be relevant for a certain subgroup but not for the whole population; thus an analysis of the whole sample may not reveal this prognostic factor. To address these problems we investigate methods to identify and assess clinically interesting subgroups in a heterogeneous population. The identification step uses a clustering algorithm and to assess significance we use a false discovery rate- (FDR-) based measure. Under the heterogeneity condition the standard FDR estimate is shown to overestimate the true FDR value, but this is remedied by an improved FDR estimation procedure. As illustrations, two real data examples from gene expression studies of lung cancer are provided.
Wu, Yanhua; Yu, Yaqin; Zhao, Tiancheng; Wang, Shibin; Fu, Yingli; Qi, Yue; Yang, Guang; Yao, Wenwang; Su, Yingying; Ma, Yue; Shi, Jieping; Jiang, Jing; Kou, Changgui
2016-01-01
The present study investigated the prevalence and risk factors for Metabolic syndrome. We evaluated the association between single nucleotide polymorphisms (SNPs) in the apolipoprotein APOA1/C3/A4/A5 gene cluster and the MetS risk and analyzed the interactions of environmental factors and APOA1/C3/A4/A5 gene cluster polymorphisms with MetS. A study on the prevalence and risk factors for MetS was conducted using data from a large cross-sectional survey representative of the population of Jilin Province situated in northeastern China. A total of 16,831 participations were randomly chosen by multistage stratified cluster sampling of residents aged from 18 to 79 years in all nine administrative areas of the province. Environmental factors associated with MetS were examined using univariate and multivariate logistic regression analyses based on the weighted sample data. A sub-sample of 1813 survey subjects who met the criteria for MetS patients and 2037 controls from this case-control study were used to evaluate the association between SNPs and MetS risk. Genomic DNA was extracted from peripheral blood lymphocytes, and SNP genotyping was determined by MALDI-TOF-MS. The associations between SNPs and MetS were examined using a case-control study design. The interactions of environmental factors and APOA1/C3/A4/A5 gene cluster polymorphisms with MetS were assessed using multivariate logistic regression analysis. The overall adjusted prevalence of MetS was 32.86% in Jilin province. The prevalence of MetS in men was 36.64%, which was significantly higher than the prevalence in women (29.66%). MetS was more common in urban areas (33.86%) than in rural areas (31.80%). The prevalence of MetS significantly increased with age (OR = 8.621, 95%CI = 6.594-11.272). Mental labor (OR = 1.098, 95%CI = 1.008-1.195), current smoking (OR = 1.259, 95%CI = 1.108-1.429), excess salt intake (OR = 1.252, 95%CI = 1.149-1.363), and a fruit and dairy intake less than 2 servings a week were positively associated with MetS (P<0.05). A family history of diabetes (OR = 1.630, 95%CI = 1.484-1.791), cardiovascular disease or cerebral diseases (OR = 1.297, 95%CI = 1.211-1.389) was associated with MetS. APOA1 rs670, APOA5 rs662799 and rs651821 revealed significant differences in genotype distributions between the MetS patients and control subjects. The minor alleles of APOA1 rs670, APOA5 rs662799 and rs651821, and APOA5 rs2075291 were associated with MetS (P<0.0016). APOA1 rs5072 and APOC3 rs5128, APOA5 rs651821 and rs662799 were in strong linkage disequilibrium to each other with r2 greater than 0.8. Five haplotypes were associated with an increased risk of MetS (OR = 1.23, 1.58, 1.80, 1.90, and 1.98). When we investigated the interactions of environmental factors and APOA1/C3/A4/A5 gene cluster gene polymorphisms, we found that APOA5 rs662799 had interactions with tobacco use and alcohol consumption (PGE<0.05). There was a high prevalence of MetS in the northeast of China. Male gender, increasing age, mental labor, family history of diabetes, cardiovascular disease or cerebral diseases, current smoking, excess salt intake, fruit and dairy intake less than 2 servings a week, and drinking were associated with MetS. The APOA1/C3/A4/A5 gene cluster was associated with MetS in the Han Chinese. APOA5 rs662799 had interactions with the environmental factors associated with MetS.
The WRKY Transcription Factor Genes in Lotus japonicus
Wang, Pengfei; Wang, Xingjun
2014-01-01
WRKY transcription factor genes play critical roles in plant growth and development, as well as stress responses. WRKY genes have been examined in various higher plants, but they have not been characterized in Lotus japonicus. The recent release of the L. japonicus whole genome sequence provides an opportunity for a genome wide analysis of WRKY genes in this species. In this study, we identified 61 WRKY genes in the L. japonicus genome. Based on the WRKY protein structure, L. japonicus WRKY (LjWRKY) genes can be classified into three groups (I–III). Investigations of gene copy number and gene clusters indicate that only one gene duplication event occurred on chromosome 4 and no clustered genes were detected on chromosomes 3 or 6. Researchers previously believed that group II and III WRKY domains were derived from the C-terminal WRKY domain of group I. Our results suggest that some WRKY genes in group II originated from the N-terminal domain of group I WRKY genes. Additional evidence to support this hypothesis was obtained by Medicago truncatula WRKY (MtWRKY) protein motif analysis. We found that LjWRKY and MtWRKY group III genes are under purifying selection, suggesting that WRKY genes will become increasingly structured and functionally conserved. PMID:24745006
Functional Organization of hsp70 Cluster in Camel (Camelus dromedarius) and Other Mammals
Garbuz, David G.; Astakhova, Lubov N.; Zatsepina, Olga G.; Arkhipova, Irina R.; Nudler, Eugene; Evgen'ev, Michael B.
2011-01-01
Heat shock protein 70 (Hsp70) is a molecular chaperone providing tolerance to heat and other challenges at the cellular and organismal levels. We sequenced a genomic cluster containing three hsp70 family genes linked with major histocompatibility complex (MHC) class III region from an extremely heat tolerant animal, camel (Camelus dromedarius). Two hsp70 family genes comprising the cluster contain heat shock elements (HSEs), while the third gene lacks HSEs and should not be induced by heat shock. Comparison of the camel hsp70 cluster with the corresponding regions from several mammalian species revealed similar organization of genes forming the cluster. Specifically, the two heat inducible hsp70 genes are arranged in tandem, while the third constitutively expressed hsp70 family member is present in inverted orientation. Comparison of regulatory regions of hsp70 genes from camel and other mammals demonstrates that transcription factor matches with highest significance are located in the highly conserved 250-bp upstream region and correspond to HSEs followed by NF-Y and Sp1 binding sites. The high degree of sequence conservation leaves little room for putative camel-specific regulatory elements. Surprisingly, RT-PCR and 5′/3′-RACE analysis demonstrated that all three hsp70 genes are expressed in camel's muscle and blood cells not only after heat shock, but under normal physiological conditions as well, and may account for tolerance of camel cells to extreme environmental conditions. A high degree of evolutionary conservation observed for the hsp70 cluster always linked with MHC locus in mammals suggests an important role of such organization for coordinated functioning of these vital genes. PMID:22096537
Familial cancer syndromes and clusters.
Birch, J M
1994-07-01
The study of rare families in which a variety of cancers occur, usually at an early age and with patterns consistent with a common hereditary mechanism, has contributed much to our understanding of the process of carcinogenesis. So far, genes identified as having a role in cancer predisposition in these families have also been important in the histogenesis of sporadic cancers. In the two most clearly defined cancer family syndromes, the Li-Fraumeni syndrome and Lynch syndrome II, the genes involved predispose to diverse but specific constellations of cancers. Genes associated with site-specific familial cancer clusters may also give rise to increased susceptibility to other cancers, and site-specific clusters may represent one end of a spectrum. A consistent feature of familial cancer syndromes is the variable expression within and between families. A challenge for the future will be to determine other factors which may interact with the principal genes involved, giving rise to this variability.
Scheps, Karen G; Varela, Viviana
Different hemoglobin isoforms are expressed during the embryonic, fetal and postnatal stages. They are formed by combination of polypeptide chains synthesized from the α- and β-globin gene clusters. Based on the fact that the presence of high hemoglobin F levels is beneficial in both sickle cell disease and severe thalassemic syndromes, a revision of the regulation of the β-globin cluster expression is proposed, especially regarding the genes encoding the y-globin chains (HBG1 and HBG2). In this review we describe the current knowledge about transcription factors and epigenetic regulators involved in the switches of the β-globin cluster. It is expected that the consolidation of knowledge in this field will allow finding new therapeutic targets for the treatment of hemoglobinopathies.
Wolf, Timo; Droste, Julian; Gren, Tetiana; Ortseifen, Vera; Schneiker-Bekel, Susanne; Zemke, Till; Pühler, Alfred; Kalinowski, Jörn
2017-07-25
Acarbose is used in the treatment of diabetes mellitus type II and is produced by Actinoplanes sp. SE50/110. Although the biosynthesis of acarbose has been intensively studied, profound knowledge about transcription factors involved in acarbose biosynthesis and their binding sites has been missing until now. In contrast to acarbose biosynthetic gene clusters in Streptomyces spp., the corresponding gene cluster of Actinoplanes sp. SE50/110 lacks genes for transcriptional regulators. The acarbose regulator C (AcrC) was identified through an in silico approach by aligning the LacI family regulators of acarbose biosynthetic gene clusters in Streptomyces spp. with the Actinoplanes sp. SE50/110 genome. The gene for acrC, located in a head-to-head arrangement with the maltose/maltodextrin ABC transporter malEFG operon, was deleted by introducing PCR targeting for Actinoplanes sp. SE50/110. Characterization was carried out through cultivation experiments, genome-wide microarray hybridizations, and RT-qPCR as well as electrophoretic mobility shift assays for the elucidation of binding motifs. The results show that AcrC binds to the intergenic region between acbE and acbD in Actinoplanes sp. SE50/110 and acts as a transcriptional repressor on these genes. The transcriptomic profile of the wild type was reconstituted through a complementation of the deleted acrC gene. Additionally, regulatory sequence motifs for the binding of AcrC were identified in the intergenic region of acbE and acbD. It was shown that AcrC expression influences acarbose formation in the early growth phase. Interestingly, AcrC does not regulate the malEFG operon. This study characterizes the first known transcription factor of the acarbose biosynthetic gene cluster in Actinoplanes sp. SE50/110. It therefore represents an important step for understanding the regulatory network of this organism. Based on this work, rational strain design for improving the biotechnological production of acarbose can now be implemented.
Strakova, Eva; Zikova, Alice; Vohradsky, Jiri
2014-01-01
A computational model of gene expression was applied to a novel test set of microarray time series measurements to reveal regulatory interactions between transcriptional regulators represented by 45 sigma factors and the genes expressed during germination of a prokaryote Streptomyces coelicolor. Using microarrays, the first 5.5 h of the process was recorded in 13 time points, which provided a database of gene expression time series on genome-wide scale. The computational modeling of the kinetic relations between the sigma factors, individual genes and genes clustered according to the similarity of their expression kinetics identified kinetically plausible sigma factor-controlled networks. Using genome sequence annotations, functional groups of genes that were predominantly controlled by specific sigma factors were identified. Using external binding data complementing the modeling approach, specific genes involved in the control of the studied process were identified and their function suggested.
Li, Xi-Hong; Wu, Mao-Yu; Wang, Ai-Li; Jiang, Yu-Qian; Jiang, Yun-Hong
2012-01-01
Anthocyanin biosynthesis in various plants is affected by environmental conditions and controlled by the transcription level of the corresponding genes. In pears (Pyrus communis cv. ‘Wujiuxiang’), anthocyanin biosynthesis is significantly induced during low temperature storage compared with that at room temperature. We further examined the transcriptional levels of anthocyanin biosynthetic genes in ‘Wujiuxiang’ pears during developmental ripening and temperature-induced storage. The expression of genes that encode flavanone 3-hydroxylase, dihydroflavonol 4-reductase, anthocyanidin synthase, UDP-glucose: flavonoid 3-O-glucosyltransferase, and R2R3 MYB transcription factor (PcMYB10) was strongly positively correlated with anthocyanin accumulation in ‘Wujiuxiang’ pears in response to both developmental and cold-temperature induction. Hierarchical clustering analysis revealed the expression patterns of the set of target genes, of which PcMYB10 and most anthocyanin biosynthetic genes were related to the same cluster. The present work may help explore the molecular mechanism that regulates anthocyanin biosynthesis and its response to abiotic stress at the transcriptional level in plants. PMID:23029391
Assessment of stem cell differentiation based on genome-wide expression profiles.
Godoy, Patricio; Schmidt-Heck, Wolfgang; Hellwig, Birte; Nell, Patrick; Feuerborn, David; Rahnenführer, Jörg; Kattler, Kathrin; Walter, Jörn; Blüthgen, Nils; Hengstler, Jan G
2018-07-05
In recent years, protocols have been established to differentiate stem and precursor cells into more mature cell types. However, progress in this field has been hampered by difficulties to assess the differentiation status of stem cell-derived cells in an unbiased manner. Here, we present an analysis pipeline based on published data and methods to quantify the degree of differentiation and to identify transcriptional control factors explaining differences from the intended target cells or tissues. The pipeline requires RNA-Seq or gene array data of the stem cell starting population, derived 'mature' cells and primary target cells or tissue. It consists of a principal component analysis to represent global expression changes and to identify possible problems of the dataset that require special attention, such as: batch effects; clustering techniques to identify gene groups with similar features; over-representation analysis to characterize biological motifs and transcriptional control factors of the identified gene clusters; and metagenes as well as gene regulatory networks for quantitative cell-type assessment and identification of influential transcription factors. Possibilities and limitations of the analysis pipeline are illustrated using the example of human embryonic stem cell and human induced pluripotent cells to generate 'hepatocyte-like cells'. The pipeline quantifies the degree of incomplete differentiation as well as remaining stemness and identifies unwanted features, such as colon- and fibroblast-associated gene clusters that are absent in real hepatocytes but typically induced by currently available differentiation protocols. Finally, transcription factors responsible for incomplete and unwanted differentiation are identified. The proposed method is widely applicable and allows an unbiased and quantitative assessment of stem cell-derived cells.This article is part of the theme issue 'Designer human tissue: coming to a lab near you'. © 2018 The Author(s).
Fungal secondary metabolites - strategies to activate silent gene clusters.
Brakhage, Axel A; Schroeckh, Volker
2011-01-01
Filamentous fungi produce a multitude of low molecular weight bioactive compounds. The increasing number of fungal genome sequences impressively demonstrated that their biosynthetic potential is far from being exploited. In fungi, the genes required for the biosynthesis of a secondary metabolite are clustered. Many of these bioinformatically newly discovered secondary metabolism gene clusters are silent under standard laboratory conditions. Consequently, no product can be found. This review summarizes the current strategies that have been successfully applied during the last years to activate these silent gene clusters in filamentous fungi, especially in the genus Aspergillus. The techniques take advantage of genome mining, vary from the simple search for compounds with bioinformatically predicted physicochemical properties up to methods that exploit a probable interaction of microorganisms. Until now, the majority of successful approaches have been based on molecular biology like the generation of gene "knock outs", promoter exchange, overexpression of transcription factors or other pleiotropic regulators. Moreover, strategies based on epigenetics opened a new avenue for the elucidation of the regulation of secondary metabolite formation and will certainly continue to play a significant role for the elucidation of cryptic natural products. The conditions under which a given gene cluster is naturally expressed are largely unknown. One technique is to attempt to simulate the natural habitat by co-cultivation of microorganisms from the same ecosystem. This has already led to the activation of silent gene clusters and the identification of novel compounds in Aspergillus nidulans. These simulation strategies will help discover new natural products in the future, and may also provide fundamental new insights into microbial communication. Copyright © 2010 Elsevier Inc. All rights reserved.
Banelli, Barbara; Brigati, Claudio; Di Vinci, Angela; Casciano, Ida; Forlani, Alessandra; Borzì, Luana; Allemanni, Giorgio; Romani, Massimo
2012-03-01
Epigenetic alterations are hallmarks of cancer and powerful biomarkers, whose clinical utilization is made difficult by the absence of standardization and of common methods of data interpretation. The coordinate methylation of many loci in cancer is defined as 'CpG island methylator phenotype' (CIMP) and identifies clinically distinct groups of patients. In neuroblastoma (NB), CIMP is defined by a methylation signature, which includes different loci, but its predictive power on outcome is entirely recapitulated by the PCDHB cluster only. We have developed a robust and cost-effective pyrosequencing-based assay that could facilitate the clinical application of CIMP in NB. This assay permits the unbiased simultaneous amplification and sequencing of 17 out of 19 genes of the PCDHB cluster for quantitative methylation analysis, taking into account all the sequence variations. As some of these variations were at CpG doublets, we bypassed the data interpretation conducted by the methylation analysis software to assign the corrected methylation value at these sites. The final result of the assay is the mean methylation level of 17 gene fragments in the protocadherin B cluster (PCDHB) cluster. We have utilized this assay to compare the methylation levels of the PCDHB cluster between high-risk and very low-risk NB patients, confirming the predictive value of CIMP. Our results demonstrate that the pyrosequencing-based assay herein described is a powerful instrument for the analysis of this gene cluster that may simplify the data comparison between different laboratories and, in perspective, could facilitate its clinical application. Furthermore, our results demonstrate that, in principle, pyrosequencing can be efficiently utilized for the methylation analysis of gene clusters with high internal homologies.
Calcisponges have a ParaHox gene and dynamic expression of dispersed NK homeobox genes.
Fortunato, Sofia A V; Adamski, Marcin; Ramos, Olivia Mendivil; Leininger, Sven; Liu, Jing; Ferrier, David E K; Adamska, Maja
2014-10-30
Sponges are simple animals with few cell types, but their genomes paradoxically contain a wide variety of developmental transcription factors, including homeobox genes belonging to the Antennapedia (ANTP) class, which in bilaterians encompass Hox, ParaHox and NK genes. In the genome of the demosponge Amphimedon queenslandica, no Hox or ParaHox genes are present, but NK genes are linked in a tight cluster similar to the NK clusters of bilaterians. It has been proposed that Hox and ParaHox genes originated from NK cluster genes after divergence of sponges from the lineage leading to cnidarians and bilaterians. On the other hand, synteny analysis lends support to the notion that the absence of Hox and ParaHox genes in Amphimedon is a result of secondary loss (the ghost locus hypothesis). Here we analysed complete suites of ANTP-class homeoboxes in two calcareous sponges, Sycon ciliatum and Leucosolenia complicata. Our phylogenetic analyses demonstrate that these calcisponges possess orthologues of bilaterian NK genes (Hex, Hmx and Msx), a varying number of additional NK genes and one ParaHox gene, Cdx. Despite the generation of scaffolds spanning multiple genes, we find no evidence of clustering of Sycon NK genes. All Sycon ANTP-class genes are developmentally expressed, with patterns suggesting their involvement in cell type specification in embryos and adults, metamorphosis and body plan patterning. These results demonstrate that ParaHox genes predate the origin of sponges, thus confirming the ghost locus hypothesis, and highlight the need to analyse the genomes of multiple sponge lineages to obtain a complete picture of the ancestral composition of the first animal genome.
Dopstadt, Julian; Neubauer, Lisa; Tudzynski, Paul; Humpf, Hans-Ulrich
2016-01-01
Claviceps purpurea is an important food contaminant and well known for the production of the toxic ergot alkaloids. Apart from that, little is known about its secondary metabolism and not all toxic substances going along with the food contamination with Claviceps are known yet. We explored the metabolite profile of a gene cluster in C. purpurea with a high homology to gene clusters, which are responsible for the formation of epipolythiodiketopiperazine (ETP) toxins in other fungi. By overexpressing the transcription factor, we were able to activate the cluster in the standard C. purpurea strain 20.1. Although all necessary genes for the formation of the characteristic disulfide bridge were expressed in the overexpression mutants, the fungus did not produce any ETPs. Isolation of pathway intermediates showed that the common biosynthetic pathway stops after the first steps. Our results demonstrate that hydroxylation of the diketopiperazine backbone is the critical step during the ETP biosynthesis. Due to a dysfunctional enzyme, the fungus is not able to produce toxic ETPs. Instead, the pathway end-products are new unusual metabolites with a unique nitrogen-sulfur bond. By heterologous expression of the Leptosphaeria maculans cytochrome P450 encoding gene sirC, we were able to identify the end-products of the ETP cluster in C. purpurea. The thioclapurines are so far unknown ETPs, which might contribute to the toxicity of other C. purpurea strains with a potentially intact ETP cluster.
Tudzynski, Paul; Humpf, Hans-Ulrich
2016-01-01
Claviceps purpurea is an important food contaminant and well known for the production of the toxic ergot alkaloids. Apart from that, little is known about its secondary metabolism and not all toxic substances going along with the food contamination with Claviceps are known yet. We explored the metabolite profile of a gene cluster in C. purpurea with a high homology to gene clusters, which are responsible for the formation of epipolythiodiketopiperazine (ETP) toxins in other fungi. By overexpressing the transcription factor, we were able to activate the cluster in the standard C. purpurea strain 20.1. Although all necessary genes for the formation of the characteristic disulfide bridge were expressed in the overexpression mutants, the fungus did not produce any ETPs. Isolation of pathway intermediates showed that the common biosynthetic pathway stops after the first steps. Our results demonstrate that hydroxylation of the diketopiperazine backbone is the critical step during the ETP biosynthesis. Due to a dysfunctional enzyme, the fungus is not able to produce toxic ETPs. Instead, the pathway end-products are new unusual metabolites with a unique nitrogen-sulfur bond. By heterologous expression of the Leptosphaeria maculans cytochrome P450 encoding gene sirC, we were able to identify the end-products of the ETP cluster in C. purpurea. The thioclapurines are so far unknown ETPs, which might contribute to the toxicity of other C. purpurea strains with a potentially intact ETP cluster. PMID:27390873
The Chloroplast atpA Gene Cluster in Chlamydomonas reinhardtii1
Drapier, Dominique; Suzuki, Hideki; Levy, Haim; Rimbault, Blandine; Kindle, Karen L.; Stern, David B.; Wollman, Francis-André
1998-01-01
Most chloroplast genes in vascular plants are organized into polycistronic transcription units, which generate a complex pattern of mono-, di-, and polycistronic transcripts. In contrast, most Chlamydomonas reinhardtii chloroplast transcripts characterized to date have been monocistronic. This paper describes the atpA gene cluster in the C. reinhardtii chloroplast genome, which includes the atpA, psbI, cemA, and atpH genes, encoding the α-subunit of the coupling-factor-1 (CF1) ATP synthase, a small photosystem II polypeptide, a chloroplast envelope membrane protein, and subunit III of the CF0 ATP synthase, respectively. We show that promoters precede the atpA, psbI, and atpH genes, but not the cemA gene, and that cemA mRNA is present only as part of di-, tri-, or tetracistronic transcripts. Deletions introduced into the gene cluster reveal, first, that CF1-α can be translated from di- or polycistronic transcripts, and, second, that substantial reductions in mRNA quantity have minimal effects on protein synthesis rates. We suggest that posttranscriptional mRNA processing is common in C. reinhardtii chloroplasts, permitting the expression of multiple genes from a single promoter. PMID:9625716
Ko, Yi-An; Mukherjee, Bhramar; Smith, Jennifer A; Kardia, Sharon L R; Allison, Matthew; Diez Roux, Ana V
2016-11-01
There has been an increased interest in identifying gene-environment interaction (G × E) in the context of multiple environmental exposures. Most G × E studies analyze one exposure at a time, but we are exposed to multiple exposures in reality. Efficient analysis strategies for complex G × E with multiple environmental factors in a single model are still lacking. Using the data from the Multiethnic Study of Atherosclerosis, we illustrate a two-step approach for modeling G × E with multiple environmental factors. First, we utilize common clustering and classification strategies (e.g., k-means, latent class analysis, classification and regression trees, Bayesian clustering using Dirichlet Process) to define subgroups corresponding to distinct environmental exposure profiles. Second, we illustrate the use of an additive main effects and multiplicative interaction model, instead of the conventional saturated interaction model using product terms of factors, to study G × E with the data-driven exposure subgroups defined in the first step. We demonstrate useful analytical approaches to translate multiple environmental exposures into one summary class. These tools not only allow researchers to consider several environmental exposures in G × E analysis but also provide some insight into how genes modify the effect of a comprehensive exposure profile instead of examining effect modification for each exposure in isolation.
Identifying and Assessing Interesting Subgroups in a Heterogeneous Population
Lee, Woojoo; Alexeyenko, Andrey; Pernemalm, Maria; Guegan, Justine; Dessen, Philippe; Lazar, Vladimir; Lehtiö, Janne; Pawitan, Yudi
2015-01-01
Biological heterogeneity is common in many diseases and it is often the reason for therapeutic failures. Thus, there is great interest in classifying a disease into subtypes that have clinical significance in terms of prognosis or therapy response. One of the most popular methods to uncover unrecognized subtypes is cluster analysis. However, classical clustering methods such as k-means clustering or hierarchical clustering are not guaranteed to produce clinically interesting subtypes. This could be because the main statistical variability—the basis of cluster generation—is dominated by genes not associated with the clinical phenotype of interest. Furthermore, a strong prognostic factor might be relevant for a certain subgroup but not for the whole population; thus an analysis of the whole sample may not reveal this prognostic factor. To address these problems we investigate methods to identify and assess clinically interesting subgroups in a heterogeneous population. The identification step uses a clustering algorithm and to assess significance we use a false discovery rate- (FDR-) based measure. Under the heterogeneity condition the standard FDR estimate is shown to overestimate the true FDR value, but this is remedied by an improved FDR estimation procedure. As illustrations, two real data examples from gene expression studies of lung cancer are provided. PMID:26339613
Fu, Qiang; Su, Zhixin; Cheng, Yuqiang; Wang, Zhaofei; Li, Shiyu; Wang, Heng'an; Sun, Jianhe; Yan, Yaxian
In order to investigate the diverse characteristics of clustered, regularly interspaced short palindromic repeat (CRISPR) arrays and the distribution of virulence factor genes in avian Escherichia coli, 80 E. coli isolates obtained from chickens with avian pathogenic E. coli (APEC) or avian fecal commensal E. coli (AFEC) were identified. Using the multiplex polymerase chain reaction (PCR), five genes were subjected to phylogenetic typing and examined for CRISPR arrays to study genetic relatedness among the strains. The strains were further analyzed for CRISPR loci and virulence factor genes to determine a possible association between their CRISPR elements and their potential virulence. The strains were divided into five phylogenetic groups: A, B1, B2, D and E. It was confirmed that two types of CRISPR arrays, CRISPR1 and CRISPR2, which contain up to 246 distinct spacers, were amplified in most of the strains. Further classification of the isolates was achieved by sorting them into nine CRISPR clusters based on their spacer profiles, which indicates a candidate typing method for E. coli. Several significant differences in invasion-associated gene distribution were found between the APEC isolates and the AFEC isolates. Our results identified the distribution of 11 virulence genes and CRISPR diversity in 80 strains. It was demonstrated that, with the exception of iucD and aslA, there was no sharp demarcation in the gene distribution between the pathogenic (APEC) and commensal (AFEC) strains, while the total number of indicated CRISPR spacers may have a positive correlation with the potential pathogenicity of the E. coli isolates. Copyright © 2016. Published by Elsevier Masson SAS.
Horizontal acquisition of toxic alkaloid synthesis in a clade of plant associated fungi.
Marcet-Houben, Marina; Gabaldón, Toni
2016-01-01
Clavicipitaceae is a fungal group that comprises species that closely interact with plants as pathogens, parasites or symbionts. A key factor in these interactions is the ability of these fungi to synthesize toxic alkaloid compounds that contribute to the protection of the plant host against herbivores. Some of these compounds such as ergot alkaloids are toxic to humans and have caused important epidemics throughout history. The gene clusters encoding the proteins responsible for the synthesis of ergot alkaloids and lolines in Clavicipitaceae have been elucidated. Notably, homologs to these gene clusters can be found in distantly related species such as Aspergillus fumigatus and Penicillium expansum, which diverged from Clavicipitaceae more than 400 million years ago. We here use a phylogenetic approach to analyze the evolution of these gene clusters. We found that the gene clusters conferring the ability to synthesize ergot alkaloids and loline emerged first in Eurotiomycetes and were then likely transferred horizontally to Clavicipitaceae. Horizontal gene transfer is known to play a role in shaping the distribution of secondary metabolism clusters across distantly related fungal species. We propose that HGT events have played an important role in the capability of Clavicipitaceae to produce two key secondary metabolites that have enhanced the ability of these species to protect their plant hosts, therefore favoring their interactions. Copyright © 2015 The Authors. Published by Elsevier Inc. All rights reserved.
The genetic epidemiology of personality disorders
Reichborn-Kjennerud, Ted
2010-01-01
Genetic epidemiologic studies indicate that all ten personality disorders (PDs) classified on the DSM-IV axis II are modestly to moderately heritable. Shared environmental and nonadditive genetic factors are of minor or no importance. No sex differences have been identified. Multivariate studies suggest that the extensive comorbidity between the PDs can be explained by three common genetic and environmental risk factors. The genetic factors do not reflect the DSM-IV cluster structure, but rather: i) broad vulnerability to PD pathology or negative emotionality; ii) high impulsivity/low agreeableness; and iii) introversion. Common genetic and environmental liability factors contribute to comorbidity between pairs or clusters of axis I and axis II disorders. Molecular genetic studies of PDs, mostly candidate gene association studies, indicate that genes linked to neurotransmitter pathways, especially in the serotonergic and dopaminergic systems, are involved. Future studies, using newer methods like genome-wide association, might take advantage of the use of endophenotypes. PMID:20373672
Transcriptional Regulatory Network Analysis of MYB Transcription Factor Family Genes in Rice.
Smita, Shuchi; Katiyar, Amit; Chinnusamy, Viswanathan; Pandey, Dev M; Bansal, Kailash C
2015-01-01
MYB transcription factor (TF) is one of the largest TF families and regulates defense responses to various stresses, hormone signaling as well as many metabolic and developmental processes in plants. Understanding these regulatory hierarchies of gene expression networks in response to developmental and environmental cues is a major challenge due to the complex interactions between the genetic elements. Correlation analyses are useful to unravel co-regulated gene pairs governing biological process as well as identification of new candidate hub genes in response to these complex processes. High throughput expression profiling data are highly useful for construction of co-expression networks. In the present study, we utilized transcriptome data for comprehensive regulatory network studies of MYB TFs by "top-down" and "guide-gene" approaches. More than 50% of OsMYBs were strongly correlated under 50 experimental conditions with 51 hub genes via "top-down" approach. Further, clusters were identified using Markov Clustering (MCL). To maximize the clustering performance, parameter evaluation of the MCL inflation score (I) was performed in terms of enriched GO categories by measuring F-score. Comparison of co-expressed cluster and clads analyzed from phylogenetic analysis signifies their evolutionarily conserved co-regulatory role. We utilized compendium of known interaction and biological role with Gene Ontology enrichment analysis to hypothesize function of coexpressed OsMYBs. In the other part, the transcriptional regulatory network analysis by "guide-gene" approach revealed 40 putative targets of 26 OsMYB TF hubs with high correlation value utilizing 815 microarray data. The putative targets with MYB-binding cis-elements enrichment in their promoter region, functional co-occurrence as well as nuclear localization supports our finding. Specially, enrichment of MYB binding regions involved in drought-inducibility implying their regulatory role in drought response in rice. Thus, the co-regulatory network analysis facilitated the identification of complex OsMYB regulatory networks, and candidate target regulon genes of selected guide MYB genes. The results contribute to the candidate gene screening, and experimentally testable hypotheses for potential regulatory MYB TFs, and their targets under stress conditions.
Clustering cancer gene expression data by projective clustering ensemble
Yu, Xianxue; Yu, Guoxian
2017-01-01
Gene expression data analysis has paramount implications for gene treatments, cancer diagnosis and other domains. Clustering is an important and promising tool to analyze gene expression data. Gene expression data is often characterized by a large amount of genes but with limited samples, thus various projective clustering techniques and ensemble techniques have been suggested to combat with these challenges. However, it is rather challenging to synergy these two kinds of techniques together to avoid the curse of dimensionality problem and to boost the performance of gene expression data clustering. In this paper, we employ a projective clustering ensemble (PCE) to integrate the advantages of projective clustering and ensemble clustering, and to avoid the dilemma of combining multiple projective clusterings. Our experimental results on publicly available cancer gene expression data show PCE can improve the quality of clustering gene expression data by at least 4.5% (on average) than other related techniques, including dimensionality reduction based single clustering and ensemble approaches. The empirical study demonstrates that, to further boost the performance of clustering cancer gene expression data, it is necessary and promising to synergy projective clustering with ensemble clustering. PCE can serve as an effective alternative technique for clustering gene expression data. PMID:28234920
Denapaite, Dalia; Rieger, Martin; Köndgen, Sophie; Brückner, Reinhold; Ochigava, Irma; Kappeler, Peter; Mätz-Rensing, Kerstin; Leendertz, Fabian; Hakenbeck, Regine
2016-01-01
Viridans streptococci were obtained from primates (great apes, rhesus monkeys, and ring-tailed lemurs) held in captivity, as well as from free-living animals (chimpanzees and lemurs) for whom contact with humans is highly restricted. Isolates represented a variety of viridans streptococci, including unknown species. Streptococcus oralis was frequently isolated from samples from great apes. Genotypic methods revealed that most of the strains clustered on separate lineages outside the main cluster of human S. oralis strains. This suggests that S. oralis is part of the commensal flora in higher primates and evolved prior to humans. Many genes described as virulence factors in Streptococcus pneumoniae were present also in other viridans streptococcal genomes. Unlike in S. pneumoniae, clustered regularly interspaced short palindromic repeat (CRISPR)-CRISPR-associated protein (Cas) gene clusters were common among viridans streptococci, and many S. oralis strains were type PI-2 (pilus islet 2) variants. S. oralis displayed a remarkable diversity of genes involved in the biosynthesis of peptidoglycan (penicillin-binding proteins and MurMN) and choline-containing teichoic acid. The small noncoding cia-dependent small RNAs (csRNAs) controlled by the response regulator CiaR might contribute to the genomic diversity, since we observed novel genomic islands between duplicated csRNAs, variably present in some isolates. All S. oralis genomes contained a β-N-acetyl-hexosaminidase gene absent in S. pneumoniae, which in contrast frequently harbors the neuraminidases NanB/C, which are absent in S. oralis. The identification of S. oralis-specific genes will help us to understand their adaptation to diverse habitats. IMPORTANCE Streptococcus pneumoniae is a rare example of a human-pathogenic bacterium among viridans streptococci, which consist of commensal symbionts, such as the close relatives Streptococcus mitis and S. oralis. We have shown that S. oralis can frequently be isolated from primates and a variety of other viridans streptococci as well. Genes and genomic islands which are known pneumococcal virulence factors are present in S. oralis and S. mitis, documenting the widespread occurrence of these compounds, which encode surface and secreted proteins. The frequent occurrence of CRISP-Cas gene clusters and a surprising variation of a set of small noncoding RNAs are factors to be considered in future research to further our understanding of mechanisms involved in the genomic diversity driven by horizontal gene transfer among viridans streptococci.
Denapaite, Dalia; Rieger, Martin; Köndgen, Sophie; Brückner, Reinhold; Ochigava, Irma; Kappeler, Peter; Mätz-Rensing, Kerstin; Leendertz, Fabian
2016-01-01
ABSTRACT Viridans streptococci were obtained from primates (great apes, rhesus monkeys, and ring-tailed lemurs) held in captivity, as well as from free-living animals (chimpanzees and lemurs) for whom contact with humans is highly restricted. Isolates represented a variety of viridans streptococci, including unknown species. Streptococcus oralis was frequently isolated from samples from great apes. Genotypic methods revealed that most of the strains clustered on separate lineages outside the main cluster of human S. oralis strains. This suggests that S. oralis is part of the commensal flora in higher primates and evolved prior to humans. Many genes described as virulence factors in Streptococcus pneumoniae were present also in other viridans streptococcal genomes. Unlike in S. pneumoniae, clustered regularly interspaced short palindromic repeat (CRISPR)–CRISPR-associated protein (Cas) gene clusters were common among viridans streptococci, and many S. oralis strains were type PI-2 (pilus islet 2) variants. S. oralis displayed a remarkable diversity of genes involved in the biosynthesis of peptidoglycan (penicillin-binding proteins and MurMN) and choline-containing teichoic acid. The small noncoding cia-dependent small RNAs (csRNAs) controlled by the response regulator CiaR might contribute to the genomic diversity, since we observed novel genomic islands between duplicated csRNAs, variably present in some isolates. All S. oralis genomes contained a β-N-acetyl-hexosaminidase gene absent in S. pneumoniae, which in contrast frequently harbors the neuraminidases NanB/C, which are absent in S. oralis. The identification of S. oralis-specific genes will help us to understand their adaptation to diverse habitats. IMPORTANCE Streptococcus pneumoniae is a rare example of a human-pathogenic bacterium among viridans streptococci, which consist of commensal symbionts, such as the close relatives Streptococcus mitis and S. oralis. We have shown that S. oralis can frequently be isolated from primates and a variety of other viridans streptococci as well. Genes and genomic islands which are known pneumococcal virulence factors are present in S. oralis and S. mitis, documenting the widespread occurrence of these compounds, which encode surface and secreted proteins. The frequent occurrence of CRISP-Cas gene clusters and a surprising variation of a set of small noncoding RNAs are factors to be considered in future research to further our understanding of mechanisms involved in the genomic diversity driven by horizontal gene transfer among viridans streptococci. PMID:27303717
Xiang, Ruidong; McNally, Jody; Rowe, Suzanne; Jonker, Arjan; Pinares-Patino, Cesar S.; Oddy, V. Hutton; Vercoe, Phil E.; McEwan, John C.; Dalrymple, Brian P.
2016-01-01
Ruminants obtain nutrients from microbial fermentation of plant material, primarily in their rumen, a multilayered forestomach. How the different layers of the rumen wall respond to diet and influence microbial fermentation, and how these process are regulated, is not well understood. Gene expression correlation networks were constructed from full thickness rumen wall transcriptomes of 24 sheep fed two different amounts and qualities of a forage and measured for methane production. The network contained two major negatively correlated gene sub-networks predominantly representing the epithelial and muscle layers of the rumen wall. Within the epithelium sub-network gene clusters representing lipid/oxo-acid metabolism, general metabolism and proliferating and differentiating cells were identified. The expression of cell cycle and metabolic genes was positively correlated with dry matter intake, ruminal short chain fatty acid concentrations and methane production. A weak correlation between lipid/oxo-acid metabolism genes and methane yield was observed. Feed consumption level explained the majority of gene expression variation, particularly for the cell cycle genes. Many known stratified epithelium transcription factors had significantly enriched targets in the epithelial gene clusters. The expression patterns of the transcription factors and their targets in proliferating and differentiating skin is mirrored in the rumen, suggesting conservation of regulatory systems. PMID:27966600
Porcine Tissue-Specific Regulatory Networks Derived from Meta-Analysis of the Transcriptome
Pérez-Montarelo, Dafne; Hudson, Nicholas J.; Fernández, Ana I.; Ramayo-Caldas, Yuliaxis; Dalrymple, Brian P.; Reverter, Antonio
2012-01-01
The processes that drive tissue identity and differentiation remain unclear for most tissue types. So are the gene networks and transcription factors (TF) responsible for the differential structure and function of each particular tissue, and this is particularly true for non model species with incomplete genomic resources. To better understand the regulation of genes responsible for tissue identity in pigs, we have inferred regulatory networks from a meta-analysis of 20 gene expression studies spanning 480 Porcine Affymetrix chips for 134 experimental conditions on 27 distinct tissues. We developed a mixed-model normalization approach with a covariance structure that accommodated the disparity in the origin of the individual studies, and obtained the normalized expression of 12,320 genes across the 27 tissues. Using this resource, we constructed a network, based on the co-expression patterns of 1,072 TF and 1,232 tissue specific genes. The resulting network is consistent with the known biology of tissue development. Within the network, genes clustered by tissue and tissues clustered by site of embryonic origin. These clusters were significantly enriched for genes annotated in key relevant biological processes and confirm gene functions and interactions from the literature. We implemented a Regulatory Impact Factor (RIF) metric to identify the key regulators in skeletal muscle and tissues from the central nervous systems. The normalization of the meta-analysis, the inference of the gene co-expression network and the RIF metric, operated synergistically towards a successful search for tissue-specific regulators. Novel among these findings are evidence suggesting a novel key role of ERCC3 as a muscle regulator. Together, our results recapitulate the known biology behind tissue specificity and provide new valuable insights in a less studied but valuable model species. PMID:23049964
Richards, Neil; Parker, David S.; Johnson, Lisa A.; Allen, Benjamin L.; Barolo, Scott; Gumucio, Deborah L.
2015-01-01
The Hedgehog (Hh) signaling pathway directs a multitude of cellular responses during embryogenesis and adult tissue homeostasis. Stimulation of the pathway results in activation of Hh target genes by the transcription factor Ci/Gli, which binds to specific motifs in genomic enhancers. In Drosophila, only a few enhancers (patched, decapentaplegic, wingless, stripe, knot, hairy, orthodenticle) have been shown by in vivo functional assays to depend on direct Ci/Gli regulation. All but one (orthodenticle) contain more than one Ci/Gli site, prompting us to directly test whether homotypic clustering of Ci/Gli binding sites is sufficient to define a Hh-regulated enhancer. We therefore developed a computational algorithm to identify Ci/Gli clusters that are enriched over random expectation, within a given region of the genome. Candidate genomic regions containing Ci/Gli clusters were functionally tested in chicken neural tube electroporation assays and in transgenic flies. Of the 22 Ci/Gli clusters tested, seven novel enhancers (and the previously known patched enhancer) were identified as Hh-responsive and Ci/Gli-dependent in one or both of these assays, including: Cuticular protein 100A (Cpr100A); invected (inv), which encodes an engrailed-related transcription factor expressed at the anterior/posterior wing disc boundary; roadkill (rdx), the fly homolog of vertebrate Spop; the segment polarity gene gooseberry (gsb); and two previously untested regions of the Hh receptor-encoding patched (ptc) gene. We conclude that homotypic Ci/Gli clustering is not sufficient information to ensure Hh-responsiveness; however, it can provide a clue for enhancer recognition within putative Hedgehog target gene loci. PMID:26710299
Architectural roles of multiple chromatin insulators at the human apolipoprotein gene cluster
Mishiro, Tsuyoshi; Ishihara, Ko; Hino, Shinjiro; Tsutsumi, Shuichi; Aburatani, Hiroyuki; Shirahige, Katsuhiko; Kinoshita, Yoshikazu; Nakao, Mitsuyoshi
2009-01-01
Long-range regulatory elements and higher-order chromatin structure coordinate the expression of multiple genes in cluster, and CTCF/cohesin-mediated chromatin insulator may be a key in this regulation. The human apolipoprotein (APO) A1/C3/A4/A5 gene region, whose alterations increase the risk of dyslipidemia and atherosclerosis, is partitioned at least by three CTCF-enriched sites and three cohesin protein RAD21-enriched sites (two overlap with the CTCF sites), resulting in the formation of two transcribed chromatin loops by interactions between insulators. The C3 enhancer and APOC3/A4/A5 promoters reside in the same loop, where the APOC3/A4 promoters are pointed towards the C3 enhancer, whereas the APOA1 promoter is present in the different loop. The depletion of either CTCF or RAD21 disrupts the chromatin loop structure, together with significant changes in the APO expression and the localization of transcription factor hepatocyte nuclear factor (HNF)-4α and transcriptionally active form of RNA polymerase II at the APO promoters. Thus, CTCF/cohesin-mediated insulators maintain the chromatin loop formation and the localization of transcriptional apparatus at the promoters, suggesting an essential role of chromatin insulation in controlling the expression of clustered genes. PMID:19322193
NFκB-mediated activation of the cellular FUT3, 5 and 6 gene cluster by herpes simplex virus type 1.
Nordén, Rickard; Samuelsson, Ebba; Nyström, Kristina
2017-11-01
Herpes simplex virus type 1 has the ability to induce expression of a human gene cluster located on chromosome 19 upon infection. This gene cluster contains three fucosyltransferases (encoded by FUT3, FUT5 and FUT6) with the ability to add a fucose to an N-acetylglucosamine residue. Little is known regarding the transcriptional activation of these three genes in human cells. Intriguingly, herpes simplex virus type 1 activates all three genes simultaneously during infection, a situation not observed in uninfected tissue, pointing towards a virus specific mechanism for transcriptional activation. The aim of this study was to define the underlying mechanism for the herpes simplex virus type 1 activation of FUT3, FUT5 and FUT6 transcription. The transcriptional activation of the FUT-gene cluster on chromosome 19 in fibroblasts was specific, not involving adjacent genes. Moreover, inhibition of NFκB signaling through panepoxydone treatment significantly decreased the induction of FUT3, FUT5 and FUT6 transcriptional activation, as did siRNA targeting of p65, in herpes simplex virus type 1 infected fibroblasts. NFκB and p65 signaling appears to play an important role in the regulation of FUT3, FUT5 and FUT6 transcriptional activation by herpes simplex virus type 1 although additional, unidentified, viral factors might account for part of the mechanism as direct interferon mediated stimulation of NFκB was not sufficient to induce the fucosyltransferase encoding gene cluster in uninfected cells. © The Author 2017. Published by Oxford University Press. All rights reserved. For permissions, please e-mail: journals.permissions@oup.com.
D'Andrea, M; Dal Monego, S; Pallavicini, A; Modonut, M; Dreos, R; Stefanon, B; Pilla, F
2011-10-01
Using an array consisting of 10 665 70-mer oligonucleotide probes, the longissimus dorsi muscle tissue expression during growth in nine pigs belonging to Casertana (CT), an autochthonous breed characterized by slow growth and a massive accumulation of backfat, was compared with that of two cosmopolitan breeds, Large White (LW) and a crossbreed (CB; Duroc × Landrace × Large White). The results were validated by real-time PCR. All animals were of the same age and were raised under the same environmental conditions. Muscle tissues were collected at 3, 6, 9 and 11 months of age, and a total of 173 genes showed significant differential expression between CT and the cosmopolitan genetic types at 3 months of age. Time series cluster analysis indicated that the CT breed had a different pattern of gene expression compared with that of the LW and the CB. Four of the eight clusters highlighted the gene differences between CT and the other two breeds, which were further supported by statistical analyses: clusters 4 and 5 contained a total of 71 genes that were underexpressed at 3 months of age, and cluster 3 and cluster 7 included 28 and 42 genes respectively that were overexpressed at 3 months of age. As expected, differentially expressed genes belonged to the category of genes coding for contractile fibres and transcription factors involved in muscle development and differentiation. These findings highlight muscle expression genes during pig growth and are useful to understand the genetic meaning of the different developmental rates. © 2011 The Authors, Animal Genetics © 2011 Stichting International Foundation for Animal Genetics.
Genomic insight into pathogenicity of dematiaceous fungus Corynespora cassiicola
Looi, Hong Keat; Toh, Yue Fen; Yew, Su Mei; Na, Shiang Ling; Tan, Yung-Chie; Chong, Pei-Sin; Khoo, Jia-Shiun; Yee, Wai-Yan; Ng, Kee Peng
2017-01-01
Corynespora cassiicola is a common plant pathogen that causes leaf spot disease in a broad range of crop, and it heavily affect rubber trees in Malaysia (Hsueh, 2011; Nghia et al., 2008). The isolation of UM 591 from a patient’s contact lens indicates the pathogenic potential of this dematiaceous fungus in human. However, the underlying factors that contribute to the opportunistic cross-infection have not been fully studied. We employed genome sequencing and gene homology annotations in attempt to identify these factors in UM 591 using data obtained from publicly available bioinformatics databases. The assembly size of UM 591 genome is 41.8 Mbp, and a total of 13,531 (≥99 bp) genes have been predicted. UM 591 is enriched with genes that encode for glycoside hydrolases, carbohydrate esterases, auxiliary activity enzymes and cell wall degrading enzymes. Virulent genes comprising of CAZymes, peptidases, and hypervirulence-associated cutinases were found to be present in the fungal genome. Comparative analysis result shows that UM 591 possesses higher number of carbohydrate esterases family 10 (CE10) CAZymes compared to other species of fungi in this study, and these enzymes hydrolyses wide range of carbohydrate and non-carbohydrate substrates. Putative melanin, siderophore, ent-kaurene, and lycopene biosynthesis gene clusters are predicted, and these gene clusters denote that UM 591 are capable of protecting itself from the UV and chemical stresses, allowing it to adapt to different environment. Putative sterigmatocystin, HC-toxin, cercosporin, and gliotoxin biosynthesis gene cluster are predicted. This finding have highlighted the necrotrophic and invasive nature of UM 591. PMID:28149676
2012-01-01
Background The expression of genes in Corynebacterium glutamicum, a Gram-positive non-pathogenic bacterium used mainly for the industrial production of amino acids, is regulated by seven different sigma factors of RNA polymerase, including the stress-responsive ECF-sigma factor SigH. The sigH gene is located in a gene cluster together with the rshA gene, putatively encoding an anti-sigma factor. The aim of this study was to analyze the transcriptional regulation of the sigH and rshA gene cluster and the effects of RshA on the SigH regulon, in order to refine the model describing the role of SigH and RshA during stress response. Results Transcription analyses revealed that the sigH gene and rshA gene are cotranscribed from four sigH housekeeping promoters in C. glutamicum. In addition, a SigH-controlled rshA promoter was found to only drive the transcription of the rshA gene. To test the role of the putative anti-sigma factor gene rshA under normal growth conditions, a C. glutamicum rshA deletion strain was constructed and used for genome-wide transcription profiling with DNA microarrays. In total, 83 genes organized in 61 putative transcriptional units, including those previously detected using sigH mutant strains, exhibited increased transcript levels in the rshA deletion mutant compared to its parental strain. The genes encoding proteins related to disulphide stress response, heat stress proteins, components of the SOS-response to DNA damage and proteasome components were the most markedly upregulated gene groups. Altogether six SigH-dependent promoters upstream of the identified genes were determined by primer extension and a refined consensus promoter consisting of 45 original promoter sequences was constructed. Conclusions The rshA gene codes for an anti-sigma factor controlling the function of the stress-responsive sigma factor SigH in C. glutamicum. Transcription of rshA from a SigH-dependent promoter may serve to quickly shutdown the SigH-dependent stress response after the cells have overcome the stress condition. Here we propose a model of the regulation of oxidative and heat stress response including redox homeostasis by SigH, RshA and the thioredoxin system. PMID:22943411
Busche, Tobias; Silar, Radoslav; Pičmanová, Martina; Pátek, Miroslav; Kalinowski, Jörn
2012-09-03
The expression of genes in Corynebacterium glutamicum, a Gram-positive non-pathogenic bacterium used mainly for the industrial production of amino acids, is regulated by seven different sigma factors of RNA polymerase, including the stress-responsive ECF-sigma factor SigH. The sigH gene is located in a gene cluster together with the rshA gene, putatively encoding an anti-sigma factor. The aim of this study was to analyze the transcriptional regulation of the sigH and rshA gene cluster and the effects of RshA on the SigH regulon, in order to refine the model describing the role of SigH and RshA during stress response. Transcription analyses revealed that the sigH gene and rshA gene are cotranscribed from four sigH housekeeping promoters in C. glutamicum. In addition, a SigH-controlled rshA promoter was found to only drive the transcription of the rshA gene. To test the role of the putative anti-sigma factor gene rshA under normal growth conditions, a C. glutamicum rshA deletion strain was constructed and used for genome-wide transcription profiling with DNA microarrays. In total, 83 genes organized in 61 putative transcriptional units, including those previously detected using sigH mutant strains, exhibited increased transcript levels in the rshA deletion mutant compared to its parental strain. The genes encoding proteins related to disulphide stress response, heat stress proteins, components of the SOS-response to DNA damage and proteasome components were the most markedly upregulated gene groups. Altogether six SigH-dependent promoters upstream of the identified genes were determined by primer extension and a refined consensus promoter consisting of 45 original promoter sequences was constructed. The rshA gene codes for an anti-sigma factor controlling the function of the stress-responsive sigma factor SigH in C. glutamicum. Transcription of rshA from a SigH-dependent promoter may serve to quickly shutdown the SigH-dependent stress response after the cells have overcome the stress condition. Here we propose a model of the regulation of oxidative and heat stress response including redox homeostasis by SigH, RshA and the thioredoxin system.
Identification of Common Differentially Expressed Genes in Urinary Bladder Cancer
Zaravinos, Apostolos; Lambrou, George I.; Boulalas, Ioannis; Delakas, Dimitris; Spandidos, Demetrios A.
2011-01-01
Background Current diagnosis and treatment of urinary bladder cancer (BC) has shown great progress with the utilization of microarrays. Purpose Our goal was to identify common differentially expressed (DE) genes among clinically relevant subclasses of BC using microarrays. Methodology/Principal Findings BC samples and controls, both experimental and publicly available datasets, were analyzed by whole genome microarrays. We grouped the samples according to their histology and defined the DE genes in each sample individually, as well as in each tumor group. A dual analysis strategy was followed. First, experimental samples were analyzed and conclusions were formulated; and second, experimental sets were combined with publicly available microarray datasets and were further analyzed in search of common DE genes. The experimental dataset identified 831 genes that were DE in all tumor samples, simultaneously. Moreover, 33 genes were up-regulated and 85 genes were down-regulated in all 10 BC samples compared to the 5 normal tissues, simultaneously. Hierarchical clustering partitioned tumor groups in accordance to their histology. K-means clustering of all genes and all samples, as well as clustering of tumor groups, presented 49 clusters. K-means clustering of common DE genes in all samples revealed 24 clusters. Genes manifested various differential patterns of expression, based on PCA. YY1 and NFκB were among the most common transcription factors that regulated the expression of the identified DE genes. Chromosome 1 contained 32 DE genes, followed by chromosomes 2 and 11, which contained 25 and 23 DE genes, respectively. Chromosome 21 had the least number of DE genes. GO analysis revealed the prevalence of transport and binding genes in the common down-regulated DE genes; the prevalence of RNA metabolism and processing genes in the up-regulated DE genes; as well as the prevalence of genes responsible for cell communication and signal transduction in the DE genes that were down-regulated in T1-Grade III tumors and up-regulated in T2/T3-Grade III tumors. Combination of samples from all microarray platforms revealed 17 common DE genes, (BMP4, CRYGD, DBH, GJB1, KRT83, MPZ, NHLH1, TACR3, ACTC1, MFAP4, SPARCL1, TAGLN, TPM2, CDC20, LHCGR, TM9SF1 and HCCS) 4 of which participate in numerous pathways. Conclusions/Significance The identification of the common DE genes among BC samples of different histology can provide further insight into the discovery of new putative markers. PMID:21483740
Symmetric nonnegative matrix factorization: algorithms and applications to probabilistic clustering.
He, Zhaoshui; Xie, Shengli; Zdunek, Rafal; Zhou, Guoxu; Cichocki, Andrzej
2011-12-01
Nonnegative matrix factorization (NMF) is an unsupervised learning method useful in various applications including image processing and semantic analysis of documents. This paper focuses on symmetric NMF (SNMF), which is a special case of NMF decomposition. Three parallel multiplicative update algorithms using level 3 basic linear algebra subprograms directly are developed for this problem. First, by minimizing the Euclidean distance, a multiplicative update algorithm is proposed, and its convergence under mild conditions is proved. Based on it, we further propose another two fast parallel methods: α-SNMF and β -SNMF algorithms. All of them are easy to implement. These algorithms are applied to probabilistic clustering. We demonstrate their effectiveness for facial image clustering, document categorization, and pattern clustering in gene expression.
Persson, Tomas; Battenberg, Kai; Demina, Irina V.; Vigil-Stenman, Theoden; Vanden Heuvel, Brian; Pujic, Petar; Facciotti, Marc T.; Wilbanks, Elizabeth G.; O'Brien, Anna; Fournier, Pascale; Cruz Hernandez, Maria Antonia; Mendoza Herrera, Alberto; Médigue, Claudine; Normand, Philippe; Pawlowski, Katharina; Berry, Alison M.
2015-01-01
Frankia strains are nitrogen-fixing soil actinobacteria that can form root symbioses with actinorhizal plants. Phylogenetically, symbiotic frankiae can be divided into three clusters, and this division also corresponds to host specificity groups. The strains of cluster II which form symbioses with actinorhizal Rosales and Cucurbitales, thus displaying a broad host range, show suprisingly low genetic diversity and to date can not be cultured. The genome of the first representative of this cluster, Candidatus Frankia datiscae Dg1 (Dg1), a microsymbiont of Datisca glomerata, was recently sequenced. A phylogenetic analysis of 50 different housekeeping genes of Dg1 and three published Frankia genomes showed that cluster II is basal among the symbiotic Frankia clusters. Detailed analysis showed that nodules of D. glomerata, independent of the origin of the inoculum, contain several closely related cluster II Frankia operational taxonomic units. Actinorhizal plants and legumes both belong to the nitrogen-fixing plant clade, and bacterial signaling in both groups involves the common symbiotic pathway also used by arbuscular mycorrhizal fungi. However, so far, no molecules resembling rhizobial Nod factors could be isolated from Frankia cultures. Alone among Frankia genomes available to date, the genome of Dg1 contains the canonical nod genes nodA, nodB and nodC known from rhizobia, and these genes are arranged in two operons which are expressed in D. glomerata nodules. Furthermore, Frankia Dg1 nodC was able to partially complement a Rhizobium leguminosarum A34 nodC::Tn5 mutant. Phylogenetic analysis showed that Dg1 Nod proteins are positioned at the root of both α- and β-rhizobial NodABC proteins. NodA-like acyl transferases were found across the phylum Actinobacteria, but among Proteobacteria only in nodulators. Taken together, our evidence indicates an Actinobacterial origin of rhizobial Nod factors. PMID:26020781
Pajoohesh-Ganji, Ahdeah; Knoblach, Susan M.; Faden, Alan I.; Byrnes, Kimberly R.
2012-01-01
Inflammation has long been implicated in secondary tissue damage after spinal cord injury (SCI). Our previous studies of inflammatory gene expression in rats after SCI revealed two temporally correlated clusters: the first was expressed early after injury and the second was up-regulated later, with peak expression at 1–2 weeks and persistent up-regulation through 6 months. To further address the role of inflammation after SCI, we examined inflammatory genes in a second species, mice, through 28 days after SCI. Using anchor gene clustering analysis, we found similar expression patterns for both the acute and chronic gene clusters previously identified after rat SCI. The acute group returned to normal expression levels by 7 days post-injury. The chronic group, which included C1qB, p22phox and galectin-3, showed peak expression at 7 days and remained up-regulated through 28 days. Immunohistochemistry and western blot analysis showed that the protein expression of these genes was consistent with the mRNA expression. Further exploration of the role of one of these genes, galectin-3, suggests that galectin-3 may contribute to secondary injury. In summary, our findings extend our prior gene profiling data by demonstrating the chronic expression of a cluster of microglial associated inflammatory genes after SCI in mice. Moreover, by demonstrating that inhibition of one such factor improves recovery, the findings suggest that such chronic up-regulation of inflammatory processes may contribute to secondary tissue damage after SCI, and that there may be a broader therapeutic window for neuroprotection than generally accepted. PMID:22884909
Wada, Masayoshi; Takahashi, Hiroki; Altaf-Ul-Amin, Md; Nakamura, Kensuke; Hirai, Masami Y; Ohta, Daisaku; Kanaya, Shigehiko
2012-07-15
Operon-like arrangements of genes occur in eukaryotes ranging from yeasts and filamentous fungi to nematodes, plants, and mammals. In plants, several examples of operon-like gene clusters involved in metabolic pathways have recently been characterized, e.g. the cyclic hydroxamic acid pathways in maize, the avenacin biosynthesis gene clusters in oat, the thalianol pathway in Arabidopsis thaliana, and the diterpenoid momilactone cluster in rice. Such operon-like gene clusters are defined by their co-regulation or neighboring positions within immediate vicinity of chromosomal regions. A comprehensive analysis of the expression of neighboring genes therefore accounts a crucial step to reveal the complete set of operon-like gene clusters within a genome. Genome-wide prediction of operon-like gene clusters should contribute to functional annotation efforts and provide novel insight into evolutionary aspects acquiring certain biological functions as well. We predicted co-expressed gene clusters by comparing the Pearson correlation coefficient of neighboring genes and randomly selected gene pairs, based on a statistical method that takes false discovery rate (FDR) into consideration for 1469 microarray gene expression datasets of A. thaliana. We estimated that A. thaliana contains 100 operon-like gene clusters in total. We predicted 34 statistically significant gene clusters consisting of 3 to 22 genes each, based on a stringent FDR threshold of 0.1. Functional relationships among genes in individual clusters were estimated by sequence similarity and functional annotation of genes. Duplicated gene pairs (determined based on BLAST with a cutoff of E<10(-5)) are included in 27 clusters. Five clusters are associated with metabolism, containing P450 genes restricted to the Brassica family and predicted to be involved in secondary metabolism. Operon-like clusters tend to include genes encoding bio-machinery associated with ribosomes, the ubiquitin/proteasome system, secondary metabolic pathways, lipid and fatty-acid metabolism, and the lipid transfer system. Copyright © 2012 Elsevier B.V. All rights reserved.
Pan, Yufang; Li, Qiaofeng; Wang, Zhizheng; Wang, Yang; Ma, Rui; Zhu, Lili; He, Guangcun; Chen, Rongzhi
2014-12-16
Thermosensitive genic male sterile (TGMS) lines and photoperiod-sensitive genic male sterile (PGMS) lines have been successfully used in hybridization to improve rice yields. However, the molecular mechanisms underlying male sterility transitions in most PGMS/TGMS rice lines are unclear. In the recently developed TGMS-Co27 line, the male sterility is based on co-suppression of a UDP-glucose pyrophosphorylase gene (Ugp1), but further study is needed to fully elucidate the molecular mechanisms involved. Microarray-based transcriptome profiling of TGMS-Co27 and wild-type Hejiang 19 (H1493) plants grown at high and low temperatures revealed that 15462 probe sets representing 8303 genes were differentially expressed in the two lines, under the two conditions, or both. Environmental factors strongly affected global gene expression. Some genes important for pollen development were strongly repressed in TGMS-Co27 at high temperature. More significantly, series-cluster analysis of differentially expressed genes (DEGs) between TGMS-Co27 plants grown under the two conditions showed that low temperature induced the expression of a gene cluster. This cluster was found to be essential for sterility transition. It includes many meiosis stage-related genes that are probably important for thermosensitive male sterility in TGMS-Co27, inter alia: Arg/Ser-rich domain (RS)-containing zinc finger proteins, polypyrimidine tract-binding proteins (PTBs), DEAD/DEAH box RNA helicases, ZOS (C2H2 zinc finger proteins of Oryza sativa), at least one polyadenylate-binding protein and some other RNA recognition motif (RRM) domain-containing proteins involved in post-transcriptional processes, eukaryotic initiation factor 5B (eIF5B), ribosomal proteins (L37, L1p/L10e, L27 and L24), aminoacyl-tRNA synthetases (ARSs), eukaryotic elongation factor Tu (eEF-Tu) and a peptide chain release factor protein involved in translation. The differential expression of 12 DEGs that are important for pollen development, low temperature responses or TGMS was validated by quantitative RT-PCR (qRT-PCR). Temperature strongly affects global gene expression and may be the common regulator of fertility in PGMS/TGMS rice lines. The identified expression changes reflect perturbations in the transcriptomic regulation of pollen development networks in TGMS-Co27. Findings from this and previous studies indicate that sets of genes involved in post-transcriptional and translation processes are involved in thermosensitive male sterility transitions in TGMS-Co27.
Peterson, Leif E
2002-01-01
CLUSFAVOR (CLUSter and Factor Analysis with Varimax Orthogonal Rotation) 5.0 is a Windows-based computer program for hierarchical cluster and principal-component analysis of microarray-based transcriptional profiles. CLUSFAVOR 5.0 standardizes input data; sorts data according to gene-specific coefficient of variation, standard deviation, average and total expression, and Shannon entropy; performs hierarchical cluster analysis using nearest-neighbor, unweighted pair-group method using arithmetic averages (UPGMA), or furthest-neighbor joining methods, and Euclidean, correlation, or jack-knife distances; and performs principal-component analysis. PMID:12184816
Intact cluster and chordate-like expression of ParaHox genes in a sea star
2013-01-01
Background The ParaHox genes are thought to be major players in patterning the gut of several bilaterian taxa. Though this is a fundamental role that these transcription factors play, their activities are not limited to the endoderm and extend to both ectodermal and mesodermal tissues. Three genes compose the ParaHox group: Gsx, Xlox and Cdx. In some taxa (mostly chordates but to some degree also in protostomes) the three genes are arranged into a genomic cluster, in a similar fashion to what has been shown for the better-known Hox genes. Sea urchins possess the full complement of ParaHox genes but they are all dispersed throughout the genome, an arrangement that, perhaps, represented the primitive condition for all echinoderms. In order to understand the evolutionary history of this group of genes we cloned and characterized all ParaHox genes, studied their expression patterns and identified their genomic loci in a member of an earlier branching group of echinoderms, the asteroid Patiria miniata. Results We identified the three ParaHox orthologs in the genome of P. miniata. While one of them, PmGsx is provided as maternal message, with no zygotic activation afterwards, the other two, PmLox and PmCdx are expressed during embryogenesis, within restricted domains of both endoderm and ectoderm. Screening of a Patiria bacterial artificial chromosome (BAC) library led to the identification of a clone containing the three genes. The transcriptional directions of PmGsx and PmLox are opposed to that of the PmCdx gene within the cluster. Conclusions The identification of P. miniata ParaHox genes has revealed the fact that these genes are clustered in the genome, in contrast to what has been reported for echinoids. Since the presence of an intact cluster, or at least a partial cluster, has been reported in chordates and polychaetes respectively, it becomes clear that within echinoderms, sea urchins have modified the original bilaterian arrangement. Moreover, the sea star ParaHox domains of expression show chordate-like features not found in the sea urchin, confirming that the dynamics of gene expression for the respective genes and their putative regulatory interactions have clearly changed over evolutionary time within the echinoid lineage. PMID:23803323
Bioinformatic prediction of leader genes in human periodontitis.
Covani, Ugo; Marconcini, Simone; Giacomelli, Luca; Sivozhelevov, Victor; Barone, Antonio; Nicolini, Claudio
2008-10-01
Genes involved in different biologic processes form complex interaction networks. However, only a few have a high number of interactions with the other genes in the network. In previous bioinformatics and experimental studies concerning the T lymphocyte cell cycle, these genes were identified and termed "leader genes." In this work, genes involved in human periodontitis were tentatively identified and ranked according to their number of interactions to obtain a preliminary, broader view of molecular mechanisms of periodontitis and plan targeted experimentation. Genes were identified with interrelated queries of several databases. The interactions among these genes were mapped and given a significance score. The weighted number of links (weighted sum of scores for every interaction in which the given gene is involved) was calculated for each gene. Genes were clustered according to this parameter. The genes in the highest cluster were termed leader genes. Sixty-one genes involved or potentially involved in periodontitis were identified. Only five were identified as leader genes, whereas 12 others were ranked in an immediately lower cluster. For 10 of 17 genes there is evidence of involvement in periodontitis; seven new genes that are potentially involved in this disease were identified. The involvement in periodontitis has been completely established for only two leader genes. We applied a validated bioinformatics algorithm to increase our knowledge of molecular mechanisms of periodontitis. Even with the limitations of this ab initio analysis, this theoretical study can suggest ad hoc experimentation targeted on significant genes and, therefore, simpler than mass-scale molecular genomics. Moreover, the identification of leader genes might suggest new potential risk factors and therapeutic targets.
Molecular analysis of SCARECROW genes expressed in white lupin cluster roots
Sbabou, Laila; Bucciarelli, Bruna; Miller, Susan; Liu, Junqi; Berhada, Fatiha; Filali-Maltouf, Abdelkarim; Allan, Deborah; Vance, Carroll
2010-01-01
The Scarecrow (SCR) transcription factor plays a crucial role in root cell radial patterning and is required for maintenance of the quiescent centre and differentiation of the endodermis. In response to phosphorus (P) deficiency, white lupin (Lupinus albus L.) root surface area increases some 50-fold to 70-fold due to the development of cluster (proteoid) roots. Previously it was reported that SCR-like expressed sequence tags (ESTs) were expressed during early cluster root development. Here the cloning of two white lupin SCR genes, LaSCR1 and LaSCR2, is reported. The predicted amino acid sequences of both LaSCR gene products are highly similar to AtSCR and contain C-terminal conserved GRAS family domains. LaSCR1 and LaSCR2 transcript accumulation localized to the endodermis of both normal and cluster roots as shown by in situ hybridization and gene promoter::reporter staining. Transcript analysis as evaluated by quantitative real-time-PCR (qRT-PCR) and RNA gel hybridization indicated that the two LaSCR genes are expressed predominantly in roots. Expression of LaSCR genes was not directly responsive to the P status of the plant but was a function of cluster root development. Suppression of LaSCR1 in transformed roots of lupin and Medicago via RNAi (RNA interference) delivered through Agrobacterium rhizogenes resulted in decreased root numbers, reflecting the potential role of LaSCR1 in maintaining root growth in these species. The results suggest that the functional orthologues of AtSCR have been characterized. PMID:20167612
Weighted gene co-expression network analysis of gene modules for the prognosis of esophageal cancer.
Zhang, Cong; Sun, Qian
2017-06-01
Esophageal cancer is a common malignant tumor, whose pathogenesis and prognosis factors are not fully understood. This study aimed to discover the gene clusters that have similar functions and can be used to predict the prognosis of esophageal cancer. The matched microarray and RNA sequencing data of 185 patients with esophageal cancer were downloaded from The Cancer Genome Atlas (TCGA), and gene co-expression networks were built without distinguishing between squamous carcinoma and adenocarcinoma. The result showed that 12 modules were associated with one or more survival data such as recurrence status, recurrence time, vital status or vital time. Furthermore, survival analysis showed that 5 out of the 12 modules were related to progression-free survival (PFS) or overall survival (OS). As the most important module, the midnight blue module with 82 genes was related to PFS, apart from the patient age, tumor grade, primary treatment success, and duration of smoking and tumor histological type. Gene ontology enrichment analysis revealed that "glycoprotein binding" was the top enriched function of midnight blue module genes. Additionally, the blue module was the exclusive gene clusters related to OS. Platelet activating factor receptor (PTAFR) and feline Gardner-Rasheed (FGR) were the top hub genes in both modeling datasets and the STRING protein interaction database. In conclusion, our study provides novel insights into the prognosis-associated genes and screens out candidate biomarkers for esophageal cancer.
NASA Astrophysics Data System (ADS)
Guo, Jingyu; Tian, Dehua; McKinney, Brett A.; Hartman, John L.
2010-06-01
Interactions between genetic and/or environmental factors are ubiquitous, affecting the phenotypes of organisms in complex ways. Knowledge about such interactions is becoming rate-limiting for our understanding of human disease and other biological phenomena. Phenomics refers to the integrative analysis of how all genes contribute to phenotype variation, entailing genome and organism level information. A systems biology view of gene interactions is critical for phenomics. Unfortunately the problem is intractable in humans; however, it can be addressed in simpler genetic model systems. Our research group has focused on the concept of genetic buffering of phenotypic variation, in studies employing the single-cell eukaryotic organism, S. cerevisiae. We have developed a methodology, quantitative high throughput cellular phenotyping (Q-HTCP), for high-resolution measurements of gene-gene and gene-environment interactions on a genome-wide scale. Q-HTCP is being applied to the complete set of S. cerevisiae gene deletion strains, a unique resource for systematically mapping gene interactions. Genetic buffering is the idea that comprehensive and quantitative knowledge about how genes interact with respect to phenotypes will lead to an appreciation of how genes and pathways are functionally connected at a systems level to maintain homeostasis. However, extracting biologically useful information from Q-HTCP data is challenging, due to the multidimensional and nonlinear nature of gene interactions, together with a relative lack of prior biological information. Here we describe a new approach for mining quantitative genetic interaction data called recursive expectation-maximization clustering (REMc). We developed REMc to help discover phenomic modules, defined as sets of genes with similar patterns of interaction across a series of genetic or environmental perturbations. Such modules are reflective of buffering mechanisms, i.e., genes that play a related role in the maintenance of physiological homeostasis. To develop the method, 297 gene deletion strains were selected based on gene-drug interactions with hydroxyurea, an inhibitor of ribonucleotide reductase enzyme activity, which is critical for DNA synthesis. To partition the gene functions, these 297 deletion strains were challenged with growth inhibitory drugs known to target different genes and cellular pathways. Q-HTCP-derived growth curves were used to quantify all gene interactions, and the data were used to test the performance of REMc. Fundamental advantages of REMc include objective assessment of total number of clusters and assignment to each cluster a log-likelihood value, which can be considered an indicator of statistical quality of clusters. To assess the biological quality of clusters, we developed a method called gene ontology information divergence z-score (GOid_z). GOid_z summarizes total enrichment of GO attributes within individual clusters. Using these and other criteria, we compared the performance of REMc to hierarchical and K-means clustering. The main conclusion is that REMc provides distinct efficiencies for mining Q-HTCP data. It facilitates identification of phenomic modules, which contribute to buffering mechanisms that underlie cellular homeostasis and the regulation of phenotypic expression.
Fang, Wei; Si, Yaqing; Douglass, Stephen; Casero, David; Merchant, Sabeeha S.; Pellegrini, Matteo; Ladunga, Istvan; Liu, Peng; Spalding, Martin H.
2012-01-01
We used RNA sequencing to query the Chlamydomonas reinhardtii transcriptome for regulation by CO2 and by the transcription regulator CIA5 (CCM1). Both CO2 and CIA5 are known to play roles in acclimation to low CO2 and in induction of an essential CO2-concentrating mechanism (CCM), but less is known about their interaction and impact on the whole transcriptome. Our comparison of the transcriptome of a wild type versus a cia5 mutant strain under three different CO2 conditions, high CO2 (5%), low CO2 (0.03 to 0.05%), and very low CO2 (<0.02%), provided an entry into global changes in the gene expression patterns occurring in response to the interaction between CO2 and CIA5. We observed a massive impact of CIA5 and CO2 on the transcriptome, affecting almost 25% of all Chlamydomonas genes, and we discovered an array of gene clusters with distinctive expression patterns that provide insight into the regulatory interaction between CIA5 and CO2. Several individual clusters respond primarily to either CIA5 or CO2, providing access to genes regulated by one factor but decoupled from the other. Three distinct clusters clearly associated with CCM-related genes may represent a rich source of candidates for new CCM components, including a small cluster of genes encoding putative inorganic carbon transporters. PMID:22634760
Chromatin-Specific Regulation of Mammalian rDNA Transcription by Clustered TTF-I Binding Sites
Diermeier, Sarah D.; Németh, Attila; Rehli, Michael; Grummt, Ingrid; Längst, Gernot
2013-01-01
Enhancers and promoters often contain multiple binding sites for the same transcription factor, suggesting that homotypic clustering of binding sites may serve a role in transcription regulation. Here we show that clustering of binding sites for the transcription termination factor TTF-I downstream of the pre-rRNA coding region specifies transcription termination, increases the efficiency of transcription initiation and affects the three-dimensional structure of rRNA genes. On chromatin templates, but not on free rDNA, clustered binding sites promote cooperative binding of TTF-I, loading TTF-I to the downstream terminators before it binds to the rDNA promoter. Interaction of TTF-I with target sites upstream and downstream of the rDNA transcription unit connects these distal DNA elements by forming a chromatin loop between the rDNA promoter and the terminators. The results imply that clustered binding sites increase the binding affinity of transcription factors in chromatin, thus influencing the timing and strength of DNA-dependent processes. PMID:24068958
A systematic analysis of genomic changes in Tg2576 mice.
Tan, Lu; Wang, Xiong; Ni, Zhong-Fei; Zhu, Xiuming; Wu, Wei; Zhu, Ling-Qiang; Liu, Dan
2013-06-01
Alzheimer's disease (AD) is an age-related neurodegenerative disorder characterized by intelligence decline, behavioral disorders and cognitive disability. The purpose of this study was to investigate gene expression in AD, based on published microarray data on Tg2576 mice. Hierarchical Cluster Analysis and Gene Ontology were employed to group genes together on the basis of their product characteristics and annotation data. Genes with prominent alterations were clustered into apoptosis and axon guidance pathways. Based on our findings and those of previous studies, we propose that the mitochondria-mediated apoptotic pathway plays a crucial role in the neuronal loss and synaptic dysfunction associated with AD. Furthermore, based on the findings of Positional Gene Enrichment analysis and Gene Set Enrichment analysis, we propose that the regulation of transcription of AD genes may be an important pathogenic factor in this neurodegenerative disease. Our results highlight the importance of genes that could subsequently be examined for their potential as prognostic markers for AD.
Omura, S; Ikeda, H; Ishikawa, J; Hanamoto, A; Takahashi, C; Shinose, M; Takahashi, Y; Horikawa, H; Nakazawa, H; Osonoe, T; Kikuchi, H; Shiba, T; Sakaki, Y; Hattori, M
2001-10-09
Streptomyces avermitilis is a soil bacterium that carries out not only a complex morphological differentiation but also the production of secondary metabolites, one of which, avermectin, is commercially important in human and veterinary medicine. The major interest in this genus Streptomyces is the diversity of its production of secondary metabolites as an industrial microorganism. A major factor in its prominence as a producer of the variety of secondary metabolites is its possession of several metabolic pathways for biosynthesis. Here we report sequence analysis of S. avermitilis, covering 99% of its genome. At least 8.7 million base pairs exist in the linear chromosome; this is the largest bacterial genome sequence, and it provides insights into the intrinsic diversity of the production of the secondary metabolites of Streptomyces. Twenty-five kinds of secondary metabolite gene clusters were found in the genome of S. avermitilis. Four of them are concerned with the biosyntheses of melanin pigments, in which two clusters encode tyrosinase and its cofactor, another two encode an ochronotic pigment derived from homogentiginic acid, and another polyketide-derived melanin. The gene clusters for carotenoid and siderophore biosyntheses are composed of seven and five genes, respectively. There are eight kinds of gene clusters for type-I polyketide compound biosyntheses, and two clusters are involved in the biosyntheses of type-II polyketide-derived compounds. Furthermore, a polyketide synthase that resembles phloroglucinol synthase was detected. Eight clusters are involved in the biosyntheses of peptide compounds that are synthesized by nonribosomal peptide synthetases. These secondary metabolite clusters are widely located in the genome but half of them are near both ends of the genome. The total length of these clusters occupies about 6.4% of the genome.
Ramamoorthy, Vellaisamy; Dhingra, Sourabh; Kincaid, Alexander; Shantappa, Sourabha; Feng, Xuehuan; Calvo, Ana M.
2013-01-01
Secondary metabolism in the model fungus Aspergillus nidulans is controlled by the conserved global regulator VeA, which also governs morphological differentiation. Among the secondary metabolites regulated by VeA is the mycotoxin sterigmatocystin (ST). The presence of VeA is necessary for the biosynthesis of this carcinogenic compound. We identified a revertant mutant able to synthesize ST intermediates in the absence of VeA. The point mutation occurred at the coding region of a gene encoding a novel putative C2H2 zinc finger domain transcription factor that we denominated mtfA. The A. nidulans mtfA gene product localizes at nuclei independently of the illumination regime. Deletion of the mtfA gene restores mycotoxin biosynthesis in the absence of veA, but drastically reduced mycotoxin production when mtfA gene expression was altered, by deletion or overexpression, in A. nidulans strains with a veA wild-type allele. Our study revealed that mtfA regulates ST production by affecting the expression of the specific ST gene cluster activator aflR. Importantly, mtfA is also a regulator of other secondary metabolism gene clusters, such as genes responsible for the synthesis of terrequinone and penicillin. As in the case of ST, deletion or overexpression of mtfA was also detrimental for the expression of terrequinone genes. Deletion of mtfA also decreased the expression of the genes in the penicillin gene cluster, reducing penicillin production. However, in this case, over-expression of mtfA enhanced the transcription of penicillin genes, increasing penicillin production more than 5 fold with respect to the control. Importantly, in addition to its effect on secondary metabolism, mtfA also affects asexual and sexual development in A. nidulans. Deletion of mtfA results in a reduction of conidiation and sexual stage. We found mtfA putative orthologs conserved in other fungal species. PMID:24066102
Fractal Clustering and Knowledge-driven Validation Assessment for Gene Expression Profiling.
Wang, Lu-Yong; Balasubramanian, Ammaiappan; Chakraborty, Amit; Comaniciu, Dorin
2005-01-01
DNA microarray experiments generate a substantial amount of information about the global gene expression. Gene expression profiles can be represented as points in multi-dimensional space. It is essential to identify relevant groups of genes in biomedical research. Clustering is helpful in pattern recognition in gene expression profiles. A number of clustering techniques have been introduced. However, these traditional methods mainly utilize shape-based assumption or some distance metric to cluster the points in multi-dimension linear Euclidean space. Their results shows poor consistence with the functional annotation of genes in previous validation study. From a novel different perspective, we propose fractal clustering method to cluster genes using intrinsic (fractal) dimension from modern geometry. This method clusters points in such a way that points in the same clusters are more self-affine among themselves than to the points in other clusters. We assess this method using annotation-based validation assessment for gene clusters. It shows that this method is superior in identifying functional related gene groups than other traditional methods.
Sá-Pinto, Alexandra; Branco, Madalena S.; Alexandrino, Paulo B.; Fontaine, Michaël C.; Baird, Stuart J. E.
2012-01-01
Knowledge of the scale of dispersal and the mechanisms governing gene flow in marine environments remains fragmentary despite being essential for understanding evolution of marine biota and to design management plans. We use the limpets Patella ulyssiponensis and Patella rustica as models for identifying factors affecting gene flow in marine organisms across the North-East Atlantic and the Mediterranean Sea. A set of allozyme loci and a fragment of the mitochondrial gene cytochrome C oxidase subunit I were screened for genetic variation through starch gel electrophoresis and DNA sequencing, respectively. An approach combining clustering algorithms with clinal analyses was used to test for the existence of barriers to gene flow and estimate their geographic location and abruptness. Sharp breaks in the genetic composition of individuals were observed in the transitions between the Atlantic and the Mediterranean and across southern Italian shores. An additional break within the Atlantic cluster separates samples from the Alboran Sea and Atlantic African shores from those of the Iberian Atlantic shores. The geographic congruence of the genetic breaks detected in these two limpet species strongly supports the existence of transpecific barriers to gene flow in the Mediterranean Sea and Northeastern Atlantic. This leads to testable hypotheses regarding factors restricting gene flow across the study area. PMID:23239977
Microarray analysis of gene expression in West Nile virus–infected human retinal pigment epithelium
Munoz-Erazo, Luis; Natoli, Ricardo; Provis, Jan Marie; Madigan, Michelle Catherine
2012-01-01
Purpose To identify key genes differentially expressed in the human retinal pigment epithelium (hRPE) following low-level West Nile virus (WNV) infection. Methods Primary hRPE and retinal pigment epithelium cell line (ARPE-19) cells were infected with WNV (multiplicity of infection 1). RNA extracted from mock-infected and WNV-infected cells was assessed for differential expression of genes using Affymetrix microarray. Quantitative real-time PCR analysis of 23 genes was used to validate the microarray results. Results Functional annotation clustering of the microarray data showed that gene clusters involved in immune and antiviral responses ranked highly, involving genes such as chemokine (C-C motif) ligand 2 (CCL2), chemokine (C-C motif) ligand 5 (CCL5), chemokine (C-X-C motif) ligand 10 (CXCL10), and toll like receptor 3 (TLR3). In conjunction with the quantitative real-time PCR analysis, other novel genes regulated by WNV infection included indoleamine 2,3-dioxygenase (IDO1), genes involved in the transforming growth factor–β pathway (bone morphogenetic protein and activin membrane-bound inhibitor homolog [BAMBI] and activating transcription factor 3 [ATF3]), and genes involved in apoptosis (tumor necrosis factor receptor superfamily, member 10d [TNFRSF10D]). WNV-infected RPE did not produce any interferon-γ, suggesting that IDO1 is induced by other soluble factors, by the virus alone, or both. Conclusions Low-level WNV infection of hRPE cells induced expression of genes that are typically associated with the host cell response to virus infection. We also identified other genes, including IDO1 and BAMBI, that may influence the RPE and therefore outer blood-retinal barrier integrity during ocular infection and inflammation, or are associated with degeneration, as seen for example in aging. PMID:22509103
Microarray characterization of gene expression changes in blood during acute ethanol exposure
2013-01-01
Background As part of the civil aviation safety program to define the adverse effects of ethanol on flying performance, we performed a DNA microarray analysis of human whole blood samples from a five-time point study of subjects administered ethanol orally, followed by breathalyzer analysis, to monitor blood alcohol concentration (BAC) to discover significant gene expression changes in response to the ethanol exposure. Methods Subjects were administered either orange juice or orange juice with ethanol. Blood samples were taken based on BAC and total RNA was isolated from PaxGene™ blood tubes. The amplified cDNA was used in microarray and quantitative real-time polymerase chain reaction (RT-qPCR) analyses to evaluate differential gene expression. Microarray data was analyzed in a pipeline fashion to summarize and normalize and the results evaluated for relative expression across time points with multiple methods. Candidate genes showing distinctive expression patterns in response to ethanol were clustered by pattern and further analyzed for related function, pathway membership and common transcription factor binding within and across clusters. RT-qPCR was used with representative genes to confirm relative transcript levels across time to those detected in microarrays. Results Microarray analysis of samples representing 0%, 0.04%, 0.08%, return to 0.04%, and 0.02% wt/vol BAC showed that changes in gene expression could be detected across the time course. The expression changes were verified by qRT-PCR. The candidate genes of interest (GOI) identified from the microarray analysis and clustered by expression pattern across the five BAC points showed seven coordinately expressed groups. Analysis showed function-based networks, shared transcription factor binding sites and signaling pathways for members of the clusters. These include hematological functions, innate immunity and inflammation functions, metabolic functions expected of ethanol metabolism, and pancreatic and hepatic function. Five of the seven clusters showed links to the p38 MAPK pathway. Conclusions The results of this study provide a first look at changing gene expression patterns in human blood during an acute rise in blood ethanol concentration and its depletion because of metabolism and excretion, and demonstrate that it is possible to detect changes in gene expression using total RNA isolated from whole blood. The analysis approach for this study serves as a workflow to investigate the biology linked to expression changes across a time course and from these changes, to identify target genes that could serve as biomarkers linked to pilot performance. PMID:23883607
Temporal Changes in Gene Expression after Injury in the Rat Retina
Vázquez-Chona, Félix; Song, Bong K.; Geisert, Eldon E.
2010-01-01
Purpose The goal of this study was to define the temporal changes in gene expression after retinal injury and to relate these changes to the inflammatory and reactive response. A specific emphasis was placed on the tetraspanin family of proteins and their relationship with markers of reactive gliosis. Methods Retinal tears were induced in adult rats by scraping the retina with a needle. After different survival times (4 hours, and 1, 3, 7, and 30 days), the retinas were removed, and mRNA was isolated, prepared, and hybridized to the Affymatrix RGU34A microarray (Santa Clara, CA). Microarray results were confirmed by using RT-PCR and correlation to protein levels was determined. Results Of the 8750 genes analyzed, approximately 393 (4.5%) were differentially expressed. Clustering analysis revealed three major profiles: (1) The early response was characterized by the upregulation of transcription factors; (2) the delayed response included a high percentage of genes related to cell cycle and cell death; and (3) the late, sustained profile clustered a significant number of genes involved in retinal gliosis. The late, sustained cluster also contained the upregulated crystallin genes. The tetraspanins Cd9, Cd81, and Cd82 were also associated with the late, sustained response. Conclusions The use of microarray technology enables definition of complex genetic changes underlying distinct phases of the cellular response to retinal injury. The early response clusters genes associate with the transcriptional regulation of the wound-healing process and cell death. Most of the genes in the late, sustained response appear to be associated with reactive gliosis. PMID:15277499
Finding gene clusters for a replicated time course study
2014-01-01
Background Finding genes that share similar expression patterns across samples is an important question that is frequently asked in high-throughput microarray studies. Traditional clustering algorithms such as K-means clustering and hierarchical clustering base gene clustering directly on the observed measurements and do not take into account the specific experimental design under which the microarray data were collected. A new model-based clustering method, the clustering of regression models method, takes into account the specific design of the microarray study and bases the clustering on how genes are related to sample covariates. It can find useful gene clusters for studies from complicated study designs such as replicated time course studies. Findings In this paper, we applied the clustering of regression models method to data from a time course study of yeast on two genotypes, wild type and YOX1 mutant, each with two technical replicates, and compared the clustering results with K-means clustering. We identified gene clusters that have similar expression patterns in wild type yeast, two of which were missed by K-means clustering. We further identified gene clusters whose expression patterns were changed in YOX1 mutant yeast compared to wild type yeast. Conclusions The clustering of regression models method can be a valuable tool for identifying genes that are coordinately transcribed by a common mechanism. PMID:24460656
Dolezal, Tomas; Gazi, Michal; Zurovec, Michal; Bryant, Peter J
2003-10-01
Many Drosophila genes exist as members of multigene families and within each family the members can be functionally redundant, making it difficult to identify them by classical mutagenesis techniques based on phenotypic screening. We have addressed this problem in a genetic analysis of a novel family of six adenosine deaminase-related growth factors (ADGFs). We used ends-in targeting to introduce mutations into five of the six ADGF genes, taking advantage of the fact that five of the family members are encoded by a three-gene cluster and a two-gene cluster. We used two targeting constructs to introduce loss-of-function mutations into all five genes, as well as to isolate different combinations of multiple mutations, independent of phenotypic consequences. The results show that (1) it is possible to use ends-in targeting to disrupt gene clusters; (2) gene conversion, which is usually considered a complication in gene targeting, can be used to help recover different mutant combinations in a single screening procedure; (3) the reduction of duplication to a single copy by induction of a double-strand break is better explained by the single-strand annealing mechanism than by simple crossing over between repeats; and (4) loss of function of the most abundantly expressed family member (ADGF-A) leads to disintegration of the fat body and the development of melanotic tumors in mutant larvae.
Katoch, Meenu; Mazmouz, Rabia; Chau, Rocky; Pearson, Leanne A.; Pickford, Russell
2016-01-01
ABSTRACT Mycosporine-like amino acids (MAAs) are an important class of secondary metabolites known for their protection against UV radiation and other stress factors. Cyanobacteria produce a variety of MAAs, including shinorine, the active ingredient in many sunscreen creams. Bioinformatic analysis of the genome of the soil-dwelling cyanobacterium Cylindrospermum stagnale PCC 7417 revealed a new gene cluster with homology to MAA synthase from Nostoc punctiforme. This newly identified gene cluster is unusual because it has five biosynthesis genes (mylA to mylE), compared to the four found in other MAA gene clusters. Heterologous expression of mylA to mylE in Escherichia coli resulted in the production of mycosporine-lysine and the novel compound mycosporine-ornithine. To our knowledge, this is the first time these compounds have been heterologously produced in E. coli and structurally characterized via direct spectral guidance. This study offers insight into the diversity, biosynthesis, and structure of cyanobacterial MAAs and highlights their amenability to heterologous production methods. IMPORTANCE Mycosporine-like amino acids (MAAs) are significant from an environmental microbiological perspective as they offer microbes protection against a variety of stress factors, including UV radiation. The heterologous expression of MAAs in E. coli is also significant from a biotechnological perspective as MAAs are the active ingredient in next-generation sunscreens. PMID:27520810
Multiconstrained gene clustering based on generalized projections
2010-01-01
Background Gene clustering for annotating gene functions is one of the fundamental issues in bioinformatics. The best clustering solution is often regularized by multiple constraints such as gene expressions, Gene Ontology (GO) annotations and gene network structures. How to integrate multiple pieces of constraints for an optimal clustering solution still remains an unsolved problem. Results We propose a novel multiconstrained gene clustering (MGC) method within the generalized projection onto convex sets (POCS) framework used widely in image reconstruction. Each constraint is formulated as a corresponding set. The generalized projector iteratively projects the clustering solution onto these sets in order to find a consistent solution included in the intersection set that satisfies all constraints. Compared with previous MGC methods, POCS can integrate multiple constraints from different nature without distorting the original constraints. To evaluate the clustering solution, we also propose a new performance measure referred to as Gene Log Likelihood (GLL) that considers genes having more than one function and hence in more than one cluster. Comparative experimental results show that our POCS-based gene clustering method outperforms current state-of-the-art MGC methods. Conclusions The POCS-based MGC method can successfully combine multiple constraints from different nature for gene clustering. Also, the proposed GLL is an effective performance measure for the soft clustering solutions. PMID:20356386
A novel sodium bicarbonate cotransporter-like gene in an ancient duplicated region: SLC4A9 at 5q31
Lipovich, Leonard; Lynch, Eric D; Lee, Ming K; King, Mary-Claire
2001-01-01
Background: Sodium bicarbonate cotransporter (NBC) genes encode proteins that execute coupled Na+ and HCO3- transport across epithelial cell membranes. We report the discovery, characterization, and genomic context of a novel human NBC-like gene, SLC4A9, on chromosome 5q31. Results: SLC4A9 was initially discovered by genomic sequence annotation and further characterized by sequencing of long-insert cDNA library clones. The predicted protein of 990 amino acids has 12 transmembrane domains and high sequence similarity to other NBCs. The 23-exon gene has 14 known mRNA isoforms. In three regions, mRNA sequence variation is generated by the inclusion or exclusion of portions of an exon. Noncoding SLC4A9 cDNAs were recovered multiple times from different libraries. The 3' untranslated region is fragmented into six alternatively spliced exons and contains expressed Alu, LINE and MER repeats. SLC4A9 has two alternative stop codons and six polyadenylation sites. Its expression is largely restricted to the kidney. In silico approaches were used to characterize two additional novel SLC4A genes and to place SLC4A9 within the context of multiple paralogous gene clusters containing members of the epidermal growth factor (EGF), ankyrin (ANK) and fibroblast growth factor (FGF) families. Seven human EGF-SLC4A-ANK-FGF clusters were found. Conclusion: The novel sodium bicarbonate cotransporter-like gene SLC4A9 demonstrates abundant alternative mRNA processing. It belongs to a growing class of functionally diverse genes characterized by inefficient highly variable splicing. The evolutionary history of the EGF-SLC4A-ANK-FGF gene clusters involves multiple rounds of duplication, apparently followed by large insertions and deletions at paralogous loci and genome-wide gene shuffling. PMID:11305939
Takeda, Itaru; Umemura, Myco; Koike, Hideaki; Asai, Kiyoshi; Machida, Masayuki
2014-08-01
Despite their biological importance, a significant number of genes for secondary metabolite biosynthesis (SMB) remain undetected due largely to the fact that they are highly diverse and are not expressed under a variety of cultivation conditions. Several software tools including SMURF and antiSMASH have been developed to predict fungal SMB gene clusters by finding core genes encoding polyketide synthase, nonribosomal peptide synthetase and dimethylallyltryptophan synthase as well as several others typically present in the cluster. In this work, we have devised a novel comparative genomics method to identify SMB gene clusters that is independent of motif information of the known SMB genes. The method detects SMB gene clusters by searching for a similar order of genes and their presence in nonsyntenic blocks. With this method, we were able to identify many known SMB gene clusters with the core genes in the genomic sequences of 10 filamentous fungi. Furthermore, we have also detected SMB gene clusters without core genes, including the kojic acid biosynthesis gene cluster of Aspergillus oryzae. By varying the detection parameters of the method, a significant difference in the sequence characteristics was detected between the genes residing inside the clusters and those outside the clusters. © The Author 2014. Published by Oxford University Press on behalf of Kazusa DNA Research Institute.
Reimegård, Johan; Kundu, Snehangshu; Pendle, Ali; Irish, Vivian F.; Shaw, Peter
2017-01-01
Abstract Co-expression of physically linked genes occurs surprisingly frequently in eukaryotes. Such chromosomal clustering may confer a selective advantage as it enables coordinated gene regulation at the chromatin level. We studied the chromosomal organization of genes involved in male reproductive development in Arabidopsis thaliana. We developed an in-silico tool to identify physical clusters of co-regulated genes from gene expression data. We identified 17 clusters (96 genes) involved in stamen development and acting downstream of the transcriptional activator MS1 (MALE STERILITY 1), which contains a PHD domain associated with chromatin re-organization. The clusters exhibited little gene homology or promoter element similarity, and largely overlapped with reported repressive histone marks. Experiments on a subset of the clusters suggested a link between expression activation and chromatin conformation: qRT-PCR and mRNA in situ hybridization showed that the clustered genes were up-regulated within 48 h after MS1 induction; out of 14 chromatin-remodeling mutants studied, expression of clustered genes was consistently down-regulated only in hta9/hta11, previously associated with metabolic cluster activation; DNA fluorescence in situ hybridization confirmed that transcriptional activation of the clustered genes was correlated with open chromatin conformation. Stamen development thus appears to involve transcriptional activation of physically clustered genes through chromatin de-condensation. PMID:28175342
Constrained clusters of gene expression profiles with pathological features.
Sese, Jun; Kurokawa, Yukinori; Monden, Morito; Kato, Kikuya; Morishita, Shinichi
2004-11-22
Gene expression profiles should be useful in distinguishing variations in disease, since they reflect accurately the status of cells. The primary clustering of gene expression reveals the genotypes that are responsible for the proximity of members within each cluster, while further clustering elucidates the pathological features of the individual members of each cluster. However, since the first clustering process and the second classification step, in which the features are associated with clusters, are performed independently, the initial set of clusters may omit genes that are associated with pathologically meaningful features. Therefore, it is important to devise a way of identifying gene expression clusters that are associated with pathological features. We present the novel technique of 'itemset constrained clustering' (IC-Clustering), which computes the optimal cluster that maximizes the interclass variance of gene expression between groups, which are divided according to the restriction that only divisions that can be expressed using common features are allowed. This constraint automatically labels each cluster with a set of pathological features which characterize that cluster. When applied to liver cancer datasets, IC-Clustering revealed informative gene expression clusters, which could be annotated with various pathological features, such as 'tumor' and 'man', or 'except tumor' and 'normal liver function'. In contrast, the k-means method overlooked these clusters.
Whole blood genome-wide expression profiling and network analysis suggest MELAS master regulators.
Mende, Susanne; Royer, Loic; Herr, Alexander; Schmiedel, Janet; Deschauer, Marcus; Klopstock, Thomas; Kostic, Vladimir S; Schroeder, Michael; Reichmann, Heinz; Storch, Alexander
2011-07-01
The heteroplasmic mitochondrial DNA (mtDNA) mutation A3243G causes the mitochondrial encephalomyopathy, lactic acidosis, and stroke-like episodes (MELAS) syndrome as one of the most frequent mitochondrial diseases. The process of reconfiguration of nuclear gene expression profile to accommodate cellular processes to the functional status of mitochondria might be a key to MELAS disease manifestation and could contribute to its diverse phenotypic presentation. To determine master regulatory protein networks and disease-modifying genes in MELAS syndrome. Analyses of whole blood transcriptomes from 10 MELAS patients using a novel strategy by combining classic Affymetrix oligonucleotide microarray profiling with regulatory and protein interaction network analyses. Hierarchical cluster analysis elucidated that the relative abundance of mutant mtDNA molecules is decisive for the nuclear gene expression response. Further analyses confirmed not only transcription factors already known to be involved in mitochondrial diseases (such as TFAM), but also detected the hypoxia-inducible factor 1 complex, nuclear factor Y and cAMP responsive element-binding protein-related transcription factors as novel master regulators for reconfiguration of nuclear gene expression in response to the MELAS mutation. Correlation analyses of gene alterations and clinico-genetic data detected significant correlations between A3243G-induced nuclear gene expression changes and mutant mtDNA load as well as disease characteristics. These potential disease-modifying genes influencing the expression of the MELAS phenotype are mainly related to clusters primarily unrelated to cellular energy metabolism, but important for nucleic acid and protein metabolism, and signal transduction. Our data thus provide a framework to search for new pathogenetic concepts and potential therapeutic approaches to treat the MELAS syndrome.
Diametrical clustering for identifying anti-correlated gene clusters.
Dhillon, Inderjit S; Marcotte, Edward M; Roshan, Usman
2003-09-01
Clustering genes based upon their expression patterns allows us to predict gene function. Most existing clustering algorithms cluster genes together when their expression patterns show high positive correlation. However, it has been observed that genes whose expression patterns are strongly anti-correlated can also be functionally similar. Biologically, this is not unintuitive-genes responding to the same stimuli, regardless of the nature of the response, are more likely to operate in the same pathways. We present a new diametrical clustering algorithm that explicitly identifies anti-correlated clusters of genes. Our algorithm proceeds by iteratively (i). re-partitioning the genes and (ii). computing the dominant singular vector of each gene cluster; each singular vector serving as the prototype of a 'diametric' cluster. We empirically show the effectiveness of the algorithm in identifying diametrical or anti-correlated clusters. Testing the algorithm on yeast cell cycle data, fibroblast gene expression data, and DNA microarray data from yeast mutants reveals that opposed cellular pathways can be discovered with this method. We present systems whose mRNA expression patterns, and likely their functions, oppose the yeast ribosome and proteosome, along with evidence for the inverse transcriptional regulation of a number of cellular systems.
Disentangling the multigenic and pleiotropic nature of molecular function
2015-01-01
Background Biological processes at the molecular level are usually represented by molecular interaction networks. Function is organised and modularity identified based on network topology, however, this approach often fails to account for the dynamic and multifunctional nature of molecular components. For example, a molecule engaging in spatially or temporally independent functions may be inappropriately clustered into a single functional module. To capture biologically meaningful sets of interacting molecules, we use experimentally defined pathways as spatial/temporal units of molecular activity. Results We defined functional profiles of Saccharomyces cerevisiae based on a minimal set of Gene Ontology terms sufficient to represent each pathway's genes. The Gene Ontology terms were used to annotate 271 pathways, accounting for pathway multi-functionality and gene pleiotropy. Pathways were then arranged into a network, linked by shared functionality. Of the genes in our data set, 44% appeared in multiple pathways performing a diverse set of functions. Linking pathways by overlapping functionality revealed a modular network with energy metabolism forming a sparse centre, surrounded by several denser clusters comprised of regulatory and metabolic pathways. Signalling pathways formed a relatively discrete cluster connected to the centre of the network. Genetic interactions were enriched within the clusters of pathways by a factor of 5.5, confirming the organisation of our pathway network is biologically significant. Conclusions Our representation of molecular function according to pathway relationships enables analysis of gene/protein activity in the context of specific functional roles, as an alternative to typical molecule-centric graph-based methods. The pathway network demonstrates the cooperation of multiple pathways to perform biological processes and organises pathways into functionally related clusters with interdependent outcomes. PMID:26678917
Valiante, Vito; Baldin, Clara; Hortschansky, Peter; Jain, Radhika; Thywißen, Andreas; Straßburger, Maria; Shelest, Ekaterina; Heinekamp, Thorsten; Brakhage, Axel A
2016-10-01
Melanins play a crucial role in defending organisms against external stressors. In several pathogenic fungi, including the human pathogen Aspergillus fumigatus, melanin production was shown to contribute to virulence. A. fumigatus produces two different types of melanins, i.e., pyomelanin and dihydroxynaphthalene (DHN)-melanin. DHN-melanin forms the gray-green pigment characteristic for conidia, playing an important role in immune evasion of conidia and thus for fungal virulence. The DHN-melanin biosynthesis pathway is encoded by six genes organized in a cluster with the polyketide synthase gene pksP as a core element. Here, cross-species promoter analysis identified specific DNA binding sites in the DHN-melanin biosynthesis genes pksP-arp1 intergenic region that can be recognized by bHLH and MADS-box transcriptional regulators. Independent deletion of two genes coding for the transcription factors DevR (bHLH) and RlmA (MADS-box) interfered with sporulation and reduced the expression of the DHN-melanin gene cluster. In vitro and in vivo experiments proved that these transcription factors cooperatively regulate pksP expression acting both as repressors and activators in a mutually exclusive manner. The dual role executed by each regulator depends on specific DNA motifs recognized in the pksP promoter region. © 2016 John Wiley & Sons Ltd.
Pang, Xiuhua; Aigle, Bertrand; Girardet, Jean-Michel; Mangenot, Sophie; Pernodet, Jean-Luc; Decaris, Bernard; Leblond, Pierre
2004-01-01
Streptomyces ambofaciens has an 8-Mb linear chromosome ending in 200-kb terminal inverted repeats. Analysis of the F6 cosmid overlapping the terminal inverted repeats revealed a locus similar to type II polyketide synthase (PKS) gene clusters. Sequence analysis identified 26 open reading frames, including genes encoding the β-ketoacyl synthase (KS), chain length factor (CLF), and acyl carrier protein (ACP) that make up the minimal PKS. These KS, CLF, and ACP subunits are highly homologous to minimal PKS subunits involved in the biosynthesis of angucycline antibiotics. The genes encoding the KS and ACP subunits are transcribed constitutively but show a remarkable increase in expression after entering transition phase. Five genes, including those encoding the minimal PKS, were replaced by resistance markers to generate single and double mutants (replacement in one and both terminal inverted repeats). Double mutants were unable to produce either diffusible orange pigment or antibacterial activity against Bacillus subtilis. Single mutants showed an intermediate phenotype, suggesting that each copy of the cluster was functional. Transformation of double mutants with a conjugative and integrative form of F6 partially restored both phenotypes. The pigmented and antibacterial compounds were shown to be two distinct molecules produced from the same biosynthetic pathway. High-pressure liquid chromatography analysis of culture extracts from wild-type and double mutants revealed a peak with an associated bioactivity that was absent from the mutants. Two additional genes encoding KS and CLF were present in the cluster. However, disruption of the second KS gene had no effect on either pigment or antibiotic production. PMID:14742212
Mustafa, Saima; Fatima, Hira; Fatima, Sadia; Khosa, Tafheem; Akbar, Atif; Shaikh, Rehan Sadiq; Iqbal, Furhan
2018-01-01
To find out a correlation between the single nucleotide polymorphisms in cluster of differentiation 28 and cluster of differentiation 40 genes with Graves' disease, if any. This case-control study was conducted at the Multan Institute of Nuclear Medicine and Radiotherapy, Multan, Pakistan, and comprised blood samples of Graves' disease patients and controls. Various risk factors were also correlated either with the genotype at each single-nucleotide polymorphism or with various combinations of genotypes studied during present investigation. Of the 160 samples, there were 80(50%) each from patients and controls. Risk factor analysis revealed that gender (p=0.008), marital status (p<0.001), education (p<0.001), smoking (p<0.001), tri-iodothyronine (P <0.001), thyroxin (p<0.001) and thyroid-stimulating hormone (p<0.000) levels in blood were associated with Graves' disease. Both single-nucleotide polymorphisms in both genes were not associated with Graves' disease, either individually or in any combined form.
Liu, Xuewu; Wang, Yuanyuan; Liang, Jiao; Wang, Luojun; Qin, Na; Zhao, Ya; Zhao, Gang
2018-05-02
Plasmodium falciparum is the most virulent malaria parasite capable of parasitizing human erythrocytes. The identification of genes related to this capability can enhance our understanding of the molecular mechanisms underlying human malaria and lead to the development of new therapeutic strategies for malaria control. With the availability of several malaria parasite genome sequences, performing computational analysis is now a practical strategy to identify genes contributing to this disease. Here, we developed and used a virtual genome method to assign 33,314 genes from three human malaria parasites, namely, P. falciparum, P. knowlesi and P. vivax, and three rodent malaria parasites, namely, P. berghei, P. chabaudi and P. yoelii, to 4605 clusters. Each cluster consisted of genes whose protein sequences were significantly similar and was considered as a virtual gene. Comparing the enriched values of all clusters in human malaria parasites with those in rodent malaria parasites revealed 115 P. falciparum genes putatively responsible for parasitizing human erythrocytes. These genes are mainly located in the chromosome internal regions and participate in many biological processes, including membrane protein trafficking and thiamine biosynthesis. Meanwhile, 289 P. berghei genes were included in the rodent parasite-enriched clusters. Most are located in subtelomeric regions and encode erythrocyte surface proteins. Comparing cluster values in P. falciparum with those in P. vivax and P. knowlesi revealed 493 candidate genes linked to virulence. Some of them encode proteins present on the erythrocyte surface and participate in cytoadhesion, virulence factor trafficking, or erythrocyte invasion, but many genes with unknown function were also identified. Cerebral malaria is characterized by accumulation of infected erythrocytes at trophozoite stage in brain microvascular. To discover cerebral malaria-related genes, fast Fourier transformation (FFT) was introduced to extract genes highly transcribed at the trophozoite stage. Finally, 55 candidate genes were identified. Considering that parasite-infected erythrocyte surface protein 2 (PIESP2) contains gap-junction-related Neuromodulin_N domain and that anti-PIESP2 might provide protection against malaria, we chose PIESP2 for further experimental study. Our analysis revealed a limited number of genes linked to human disease in P. falciparum genome. These genes could be interesting targets for further functional characterization.
Breakup of a homeobox cluster after genome duplication in teleosts
Mulley, John F.; Chiu, Chi-hua; Holland, Peter W. H.
2006-01-01
Several families of homeobox genes are arranged in genomic clusters in metazoan genomes, including the Hox, ParaHox, NK, Rhox, and Iroquois gene clusters. The selective pressures responsible for maintenance of these gene clusters are poorly understood. The ParaHox gene cluster is evolutionarily conserved between amphioxus and human but is fragmented in teleost fishes. We show that two basal ray-finned fish, Polypterus and Amia, each possess an intact ParaHox cluster; this implies that the selective pressure maintaining clustering was lost after whole-genome duplication in teleosts. Cluster breakup is because of gene loss, not transposition or inversion, and the total number of ParaHox genes is the same in teleosts, human, mouse, and frog. We propose that this homeobox gene cluster is held together in chordates by the existence of interdigitated control regions that could be separated after locus duplication in the teleost fish. PMID:16801555
Gregg, Christina M.; Goetzl, Sebastian; Jeoung, Jae-Hun
2016-01-01
Acetyl-CoA synthase (ACS) catalyzes the reversible condensation of CO, CoA, and a methyl-cation to form acetyl-CoA at a unique Ni,Ni-[4Fe4S] cluster (the A-cluster). However, it was unknown which proteins support the assembly of the A-cluster. We analyzed the product of a gene from the cluster containing the ACS gene, cooC2 from Carboxydothermus hydrogenoformans, named AcsFCh, and showed that it acts as a maturation factor of ACS. AcsFCh and inactive ACS form a stable 2:1 complex that binds two nickel ions with higher affinity than the individual components. The nickel-bound ACS-AcsFCh complex remains inactive until MgATP is added, thereby converting inactive to active ACS. AcsFCh is a MinD-type ATPase and belongs to the CooC protein family, which can be divided into homologous subgroups. We propose that proteins of one subgroup are responsible for assembling the Ni,Ni-[4Fe4S] cluster of ACS, whereas proteins of a second subgroup mature the [Ni4Fe4S] cluster of carbon monoxide dehydrogenases. PMID:27382049
Yoshikawa, Mamoru; Kojima, Hiromi; Wada, Kota; Tsukidate, Toshiharu; Okada, Naoko; Saito, Hirohisa; Moriyama, Hiroshi
2006-07-01
To investigate the role of fibroblasts in the pathogenesis of cholesteatoma. Tissue specimens were obtained from our patients. Middle ear cholesteatoma-derived fibroblasts (MECFs) and postauricular skin-derived fibroblasts (SFs) as controls were then cultured for a few weeks. These fibroblasts were stimulated with interleukin (IL) 1alpha and/or IL-1beta before gene expression assays. We used the human genome U133A probe array (GeneChip) and real-time polymerase chain reaction to examine and compare the gene expression profiles of the MECFs and SFs. Six patients who had undergone tympanoplasty. The IL-1alpha-regulated genes were classified into 4 distinct clusters on the basis of profiles differentially regulated by SF and MECF using a hierarchical clustering analysis. The messenger RNA expressions of LARC (liver and activation-regulated chemokine), GMCSF (granulocyte-macrophage colony-stimulating factor), epiregulin, ICAM1 (intercellular adhesion molecule 1), and TGFA (transforming growth factor alpha) were more strongly up-regulated by IL-1alpha and/or IL-1beta in MECF than in SF, suggesting that these fibroblasts derived from different tissues retained their typical gene expression profiles. Fibroblasts may play a role in hyperkeratosis of middle ear cholesteatoma by releasing molecules involved in inflammation and epidermal growth. These fibroblasts may retain tissue-specific characteristics presumably controlled by epigenetic mechanisms.
Carvalho, Fabíola M; Souza, Rangel C; Barcellos, Fernando G; Hungria, Mariangela; Vasconcelos, Ana Tereza R
2010-02-08
Species belonging to the Rhizobiales are intriguing and extensively researched for including both bacteria with the ability to fix nitrogen when in symbiosis with leguminous plants and pathogenic bacteria to animals and plants. Similarities between the strategies adopted by pathogenic and symbiotic Rhizobiales have been described, as well as high variability related to events of horizontal gene transfer. Although it is well known that chromosomal rearrangements, mutations and horizontal gene transfer influence the dynamics of bacterial genomes, in Rhizobiales, the scenario that determine pathogenic or symbiotic lifestyle are not clear and there are very few studies of comparative genomic between these classes of prokaryotic microorganisms trying to delineate the evolutionary characterization of symbiosis and pathogenesis. Non-symbiotic nitrogen-fixing bacteria and bacteria involved in bioremediation closer to symbionts and pathogens in study may assist in the origin and ancestry genes and the gene flow occurring in Rhizobiales. The genomic comparisons of 19 species of Rhizobiales, including nitrogen-fixing, bioremediators and pathogens resulted in 33 common clusters to biological nitrogen fixation and pathogenesis, 15 clusters exclusive to all nitrogen-fixing bacteria and bacteria involved in bioremediation, 13 clusters found in only some nitrogen-fixing and bioremediation bacteria, 01 cluster exclusive to some symbionts, and 01 cluster found only in some pathogens analyzed. In BBH performed to all strains studied, 77 common genes were obtained, 17 of which were related to biological nitrogen fixation and pathogenesis. Phylogenetic reconstructions for Fix, Nif, Nod, Vir, and Trb showed possible horizontal gene transfer events, grouping species of different phenotypes. The presence of symbiotic and virulence genes in both pathogens and symbionts does not seem to be the only determinant factor for lifestyle evolution in these microorganisms, although they may act in common stages of host infection. The phylogenetic analysis for many distinct operons involved in these processes emphasizes the relevance of horizontal gene transfer events in the symbiotic and pathogenic similarity.
Medina, Angel; Schmidt-Heydt, Markus; Cárdenas-Chávez, Diana L.; Parra, Roberto; Geisen, Rolf; Magan, Naresh
2013-01-01
The objective of this study was to integrate data on the effect of water activity (aw; 0.995–0.93) and temperature (20–35°C) on activation of the biosynthetic FUM genes, growth and the mycotoxins fumonisin (FB1, FB2) by Fusarium verticillioides in vitro. The relative expression of nine biosynthetic cluster genes (FUM1, FUM7, FUM10, FUM11, FUM12, FUM13, FUM14, FUM16 and FUM19) in relation to the environmental factors was determined using a microarray analysis. The expression was related to growth and phenotypic FB1 and FB2 production. These data were used to develop a mixed-growth-associated product formation model and link this to a linear combination of the expression data for the nine genes. The model was then validated by examining datasets outside the model fitting conditions used (35°C). The relationship between the key gene (FUM1) and other genes in the cluster (FUM11, FUM13, FUM9, FUM14) were examined in relation to aw, temperature, FB1 and FB2 production by developing ternary diagrams of relative expression. This model is important in developing an integrated systems approach to develop prevention strategies to control fumonisin biosynthesis in staple food commodities and could also be used to predict the potential impact that climate change factors may have on toxin production. PMID:23697716
Tago, Kanako; Okubo, Takashi; Shimomura, Yumi; Kikuchi, Yoshitomo; Hori, Tomoyuki; Nagayama, Atsushi; Hayatsu, Masahito
2015-01-01
The effects of environmental factors such as pH and nutrient content on the ecology of ammonia-oxidizing bacteria (AOB) and archaea (AOA) in soil has been extensively studied using experimental fields. However, how these environmental factors intricately influence the community structure of AOB and AOA in soil from farmers' fields is unclear. In the present study, the abundance and diversity of AOB and AOA in soils collected from farmers' sugarcane fields were investigated using quantitative PCR and barcoded pyrosequencing targeting the ammonia monooxygenase alpha subunit (amoA) gene. The abundances of AOB and AOA amoA genes were estimated to be in the range of 1.8 × 10(5)-9.2 × 10(6) and 1.7 × 10(6)-5.3 × 10(7) gene copies g dry soil(-1), respectively. The abundance of both AOB and AOA positively correlated with the potential nitrification rate. The dominant sequence reads of AOB and AOA were placed in Nitrosospira-related and Nitrososphaera-related clusters in all soils, respectively, which varied at the level of their sub-clusters in each soil. The relationship between these ammonia-oxidizing community structures and soil pH was shown to be significant by the Mantel test. The relative abundances of the OTU1 of Nitrosospira cluster 3 and Nitrososphaera subcluster 7.1 negatively correlated with soil pH. These results indicated that soil pH was the most important factor shaping the AOB and AOA community structures, and that certain subclusters of AOB and AOA adapted to and dominated the acidic soil of agricultural sugarcane fields.
Tago, Kanako; Okubo, Takashi; Shimomura, Yumi; Kikuchi, Yoshitomo; Hori, Tomoyuki; Nagayama, Atsushi; Hayatsu, Masahito
2015-01-01
The effects of environmental factors such as pH and nutrient content on the ecology of ammonia-oxidizing bacteria (AOB) and archaea (AOA) in soil has been extensively studied using experimental fields. However, how these environmental factors intricately influence the community structure of AOB and AOA in soil from farmers’ fields is unclear. In the present study, the abundance and diversity of AOB and AOA in soils collected from farmers’ sugarcane fields were investigated using quantitative PCR and barcoded pyrosequencing targeting the ammonia monooxygenase alpha subunit (amoA) gene. The abundances of AOB and AOA amoA genes were estimated to be in the range of 1.8 × 105–9.2 × 106 and 1.7 × 106–5.3 × 107 gene copies g dry soil−1, respectively. The abundance of both AOB and AOA positively correlated with the potential nitrification rate. The dominant sequence reads of AOB and AOA were placed in Nitrosospira-related and Nitrososphaera-related clusters in all soils, respectively, which varied at the level of their sub-clusters in each soil. The relationship between these ammonia-oxidizing community structures and soil pH was shown to be significant by the Mantel test. The relative abundances of the OTU1 of Nitrosospira cluster 3 and Nitrososphaera subcluster 7.1 negatively correlated with soil pH. These results indicated that soil pH was the most important factor shaping the AOB and AOA community structures, and that certain subclusters of AOB and AOA adapted to and dominated the acidic soil of agricultural sugarcane fields. PMID:25736866
Bessonov, Kyrylo; Walkey, Christopher J.; Shelp, Barry J.; van Vuuren, Hennie J. J.; Chiu, David; van der Merwe, George
2013-01-01
Analyzing time-course expression data captured in microarray datasets is a complex undertaking as the vast and complex data space is represented by a relatively low number of samples as compared to thousands of available genes. Here, we developed the Interdependent Correlation Clustering (ICC) method to analyze relationships that exist among genes conditioned on the expression of a specific target gene in microarray data. Based on Correlation Clustering, the ICC method analyzes a large set of correlation values related to gene expression profiles extracted from given microarray datasets. ICC can be applied to any microarray dataset and any target gene. We applied this method to microarray data generated from wine fermentations and selected NSF1, which encodes a C2H2 zinc finger-type transcription factor, as the target gene. The validity of the method was verified by accurate identifications of the previously known functional roles of NSF1. In addition, we identified and verified potential new functions for this gene; specifically, NSF1 is a negative regulator for the expression of sulfur metabolism genes, the nuclear localization of Nsf1 protein (Nsf1p) is controlled in a sulfur-dependent manner, and the transcription of NSF1 is regulated by Met4p, an important transcriptional activator of sulfur metabolism genes. The inter-disciplinary approach adopted here highlighted the accuracy and relevancy of the ICC method in mining for novel gene functions using complex microarray datasets with a limited number of samples. PMID:24130853
Expressed sequence tags from poplar wood tissues--a comparative analysis from multiple libraries.
Déjardin, A; Leplé, J-C; Lesage-Descauses, M-C; Costa, G; Pilate, G
2004-01-01
Xylogenesis involves successive developmental processes--cambial division, cell expansion and differentiation, cell death--each occurring along a gradient from the cambium to the pith of the stem. Taking advantage of the high level of organisation of wood tissues, we isolated cambial zone (CZ), differentiating xylem (DX) and mature xylem (MX) from both tension wood (TW) and opposite wood (OW) of bent poplars. Four different cDNA libraries were then constructed and used to generate 10,062 EST, reflecting the genes expressed in the different wood tissues. For the most abundant clusters, the EST distributions were compared between libraries in order to identify genes specific or over-represented at some specific developmental stages. They clearly showed a developmental shift between CZ and DX, whereas there is a continuity of development between DX and MX. CZ was mainly characterized by clusters of genes involved in cell cycle, protein synthesis and fate. Interestingly, two clusters with no assigned function were found specific to the cambial zone. In DX and MX, clusters were mostly involved in methylation of lignin precursors and microtubule cytoskeleton. In addition, in DX, EST from TW and OW were compared: five clusters of arabinogalactan proteins, one for sucrose synthase and one for fructokinase were specific or over-represented in TW. Moreover, a putative transcription factor and a cluster of unknown function were also identified in DX-TW. The informative comparison of multiple libraries prepared from wood tissues led to the identification of genes--some with still unknown functions--putatively involved in xylogenesis and tension wood formation.
Charles, J. P.; Chihara, C.; Nejad, S.; Riddiford, L. M.
1997-01-01
A 36-kb genomic DNA segment of the Drosophila melanogaster genome containing 12 clustered cuticle genes has been mapped and partially sequenced. The cluster maps at 65A 5-6 on the left arm of the third chromosome, in agreement with the previously determined location of a putative cluster encompassing the genes for the third instar larval cuticle proteins LCP5, LCP6 and LCP8. This cluster is the largest cuticle gene cluster discovered to date and shows a number of surprising features that explain in part the genetic complexity of the LCP5, LCP6 and LCP8 loci. The genes encoding LCP5 and LCP8 are multiple copy genes and the presence of extensive similarity in their coding regions gives the first evidence for gene conversion in cuticle genes. In addition, five genes in the cluster are intronless. Four of these five have arisen by retroposition. The other genes in the cluster have a single intron located at an unusual location for insect cuticle genes. PMID:9383064
Summerfield, Taryn L.; Yu, Lianbo; Gulati, Parul; Zhang, Jie; Huang, Kun; Romero, Roberto; Kniss, Douglas A.
2011-01-01
A majority of the studies examining the molecular regulation of human labor have been conducted using single gene approaches. While the technology to produce multi-dimensional datasets is readily available, the means for facile analysis of such data are limited. The objective of this study was to develop a systems approach to infer regulatory mechanisms governing global gene expression in cytokine-challenged cells in vitro, and to apply these methods to predict gene regulatory networks (GRNs) in intrauterine tissues during term parturition. To this end, microarray analysis was applied to human amnion mesenchymal cells (AMCs) stimulated with interleukin-1β, and differentially expressed transcripts were subjected to hierarchical clustering, temporal expression profiling, and motif enrichment analysis, from which a GRN was constructed. These methods were then applied to fetal membrane specimens collected in the absence or presence of spontaneous term labor. Analysis of cytokine-responsive genes in AMCs revealed a sterile immune response signature, with promoters enriched in response elements for several inflammation-associated transcription factors. In comparison to the fetal membrane dataset, there were 34 genes commonly upregulated, many of which were part of an acute inflammation gene expression signature. Binding motifs for nuclear factor-κB were prominent in the gene interaction and regulatory networks for both datasets; however, we found little evidence to support the utilization of pathogen-associated molecular pattern (PAMP) signaling. The tissue specimens were also enriched for transcripts governed by hypoxia-inducible factor. The approach presented here provides an uncomplicated means to infer global relationships among gene clusters involved in cellular responses to labor-associated signals. PMID:21655103
Identification and characterization of Rhox13, a novel X-linked mouse homeobox gene
Geyer, Christopher B.; Eddy, Edward M.
2008-01-01
Homeobox genes encode transcription factors whose expression organizes programs of development. A number of homeobox genes expressed in reproductive tissues have been identified recently, including a colinear cluster on the X chromosome in mice. This has led to an increased interest in understanding the role(s) of homeobox genes in regulating development of reproductive tissues including the testis, ovary, and placenta. Here we report the identification and characterization of a novel homeobox gene of the paired-like class on the X chromosome distal to the reproductive homeobox (Rhox) cluster in mice. Transcripts are found in the testis and ovary as early as 13.5 days post-coitum (dpc). Transcription ceases in the ovary by 3 days post-partum (dpp), but continues in the testis through adulthood. The Rhox13 gene encodes a 25.3 kDa protein expressed in the adult testis in germ cells at the basal aspect of the seminiferous epithelium. PMID:18675325
Su, Hongyan; Zhang, Shizhong; Yuan, Xiaowei; Chen, Changtian; Wang, Xiao-Fei; Hao, Yu-Jin
2013-10-01
NAC (NAM, ATAF1,2, and CUC2) proteins constitute one of the largest families of plant-specific transcription factors. To date, little is known about the NAC genes in the apple (Malus domestica). In this study, a total of 180 NAC genes were identified in the apple genome and were phylogenetically clustered into six groups (I-VI) with the NAC genes from Arabidopsis and rice. The predicted apple NAC genes were distributed across all of 17 chromosomes at various densities. Additionally, the gene structure and motif compositions of the apple NAC genes were analyzed. Moreover, the expression of 29 selected apple NAC genes was analyzed in different tissues and under different abiotic stress conditions. All of the selected genes, with the exception of four genes, were expressed in at least one of the tissues tested, which indicates that the NAC genes are involved in various aspects of the physiological and developmental processes of the apple. Encouragingly, 17 of the selected genes were found to respond to one or more of the abiotic stress treatments, and these 17 genes included not only the expected 7 genes that were clustered with the well-known stress-related marker genes in group IV but also 10 genes located in other subgroups, none of which contains members that have been reported to be stress-related. To the best of our knowledge, this report describes the first genome-wide analysis of the apple NAC gene family, and the results should provide valuable information for understanding the classification and putative functions of this family. Copyright © 2013 Elsevier Masson SAS. All rights reserved.
Identification of lethal cluster of genes in the yeast transcription network
NASA Astrophysics Data System (ADS)
Rho, K.; Jeong, H.; Kahng, B.
2006-05-01
Identification of essential or lethal genes would be one of the ultimate goals in drug designs. Here we introduce an in silico method to select the cluster with a high population of lethal genes, called lethal cluster, through microarray assay. We construct a gene transcription network based on the microarray expression level. Links are added one by one in the descending order of the Pearson correlation coefficients between two genes. As the link density p increases, two meaningful link densities pm and ps are observed. At pm, which is smaller than the percolation threshold, the number of disconnected clusters is maximum, and the lethal genes are highly concentrated in a certain cluster that needs to be identified. Thus the deletion of all genes in that cluster could efficiently lead to a lethal inviable mutant. This lethal cluster can be identified by an in silico method. As p increases further beyond the percolation threshold, the power law behavior in the degree distribution of a giant cluster appears at ps. We measure the degree of each gene at ps. With the information pertaining to the degrees of each gene at ps, we return to the point pm and calculate the mean degree of genes of each cluster. We find that the lethal cluster has the largest mean degree.
Freyre-González, Julio A; Alonso-Pavón, José A; Treviño-Quintanilla, Luis G; Collado-Vides, Julio
2008-10-27
Previous studies have used different methods in an effort to extract the modular organization of transcriptional regulatory networks. However, these approaches are not natural, as they try to cluster strongly connected genes into a module or locate known pleiotropic transcription factors in lower hierarchical layers. Here, we unravel the transcriptional regulatory network of Escherichia coli by separating it into its key elements, thus revealing its natural organization. We also present a mathematical criterion, based on the topological features of the transcriptional regulatory network, to classify the network elements into one of two possible classes: hierarchical or modular genes. We found that modular genes are clustered into physiologically correlated groups validated by a statistical analysis of the enrichment of the functional classes. Hierarchical genes encode transcription factors responsible for coordinating module responses based on general interest signals. Hierarchical elements correlate highly with the previously studied global regulators, suggesting that this could be the first mathematical method to identify global regulators. We identified a new element in transcriptional regulatory networks never described before: intermodular genes. These are structural genes that integrate, at the promoter level, signals coming from different modules, and therefore from different physiological responses. Using the concept of pleiotropy, we have reconstructed the hierarchy of the network and discuss the role of feedforward motifs in shaping the hierarchical backbone of the transcriptional regulatory network. This study sheds new light on the design principles underpinning the organization of transcriptional regulatory networks, showing a novel nonpyramidal architecture composed of independent modules globally governed by hierarchical transcription factors, whose responses are integrated by intermodular genes.
Crnovčić, Ivana; Rückert, Christian; Semsary, Siamak; Lang, Manuel; Kalinowski, Jörn; Keller, Ullrich
2017-01-01
Sequencing the actinomycin (acm) biosynthetic gene cluster of Streptomyces antibioticus IMRU 3720, which produces actinomycin X (Acm X), revealed 20 genes organized into a highly similar framework as in the bi-armed acm C biosynthetic gene cluster of Streptomyces chrysomallus but without an attached additional extra arm of orthologues as in the latter. Curiously, the extra arm of the S. chrysomallus gene cluster turned out to perfectly match the single arm of the S. antibioticus gene cluster in the same order of orthologues including the the presence of two pseudogenes, scacmM and scacmN, encoding a cytochrome P450 and its ferredoxin, respectively. Orthologues of the latter genes were both missing in the principal arm of the S. chrysomallus acm C gene cluster. All orthologues of the extra arm showed a G +C-contents different from that of their counterparts in the principal arm. Moreover, the similarities of translation products from the extra arm were all higher to the corresponding translation products of orthologue genes from the S. antibioticus acm X gene cluster than to those encoded by the principal arm of their own gene cluster. This suggests that the duplicated structure of the S. chrysomallus acm C biosynthetic gene cluster evolved from previous fusion between two one-armed acm gene clusters each from a different genetic background. However, while scacmM and scacmN in the extra arm of the S. chrysomallus acm C gene cluster are mutated and therefore are non-functional, their orthologues saacmM and saacmN in the S. antibioticus acm C gene cluster show no defects seemingly encoding active enzymes with functions specific for Acm X biosynthesis. Both acm biosynthetic gene clusters lack a kynurenine-3-monooxygenase gene necessary for biosynthesis of 3-hydroxy-4-methylanthranilic acid, the building block of the Acm chromophore, which suggests participation of a genome-encoded relevant monooxygenase during Acm biosynthesis in both S. chrysomallus and S. antibioticus. PMID:28435299
Crnovčić, Ivana; Rückert, Christian; Semsary, Siamak; Lang, Manuel; Kalinowski, Jörn; Keller, Ullrich
2017-01-01
Sequencing the actinomycin ( acm ) biosynthetic gene cluster of Streptomyces antibioticus IMRU 3720, which produces actinomycin X (Acm X), revealed 20 genes organized into a highly similar framework as in the bi-armed acm C biosynthetic gene cluster of Streptomyces chrysomallus but without an attached additional extra arm of orthologues as in the latter. Curiously, the extra arm of the S. chrysomallus gene cluster turned out to perfectly match the single arm of the S. antibioticus gene cluster in the same order of orthologues including the the presence of two pseudogenes, scacmM and scacmN , encoding a cytochrome P450 and its ferredoxin, respectively. Orthologues of the latter genes were both missing in the principal arm of the S. chrysomallus acm C gene cluster. All orthologues of the extra arm showed a G +C-contents different from that of their counterparts in the principal arm. Moreover, the similarities of translation products from the extra arm were all higher to the corresponding translation products of orthologue genes from the S. antibioticus acm X gene cluster than to those encoded by the principal arm of their own gene cluster. This suggests that the duplicated structure of the S. chrysomallus acm C biosynthetic gene cluster evolved from previous fusion between two one-armed acm gene clusters each from a different genetic background. However, while scacmM and scacmN in the extra arm of the S. chrysomallus acm C gene cluster are mutated and therefore are non-functional, their orthologues saacmM and saacmN in the S. antibioticus acm C gene cluster show no defects seemingly encoding active enzymes with functions specific for Acm X biosynthesis. Both acm biosynthetic gene clusters lack a kynurenine-3-monooxygenase gene necessary for biosynthesis of 3-hydroxy-4-methylanthranilic acid, the building block of the Acm chromophore, which suggests participation of a genome-encoded relevant monooxygenase during Acm biosynthesis in both S. chrysomallus and S. antibioticus .
Unsupervised deep learning reveals prognostically relevant subtypes of glioblastoma.
Young, Jonathan D; Cai, Chunhui; Lu, Xinghua
2017-10-03
One approach to improving the personalized treatment of cancer is to understand the cellular signaling transduction pathways that cause cancer at the level of the individual patient. In this study, we used unsupervised deep learning to learn the hierarchical structure within cancer gene expression data. Deep learning is a group of machine learning algorithms that use multiple layers of hidden units to capture hierarchically related, alternative representations of the input data. We hypothesize that this hierarchical structure learned by deep learning will be related to the cellular signaling system. Robust deep learning model selection identified a network architecture that is biologically plausible. Our model selection results indicated that the 1st hidden layer of our deep learning model should contain about 1300 hidden units to most effectively capture the covariance structure of the input data. This agrees with the estimated number of human transcription factors, which is approximately 1400. This result lends support to our hypothesis that the 1st hidden layer of a deep learning model trained on gene expression data may represent signals related to transcription factor activation. Using the 3rd hidden layer representation of each tumor as learned by our unsupervised deep learning model, we performed consensus clustering on all tumor samples-leading to the discovery of clusters of glioblastoma multiforme with differential survival. One of these clusters contained all of the glioblastoma samples with G-CIMP, a known methylation phenotype driven by the IDH1 mutation and associated with favorable prognosis, suggesting that the hidden units in the 3rd hidden layer representations captured a methylation signal without explicitly using methylation data as input. We also found differentially expressed genes and well-known mutations (NF1, IDH1, EGFR) that were uniquely correlated with each of these clusters. Exploring these unique genes and mutations will allow us to further investigate the disease mechanisms underlying each of these clusters. In summary, we show that a deep learning model can be trained to represent biologically and clinically meaningful abstractions of cancer gene expression data. Understanding what additional relationships these hidden layer abstractions have with the cancer cellular signaling system could have a significant impact on the understanding and treatment of cancer.
Lukashin, A V; Fuchs, R
2001-05-01
Cluster analysis of genome-wide expression data from DNA microarray hybridization studies has proved to be a useful tool for identifying biologically relevant groupings of genes and samples. In the present paper, we focus on several important issues related to clustering algorithms that have not yet been fully studied. We describe a simple and robust algorithm for the clustering of temporal gene expression profiles that is based on the simulated annealing procedure. In general, this algorithm guarantees to eventually find the globally optimal distribution of genes over clusters. We introduce an iterative scheme that serves to evaluate quantitatively the optimal number of clusters for each specific data set. The scheme is based on standard approaches used in regular statistical tests. The basic idea is to organize the search of the optimal number of clusters simultaneously with the optimization of the distribution of genes over clusters. The efficiency of the proposed algorithm has been evaluated by means of a reverse engineering experiment, that is, a situation in which the correct distribution of genes over clusters is known a priori. The employment of this statistically rigorous test has shown that our algorithm places greater than 90% genes into correct clusters. Finally, the algorithm has been tested on real gene expression data (expression changes during yeast cell cycle) for which the fundamental patterns of gene expression and the assignment of genes to clusters are well understood from numerous previous studies.
Cary, J. W.; Han, Z.; Yin, Y.; Lohmar, J. M.; Shantappa, S.; Harris-Coward, P. Y.; Mack, B.; Ehrlich, K. C.; Wei, Q.; Arroyo-Manzanares, N.; Uka, V.; Vanhaecke, L.; Bhatnagar, D.; Yu, J.; Nierman, W. C.; Johns, M. A.; Sorensen, D.; Shen, H.; De Saeger, S.; Diana Di Mavungu, J.
2015-01-01
The global regulatory veA gene governs development and secondary metabolism in numerous fungal species, including Aspergillus flavus. This is especially relevant since A. flavus infects crops of agricultural importance worldwide, contaminating them with potent mycotoxins. The most well-known are aflatoxins, which are cytotoxic and carcinogenic polyketide compounds. The production of aflatoxins and the expression of genes implicated in the production of these mycotoxins are veA dependent. The genes responsible for the synthesis of aflatoxins are clustered, a signature common for genes involved in fungal secondary metabolism. Studies of the A. flavus genome revealed many gene clusters possibly connected to the synthesis of secondary metabolites. Many of these metabolites are still unknown, or the association between a known metabolite and a particular gene cluster has not yet been established. In the present transcriptome study, we show that veA is necessary for the expression of a large number of genes. Twenty-eight out of the predicted 56 secondary metabolite gene clusters include at least one gene that is differentially expressed depending on presence or absence of veA. One of the clusters under the influence of veA is cluster 39. The absence of veA results in a downregulation of the five genes found within this cluster. Interestingly, our results indicate that the cluster is expressed mainly in sclerotia. Chemical analysis of sclerotial extracts revealed that cluster 39 is responsible for the production of aflavarin. PMID:26209694
A tripartite clustering analysis on microRNA, gene and disease model.
Shen, Chengcheng; Liu, Ying
2012-02-01
Alteration of gene expression in response to regulatory molecules or mutations could lead to different diseases. MicroRNAs (miRNAs) have been discovered to be involved in regulation of gene expression and a wide variety of diseases. In a tripartite biological network of human miRNAs, their predicted target genes and the diseases caused by altered expressions of these genes, valuable knowledge about the pathogenicity of miRNAs, involved genes and related disease classes can be revealed by co-clustering miRNAs, target genes and diseases simultaneously. Tripartite co-clustering can lead to more informative results than traditional co-clustering with only two kinds of members and pass the hidden relational information along the relation chain by considering multi-type members. Here we report a spectral co-clustering algorithm for k-partite graph to find clusters with heterogeneous members. We use the method to explore the potential relationships among miRNAs, genes and diseases. The clusters obtained from the algorithm have significantly higher density than randomly selected clusters, which means members in the same cluster are more likely to have common connections. Results also show that miRNAs in the same family based on the hairpin sequences tend to belong to the same cluster. We also validate the clustering results by checking the correlation of enriched gene functions and disease classes in the same cluster. Finally, widely studied miR-17-92 and its paralogs are analyzed as a case study to reveal that genes and diseases co-clustered with the miRNAs are in accordance with current research findings.
Pyeon, Hye-Rim; Nah, Hee-Ju; Kang, Seung-Hoon; Choi, Si-Sun; Kim, Eung-Soo
2017-05-31
Heterologous expression of biosynthetic gene clusters of natural microbial products has become an essential strategy for titer improvement and pathway engineering of various potentially-valuable natural products. A Streptomyces artificial chromosomal conjugation vector, pSBAC, was previously successfully applied for precise cloning and tandem integration of a large polyketide tautomycetin (TMC) biosynthetic gene cluster (Nah et al. in Microb Cell Fact 14(1):1, 2015), implying that this strategy could be employed to develop a custom overexpression scheme of natural product pathway clusters present in actinomycetes. To validate the pSBAC system as a generally-applicable heterologous overexpression system for a large-sized polyketide biosynthetic gene cluster in Streptomyces, another model polyketide compound, the pikromycin biosynthetic gene cluster, was preciously cloned and heterologously expressed using the pSBAC system. A unique HindIII restriction site was precisely inserted at one of the border regions of the pikromycin biosynthetic gene cluster within the chromosome of Streptomyces venezuelae, followed by site-specific recombination of pSBAC into the flanking region of the pikromycin gene cluster. Unlike the previous cloning process, one HindIII site integration step was skipped through pSBAC modification. pPik001, a pSBAC containing the pikromycin biosynthetic gene cluster, was directly introduced into two heterologous hosts, Streptomyces lividans and Streptomyces coelicolor, resulting in the production of 10-deoxymethynolide, a major pikromycin derivative. When two entire pikromycin biosynthetic gene clusters were tandemly introduced into the S. lividans chromosome, overproduction of 10-deoxymethynolide and the presence of pikromycin, which was previously not detected, were both confirmed. Moreover, comparative qRT-PCR results confirmed that the transcription of pikromycin biosynthetic genes was significantly upregulated in S. lividans containing tandem clusters of pikromycin biosynthetic gene clusters. The 60 kb pikromycin biosynthetic gene cluster was isolated in a single integration pSBAC vector. Introduction of the pikromycin biosynthetic gene cluster into the pikromycin non-producing strains resulted in higher pikromycin production. The utility of the pSBAC system as a precise cloning tool for large-sized biosynthetic gene clusters was verified through heterologous expression of the pikromycin biosynthetic gene cluster. Moreover, this pSBAC-driven heterologous expression strategy was confirmed to be an ideal approach for production of low and inconsistent natural products such as pikromycin in S. venezuelae, implying that this strategy could be employed for development of a custom overexpression scheme of natural product biosynthetic gene clusters in actinomycetes.
From hormones to secondary metabolism: the emergence of metabolic gene clusters in plants.
Chu, Hoi Yee; Wegel, Eva; Osbourn, Anne
2011-04-01
Gene clusters for the synthesis of secondary metabolites are a common feature of microbial genomes. Well-known examples include clusters for the synthesis of antibiotics in actinomycetes, and also for the synthesis of antibiotics and toxins in filamentous fungi. Until recently it was thought that genes for plant metabolic pathways were not clustered, and this is certainly true in many cases; however, five plant secondary metabolic gene clusters have now been discovered, all of them implicated in synthesis of defence compounds. An obvious assumption might be that these eukaryotic gene clusters have arisen by horizontal gene transfer from microbes, but there is compelling evidence to indicate that this is not the case. This raises intriguing questions about how widespread such clusters are, what the significance of clustering is, why genes for some metabolic pathways are clustered and those for others are not, and how these clusters form. In answering these questions we may hope to learn more about mechanisms of genome plasticity and adaptive evolution in plants. It is noteworthy that for the five plant secondary metabolic gene clusters reported so far, the enzymes for the first committed steps all appear to have been recruited directly or indirectly from primary metabolic pathways involved in hormone synthesis. This may or may not turn out to be a common feature of plant secondary metabolic gene clusters as new clusters emerge. © 2011 The Authors. The Plant Journal © 2011 Blackwell Publishing Ltd.
Zhong, Xingyu; Tian, Yuqing; Niu, Guoqing; Tan, Huarong
2013-07-01
A draft genome sequence of Streptomyces ansochromogenes 7100 was generated using 454 sequencing technology. In combination with local BLAST searches and gap filling techniques, a comprehensive antiSMASH-based method was adopted to assemble the secondary metabolite biosynthetic gene clusters in the draft genome of S. ansochromogenes. A total of at least 35 putative gene clusters were identified and assembled. Transcriptional analysis showed that 20 of the 35 gene clusters were expressed in either or all of the three different media tested, whereas the other 15 gene clusters were silent in all three different media. This study provides a comprehensive method to identify and assemble secondary metabolite biosynthetic gene clusters in draft genomes of Streptomyces, and will significantly promote functional studies of these secondary metabolite biosynthetic gene clusters.
Supervised group Lasso with applications to microarray data analysis
Ma, Shuangge; Song, Xiao; Huang, Jian
2007-01-01
Background A tremendous amount of efforts have been devoted to identifying genes for diagnosis and prognosis of diseases using microarray gene expression data. It has been demonstrated that gene expression data have cluster structure, where the clusters consist of co-regulated genes which tend to have coordinated functions. However, most available statistical methods for gene selection do not take into consideration the cluster structure. Results We propose a supervised group Lasso approach that takes into account the cluster structure in gene expression data for gene selection and predictive model building. For gene expression data without biological cluster information, we first divide genes into clusters using the K-means approach and determine the optimal number of clusters using the Gap method. The supervised group Lasso consists of two steps. In the first step, we identify important genes within each cluster using the Lasso method. In the second step, we select important clusters using the group Lasso. Tuning parameters are determined using V-fold cross validation at both steps to allow for further flexibility. Prediction performance is evaluated using leave-one-out cross validation. We apply the proposed method to disease classification and survival analysis with microarray data. Conclusion We analyze four microarray data sets using the proposed approach: two cancer data sets with binary cancer occurrence as outcomes and two lymphoma data sets with survival outcomes. The results show that the proposed approach is capable of identifying a small number of influential gene clusters and important genes within those clusters, and has better prediction performance than existing methods. PMID:17316436
Mugford, Sam T.; Louveau, Thomas; Melton, Rachel; Qi, Xiaoquan; Bakht, Saleha; Hill, Lionel; Tsurushima, Tetsu; Honkanen, Suvi; Rosser, Susan J.; Lomonossoff, George P.; Osbourn, Anne
2013-01-01
Operon-like gene clusters are an emerging phenomenon in the field of plant natural products. The genes encoding some of the best-characterized plant secondary metabolite biosynthetic pathways are scattered across plant genomes. However, an increasing number of gene clusters encoding the synthesis of diverse natural products have recently been reported in plant genomes. These clusters have arisen through the neo-functionalization and relocation of existing genes within the genome, and not by horizontal gene transfer from microbes. The reasons for clustering are not yet clear, although this form of gene organization is likely to facilitate co-inheritance and co-regulation. Oats (Avena spp) synthesize antimicrobial triterpenoids (avenacins) that provide protection against disease. The synthesis of these compounds is encoded by a gene cluster. Here we show that a module of three adjacent genes within the wider biosynthetic gene cluster is required for avenacin acylation. Through the characterization of these genes and their encoded proteins we present a model of the subcellular organization of triterpenoid biosynthesis. PMID:23532069
2009-01-01
Background Soybeans grown in the upper Midwestern United States often suffer from iron deficiency chlorosis, which results in yield loss at the end of the season. To better understand the effect of iron availability on soybean yield, we identified genes in two near isogenic lines with changes in expression patterns when plants were grown in iron sufficient and iron deficient conditions. Results Transcriptional profiles of soybean (Glycine max, L. Merr) near isogenic lines Clark (PI548553, iron efficient) and IsoClark (PI547430, iron inefficient) grown under Fe-sufficient and Fe-limited conditions were analyzed and compared using the Affymetrix® GeneChip® Soybean Genome Array. There were 835 candidate genes in the Clark (PI548553) genotype and 200 candidate genes in the IsoClark (PI547430) genotype putatively involved in soybean's iron stress response. Of these candidate genes, fifty-eight genes in the Clark genotype were identified with a genetic location within known iron efficiency QTL and 21 in the IsoClark genotype. The arrays also identified 170 single feature polymorphisms (SFPs) specific to either Clark or IsoClark. A sliding window analysis of the microarray data and the 7X genome assembly coupled with an iterative model of the data showed the candidate genes are clustered in the genome. An analysis of 5' untranslated regions in the promoter of candidate genes identified 11 conserved motifs in 248 differentially expressed genes, all from the Clark genotype, representing 129 clusters identified earlier, confirming the cluster analysis results. Conclusion These analyses have identified the first genes with expression patterns that are affected by iron stress and are located within QTL specific to iron deficiency stress. The genetic location and promoter motif analysis results support the hypothesis that the differentially expressed genes are co-regulated. The combined results of all analyses lead us to postulate iron inefficiency in soybean is a result of a mutation in a transcription factor(s), which controls the expression of genes required in inducing an iron stress response. PMID:19678937
Differentiation of human-induced pluripotent stem cells into insulin-producing clusters.
Shaer, Anahita; Azarpira, Negar; Vahdati, Akbar; Karimi, Mohammad Hosein; Shariati, Mehrdad
2015-02-01
In diabetes mellitus type 1, beta cells are mostly destroyed; while in diabetes mellitus type 2, beta cells are reduced by 40% to 60%. We hope that soon, stem cells can be used in diabetes therapy via pancreatic beta cell replacement. Induced pluripotent stem cells are a kind of stem cell taken from an adult somatic cell by "stimulating" certain genes. These induced pluripotent stem cells may be a promising source of cell therapy. This study sought to produce isletlike clusters of insulin-producing cells taken from induced pluripotent stem cells. A human-induced pluripotent stem cell line was induced into isletlike clusters via a 4-step protocol, by adding insulin, transferrin, and selenium (ITS), N2, B27, fibroblast growth factor, and nicotinamide. During differentiation, expression of pancreatic β-cell genes was evaluated by reverse transcriptase-polymerase chain reaction; the morphologic changes of induced pluripotent stem cells toward isletlike clusters were observed by a light microscope. Dithizone staining was used to stain these isletlike clusters. Insulin produced by these clusters was evaluated by radio immunosorbent assay, and the secretion capacity was analyzed with a glucose challenge test. Differentiation was evaluated by analyzing the morphology, dithizone staining, real-time quantitative polymerase chain reaction, and immunocytochemistry. Gene expression of insulin, glucagon, PDX1, NGN3, PAX4, PAX6, NKX6.1, KIR6.2, and GLUT2 were documented by analyzing real-time quantitative polymerase chain reaction. Dithizone-stained cellular clusters were observed after 23 days. The isletlike clusters significantly produced insulin. The isletlike clusters could increase insulin secretion after a glucose challenge test. This work provides a model for studying the differentiation of human-induced pluripotent stem cells to insulin-producing cells.
Scott, Barry; Young, Carolyn A.; Saikia, Sanjay; McMillan, Lisa K.; Monahan, Brendon J.; Koulman, Albert; Astin, Jonathan; Eaton, Carla J.; Bryant, Andrea; Wrenn, Ruth E.; Finch, Sarah C.; Tapper, Brian A.; Parker, Emily J.; Jameson, Geoffrey B.
2013-01-01
The indole-diterpene paxilline is an abundant secondary metabolite synthesized by Penicillium paxilli. In total, 21 genes have been identified at the PAX locus of which six have been previously confirmed to have a functional role in paxilline biosynthesis. A combination of bioinformatics, gene expression and targeted gene replacement analyses were used to define the boundaries of the PAX gene cluster. Targeted gene replacement identified seven genes, paxG, paxA, paxM, paxB, paxC, paxP and paxQ that were all required for paxilline production, with one additional gene, paxD, required for regular prenylation of the indole ring post paxilline synthesis. The two putative transcription factors, PP104 and PP105, were not co-regulated with the pax genes and based on targeted gene replacement, including the double knockout, did not have a role in paxilline production. The relationship of indole dimethylallyl transferases involved in prenylation of indole-diterpenes such as paxilline or lolitrem B, can be found as two disparate clades, not supported by prenylation type (e.g., regular or reverse). This paper provides insight into the P. paxilli indole-diterpene locus and reviews the recent advances identified in paxilline biosynthesis. PMID:23949005
Regulatory genes and their roles for improvement of antibiotic biosynthesis in Streptomyces.
Lu, Fengjuan; Hou, Yanyan; Zhang, Heming; Chu, Yiwen; Xia, Haiyang; Tian, Yongqiang
2017-08-01
The numerous secondary metabolites in Streptomyces spp. are crucial for various applications. For example, cephamycin C is used as an antibiotic, and avermectin is used as an insecticide. Specifically, antibiotic yield is closely related to many factors, such as the external environment, nutrition (including nitrogen and carbon sources), biosynthetic efficiency and the regulatory mechanisms in producing strains. There are various types of regulatory genes that work in different ways, such as pleiotropic (or global) regulatory genes, cluster-situated regulators, which are also called pathway-specific regulatory genes, and many other regulators. The study of regulatory genes that influence antibiotic biosynthesis in Streptomyces spp. not only provides a theoretical basis for antibiotic biosynthesis in Streptomyces but also helps to increase the yield of antibiotics via molecular manipulation of these regulatory genes. Currently, more and more emphasis is being placed on the regulatory genes of antibiotic biosynthetic gene clusters in Streptomyces spp., and many studies on these genes have been performed to improve the yield of antibiotics in Streptomyces. This paper lists many antibiotic biosynthesis regulatory genes in Streptomyces spp. and focuses on frequently investigated regulatory genes that are involved in pathway-specific regulation and pleiotropic regulation and their applications in genetic engineering.
Battenberg, Kai; Wren, Jannah A.; Hillman, Janell; Edwards, Joseph; Huang, Liujing
2016-01-01
ABSTRACT The actinobacterial genus Frankia establishes nitrogen-fixing root nodule symbioses with specific hosts within the nitrogen-fixing plant clade. Of four genetically distinct subgroups of Frankia, cluster I, II, and III strains are capable of forming effective nitrogen-fixing symbiotic associations, while cluster IV strains generally do not. Cluster II Frankia strains have rarely been detected in soil devoid of host plants, unlike cluster I or III strains, suggesting a stronger association with their host. To investigate the degree of host influence, we characterized the cluster II Frankia strain distribution in rhizosphere soil in three locations in northern California. The presence/absence of cluster II Frankia strains at a given site correlated significantly with the presence/absence of host plants on the site, as determined by glutamine synthetase (glnA) gene sequence analysis, and by microbiome analysis (16S rRNA gene) of a subset of host/nonhost rhizosphere soils. However, the distribution of cluster II Frankia strains was not significantly affected by other potential determinants such as host-plant species, geographical location, climate, soil pH, or soil type. Rhizosphere soil microbiome analysis showed that cluster II Frankia strains occupied only a minute fraction of the microbiome even in the host-plant-present site and further revealed no statistically significant difference in the α-diversity or in the microbiome composition between the host-plant-present or -absent sites. Taken together, these data suggest that host plants provide a factor that is specific for cluster II Frankia strains, not a general growth-promoting factor. Further, the factor accumulates or is transported at the site level, i.e., beyond the host rhizosphere. IMPORTANCE Biological nitrogen fixation is a bacterial process that accounts for a major fraction of net new nitrogen input in terrestrial ecosystems. Transfer of fixed nitrogen to plant biomass is especially efficient via root nodule symbioses, which represent evolutionarily and ecologically specialized mutualistic associations. Frankia spp. (Actinobacteria), especially cluster II Frankia spp., have an extremely broad host range, yet comparatively little is known about the soil ecology of these organisms in relation to the host plants and their rhizosphere microbiomes. This study reveals a strong influence of the host plant on soil distribution of cluster II Frankia spp. PMID:27795313
2014-01-01
Background Highly adapted plant species are able to alter their root architecture to improve nutrient uptake and thrive in environments with limited nutrient supply. Cluster roots (CRs) are specialised structures of dense lateral roots formed by several plant species for the effective mining of nutrient rich soil patches through a combination of increased surface area and exudation of carboxylates. White lupin is becoming a model-species allowing for the discovery of gene networks involved in CR development. A greater understanding of the underlying molecular mechanisms driving these developmental processes is important for the generation of smarter plants for a world with diminishing resources to improve food security. Results RNA-seq analyses for three developmental stages of the CR formed under phosphorus-limited conditions and two of non-cluster roots have been performed for white lupin. In total 133,045,174 high-quality paired-end reads were used for a de novo assembly of the root transcriptome and merged with LAGI01 (Lupinus albus gene index) to generate an improved LAGI02 with 65,097 functionally annotated contigs. This was followed by comparative gene expression analysis. We show marked differences in the transcriptional response across the various cluster root stages to adjust to phosphate limitation by increasing uptake capacity and adjusting metabolic pathways. Several transcription factors such as PLT, SCR, PHB, PHV or AUX/IAA with a known role in the control of meristem activity and developmental processes show an increased expression in the tip of the CR. Genes involved in hormonal responses (PIN, LAX, YUC) and cell cycle control (CYCA/B, CDK) are also differentially expressed. In addition, we identify primary transcripts of miRNAs with established function in the root meristem. Conclusions Our gene expression analysis shows an intricate network of transcription factors and plant hormones controlling CR initiation and formation. In addition, functional differences between the different CR developmental stages in the acclimation to phosphorus starvation have been identified. PMID:24666749
Prokaryotic Gene Clusters: A Rich Toolbox for Synthetic Biology
Fischbach, Michael; Voigt, Christopher A.
2014-01-01
Bacteria construct elaborate nanostructures, obtain nutrients and energy from diverse sources, synthesize complex molecules, and implement signal processing to react to their environment. These complex phenotypes require the coordinated action of multiple genes, which are often encoded in a contiguous region of the genome, referred to as a gene cluster. Gene clusters sometimes contain all of the genes necessary and sufficient for a particular function. As an evolutionary mechanism, gene clusters facilitate the horizontal transfer of the complete function between species. Here, we review recent work on a number of clusters whose functions are relevant to biotechnology. Engineering these clusters has been hindered by their regulatory complexity, the need to balance the expression of many genes, and a lack of tools to design and manipulate DNA at this scale. Advances in synthetic biology will enable the large-scale bottom-up engineering of the clusters to optimize their functions, wake up cryptic clusters, or to transfer them between organisms. Understanding and manipulating gene clusters will move towards an era of genome engineering, where multiple functions can be “mixed-and-matched” to create a designer organism. PMID:21154668
The WRKY transcription factor family and senescence in switchgrass.
Rinerson, Charles I; Scully, Erin D; Palmer, Nathan A; Donze-Reiner, Teresa; Rabara, Roel C; Tripathi, Prateek; Shen, Qingxi J; Sattler, Scott E; Rohila, Jai S; Sarath, Gautam; Rushton, Paul J
2015-11-09
Early aerial senescence in switchgrass (Panicum virgatum) can significantly limit biomass yields. WRKY transcription factors that can regulate senescence could be used to reprogram senescence and enhance biomass yields. All potential WRKY genes present in the version 1.0 of the switchgrass genome were identified and curated using manual and bioinformatic methods. Expression profiles of WRKY genes in switchgrass flag leaf RNA-Seq datasets were analyzed using clustering and network analyses tools to identify both WRKY and WRKY-associated gene co-expression networks during leaf development and senescence onset. We identified 240 switchgrass WRKY genes including members of the RW5 and RW6 families of resistance proteins. Weighted gene co-expression network analysis of the flag leaf transcriptomes across development readily separated clusters of co-expressed genes into thirteen modules. A visualization highlighted separation of modules associated with the early and senescence-onset phases of flag leaf growth. The senescence-associated module contained 3000 genes including 23 WRKYs. Putative promoter regions of senescence-associated WRKY genes contained several cis-element-like sequences suggestive of responsiveness to both senescence and stress signaling pathways. A phylogenetic comparison of senescence-associated WRKY genes from switchgrass flag leaf with senescence-associated WRKY genes from other plants revealed notable hotspots in Group I, IIb, and IIe of the phylogenetic tree. We have identified and named 240 WRKY genes in the switchgrass genome. Twenty three of these genes show elevated mRNA levels during the onset of flag leaf senescence. Eleven of the WRKY genes were found in hotspots of related senescence-associated genes from multiple species and thus represent promising targets for future switchgrass genetic improvement. Overall, individual WRKY gene expression profiles could be readily linked to developmental stages of flag leaves.
GraphTeams: a method for discovering spatial gene clusters in Hi-C sequencing data.
Schulz, Tizian; Stoye, Jens; Doerr, Daniel
2018-05-08
Hi-C sequencing offers novel, cost-effective means to study the spatial conformation of chromosomes. We use data obtained from Hi-C experiments to provide new evidence for the existence of spatial gene clusters. These are sets of genes with associated functionality that exhibit close proximity to each other in the spatial conformation of chromosomes across several related species. We present the first gene cluster model capable of handling spatial data. Our model generalizes a popular computational model for gene cluster prediction, called δ-teams, from sequences to graphs. Following previous lines of research, we subsequently extend our model to allow for several vertices being associated with the same label. The model, called δ-teams with families, is particular suitable for our application as it enables handling of gene duplicates. We develop algorithmic solutions for both models. We implemented the algorithm for discovering δ-teams with families and integrated it into a fully automated workflow for discovering gene clusters in Hi-C data, called GraphTeams. We applied it to human and mouse data to find intra- and interchromosomal gene cluster candidates. The results include intrachromosomal clusters that seem to exhibit a closer proximity in space than on their chromosomal DNA sequence. We further discovered interchromosomal gene clusters that contain genes from different chromosomes within the human genome, but are located on a single chromosome in mouse. By identifying δ-teams with families, we provide a flexible model to discover gene cluster candidates in Hi-C data. Our analysis of Hi-C data from human and mouse reveals several known gene clusters (thus validating our approach), but also few sparsely studied or possibly unknown gene cluster candidates that could be the source of further experimental investigations.
Finding approximate gene clusters with Gecko 3.
Winter, Sascha; Jahn, Katharina; Wehner, Stefanie; Kuchenbecker, Leon; Marz, Manja; Stoye, Jens; Böcker, Sebastian
2016-11-16
Gene-order-based comparison of multiple genomes provides signals for functional analysis of genes and the evolutionary process of genome organization. Gene clusters are regions of co-localized genes on genomes of different species. The rapid increase in sequenced genomes necessitates bioinformatics tools for finding gene clusters in hundreds of genomes. Existing tools are often restricted to few (in many cases, only two) genomes, and often make restrictive assumptions such as short perfect conservation, conserved gene order or monophyletic gene clusters. We present Gecko 3, an open-source software for finding gene clusters in hundreds of bacterial genomes, that comes with an easy-to-use graphical user interface. The underlying gene cluster model is intuitive, can cope with low degrees of conservation as well as misannotations and is complemented by a sound statistical evaluation. To evaluate the biological benefit of Gecko 3 and to exemplify our method, we search for gene clusters in a dataset of 678 bacterial genomes using Synechocystis sp. PCC 6803 as a reference. We confirm detected gene clusters reviewing the literature and comparing them to a database of operons; we detect two novel clusters, which were confirmed by publicly available experimental RNA-Seq data. The computational analysis is carried out on a laptop computer in <40 min. © The Author(s) 2016. Published by Oxford University Press on behalf of Nucleic Acids Research.
Genomic and Metabolomic Profile Associated to Clustering of Cardio-Metabolic Risk Factors
Marrachelli, Vannina G.; Rentero, Pilar; Mansego, María L.; Morales, Jose Manuel; Galan, Inma; Pardo-Tendero, Mercedes; Martinez, Fernando; Martin-Escudero, Juan Carlos; Briongos, Laisa; Chaves, Felipe Javier; Redon, Josep; Monleon, Daniel
2016-01-01
Background To identify metabolomic and genomic markers associated with the presence of clustering of cardiometabolic risk factors (CMRFs) from a general population. Methods and Findings One thousand five hundred and two subjects, Caucasian, > 18 years, representative of the general population, were included. Blood pressure measurement, anthropometric parameters and metabolic markers were measured. Subjects were grouped according the number of CMRFs (Group 1: <2; Group 2: 2; Group 3: 3 or more CMRFs). Using SNPlex, 1251 SNPs potentially associated to clustering of three or more CMRFs were analyzed. Serum metabolomic profile was assessed by 1H NMR spectra using a Brucker Advance DRX 600 spectrometer. From the total population, 1217 (mean age 54±19, 50.6% men) with high genotyping call rate were analysed. A differential metabolomic profile, which included products from mitochondrial metabolism, extra mitochondrial metabolism, branched amino acids and fatty acid signals were observed among the three groups. The comparison of metabolomic patterns between subjects of Groups 1 to 3 for each of the genotypes associated to those subjects with three or more CMRFs revealed two SNPs, the rs174577_AA of FADS2 gene and the rs3803_TT of GATA2 transcription factor gene, with minimal or no statistically significant differences. Subjects with and without three or more CMRFs who shared the same genotype and metabolomic profile differed in the pattern of CMRFS cluster. Subjects of Group 3 and the AA genotype of the rs174577 had a lower prevalence of hypertension compared to the CC and CT genotype. In contrast, subjects of Group 3 and the TT genotype of the rs3803 polymorphism had a lower prevalence of T2DM, although they were predominantly males and had higher values of plasma creatinine. Conclusions The results of the present study add information to the metabolomics profile and to the potential impact of genetic factors on the variants of clustering of cardiometabolic risk factors. PMID:27589269
Genomic and Metabolomic Profile Associated to Clustering of Cardio-Metabolic Risk Factors.
Marrachelli, Vannina G; Rentero, Pilar; Mansego, María L; Morales, Jose Manuel; Galan, Inma; Pardo-Tendero, Mercedes; Martinez, Fernando; Martin-Escudero, Juan Carlos; Briongos, Laisa; Chaves, Felipe Javier; Redon, Josep; Monleon, Daniel
2016-01-01
To identify metabolomic and genomic markers associated with the presence of clustering of cardiometabolic risk factors (CMRFs) from a general population. One thousand five hundred and two subjects, Caucasian, > 18 years, representative of the general population, were included. Blood pressure measurement, anthropometric parameters and metabolic markers were measured. Subjects were grouped according the number of CMRFs (Group 1: <2; Group 2: 2; Group 3: 3 or more CMRFs). Using SNPlex, 1251 SNPs potentially associated to clustering of three or more CMRFs were analyzed. Serum metabolomic profile was assessed by 1H NMR spectra using a Brucker Advance DRX 600 spectrometer. From the total population, 1217 (mean age 54±19, 50.6% men) with high genotyping call rate were analysed. A differential metabolomic profile, which included products from mitochondrial metabolism, extra mitochondrial metabolism, branched amino acids and fatty acid signals were observed among the three groups. The comparison of metabolomic patterns between subjects of Groups 1 to 3 for each of the genotypes associated to those subjects with three or more CMRFs revealed two SNPs, the rs174577_AA of FADS2 gene and the rs3803_TT of GATA2 transcription factor gene, with minimal or no statistically significant differences. Subjects with and without three or more CMRFs who shared the same genotype and metabolomic profile differed in the pattern of CMRFS cluster. Subjects of Group 3 and the AA genotype of the rs174577 had a lower prevalence of hypertension compared to the CC and CT genotype. In contrast, subjects of Group 3 and the TT genotype of the rs3803 polymorphism had a lower prevalence of T2DM, although they were predominantly males and had higher values of plasma creatinine. The results of the present study add information to the metabolomics profile and to the potential impact of genetic factors on the variants of clustering of cardiometabolic risk factors.
Differential expression of cysteine desulfurases in soybean
2011-01-01
Background Iron-sulfur [Fe-S] clusters are prosthetic groups required to sustain fundamental life processes including electron transfer, metabolic reactions, sensing, signaling, gene regulation and stabilization of protein structures. In plants, the biogenesis of Fe-S protein is compartmentalized and adapted to specific needs of the cell. Many environmental factors affect plant development and limit productivity and geographical distribution. The impact of these limiting factors is particularly relevant for major crops, such as soybean, which has worldwide economic importance. Results Here we analyze the transcriptional profile of the soybean cysteine desulfurases NFS1, NFS2 and ISD11 genes, involved in the biogenesis of [Fe-S] clusters, by quantitative RT-PCR. NFS1, ISD11 and NFS2 encoding two mitochondrial and one plastid located proteins, respectively, are duplicated and showed distinct transcript levels considering tissue and stress response. NFS1 and ISD11 are highly expressed in roots, whereas NFS2 showed no differential expression in tissues. Cold-treated plants showed a decrease in NFS2 and ISD11 transcript levels in roots, and an increased expression of NFS1 and ISD11 genes in leaves. Plants treated with salicylic acid exhibited increased NFS1 transcript levels in roots but lower levels in leaves. In silico analysis of promoter regions indicated the presence of different cis-elements in cysteine desulfurase genes, in good agreement with differential expression of each locus. Our data also showed that increasing of transcript levels of mitochondrial genes, NFS1/ISD11, are associated with higher activities of aldehyde oxidase and xanthine dehydrogenase, two cytosolic Fe-S proteins. Conclusions Our results suggest a relationship between gene expression pattern, biochemical effects, and transcription factor binding sites in promoter regions of cysteine desulfurase genes. Moreover, data show proportionality between NFS1 and ISD11 genes expression. PMID:22099069
Shahdoust, Maryam; Hajizadeh, Ebrahim; Mozdarani, Hossein; Chehrei, Ali
2013-01-01
Cigarette smoking is the major risk factor for development of lung cancer. Identification of effects of tobacco on airway gene expression may provide insight into the causes. This research aimed to compare gene expression of large airway epithelium cells in normal smokers (n=13) and non-smokers (n=9) in order to find genes which discriminate the two groups and assess cigarette smoking effects on large airway epithelium cells. Genes discriminating smokers from non-smokers were identified by applying a neural network clustering method, growing self-organizing maps (GSOM), to microarray data according to class discrimination scores. An index was computed based on differentiation between each mean of gene expression in the two groups. This clustering approach provided the possibility of comparing thousands of genes simultaneously. The applied approach compared the mean of 7,129 genes in smokers and non-smokers simultaneously and classified the genes of large airway epithelium cells which had differently expressed in smokers comparing with non-smokers. Seven genes were identified which had the highest different expression in smokers compared with the non-smokers group: NQO1, H19, ALDH3A1, AKR1C1, ABHD2, GPX2 and ADH7. Most (NQO1, ALDH3A1, AKR1C1, H19 and GPX2) are known to be clinically notable in lung cancer studies. Furthermore, statistical discriminate analysis showed that these genes could classify samples in smokers and non-smokers correctly with 100% accuracy. With the performed GSOM map, other nodes with high average discriminate scores included genes with alterations strongly related to the lung cancer such as AKR1C3, CYP1B1, UCHL1 and AKR1B10. This clustering by comparing expression of thousands of genes at the same time revealed alteration in normal smokers. Most of the identified genes were strongly relevant to lung cancer in the existing literature. The genes may be utilized to identify smokers with increased risk for lung cancer. A large sample study is now recommended to determine relations between the genes ABHD2 and ADH7 and smoking.
Differential expression of anti-angiogenic factors and guidance genes in the developing macula.
Kozulin, Peter; Natoli, Riccardo; O'Brien, Keely M Bumsted; Madigan, Michele C; Provis, Jan M
2009-01-01
The primate retina contains a specialized, cone-rich macula, which mediates high acuity and color vision. The spatial resolution provided by the neural retina at the macula is optimized by stereotyped retinal blood vessel and ganglion cell axon patterning, which radiate away from the macula and reduce shadowing of macular photoreceptors. However, the genes that mediate these specializations, and the reasons for the vulnerability of the macula to degenerative disease, remain obscure. The aim of this study was to identify novel genes that may influence retinal vascular patterning and definition of the foveal avascular area. We used RNA from human fetal retinas at 19-20 weeks of gestation (WG; n=4) to measure differential gene expression in the macula, a region nasal to disc (nasal) and in the surrounding retina (surround) by hybridization to 12 GeneChip microarrays (HG-U133 Plus 2.0). The raw data was subjected to quality control assessment and preprocessing, using GC-RMA. We then used ANOVA analysis (Partek) Genomic Suite 6.3) and clustering (DAVID website) to identify the most highly represented genes clustered according to "biological process." The neural retina is fully differentiated at the macula at 19-20 WG, while neuronal progenitor cells are present throughout the rest of the retina. We therefore excluded genes associated with the cell cycle, and markers of differentiated neurons, from further analyses. Significantly regulated genes (p<0.01) were then identified in a second round of clustering according to molecular/reaction (KEGG) pathway. Genes of interest were verified by quantitative PCR (QRT-PCR), and 2 genes were localized by in situ hybridization. We generated two lists of differentially regulated genes: "macula versus surround" and "macula versus nasal." KEGG pathway clustering of the filtered gene lists identified 25 axon guidance-related genes that are differentially regulated in the macula. Furthermore, we found significant upregulation of three anti-angiogenic factors in the macula: pigment epithelium derived factor (PEDF), natriuretic peptide precurusor B (NPPB), and collagen type IValpha2. Differential expression of several members of the ephrin and semaphorin axon guidance gene families, PEDF, and NPPB was verified by QRT-PCR. Localization of PEDF and Eph-A6 mRNAs in sections of macaque retina shows expression of both genes concentrates in the ganglion cell layer (GCL) at the developing fovea, consistent with an involvement in definition of the foveal avascular area. Because the axons of macular ganglion cells exit the retina from around 8 WG, we suggest that the axon guidance genes highly expressed at the macula at 19-20 WG are also involved in vascular patterning, along with PEDF and NPPB. Localization of both PEDF and Eph-A6 mRNAs to the GCL of the developing fovea supports this idea. It is possible that specialization of the macular vessels, including definition of the foveal avascular area, is mediated by processes that piggyback on axon guidance mechanisms in effect earlier in development. These findings may be useful to understand the vulnerability of the macula to degeneration and to develop new therapeutic strategies to inhibit neovascularization.
Functional clustering of time series gene expression data by Granger causality
2012-01-01
Background A common approach for time series gene expression data analysis includes the clustering of genes with similar expression patterns throughout time. Clustered gene expression profiles point to the joint contribution of groups of genes to a particular cellular process. However, since genes belong to intricate networks, other features, besides comparable expression patterns, should provide additional information for the identification of functionally similar genes. Results In this study we perform gene clustering through the identification of Granger causality between and within sets of time series gene expression data. Granger causality is based on the idea that the cause of an event cannot come after its consequence. Conclusions This kind of analysis can be used as a complementary approach for functional clustering, wherein genes would be clustered not solely based on their expression similarity but on their topological proximity built according to the intensity of Granger causality among them. PMID:23107425
Engineering synthetic TALE and CRISPR/Cas9 transcription factors for regulating gene expression.
Kabadi, Ami M; Gersbach, Charles A
2014-09-01
Engineered DNA-binding proteins that can be targeted to specific sites in the genome to manipulate gene expression have enabled many advances in biomedical research. This includes generating tools to study fundamental aspects of gene regulation and the development of a new class of gene therapies that alter the expression of endogenous genes. Designed transcription factors have entered clinical trials for the treatment of human diseases and others are in preclinical development. High-throughput and user-friendly platforms for designing synthetic DNA-binding proteins present innovative methods for deciphering cell biology and designing custom synthetic gene circuits. We review two platforms for designing synthetic transcription factors for manipulating gene expression: Transcription activator-like effectors (TALEs) and the RNA-guided clustered regularly interspaced short palindromic repeats (CRISPR)/Cas9 system. We present an overview of each technology and a guide for designing and assembling custom TALE- and CRISPR/Cas9-based transcription factors. We also discuss characteristics of each platform that are best suited for different applications. Copyright © 2014 Elsevier Inc. All rights reserved.
Swaminathan, Sivakumar; Morrone, Dana; Wang, Qiang; Fulton, D. Bruce; Peters, Reuben J.
2009-01-01
Biosynthetic gene clusters are common in microbial organisms, but rare in plants, raising questions regarding the evolutionary forces that drive their assembly in multicellular eukaryotes. Here, we characterize the biochemical function of a rice (Oryza sativa) cytochrome P450 monooxygenase, CYP76M7, which seems to act in the production of antifungal phytocassanes and defines a second diterpenoid biosynthetic gene cluster in rice. This cluster is uniquely multifunctional, containing enzymatic genes involved in the production of two distinct sets of phytoalexins, the antifungal phytocassanes and antibacterial oryzalides/oryzadiones, with the corresponding genes being subject to distinct transcriptional regulation. The lack of uniform coregulation of the genes within this multifunctional cluster suggests that this was not a primary driving force in its assembly. However, the cluster is dedicated to specialized metabolism, as all genes in the cluster are involved in phytoalexin metabolism. We hypothesize that this dedication to specialized metabolism led to the assembly of the corresponding biosynthetic gene cluster. Consistent with this hypothesis, molecular phylogenetic comparison demonstrates that the two rice diterpenoid biosynthetic gene clusters have undergone independent elaboration to their present-day forms, indicating continued evolutionary pressure for coclustering of enzymatic genes encoding components of related biosynthetic pathways. PMID:19825834
Distribution and Genetic Diversity of Bacteriocin Gene Clusters in Rumen Microbial Genomes.
Azevedo, Analice C; Bento, Cláudia B P; Ruiz, Jeronimo C; Queiroz, Marisa V; Mantovani, Hilário C
2015-10-01
Some species of ruminal bacteria are known to produce antimicrobial peptides, but the screening procedures have mostly been based on in vitro assays using standardized methods. Recent sequencing efforts have made available the genome sequences of hundreds of ruminal microorganisms. In this work, we performed genome mining of the complete and partial genome sequences of 224 ruminal bacteria and 5 ruminal archaea to determine the distribution and diversity of bacteriocin gene clusters. A total of 46 bacteriocin gene clusters were identified in 33 strains of ruminal bacteria. Twenty gene clusters were related to lanthipeptide biosynthesis, while 11 gene clusters were associated with sactipeptide production, 7 gene clusters were associated with class II bacteriocin production, and 8 gene clusters were associated with class III bacteriocin production. The frequency of strains whose genomes encode putative antimicrobial peptide precursors was 14.4%. Clusters related to the production of sactipeptides were identified for the first time among ruminal bacteria. BLAST analysis indicated that the majority of the gene clusters (88%) encoding putative lanthipeptides contained all the essential genes required for lanthipeptide biosynthesis. Most strains of Streptococcus (66.6%) harbored complete lanthipeptide gene clusters, in addition to an open reading frame encoding a putative class II bacteriocin. Albusin B-like proteins were found in 100% of the Ruminococcus albus strains screened in this study. The in silico analysis provided evidence of novel biosynthetic gene clusters in bacterial species not previously related to bacteriocin production, suggesting that the rumen microbiota represents an underexplored source of antimicrobial peptides. Copyright © 2015, American Society for Microbiology. All Rights Reserved.
Han, Xiaolong; Chakrabortti, Alolika; Zhu, Jindong; Liang, Zhao-Xun; Li, Jinming
2016-08-15
Aspergillus westerdijkiae produces ochratoxin A (OTA) in Aspergillus section Circumdati. It is responsible for the contamination of agricultural crops, fruits, and food commodities, as its secondary metabolite OTA poses a potential threat to animals and humans. As a member of the filamentous fungi family, its capacity for enzymatic catalysis and secondary metabolite production is valuable in industrial production and medicine. To understand the genetic factors underlying its pathogenicity, enzymatic degradation, and secondary metabolism, we analysed the whole genome of A. westerdijkiae and compared it with eight other sequenced Aspergillus species. We sequenced the complete genome of A. westerdijkiae and assembled approximately 36 Mb of its genomic DNA, in which we identified 10,861 putative protein-coding genes. We constructed a phylogenetic tree of A. westerdijkiae and eight other sequenced Aspergillus species and found that the sister group of A. westerdijkiae was the A. oryzae - A. flavus clade. By searching the associated databases, we identified 716 cytochrome P450 enzymes, 633 carbohydrate-active enzymes, and 377 proteases. By combining comparative analysis with Kyoto Encyclopaedia of Genes and Genomes (KEGG), Conserved Domains Database (CDD), and Pfam annotations, we predicted 228 potential carbohydrate-active enzymes related to plant polysaccharide degradation (PPD). We found a large number of secondary biosynthetic gene clusters, which suggested that A. westerdijkiae had a remarkable capacity to produce secondary metabolites. Furthermore, we obtained two more reliable and integrated gene sequences containing the reported portions of OTA biosynthesis and identified their respective secondary metabolite clusters. We also systematically annotated these two hybrid t1pks-nrps gene clusters involved in OTA biosynthesis. These two clusters were separate in the genome, and one of them encoded a couple of GH3 and AA3 enzyme genes involved in sucrose and glucose metabolism. The genomic information obtained in this study is valuable for understanding the life cycle and pathogenicity of A. westerdijkiae. We identified numerous enzyme genes that are potentially involved in host invasion and pathogenicity, and we provided a preliminary prediction for each putative secondary metabolite (SM) gene cluster. In particular, for the OTA-related SM gene clusters, we delivered their components with domain and pathway annotations. This study sets the stage for experimental verification of the biosynthetic and regulatory mechanisms of OTA and for the discovery of new secondary metabolites.
Drivers of genetic diversity in secondary metabolic gene clusters within a fungal species
Lind, Abigail L.; Wisecaver, Jennifer H.; Lameiras, Catarina; Wiemann, Philipp; Palmer, Jonathan M.; Keller, Nancy P.; Rodrigues, Fernando; Goldman, Gustavo H.
2017-01-01
Filamentous fungi produce a diverse array of secondary metabolites (SMs) critical for defense, virulence, and communication. The metabolic pathways that produce SMs are found in contiguous gene clusters in fungal genomes, an atypical arrangement for metabolic pathways in other eukaryotes. Comparative studies of filamentous fungal species have shown that SM gene clusters are often either highly divergent or uniquely present in one or a handful of species, hampering efforts to determine the genetic basis and evolutionary drivers of SM gene cluster divergence. Here, we examined SM variation in 66 cosmopolitan strains of a single species, the opportunistic human pathogen Aspergillus fumigatus. Investigation of genome-wide within-species variation revealed 5 general types of variation in SM gene clusters: nonfunctional gene polymorphisms; gene gain and loss polymorphisms; whole cluster gain and loss polymorphisms; allelic polymorphisms, in which different alleles corresponded to distinct, nonhomologous clusters; and location polymorphisms, in which a cluster was found to differ in its genomic location across strains. These polymorphisms affect the function of representative A. fumigatus SM gene clusters, such as those involved in the production of gliotoxin, fumigaclavine, and helvolic acid as well as the function of clusters with undefined products. In addition to enabling the identification of polymorphisms, the detection of which requires extensive genome-wide synteny conservation (e.g., mobile gene clusters and nonhomologous cluster alleles), our approach also implicated multiple underlying genetic drivers, including point mutations, recombination, and genomic deletion and insertion events as well as horizontal gene transfer from distant fungi. Finally, most of the variants that we uncover within A. fumigatus have been previously hypothesized to contribute to SM gene cluster diversity across entire fungal classes and phyla. We suggest that the drivers of genetic diversity operating within a fungal species shown here are sufficient to explain SM cluster macroevolutionary patterns. PMID:29149178
Liu, Ying; Ciliax, Brian J; Borges, Karin; Dasigi, Venu; Ram, Ashwin; Navathe, Shamkant B; Dingledine, Ray
2004-01-01
One of the key challenges of microarray studies is to derive biological insights from the unprecedented quatities of data on gene-expression patterns. Clustering genes by functional keyword association can provide direct information about the nature of the functional links among genes within the derived clusters. However, the quality of the keyword lists extracted from biomedical literature for each gene significantly affects the clustering results. We extracted keywords from MEDLINE that describes the most prominent functions of the genes, and used the resulting weights of the keywords as feature vectors for gene clustering. By analyzing the resulting cluster quality, we compared two keyword weighting schemes: normalized z-score and term frequency-inverse document frequency (TFIDF). The best combination of background comparison set, stop list and stemming algorithm was selected based on precision and recall metrics. In a test set of four known gene groups, a hierarchical algorithm correctly assigned 25 of 26 genes to the appropriate clusters based on keywords extracted by the TDFIDF weighting scheme, but only 23 og 26 with the z-score method. To evaluate the effectiveness of the weighting schemes for keyword extraction for gene clusters from microarray profiles, 44 yeast genes that are differentially expressed during the cell cycle were used as a second test set. Using established measures of cluster quality, the results produced from TFIDF-weighted keywords had higher purity, lower entropy, and higher mutual information than those produced from normalized z-score weighted keywords. The optimized algorithms should be useful for sorting genes from microarray lists into functionally discrete clusters.
Shoguchi, Eiichi; Beedessee, Girish; Tada, Ipputa; Hisata, Kanako; Kawashima, Takeshi; Takeuchi, Takeshi; Arakaki, Nana; Fujie, Manabu; Koyanagi, Ryo; Roy, Michael C; Kawachi, Masanobu; Hidaka, Michio; Satoh, Noriyuki; Shinzato, Chuya
2018-06-14
The marine dinoflagellate, Symbiodinium, is a well-known photosynthetic partner for coral and other diverse, non-photosynthetic hosts in subtropical and tropical shallows, where it comprises an essential component of marine ecosystems. Using molecular phylogenetics, the genus Symbiodinium has been classified into nine major clades, A-I, and one of the reported differences among phenotypes is their capacity to synthesize mycosporine-like amino acids (MAAs), which absorb UV radiation. However, the genetic basis for this difference in synthetic capacity is unknown. To understand genetics underlying Symbiodinium diversity, we report two draft genomes, one from clade A, presumed to have been the earliest branching clade, and the other from clade C, in the terminal branch. The nuclear genome of Symbiodinium clade A (SymA) has more gene families than that of clade C, with larger numbers of organelle-related genes, including mitochondrial transcription terminal factor (mTERF) and Rubisco. While clade C (SymC) has fewer gene families, it displays specific expansions of repeat domain-containing genes, such as leucine-rich repeats (LRRs) and retrovirus-related dUTPases. Interestingly, the SymA genome encodes a gene cluster for MAA biosynthesis, potentially transferred from an endosymbiotic red alga (probably of bacterial origin), while SymC has completely lost these genes. Our analysis demonstrates that SymC appears to have evolved by losing gene families, such as the MAA biosynthesis gene cluster. In contrast to the conservation of genes related to photosynthetic ability, the terminal clade has suffered more gene family losses than other clades, suggesting a possible adaptation to symbiosis. Overall, this study implies that Symbiodinium ecology drives acquisition and loss of gene families.
2010-01-01
Background Cluster analysis, and in particular hierarchical clustering, is widely used to extract information from gene expression data. The aim is to discover new classes, or sub-classes, of either individuals or genes. Performing a cluster analysis commonly involve decisions on how to; handle missing values, standardize the data and select genes. In addition, pre-processing, involving various types of filtration and normalization procedures, can have an effect on the ability to discover biologically relevant classes. Here we consider cluster analysis in a broad sense and perform a comprehensive evaluation that covers several aspects of cluster analyses, including normalization. Result We evaluated 2780 cluster analysis methods on seven publicly available 2-channel microarray data sets with common reference designs. Each cluster analysis method differed in data normalization (5 normalizations were considered), missing value imputation (2), standardization of data (2), gene selection (19) or clustering method (11). The cluster analyses are evaluated using known classes, such as cancer types, and the adjusted Rand index. The performances of the different analyses vary between the data sets and it is difficult to give general recommendations. However, normalization, gene selection and clustering method are all variables that have a significant impact on the performance. In particular, gene selection is important and it is generally necessary to include a relatively large number of genes in order to get good performance. Selecting genes with high standard deviation or using principal component analysis are shown to be the preferred gene selection methods. Hierarchical clustering using Ward's method, k-means clustering and Mclust are the clustering methods considered in this paper that achieves the highest adjusted Rand. Normalization can have a significant positive impact on the ability to cluster individuals, and there are indications that background correction is preferable, in particular if the gene selection is successful. However, this is an area that needs to be studied further in order to draw any general conclusions. Conclusions The choice of cluster analysis, and in particular gene selection, has a large impact on the ability to cluster individuals correctly based on expression profiles. Normalization has a positive effect, but the relative performance of different normalizations is an area that needs more research. In summary, although clustering, gene selection and normalization are considered standard methods in bioinformatics, our comprehensive analysis shows that selecting the right methods, and the right combinations of methods, is far from trivial and that much is still unexplored in what is considered to be the most basic analysis of genomic data. PMID:20937082
Determining Physical Mechanisms of Gene Expression Regulation from Single Cell Gene Expression Data.
Ezer, Daphne; Moignard, Victoria; Göttgens, Berthold; Adryan, Boris
2016-08-01
Many genes are expressed in bursts, which can contribute to cell-to-cell heterogeneity. It is now possible to measure this heterogeneity with high throughput single cell gene expression assays (single cell qPCR and RNA-seq). These experimental approaches generate gene expression distributions which can be used to estimate the kinetic parameters of gene expression bursting, namely the rate that genes turn on, the rate that genes turn off, and the rate of transcription. We construct a complete pipeline for the analysis of single cell qPCR data that uses the mathematics behind bursty expression to develop more accurate and robust algorithms for analyzing the origin of heterogeneity in experimental samples, specifically an algorithm for clustering cells by their bursting behavior (Simulated Annealing for Bursty Expression Clustering, SABEC) and a statistical tool for comparing the kinetic parameters of bursty expression across populations of cells (Estimation of Parameter changes in Kinetics, EPiK). We applied these methods to hematopoiesis, including a new single cell dataset in which transcription factors (TFs) involved in the earliest branchpoint of blood differentiation were individually up- and down-regulated. We could identify two unique sub-populations within a seemingly homogenous group of hematopoietic stem cells. In addition, we could predict regulatory mechanisms controlling the expression levels of eighteen key hematopoietic transcription factors throughout differentiation. Detailed information about gene regulatory mechanisms can therefore be obtained simply from high throughput single cell gene expression data, which should be widely applicable given the rapid expansion of single cell genomics.
Ortholog-based screening and identification of genes related to intracellular survival.
Yang, Xiaowen; Wang, Jiawei; Bing, Guoxia; Bie, Pengfei; De, Yanyan; Lyu, Yanli; Wu, Qingmin
2018-04-20
Bioinformatics and comparative genomics analysis methods were used to predict unknown pathogen genes based on homology with identified or functionally clustered genes. In this study, the genes of common pathogens were analyzed to screen and identify genes associated with intracellular survival through sequence similarity, phylogenetic tree analysis and the λ-Red recombination system test method. The total 38,952 protein-coding genes of common pathogens were divided into 19,775 clusters. As demonstrated through a COG analysis, information storage and processing genes might play an important role intracellular survival. Only 19 clusters were present in facultative intracellular pathogens, and not all were present in extracellular pathogens. Construction of a phylogenetic tree selected 18 of these 19 clusters. Comparisons with the DEG database and previous research revealed that seven other clusters are considered essential gene clusters and that seven other clusters are associated with intracellular survival. Moreover, this study confirmed that clusters screened by orthologs with similar function could be replaced with an approved uvrY gene and its orthologs, and the results revealed that the usg gene is associated with intracellular survival. The study improves the current understanding of intracellular pathogens characteristics and allows further exploration of the intracellular survival-related gene modules in these pathogens. Copyright © 2018. Published by Elsevier B.V.
Nowrousian, Minou
2009-04-01
During fungal fruiting body development, hyphae aggregate to form multicellular structures that protect and disperse the sexual spores. Analysis of microarray data revealed a gene cluster strongly upregulated during fruiting body development in the ascomycete Sordaria macrospora. Real time PCR analysis showed that the genes from the orthologous cluster in Neurospora crassa are also upregulated during development. The cluster encodes putative polyketide biosynthesis enzymes, including a reducing polyketide synthase. Analysis of knockout strains of a predicted dehydrogenase gene from the cluster showed that mutants in N. crassa and S. macrospora are delayed in fruiting body formation. In addition to the upregulated cluster, the N. crassa genome comprises another cluster containing a polyketide synthase gene, and five additional reducing polyketide synthase (rpks) genes that are not part of clusters. To study the role of these genes in sexual development, expression of the predicted rpks genes in S. macrospora (five genes) and N. crassa (six genes) was analyzed; all but one are upregulated during sexual development. Analysis of knockout strains for the N. crassa rpks genes showed that one of them is essential for fruiting body formation. These data indicate that polyketides produced by RPKSs are involved in sexual development in filamentous ascomycetes.
Singh, Anil Kumar; Sharma, Vishal; Pal, Awadhesh Kumar; Acharya, Vishal; Ahuja, Paramvir Singh
2013-08-01
NAC [no apical meristem (NAM), Arabidopsis thaliana transcription activation factor [ATAF1/2] and cup-shaped cotyledon (CUC2)] proteins belong to one of the largest plant-specific transcription factor (TF) families and play important roles in plant development processes, response to biotic and abiotic cues and hormone signalling. Our genome-wide analysis identified 110 StNAC genes in potato encoding for 136 proteins, including 14 membrane-bound TFs. The physical map positions of StNAC genes on 12 potato chromosomes were non-random, and 40 genes were found to be distributed in 16 clusters. The StNAC proteins were phylogenetically clustered into 12 subgroups. Phylogenetic analysis of StNACs along with their Arabidopsis and rice counterparts divided these proteins into 18 subgroups. Our comparative analysis has also identified 36 putative TNAC proteins, which appear to be restricted to Solanaceae family. In silico expression analysis, using Illumina RNA-seq transcriptome data, revealed tissue-specific, biotic, abiotic stress and hormone-responsive expression profile of StNAC genes. Several StNAC genes, including StNAC072 and StNAC101that are orthologs of known stress-responsive Arabidopsis RESPONSIVE TO DEHYDRATION 26 (RD26) were identified as highly abiotic stress responsive. Quantitative real-time polymerase chain reaction analysis largely corroborated the expression profile of StNAC genes as revealed by the RNA-seq data. Taken together, this analysis indicates towards putative functions of several StNAC TFs, which will provide blue-print for their functional characterization and utilization in potato improvement.
Poole, William; Leinonen, Kalle; Shmulevich, Ilya
2017-01-01
Cancer researchers have long recognized that somatic mutations are not uniformly distributed within genes. However, most approaches for identifying cancer mutations focus on either the entire-gene or single amino-acid level. We have bridged these two methodologies with a multiscale mutation clustering algorithm that identifies variable length mutation clusters in cancer genes. We ran our algorithm on 539 genes using the combined mutation data in 23 cancer types from The Cancer Genome Atlas (TCGA) and identified 1295 mutation clusters. The resulting mutation clusters cover a wide range of scales and often overlap with many kinds of protein features including structured domains, phosphorylation sites, and known single nucleotide variants. We statistically associated these multiscale clusters with gene expression and drug response data to illuminate the functional and clinical consequences of mutations in our clusters. Interestingly, we find multiple clusters within individual genes that have differential functional associations: these include PTEN, FUBP1, and CDH1. This methodology has potential implications in identifying protein regions for drug targets, understanding the biological underpinnings of cancer, and personalizing cancer treatments. Toward this end, we have made the mutation clusters and the clustering algorithm available to the public. Clusters and pathway associations can be interactively browsed at m2c.systemsbiology.net. The multiscale mutation clustering algorithm is available at https://github.com/IlyaLab/M2C. PMID:28170390
Poole, William; Leinonen, Kalle; Shmulevich, Ilya; Knijnenburg, Theo A; Bernard, Brady
2017-02-01
Cancer researchers have long recognized that somatic mutations are not uniformly distributed within genes. However, most approaches for identifying cancer mutations focus on either the entire-gene or single amino-acid level. We have bridged these two methodologies with a multiscale mutation clustering algorithm that identifies variable length mutation clusters in cancer genes. We ran our algorithm on 539 genes using the combined mutation data in 23 cancer types from The Cancer Genome Atlas (TCGA) and identified 1295 mutation clusters. The resulting mutation clusters cover a wide range of scales and often overlap with many kinds of protein features including structured domains, phosphorylation sites, and known single nucleotide variants. We statistically associated these multiscale clusters with gene expression and drug response data to illuminate the functional and clinical consequences of mutations in our clusters. Interestingly, we find multiple clusters within individual genes that have differential functional associations: these include PTEN, FUBP1, and CDH1. This methodology has potential implications in identifying protein regions for drug targets, understanding the biological underpinnings of cancer, and personalizing cancer treatments. Toward this end, we have made the mutation clusters and the clustering algorithm available to the public. Clusters and pathway associations can be interactively browsed at m2c.systemsbiology.net. The multiscale mutation clustering algorithm is available at https://github.com/IlyaLab/M2C.
Analysis of multiplex gene expression maps obtained by voxelation.
An, Li; Xie, Hongbo; Chin, Mark H; Obradovic, Zoran; Smith, Desmond J; Megalooikonomou, Vasileios
2009-04-29
Gene expression signatures in the mammalian brain hold the key to understanding neural development and neurological disease. Researchers have previously used voxelation in combination with microarrays for acquisition of genome-wide atlases of expression patterns in the mouse brain. On the other hand, some work has been performed on studying gene functions, without taking into account the location information of a gene's expression in a mouse brain. In this paper, we present an approach for identifying the relation between gene expression maps obtained by voxelation and gene functions. To analyze the dataset, we chose typical genes as queries and aimed at discovering similar gene groups. Gene similarity was determined by using the wavelet features extracted from the left and right hemispheres averaged gene expression maps, and by the Euclidean distance between each pair of feature vectors. We also performed a multiple clustering approach on the gene expression maps, combined with hierarchical clustering. Among each group of similar genes and clusters, the gene function similarity was measured by calculating the average gene function distances in the gene ontology structure. By applying our methodology to find similar genes to certain target genes we were able to improve our understanding of gene expression patterns and gene functions. By applying the clustering analysis method, we obtained significant clusters, which have both very similar gene expression maps and very similar gene functions respectively to their corresponding gene ontologies. The cellular component ontology resulted in prominent clusters expressed in cortex and corpus callosum. The molecular function ontology gave prominent clusters in cortex, corpus callosum and hypothalamus. The biological process ontology resulted in clusters in cortex, hypothalamus and choroid plexus. Clusters from all three ontologies combined were most prominently expressed in cortex and corpus callosum. The experimental results confirm the hypothesis that genes with similar gene expression maps might have similar gene functions. The voxelation data takes into account the location information of gene expression level in mouse brain, which is novel in related research. The proposed approach can potentially be used to predict gene functions and provide helpful suggestions to biologists.
Jarvi, S.I.; Tarr, C.L.; Mcintosh, C.E.; Atkinson, C.T.; Fleischer, R.C.
2004-01-01
The native Hawaiian honeycreepers represent a classic example of adaptive radiation and speciation, but currently face one the highest extinction rates in the world. Although multiple factors have likely influenced the fate of Hawaiian birds, the relatively recent introduction of avian malaria is thought to be a major factor limiting honeycreeper distribution and abundance. We have initiated genetic analyses of class II ?? chain Mhc genes in four species of honeycreepers using methods that eliminate the possibility of sequencing mosaic variants formed by cloning heteroduplexed polymerase chain reaction products. Phylogenetic analyses group the honeycreeper Mhc sequences into two distinct clusters. Variation within one cluster is high, with dN > d S and levels of diversity similar to other studies of Mhc (B system) genes in birds. The second cluster is nearly invariant and includes sequences from honeycreepers (Fringillidae), a sparrow (Emberizidae) and a blackbird (Emberizidae). This highly conserved cluster appears reminiscent of the independently segregating Rfp-Y system of genes defined in chickens. The notion that balancing selection operates at the Mhc in the honeycreepers is supported by transpecies polymorphism and strikingly high dN/dS ratios at codons putatively involved in peptide interaction. Mitochondrial DNA control region sequences were invariant in the i'iwi, but were highly variable in the 'amakihi. By contrast, levels of variability of class II ?? chain Mhc sequence codons that are hypothesized to be directly involved in peptide interactions appear comparable between i'iwi and 'amakihi. In the i'iwi, natural selection may have maintained variation within the Mhc, even in the face of what appears to a genetic bottleneck.
Conserved noncoding sequences (CNSs) in higher plants.
Freeling, Michael; Subramaniam, Shabarinath
2009-04-01
Plant conserved noncoding sequences (CNSs)--a specific category of phylogenetic footprint--have been shown experimentally to function. No plant CNS is conserved to the extent that ultraconserved noncoding sequences are conserved in vertebrates. Plant CNSs are enriched in known transcription factor or other cis-acting binding sites, and are usually clustered around genes. Genes that encode transcription factors and/or those that respond to stimuli are particularly CNS-rich. Only rarely could this function involve small RNA binding. Some transcribed CNSs encode short translation products as a form of negative control. Approximately 4% of Arabidopsis gene content is estimated to be both CNS-rich and occupies a relatively long stretch of chromosome: Bigfoot genes (long phylogenetic footprints). We discuss a 'DNA-templated protein assembly' idea that might help explain Bigfoot gene CNSs.
Haarmann, Thomas; Machado, Caroline; Lübbe, Yvonne; Correia, Telmo; Schardl, Christopher L; Panaccione, Daniel G; Tudzynski, Paul
2005-06-01
The genomic region of Claviceps purpurea strain P1 containing the ergot alkaloid gene cluster [Tudzynski, P., Hölter, K., Correia, T., Arntz, C., Grammel, N., Keller, U., 1999. Evidence for an ergot alkaloid gene cluster in Claviceps purpurea. Mol. Gen. Genet. 261, 133-141] was explored by chromosome walking, and additional genes probably involved in the ergot alkaloid biosynthesis have been identified. The putative cluster sequence (extending over 68.5kb) contains 4 different nonribosomal peptide synthetase (NRPS) genes and several putative oxidases. Northern analysis showed that most of the genes were co-regulated (repressed by high phosphate), and identified probable flanking genes by lack of co-regulation. Comparison of the cluster sequences of strain P1, an ergotamine producer, with that of strain ECC93, an ergocristine producer, showed high conservation of most of the cluster genes, but significant variation in the NRPS modules, strongly suggesting that evolution of these chemical races of C. purpurea is determined by evolution of NRPS module specificity.
Patel, Vidushi S; Cooper, Steven J B; Deakin, Janine E; Fulton, Bob; Graves, Tina; Warren, Wesley C; Wilson, Richard K; Graves, Jennifer A M
2008-07-25
Vertebrate alpha (alpha)- and beta (beta)-globin gene families exemplify the way in which genomes evolve to produce functional complexity. From tandem duplication of a single globin locus, the alpha- and beta-globin clusters expanded, and then were separated onto different chromosomes. The previous finding of a fossil beta-globin gene (omega) in the marsupial alpha-cluster, however, suggested that duplication of the alpha-beta cluster onto two chromosomes, followed by lineage-specific gene loss and duplication, produced paralogous alpha- and beta-globin clusters in birds and mammals. Here we analyse genomic data from an egg-laying monotreme mammal, the platypus (Ornithorhynchus anatinus), to explore haemoglobin evolution at the stem of the mammalian radiation. The platypus alpha-globin cluster (chromosome 21) contains embryonic and adult alpha- globin genes, a beta-like omega-globin gene, and the GBY globin gene with homology to cytoglobin, arranged as 5'-zeta-zeta'-alphaD-alpha3-alpha2-alpha1-omega-GBY-3'. The platypus beta-globin cluster (chromosome 2) contains single embryonic and adult globin genes arranged as 5'-epsilon-beta-3'. Surprisingly, all of these globin genes were expressed in some adult tissues. Comparison of flanking sequences revealed that all jawed vertebrate alpha-globin clusters are flanked by MPG-C16orf35 and LUC7L, whereas all bird and mammal beta-globin clusters are embedded in olfactory genes. Thus, the mammalian alpha- and beta-globin clusters are orthologous to the bird alpha- and beta-globin clusters respectively. We propose that alpha- and beta-globin clusters evolved from an ancient MPG-C16orf35-alpha-beta-GBY-LUC7L arrangement 410 million years ago. A copy of the original beta (represented by omega in marsupials and monotremes) was inserted into an array of olfactory genes before the amniote radiation (>315 million years ago), then duplicated and diverged to form orthologous clusters of beta-globin genes with different expression profiles in different lineages.
Gene amplification of the transcription factor DP1 and CTNND1 in human lung cancer.
Castillo, Sandra D; Angulo, Barbara; Suarez-Gauthier, Ana; Melchor, Lorenzo; Medina, Pedro P; Sanchez-Verde, Lydia; Torres-Lanzas, Juan; Pita, Guillermo; Benitez, Javier; Sanchez-Cespedes, Montse
2010-09-01
The search for novel oncogenes is important because they could be the target of future specific anticancer therapies. In the present paper we report the identification of novel amplified genes in lung cancer by means of global gene expression analysis. To screen for amplicons, we aligned the gene expression data according to the position of transcripts in the human genome and searched for clusters of over-expressed genes. We found several clusters with gene over-expression, suggesting an underlying genomic amplification. FISH and microarray analysis for DNA copy number in two clusters, at chromosomes 11q12 and 13q34, confirmed the presence of amplifications spanning about 0.4 and 1 Mb for 11q12 and 13q34, respectively. Amplification at these regions each occurred at a frequency of 3%. Moreover, quantitative RT-PCR of each individual transcript within the amplicons allowed us to verify the increased in gene expression of several genes. The p120ctn and DP1 proteins, encoded by two candidate oncogenes, CTNND1 and TFDP1, at 11q12 and 13q amplicons, respectively, showed very strong immunostaining in lung tumours with gene amplification. We then focused on the 13q34 amplicon and in the TFDP1 candidate oncogene. To further determine the oncogenic properties of DP1, we searched for lung cancer cell lines carrying TFDP1 amplification. Depletion of TFDP1 expression by small interference RNA in a lung cancer cell line (HCC33) with TFDP1 amplification and protein over-expression reduced cell viability by 50%. In conclusion, we report the identification of two novel amplicons, at 13q34 and 11q12, each occurring at a frequency of 3% of non-small cell lung cancers. TFDP1, which encodes the E2F-associated transcription factor DP1 is a candidate oncogene at 13q34. The data discussed in this publication have been deposited in NCBIs Gene Expression Omnibus (GEO; http://www.ncbi.nlm.nih.gov/geo/) and are accessible through GEO Series Accession No. GSE21168.
Gao, Haiyan; Yang, Mei; Zhang, Xiaolan
2018-04-01
The present study aimed to investigate potential recurrence-risk biomarkers based on significant pathways for Luminal A breast cancer through gene expression profile analysis. Initially, the gene expression profiles of Luminal A breast cancer patients were downloaded from The Cancer Genome Atlas database. The differentially expressed genes (DEGs) were identified using a Limma package and the hierarchical clustering analysis was conducted for the DEGs. In addition, the functional pathways were screened using Kyoto Encyclopedia of Genes and Genomes pathway enrichment analyses and rank ratio calculation. The multigene prognostic assay was exploited based on the statistically significant pathways and its prognostic function was tested using train set and verified using the gene expression data and survival data of Luminal A breast cancer patients downloaded from the Gene Expression Omnibus. A total of 300 DEGs were identified between good and poor outcome groups, including 176 upregulated genes and 124 downregulated genes. The DEGs may be used to effectively distinguish Luminal A samples with different prognoses verified by hierarchical clustering analysis. There were 9 pathways screened as significant pathways and a total of 18 DEGs involved in these 9 pathways were identified as prognostic biomarkers. According to the survival analysis and receiver operating characteristic curve, the obtained 18-gene prognostic assay exhibited good prognostic function with high sensitivity and specificity to both the train and test samples. In conclusion the 18-gene prognostic assay including the key genes, transcription factor 7-like 2, anterior parietal cortex and lymphocyte enhancer factor-1 may provide a new method for predicting outcomes and may be conducive to the promotion of precision medicine for Luminal A breast cancer.
CORM: An R Package Implementing the Clustering of Regression Models Method for Gene Clustering
Shi, Jiejun; Qin, Li-Xuan
2014-01-01
We report a new R package implementing the clustering of regression models (CORM) method for clustering genes using gene expression data and provide data examples illustrating each clustering function in the package. The CORM package is freely available at CRAN from http://cran.r-project.org. PMID:25452684
Kienle, Dirk; Katzenberger, Tiemo; Ott, German; Saupe, Doreen; Benner, Axel; Kohlhammer, Holger; Barth, Thomas F E; Höller, Sylvia; Kalla, Jörg; Rosenwald, Andreas; Müller-Hermelink, Hans Konrad; Möller, Peter; Lichter, Peter; Döhner, Hartmut; Stilgenbauer, Stephan
2007-07-01
There is evidence for a direct role of quantitative gene expression deregulation in mantle-cell lymphoma (MCL) pathogenesis. Our aim was to investigate gene expression associations with other pathogenic factors and the significance of gene expression in a multivariate survival analysis. Quantitative expression of 20 genes of potential relevance for MCL prognosis and pathogenesis were analyzed using real-time reverse transcriptase polymerase chain reaction and correlated with clinical and genetic factors, tumor morphology, and Ki-67 index in 65 MCL samples. Genomic losses at the loci of TP53, RB1, and P16 were associated with reduced transcript levels of the respective genes, indicating a gene-dosage effect as the pathomechanism. Analysis of gene expression correlations between the candidate genes revealed a separation into two clusters, one dominated by proliferation activators, another by proliferation inhibitors and regulators of apoptosis. Whereas only weak associations were identified between gene expression and clinical parameters or blastoid morphology, several genes were correlated closely with the Ki-67 index, including the short CCND1 variant (positive correlation) and RB1, ATM, P27, and BMI (negative correlation). In multivariate survival analysis, expression levels of MYC, MDM2, EZH2, and CCND1 were the strongest prognostic factors independently of tumor proliferation and clinical factors. These results indicate a pathogenic contribution of several gene transcript levels to the biology and clinical course of MCL. Genes can be differentiated into factors contributing to proliferation deregulation, either by enhancement or loss of inhibition, and proliferation-independent factors potentially contributing to MCL pathogenesis by apoptosis impairment.
Sakai, Kanae; Komaki, Hisayuki; Gonoi, Tohru
2015-01-01
Nocardithiocin is a thiopeptide compound isolated from the opportunistic pathogen Nocardia pseudobrasiliensis. It shows a strong activity against acid-fast bacteria and is also active against rifampicin-resistant Mycobacterium tuberculosis. Here, we report the identification of the nocardithiocin gene cluster in N. pseudobrasiliensis IFM 0761 based on conserved thiopeptide biosynthesis gene sequence and the whole genome sequence. The predicted gene cluster was confirmed by gene disruption and complementation. As expected, strains containing the disrupted gene did not produce nocardithiocin while gene complementation restored nocardithiocin production in these strains. The predicted cluster was further analyzed using RNA-seq which showed that the nocardithiocin gene cluster contains 12 genes within a 15.2-kb region. This finding will promote the improvement of nocardithiocin productivity and its derivatives production. PMID:26588225
USDA-ARS?s Scientific Manuscript database
Background: In many bacteria including E. coli, genes encoding O-antigens are clustered in the chromosome, with a 39-bp JUMPstart sequence and gnd gene located upstream and downstream of the cluster, respectively. For determining the DNA sequence of the E. coli O-antigen gene cluster, one set of P...
Marui, Junichiro; Yamane, Noriko; Ohashi-Kunihiro, Sumiko; Ando, Tomohiro; Terabayashi, Yasunobu; Sano, Motoaki; Ohashi, Shinichi; Ohshima, Eiji; Tachibana, Kuniharu; Higa, Yoshitaka; Nishimura, Marie; Koike, Hideaki; Machida, Masayuki
2011-07-01
A gene encoding the Zn(II)(2)Cys(6) transcriptional factor is clustered with two genes involved in biosynthesis of a secondary metabolite, kojic acid (KA), in Aspergillus oryzae. We determined that the gene was essential for KA production and the transcriptional activation of KA biosynthetic genes, which were triggered by the addition of KA. Copyright © 2011 The Society for Biotechnology, Japan. Published by Elsevier B.V. All rights reserved.
[siRNA-mediated tissue factor knockdown in porcine neonatal islet cell clusters in vitro].
Ji, Ming; Yi, Shounan; Yu, Deling; Wang, Wei
2011-12-01
To determine the genetic modification on neonatal porcine islet cell clusters (NICC) by small interfering RNA (siRNA)-mediated tissue factor (TF) knockdown in vitro. Porcine NICC were transfected with 5 pairs of designed siRNA respectively or in different combinations with lipofectamine 2000. Transfected NICC were analyzed for TF gene by real-time PCR to select the siRNA which worked best. Meanwhile, the viability of NICC after the TF siRNA transfection was examined by FACS. The efficiency of TF gene and protein suppression was measured by real-time PCR and and FACS respectively. Real-time PCR and FACS showed that a 60% reduction in the TF gene expression and a 50% reduction in the protien level of TF on NICC were achieved by transfecting 3 pairs of selected siRNA. The siRNA transfection had no significant effect on the viability of NICC which was analyzed by FACS. The expression of TF on porcine NICC is efficiently suppressed by 3 pairs of designed siRNA in vitro.
USDA-ARS?s Scientific Manuscript database
Seed dormancy has been associated with red grain color in cereal crops. The association was linked to the cluster of quantitative trait loci qSD7-1/qPC7 in weedy red rice. This research delimited the cluster to Os07g11020 or Rc encoding a predicted bHLH family transcription factor by intragenic reco...
Ishihama, Nobuaki; Yamada, Reiko; Yoshioka, Miki; Katou, Shinpei; Yoshioka, Hirofumi
2011-01-01
Mitogen-activated protein kinase (MAPK) cascades have pivotal roles in plant innate immunity. However, downstream signaling of plant defense-related MAPKs is not well understood. Here, we provide evidence that the Nicotiana benthamiana WRKY8 transcription factor is a physiological substrate of SIPK, NTF4, and WIPK. Clustered Pro-directed Ser residues (SP cluster), which are conserved in group I WRKY proteins, in the N-terminal region of WRKY8 were phosphorylated by these MAPKs in vitro. Antiphosphopeptide antibodies indicated that Ser residues in the SP cluster of WRKY8 are phosphorylated by SIPK, NTF4, and WIPK in vivo. The interaction of WRKY8 with MAPKs depended on its D domain, which is a MAPK-interacting motif, and this interaction was required for effective phosphorylation of WRKY8 in plants. Phosphorylation of WRKY8 increased its DNA binding activity to the cognate W-box sequence. The phospho-mimicking mutant of WRKY8 showed higher transactivation activity, and its ectopic expression induced defense-related genes, such as 3-hydroxy-3-methylglutaryl CoA reductase 2 and NADP-malic enzyme. By contrast, silencing of WRKY8 decreased the expression of defense-related genes and increased disease susceptibility to the pathogens Phytophthora infestans and Colletotrichum orbiculare. Thus, MAPK-mediated phosphorylation of WRKY8 has an important role in the defense response through activation of downstream genes. PMID:21386030
Nykyri, Johanna; Mattinen, Laura; Niemi, Outi; Adhikari, Satish; Kõiv, Viia; Somervuo, Panu; Fang, Xin; Auvinen, Petri; Mäe, Andres; Palva, E. Tapio; Pirhonen, Minna
2013-01-01
In this study, we characterized a putative Flp/Tad pilus-encoding gene cluster, and we examined its regulation at the transcriptional level and its role in the virulence of potato pathogenic enterobacteria of the genus Pectobacterium. The Flp/Tad pilus-encoding gene clusters in Pectobacterium atrosepticum, Pectobacterium wasabiae and Pectobacterium aroidearum were compared to previously characterized flp/tad gene clusters, including that of the well-studied Flp/Tad pilus model organism Aggregatibacter actinomycetemcomitans, in which this pilus is a major virulence determinant. Comparative analyses revealed substantial protein sequence similarity and open reading frame synteny between the previously characterized flp/tad gene clusters and the cluster in Pectobacterium, suggesting that the predicted flp/tad gene cluster in Pectobacterium encodes a Flp/Tad pilus-like structure. We detected genes for a novel two-component system adjacent to the flp/tad gene cluster in Pectobacterium, and mutant analysis demonstrated that this system has a positive effect on the transcription of selected Flp/Tad pilus biogenesis genes, suggesting that this response regulator regulate the flp/tad gene cluster. Mutagenesis of either the predicted regulator gene or selected Flp/Tad pilus biogenesis genes had a significant impact on the maceration ability of the bacterial strains in potato tubers, indicating that the Flp/Tad pilus-encoding gene cluster represents a novel virulence determinant in Pectobacterium. Soft-rot enterobacteria in the genera Pectobacterium and Dickeya are of great agricultural importance, and an investigation of the virulence of these pathogens could facilitate improvements in agricultural practices, thus benefiting farmers, the potato industry and consumers. PMID:24040039
Nykyri, Johanna; Mattinen, Laura; Niemi, Outi; Adhikari, Satish; Kõiv, Viia; Somervuo, Panu; Fang, Xin; Auvinen, Petri; Mäe, Andres; Palva, E Tapio; Pirhonen, Minna
2013-01-01
In this study, we characterized a putative Flp/Tad pilus-encoding gene cluster, and we examined its regulation at the transcriptional level and its role in the virulence of potato pathogenic enterobacteria of the genus Pectobacterium. The Flp/Tad pilus-encoding gene clusters in Pectobacterium atrosepticum, Pectobacterium wasabiae and Pectobacterium aroidearum were compared to previously characterized flp/tad gene clusters, including that of the well-studied Flp/Tad pilus model organism Aggregatibacter actinomycetemcomitans, in which this pilus is a major virulence determinant. Comparative analyses revealed substantial protein sequence similarity and open reading frame synteny between the previously characterized flp/tad gene clusters and the cluster in Pectobacterium, suggesting that the predicted flp/tad gene cluster in Pectobacterium encodes a Flp/Tad pilus-like structure. We detected genes for a novel two-component system adjacent to the flp/tad gene cluster in Pectobacterium, and mutant analysis demonstrated that this system has a positive effect on the transcription of selected Flp/Tad pilus biogenesis genes, suggesting that this response regulator regulate the flp/tad gene cluster. Mutagenesis of either the predicted regulator gene or selected Flp/Tad pilus biogenesis genes had a significant impact on the maceration ability of the bacterial strains in potato tubers, indicating that the Flp/Tad pilus-encoding gene cluster represents a novel virulence determinant in Pectobacterium. Soft-rot enterobacteria in the genera Pectobacterium and Dickeya are of great agricultural importance, and an investigation of the virulence of these pathogens could facilitate improvements in agricultural practices, thus benefiting farmers, the potato industry and consumers.
An effective fuzzy kernel clustering analysis approach for gene expression data.
Sun, Lin; Xu, Jiucheng; Yin, Jiaojiao
2015-01-01
Fuzzy clustering is an important tool for analyzing microarray data. A major problem in applying fuzzy clustering method to microarray gene expression data is the choice of parameters with cluster number and centers. This paper proposes a new approach to fuzzy kernel clustering analysis (FKCA) that identifies desired cluster number and obtains more steady results for gene expression data. First of all, to optimize characteristic differences and estimate optimal cluster number, Gaussian kernel function is introduced to improve spectrum analysis method (SAM). By combining subtractive clustering with max-min distance mean, maximum distance method (MDM) is proposed to determine cluster centers. Then, the corresponding steps of improved SAM (ISAM) and MDM are given respectively, whose superiority and stability are illustrated through performing experimental comparisons on gene expression data. Finally, by introducing ISAM and MDM into FKCA, an effective improved FKCA algorithm is proposed. Experimental results from public gene expression data and UCI database show that the proposed algorithms are feasible for cluster analysis, and the clustering accuracy is higher than the other related clustering algorithms.
[Breast cancer genetics. BRCA1 and BRCA2: the main genes for disease predisposition].
Ruiz-Flores, P; Calderón-Garcidueñas, A L; Barrera-Saldaña, H A
2001-01-01
Breast cancer is among the most common world cancers. In Mexico this neoplasm has been progressively increasing since 1990 and is expected to continue. The risk factors for this disease are age, some reproductive factors, ionizing radiation, contraceptives, obesity and high fat diets, among other factors. The main risk factor for BC is a positive family history. Several families, in which clustering but no mendelian inheritance exists, the BC is due probably to mutations in low penetrance genes and/or environmental factors. In families with autosomal dominant trait, the BRCA1 and BRCA2 genes are frequently mutated. These genes are the two main BC susceptibility genes. BRCA1 predispose to BC and ovarian cancer, while BRCA2 mutations predispose to BC in men and women. Both are long genes, tumor suppressors, functioning in a cell cycle dependent manner, and it is believed that both switch on the transcription of several genes, and participate in DNA repair. The mutations profile of these genes is known in developed countries, while in Latin America their search has just began. A multidisciplinary group most be responsible of the clinical management of patients with mutations in BRCA1 and BRCA2, and the risk assignment and Genetic counseling most be done carefully.
Kim, Jun-Mo; Lim, Kyu-Sang; Byun, Mijeong; Lee, Kyung-Tai; Yang, Young-Rok; Park, Mina; Lim, Dajeong; Chai, Han-Ha; Bang, Han-Tae; Hwangbo, Jong; Choi, Yang-Ho; Cho, Yong-Min; Park, Jong-Eun
2017-11-01
White Pekin duck is an important meat resource in the livestock industries. However, the temperature increase due to global warming has become a serious environmental factor in duck production, because of hyperthermia. Therefore, identifying the gene regulations and understanding the molecular mechanism for adaptation to the warmer environment will provide insightful information on the acclimation system of ducks. This study examined transcriptomic responses to heat stress treatments (3 and 6 h at 35 °C) and control (C, 25 °C) using RNA-sequencing analysis of genes from the breast muscle tissue. Based on three distinct differentially expressed gene (DEG) sets (3H/C, 6H/C, and 6H/3H), the expression patterns of significant DEGs (absolute log2 > 1.0 and false discovery rate < 0.05) were clustered into three responsive gene groups divided into upregulated and downregulated genes. Next, we analyzed the clusters that showed relatively higher expression levels in 3H/C and lower levels in 6H/C with much lower or opposite levels in 6H/3H; we referred to these clusters as the adaptable responsive gene group. These genes were significantly enriched in the ErbB signaling pathway, neuroactive ligand-receptor interaction and type II diabetes mellitus in the KEGG pathways (P < 0.01). From the functional enrichment analysis and significantly regulated genes observed in the enriched pathways, we think that the adaptable responsive genes are responsible for the acclimation mechanism of ducks and suggest that the regulation of phosphoinositide 3-kinase genes including PIK3R6, PIK3R5, and PIK3C2B has an important relationship with the mechanisms of adaptation to heat stress in ducks.
Nakaoka, Hirofumi; Tajima, Atsushi; Yoneyama, Taku; Hosomichi, Kazuyoshi; Kasuya, Hidetoshi; Mizutani, Tohru; Inoue, Ituro
2014-08-01
The rupture of intracranial aneurysm (IA) causes subarachnoid hemorrhage associated with high morbidity and mortality. We compared gene expression profiles in aneurysmal domes between unruptured IAs and ruptured IAs (RIAs) to elucidate biological mechanisms predisposing to the rupture of IA. We determined gene expression levels of 8 RIAs, 5 unruptured IAs, and 10 superficial temporal arteries with the Agilent microarrays. To explore biological heterogeneity of IAs, we classified the samples into subgroups showing similar gene expression patterns, using clustering methods. The clustering analysis identified 4 groups: superficial temporal arteries and unruptured IAs were aggregated into their own clusters, whereas RIAs segregated into 2 distinct subgroups (early and late RIAs). Comparing gene expression levels between early RIAs and unruptured IAs, we identified 430 upregulated and 617 downregulated genes in early RIAs. The upregulated genes were associated with inflammatory and immune responses and phagocytosis including S100/calgranulin genes (S100A8, S100A9, and S100A12). The downregulated genes suggest mechanical weakness of aneurysm walls. The expressions of Krüppel-like family of transcription factors (KLF2, KLF12, and KLF15), which were anti-inflammatory regulators, and CDKN2A, which was located on chromosome 9p21 that was the most consistently replicated locus in genome-wide association studies of IA, were also downregulated. We demonstrate that gene expression patterns of RIAs were different according to the age of patients. The results suggest that macrophage-mediated inflammation is a key biological pathway for IA rupture. The identified genes can be good candidates for molecular markers of rupture-prone IAs and therapeutic targets. © 2014 American Heart Association, Inc.
Li, Yaqian; Du, Xilin; Lu, Zhi John; Wu, Daqiang; Zhao, Yilei; Ren, Bin; Huang, Jiaofang; Huang, Xianqing; Xu, Yuhong; Xu, Yuquan
2011-01-01
Background Phenazines are important compounds produced by pseudomonads and other bacteria. Two phz gene clusters called phzA1-G1 and phzA2-G2, respectively, were found in the genome of Pseudomonas sp. M18, an effective biocontrol agent, which is highly homologous to the opportunistic human pathogen P. aeruginosa PAO1, however little is known about the correlation between the expressions of two phz gene clusters. Methodology/Principal Findings Two chromosomal insertion inactivated mutants for the two gene clusters were constructed respectively and the correlation between the expressions of two phz gene clusters was investigated in strain M18. Phenazine-1-carboxylic acid (PCA) molecules produced from phzA2-G2 gene cluster are able to auto-regulate expression itself and activate the expression of phzA1-G1 gene cluster in a circulated amplification pattern. However, the post-transcriptional expression of phzA1-G1 transcript was blocked principally through 5′-untranslated region (UTR). In contrast, the phzA2-G2 gene cluster was transcribed to a lesser extent and translated efficiently and was negatively regulated by the GacA signal transduction pathway, mainly at a post-transcriptional level. Conclusions/Significance A single molecule, PCA, produced in different quantities by the two phz gene clusters acted as the functional mediator and the two phz gene clusters developed a specific regulatory mechanism which acts through 5′-UTR to transfer a single, but complex bacterial signaling event in Pseudomonas sp. strain M18. PMID:21559370
DOE Office of Scientific and Technical Information (OSTI.GOV)
Zhai, Ying; Bai, Silei; Liu, Jingjing
Dithiolopyrrolone group antibiotics characterized by an electronically unique dithiolopyrrolone heterobicyclic core are known for their antibacterial, antifungal, insecticidal and antitumor activities. Recently the biosynthetic gene clusters for two dithiolopyrrolone compounds, holomycin and thiomarinol, have been identified respectively in different bacterial species. Here, we report a novel dithiolopyrrolone biosynthetic gene cluster (aut) isolated from Streptomyces thioluteus DSM 40027 which produces two pyrrothine derivatives, aureothricin and thiolutin. By comparison with other characterized dithiolopyrrolone clusters, eight genes in the aut cluster were verified to be responsible for the assembly of dithiolopyrrolone core. The aut cluster was further confirmed by heterologous expression and in-framemore » gene deletion experiments. Intriguingly, we found that the heterogenetic thioesterase HlmK derived from the holomycin (hlm) gene cluster in Streptomyces clavuligerus significantly improved heterologous biosynthesis of dithiolopyrrolones in Streptomyces albus through coexpression with the aut cluster. In the previous studies, HlmK was considered invalid because it has a Ser to Gly point mutation within the canonical Ser-His-Asp catalytic triad of thioesterases. However, gene inactivation and complementation experiments in our study unequivocally demonstrated that HlmK is an active distinctive type II thioesterase that plays a beneficial role in dithiolopyrrolone biosynthesis. - Highlights: • Cloning of the aureothricin biosynthetic gene cluster from Streptomyces thioluteus DSM 40027. • Identification of the aureothricin gene cluster by heterologous expression and in-frame gene deletion. • The heterogenetic thioesterase HlmK significantly improved dithiolopyrrolones production of the aureothricin gene cluster. • Identification of HlmK as an unusual type II thioesterase.« less
Osborne, Peter W; Benoit, Gérard; Laudet, Vincent; Schubert, Michael; Ferrier, David E K
2009-03-01
The ParaHox cluster is the evolutionary sister to the Hox cluster. Like the Hox cluster, the ParaHox cluster displays spatial and temporal regulation of the component genes along the anterior/posterior axis in a manner that correlates with the gene positions within the cluster (a feature called collinearity). The ParaHox cluster is however a simpler system to study because it is composed of only three genes. We provide a detailed analysis of the amphioxus ParaHox cluster and, for the first time in a single species, examine the regulation of the cluster in response to a single developmental signalling molecule, retinoic acid (RA). Embryos treated with either RA or RA antagonist display altered ParaHox gene expression: AmphiGsx expression shifts in the neural tube, and the endodermal boundary between AmphiXlox and AmphiCdx shifts its anterior/posterior position. We identified several putative retinoic acid response elements and in vitro assays suggest some may participate in RA regulation of the ParaHox genes. By comparison to vertebrate ParaHox gene regulation we explore the evolutionary implications. This work highlights how insights into the regulation and evolution of more complex vertebrate arrangements can be obtained through studies of a simpler, unduplicated amphioxus gene cluster.
A Stationary Wavelet Entropy-Based Clustering Approach Accurately Predicts Gene Expression
Nguyen, Nha; Vo, An; Choi, Inchan
2015-01-01
Abstract Studying epigenetic landscapes is important to understand the condition for gene regulation. Clustering is a useful approach to study epigenetic landscapes by grouping genes based on their epigenetic conditions. However, classical clustering approaches that often use a representative value of the signals in a fixed-sized window do not fully use the information written in the epigenetic landscapes. Clustering approaches to maximize the information of the epigenetic signals are necessary for better understanding gene regulatory environments. For effective clustering of multidimensional epigenetic signals, we developed a method called Dewer, which uses the entropy of stationary wavelet of epigenetic signals inside enriched regions for gene clustering. Interestingly, the gene expression levels were highly correlated with the entropy levels of epigenetic signals. Dewer separates genes better than a window-based approach in the assessment using gene expression and achieved a correlation coefficient above 0.9 without using any training procedure. Our results show that the changes of the epigenetic signals are useful to study gene regulation. PMID:25383910
de Lima-Morales, Daiana; Chaves-Moreno, Diego; Wos-Oxley, Melissa L; Jáuregui, Ruy; Vilchez-Vargas, Ramiro; Pieper, Dietmar H
2016-01-01
Pseudomonas veronii 1YdBTEX2, a benzene and toluene degrader, and Pseudomonas veronii 1YB2, a benzene degrader, have previously been shown to be key players in a benzene-contaminated site. These strains harbor unique catabolic pathways for the degradation of benzene comprising a gene cluster encoding an isopropylbenzene dioxygenase where genes encoding downstream enzymes were interrupted by stop codons. Extradiol dioxygenases were recruited from gene clusters comprising genes encoding a 2-hydroxymuconic semialdehyde dehydrogenase necessary for benzene degradation but typically absent from isopropylbenzene dioxygenase-encoding gene clusters. The benzene dihydrodiol dehydrogenase-encoding gene was not clustered with any other aromatic degradation genes, and the encoded protein was only distantly related to dehydrogenases of aromatic degradation pathways. The involvement of the different gene clusters in the degradation pathways was suggested by real-time quantitative reverse transcription PCR. Copyright © 2015, American Society for Microbiology. All Rights Reserved.
LCGbase: A Comprehensive Database for Lineage-Based Co-regulated Genes.
Wang, Dapeng; Zhang, Yubin; Fan, Zhonghua; Liu, Guiming; Yu, Jun
2012-01-01
Animal genes of different lineages, such as vertebrates and arthropods, are well-organized and blended into dynamic chromosomal structures that represent a primary regulatory mechanism for body development and cellular differentiation. The majority of genes in a genome are actually clustered, which are evolutionarily stable to different extents and biologically meaningful when evaluated among genomes within and across lineages. Until now, many questions concerning gene organization, such as what is the minimal number of genes in a cluster and what is the driving force leading to gene co-regulation, remain to be addressed. Here, we provide a user-friendly database-LCGbase (a comprehensive database for lineage-based co-regulated genes)-hosting information on evolutionary dynamics of gene clustering and ordering within animal kingdoms in two different lineages: vertebrates and arthropods. The database is constructed on a web-based Linux-Apache-MySQL-PHP framework and effective interactive user-inquiry service. Compared to other gene annotation databases with similar purposes, our database has three comprehensible advantages. First, our database is inclusive, including all high-quality genome assemblies of vertebrates and representative arthropod species. Second, it is human-centric since we map all gene clusters from other genomes in an order of lineage-ranks (such as primates, mammals, warm-blooded, and reptiles) onto human genome and start the database from well-defined gene pairs (a minimal cluster where the two adjacent genes are oriented as co-directional, convergent, and divergent pairs) to large gene clusters. Furthermore, users can search for any adjacent genes and their detailed annotations. Third, the database provides flexible parameter definitions, such as the distance of transcription start sites between two adjacent genes, which is extendable to genes that flanking the cluster across species. We also provide useful tools for sequence alignment, gene ontology (GO) annotation, promoter identification, gene expression (co-expression), and evolutionary analysis. This database not only provides a way to define lineage-specific and species-specific gene clusters but also facilitates future studies on gene co-regulation, epigenetic control of gene expression (DNA methylation and histone marks), and chromosomal structures in a context of gene clusters and species evolution. LCGbase is freely available at http://lcgbase.big.ac.cn/LCGbase.
Takahashi, Hiroki; Hotta, Kohji; Takagi, Chiyo; Ueno, Naoto; Satoh, Nori; Shoguchi, Eiichi
2010-02-01
Brachyury, a T-box transcription factor, is expressed in ascidian embryos exclusively in primordial notochord cells and plays a pivotal role in differentiation of notochord cells. Previously, we identified approximately 450 genes downstream of Ciona intestinalis Brachyury (Ci-Bra), and characterized the expression profiles of 45 of these in differentiating notochord cells. In this study, we looked for cisregulatory sequences in minimal enhancers of 20 Ci-Bra downstream genes by electroporating region within approximately 3 kb upstream of each gene fused with lacZ. Eight of the 20 reporters were expressed in notochord cells. The minimal enchancer for each of these eight genes was narrowed to a region approximately 0.5-1.0-kb long. We also explored the genome-wide and coordinate regulation of 43 Ci-Bra-downstream genes. When we determined their chromosomal localization, it became evident that they are not clustered in a given region of the genome, but rather distributed evenly over 13 of the 14 pairs of chromosomes, suggesting that gene clustering does not contribute to coordinate control of the Ci-Bra downstream gene expression. Our results might provide Insights Into the molecular mechanisms underlying notochord formation in chordates.
Romero-Campero, Francisco J; Perez-Hurtado, Ignacio; Lucas-Reina, Eva; Romero, Jose M; Valverde, Federico
2016-03-12
Chlamydomonas reinhardtii is the model organism that serves as a reference for studies in algal genomics and physiology. It is of special interest in the study of the evolution of regulatory pathways from algae to higher plants. Additionally, it has recently gained attention as a potential source for bio-fuel and bio-hydrogen production. The genome of Chlamydomonas is available, facilitating the analysis of its transcriptome by RNA-seq data. This has produced a massive amount of data that remains fragmented making necessary the application of integrative approaches based on molecular systems biology. We constructed a gene co-expression network based on RNA-seq data and developed a web-based tool, ChlamyNET, for the exploration of the Chlamydomonas transcriptome. ChlamyNET exhibits a scale-free and small world topology. Applying clustering techniques, we identified nine gene clusters that capture the structure of the transcriptome under the analyzed conditions. One of the most central clusters was shown to be involved in carbon/nitrogen metabolism and signalling, whereas one of the most peripheral clusters was involved in DNA replication and cell cycle regulation. The transcription factors and regulators in the Chlamydomonas genome have been identified in ChlamyNET. The biological processes potentially regulated by them as well as their putative transcription factor binding sites were determined. The putative light regulated transcription factors and regulators in the Chlamydomonas genome were analyzed in order to provide a case study on the use of ChlamyNET. Finally, we used an independent data set to cross-validate the predictive power of ChlamyNET. The topological properties of ChlamyNET suggest that the Chlamydomonas transcriptome posseses important characteristics related to error tolerance, vulnerability and information propagation. The central part of ChlamyNET constitutes the core of the transcriptome where most authoritative hub genes are located interconnecting key biological processes such as light response with carbon and nitrogen metabolism. Our study reveals that key elements in the regulation of carbon and nitrogen metabolism, light response and cell cycle identified in higher plants were already established in Chlamydomonas. These conserved elements are not only limited to transcription factors, regulators and their targets, but also include the cis-regulatory elements recognized by them.
Pepke, Shirley; Ver Steeg, Greg
2017-03-15
De novo inference of clinically relevant gene function relationships from tumor RNA-seq remains a challenging task. Current methods typically either partition patient samples into a few subtypes or rely upon analysis of pairwise gene correlations that will miss some groups in noisy data. Leveraging higher dimensional information can be expected to increase the power to discern targetable pathways, but this is commonly thought to be an intractable computational problem. In this work we adapt a recently developed machine learning algorithm for sensitive detection of complex gene relationships. The algorithm, CorEx, efficiently optimizes over multivariate mutual information and can be iteratively applied to generate a hierarchy of relatively independent latent factors. The learned latent factors are used to stratify patients for survival analysis with respect to both single factors and combinations. These analyses are performed and interpreted in the context of biological function annotations and protein network interactions that might be utilized to match patients to multiple therapies. Analysis of ovarian tumor RNA-seq samples demonstrates the algorithm's power to infer well over one hundred biologically interpretable gene cohorts, several times more than standard methods such as hierarchical clustering and k-means. The CorEx factor hierarchy is also informative, with related but distinct gene clusters grouped by upper nodes. Some latent factors correlate with patient survival, including one for a pathway connected with the epithelial-mesenchymal transition in breast cancer that is regulated by a microRNA that modulates epigenetics. Further, combinations of factors lead to a synergistic survival advantage in some cases. In contrast to studies that attempt to partition patients into a small number of subtypes (typically 4 or fewer) for treatment purposes, our approach utilizes subgroup information for combinatoric transcriptional phenotyping. Considering only the 66 gene expression groups that are found to both have significant Gene Ontology enrichment and are small enough to indicate specific drug targets implies a computational phenotype for ovarian cancer that allows for 3 66 possible patient profiles, enabling truly personalized treatment. The findings here demonstrate a new technique that sheds light on the complexity of gene expression dependencies in tumors and could eventually enable the use of patient RNA-seq profiles for selection of personalized and effective cancer treatments.
Wasito, Ito; Hashim, Siti Zaiton M; Sukmaningrum, Sri
2007-01-01
Gene expression profiling plays an important role in the identification of biological and clinical properties of human solid tumors such as colorectal carcinoma. Profiling is required to reveal underlying molecular features for diagnostic and therapeutic purposes. A non-parametric density-estimation-based approach called iterative local Gaussian clustering (ILGC), was used to identify clusters of expressed genes. We used experimental data from a previous study by Muro and others consisting of 1,536 genes in 100 colorectal cancer and 11 normal tissues. In this dataset, the ILGC finds three clusters, two large and one small gene clusters, similar to their results which used Gaussian mixture clustering. The correlation of each cluster of genes and clinical properties of malignancy of human colorectal cancer was analysed for the existence of tumor or normal, the existence of distant metastasis and the existence of lymph node metastasis. PMID:18305825
Wasito, Ito; Hashim, Siti Zaiton M; Sukmaningrum, Sri
2007-12-30
Gene expression profiling plays an important role in the identification of biological and clinical properties of human solid tumors such as colorectal carcinoma. Profiling is required to reveal underlying molecular features for diagnostic and therapeutic purposes. A non-parametric density-estimation-based approach called iterative local Gaussian clustering (ILGC), was used to identify clusters of expressed genes. We used experimental data from a previous study by Muro and others consisting of 1,536 genes in 100 colorectal cancer and 11 normal tissues. In this dataset, the ILGC finds three clusters, two large and one small gene clusters, similar to their results which used Gaussian mixture clustering. The correlation of each cluster of genes and clinical properties of malignancy of human colorectal cancer was analysed for the existence of tumor or normal, the existence of distant metastasis and the existence of lymph node metastasis.
A cluster merging method for time series microarray with production values.
Chira, Camelia; Sedano, Javier; Camara, Monica; Prieto, Carlos; Villar, Jose R; Corchado, Emilio
2014-09-01
A challenging task in time-course microarray data analysis is to cluster genes meaningfully combining the information provided by multiple replicates covering the same key time points. This paper proposes a novel cluster merging method to accomplish this goal obtaining groups with highly correlated genes. The main idea behind the proposed method is to generate a clustering starting from groups created based on individual temporal series (representing different biological replicates measured in the same time points) and merging them by taking into account the frequency by which two genes are assembled together in each clustering. The gene groups at the level of individual time series are generated using several shape-based clustering methods. This study is focused on a real-world time series microarray task with the aim to find co-expressed genes related to the production and growth of a certain bacteria. The shape-based clustering methods used at the level of individual time series rely on identifying similar gene expression patterns over time which, in some models, are further matched to the pattern of production/growth. The proposed cluster merging method is able to produce meaningful gene groups which can be naturally ranked by the level of agreement on the clustering among individual time series. The list of clusters and genes is further sorted based on the information correlation coefficient and new problem-specific relevant measures. Computational experiments and results of the cluster merging method are analyzed from a biological perspective and further compared with the clustering generated based on the mean value of time series and the same shape-based algorithm.
Yu, Jiujiang
2012-10-25
Traditional molecular techniques have been used in research in discovering the genes and enzymes that are involved in aflatoxin formation and genetic regulation. We cloned most, if not all, of the aflatoxin pathway genes. A consensus gene cluster for aflatoxin biosynthesis was discovered in 2005. The factors that affect aflatoxin formation have been studied. In this report, the author summarized the current status of research progress and future possibilities that may be used for solving aflatoxin contamination.
Yu, Jiujiang
2012-01-01
Traditional molecular techniques have been used in research in discovering the genes and enzymes that are involved in aflatoxin formation and genetic regulation. We cloned most, if not all, of the aflatoxin pathway genes. A consensus gene cluster for aflatoxin biosynthesis was discovered in 2005. The factors that affect aflatoxin formation have been studied. In this report, the author summarized the current status of research progress and future possibilities that may be used for solving aflatoxin contamination. PMID:23202305
Cherednichenko, A A; Trifonova, E A; Vagaitseva, K V; Bocharova, A V; Varzari, A M; Radzhabov, M O; Stepanov, V A
2015-01-01
The data on distribution of genetic diversity in gene polymorphisms associated with autoimmune and allergic diseases and with regulation of immunoglobulin E and cytokines levels in 26 populations of the Northern Eurasia is presented. Substantial correlation between the values of average expected heterozygosity by 44 gene polymorphisms with climatic and geographical factors has not been revealed. Clustering of population groups in correspondence with their geographic locations is observed. The degree of gene differentiation among populations and the selective neutrality of gene polymorphisms have been assessed. The results of our work evidence the substantial genetic diversity and differentiation of human populations by studied genes.
Liu, Ying; Navathe, Shamkant B; Pivoshenko, Alex; Dasigi, Venu G; Dingledine, Ray; Ciliax, Brian J
2006-01-01
One of the key challenges of microarray studies is to derive biological insights from the gene-expression patterns. Clustering genes by functional keyword association can provide direct information about the functional links among genes. However, the quality of the keyword lists significantly affects the clustering results. We compared two keyword weighting schemes: normalised z-score and term frequency-inverse document frequency (TFIDF). Two gene sets were tested to evaluate the effectiveness of the weighting schemes for keyword extraction for gene clustering. Using established measures of cluster quality, the results produced from TFIDF-weighted keywords outperformed those produced from normalised z-score weighted keywords. The optimised algorithms should be useful for partitioning genes from microarray lists into functionally discrete clusters.
Zhang, Xiujun; Alemany, Lawrence B.; Fiedler, Hans-Peter; Goodfellow, Michael; Parry, Ronald J.
2008-01-01
The antibiotics lactonamycin and lactonamycin Z provide attractive leads for antibacterial drug development. Both antibiotics contain a novel aglycone core called lactonamycinone. To gain insight into lactonamycinone biosynthesis, cloning and precursor incorporation experiments were undertaken. The lactonamycin gene cluster was initially cloned from Streptomyces rishiriensis. Sequencing of ca. 61 kb of S. rishiriensis DNA revealed the presence of 57 open reading frames. These included genes coding for the biosynthesis of l-rhodinose, the sugar found in lactonamycin, and genes similar to those in the tetracenomycin biosynthetic gene cluster. Since lactonamycin production by S. rishiriensis could not be sustained, additional proof for the identity of the S. rishiriensis cluster was obtained by cloning the lactonamycin Z gene cluster from Streptomyces sanglieri. Partial sequencing of the S. sanglieri cluster revealed 15 genes that exhibited a very high degree of similarity to genes within the lactonamycin cluster, as well as an identical organization. Double-crossover disruption of one gene in the S. sanglieri cluster abolished lactonamycin Z production, and production was restored by complementation. These results confirm the identity of the genetic locus cloned from S. sanglieri and indicate that the highly similar locus in S. rishiriensis encodes lactonamycin biosynthetic genes. Precursor incorporation experiments with S. sanglieri revealed that lactonamycinone is biosynthesized in an unusual manner whereby glycine or a glycine derivative serves as a starter unit that is extended by nine acetate units. Analysis of the gene clusters and of the precursor incorporation data suggested a hypothetical scheme for lactonamycinone biosynthesis. PMID:18070976
Bowen, Lizabeth; Miles, A. Keith; Ballachey, Brenda E.; Waters, Shannon C.; Bodkin, James L.
2016-01-01
Using a panel of genes stimulated by oil exposure in a laboratory study, we evaluated gene transcription in blood leukocytes sampled from sea otters captured from 2006–2012 in western Prince William Sound (WPWS), Alaska, 17–23 years after the 1989 Exxon Valdez oil spill (EVOS). We compared WPWS sea otters to reference populations (not affected by the EVOS) from the Alaska Peninsula (2009), Katmai National Park and Preserve (2009), Clam Lagoon at Adak Island (2012), Kodiak Island (2005) and captive sea otters in aquaria. Statistically, sea otter gene transcript profiles separated into three distinct clusters: Cluster 1, Kodiak and WPWS 2006–2008 (higher relative transcription); Cluster 2, Clam Lagoon and WPWS 2010–2012 (lower relative transcription); and Cluster 3, Alaska Peninsula, Katmai and captive sea otters (intermediate relative transcription). The lower transcription of the aryl hydrocarbon receptor (AHR), an established biomarker for hydrocarbon exposure, in WPWS 2010–2012 compared to earlier samples from WPWS is consistent with declining hydrocarbon exposure, but the pattern of overall low levels of transcription seen in WPWS 2010–2012 could be related to other factors, such as food limitation, pathogens or injury, and may indicate an inability to mount effective responses to stressors. Decreased transcriptional response across the entire gene panel precludes the evaluation of whether or not individual sea otters show signs of exposure to lingering oil. However, related studies on sea otter demographics indicate that by 2012, the sea otter population in WPWS had recovered, which indicates diminishing oil exposure.
Differential expression of anti-angiogenic factors and guidance genes in the developing macula
Kozulin, Peter; Natoli, Riccardo; O’Brien, Keely M. Bumsted; Madigan, Michele C.
2009-01-01
Purpose The primate retina contains a specialized, cone-rich macula, which mediates high acuity and color vision. The spatial resolution provided by the neural retina at the macula is optimized by stereotyped retinal blood vessel and ganglion cell axon patterning, which radiate away from the macula and reduce shadowing of macular photoreceptors. However, the genes that mediate these specializations, and the reasons for the vulnerability of the macula to degenerative disease, remain obscure. The aim of this study was to identify novel genes that may influence retinal vascular patterning and definition of the foveal avascular area. Methods We used RNA from human fetal retinas at 19–20 weeks of gestation (WG; n=4) to measure differential gene expression in the macula, a region nasal to disc (nasal) and in the surrounding retina (surround) by hybridization to 12 GeneChip® microarrays (HG-U133 Plus 2.0). The raw data was subjected to quality control assessment and preprocessing, using GC-RMA. We then used ANOVA analysis (Partek® Genomic Suite™ 6.3) and clustering (DAVID website) to identify the most highly represented genes clustered according to “biological process.” The neural retina is fully differentiated at the macula at 19–20 WG, while neuronal progenitor cells are present throughout the rest of the retina. We therefore excluded genes associated with the cell cycle, and markers of differentiated neurons, from further analyses. Significantly regulated genes (p<0.01) were then identified in a second round of clustering according to molecular/reaction (KEGG) pathway. Genes of interest were verified by quantitative PCR (QRT–PCR), and 2 genes were localized by in situ hybridization. Results We generated two lists of differentially regulated genes: “macula versus surround” and “macula versus nasal.” KEGG pathway clustering of the filtered gene lists identified 25 axon guidance-related genes that are differentially regulated in the macula. Furthermore, we found significant upregulation of three anti-angiogenic factors in the macula: pigment epithelium derived factor (PEDF), natriuretic peptide precurusor B (NPPB), and collagen type IVα2. Differential expression of several members of the ephrin and semaphorin axon guidance gene families, PEDF, and NPPB was verified by QRT–PCR. Localization of PEDF and Eph-A6 mRNAs in sections of macaque retina shows expression of both genes concentrates in the ganglion cell layer (GCL) at the developing fovea, consistent with an involvement in definition of the foveal avascular area. Conclusions Because the axons of macular ganglion cells exit the retina from around 8 WG, we suggest that the axon guidance genes highly expressed at the macula at 19–20 WG are also involved in vascular patterning, along with PEDF and NPPB. Localization of both PEDF and Eph-A6 mRNAs to the GCL of the developing fovea supports this idea. It is possible that specialization of the macular vessels, including definition of the foveal avascular area, is mediated by processes that piggyback on axon guidance mechanisms in effect earlier in development. These findings may be useful to understand the vulnerability of the macula to degeneration and to develop new therapeutic strategies to inhibit neovascularization. PMID:19145251
Aguiar, Bruno; Vieira, Jorge; Cunha, Ana E; Fonseca, Nuno A; Reboiro-Jato, David; Reboiro-Jato, Miguel; Fdez-Riverola, Florentino; Raspé, Olivier; Vieira, Cristina P
2013-05-01
S-RNase-based gametophytic self-incompatibility evolved once before the split of the Asteridae and Rosidae. In Prunus (tribe Amygdaloideae of Rosaceae), the self-incompatibility S-pollen is a single F-box gene that presents the expected evolutionary signatures. In Malus and Pyrus (subtribe Pyrinae of Rosaceae), however, clusters of F-box genes (called SFBBs) have been described that are expressed in pollen only and are linked to the S-RNase gene. Although polymorphic, SFBB genes present levels of diversity lower than those of the S-RNase gene. They have been suggested as putative S-pollen genes, in a system of non-self recognition by multiple factors. Subsets of allelic products of the different SFBB genes interact with non-self S-RNases, marking them for degradation, and allowing compatible pollinations. This study performed a detailed characterization of SFBB genes in Sorbus aucuparia (Pyrinae) to address three predictions of the non-self recognition by multiple factors model. As predicted, the number of SFBB genes was large to account for the many S-RNase specificities. Secondly, like the S-RNase gene, the SFBB genes were old. Thirdly, amino acids under positive selection-those that could be involved in specificity determination-were identified when intra-haplotype SFBB genes were analysed using codon models. Overall, the findings reported here support the non-self recognition by multiple factors model.
Computational gene expression profiling under salt stress reveals patterns of co-expression
Sanchita; Sharma, Ashok
2016-01-01
Plants respond differently to environmental conditions. Among various abiotic stresses, salt stress is a condition where excess salt in soil causes inhibition of plant growth. To understand the response of plants to the stress conditions, identification of the responsible genes is required. Clustering is a data mining technique used to group the genes with similar expression. The genes of a cluster show similar expression and function. We applied clustering algorithms on gene expression data of Solanum tuberosum showing differential expression in Capsicum annuum under salt stress. The clusters, which were common in multiple algorithms were taken further for analysis. Principal component analysis (PCA) further validated the findings of other cluster algorithms by visualizing their clusters in three-dimensional space. Functional annotation results revealed that most of the genes were involved in stress related responses. Our findings suggest that these algorithms may be helpful in the prediction of the function of co-expressed genes. PMID:26981411
Mihali, Troco K; Kellmann, Ralf; Neilan, Brett A
2009-03-30
Saxitoxin and its analogues collectively known as the paralytic shellfish toxins (PSTs) are neurotoxic alkaloids and are the cause of the syndrome named paralytic shellfish poisoning. PSTs are produced by a unique biosynthetic pathway, which involves reactions that are rare in microbial metabolic pathways. Nevertheless, distantly related organisms such as dinoflagellates and cyanobacteria appear to produce these toxins using the same pathway. Hypothesised explanations for such an unusual phylogenetic distribution of this shared uncommon metabolic pathway, include a polyphyletic origin, an involvement of symbiotic bacteria, and horizontal gene transfer. We describe the identification, annotation and bioinformatic characterisation of the putative paralytic shellfish toxin biosynthesis clusters in an Australian isolate of Anabaena circinalis and an American isolate of Aphanizomenon sp., both members of the Nostocales. These putative PST gene clusters span approximately 28 kb and contain genes coding for the biosynthesis and export of the toxin. A putative insertion/excision site in the Australian Anabaena circinalis AWQC131C was identified, and the organization and evolution of the gene clusters are discussed. A biosynthetic pathway leading to the formation of saxitoxin and its analogues in these organisms is proposed. The PST biosynthesis gene cluster presents a mosaic structure, whereby genes have apparently transposed in segments of varying size, resulting in different gene arrangements in all three sxt clusters sequenced so far. The gene cluster organizational structure and sequence similarity seems to reflect the phylogeny of the producer organisms, indicating that the gene clusters have an ancient origin, or that their lateral transfer was also an ancient event. The knowledge we gain from the characterisation of the PST biosynthesis gene clusters, including the identity and sequence of the genes involved in the biosynthesis, may also afford the identification of these gene clusters in dinoflagellates, the cause of human mortalities and significant financial loss to the tourism and shellfish industries.
Mihali, Troco K; Kellmann, Ralf; Neilan, Brett A
2009-01-01
Background Saxitoxin and its analogues collectively known as the paralytic shellfish toxins (PSTs) are neurotoxic alkaloids and are the cause of the syndrome named paralytic shellfish poisoning. PSTs are produced by a unique biosynthetic pathway, which involves reactions that are rare in microbial metabolic pathways. Nevertheless, distantly related organisms such as dinoflagellates and cyanobacteria appear to produce these toxins using the same pathway. Hypothesised explanations for such an unusual phylogenetic distribution of this shared uncommon metabolic pathway, include a polyphyletic origin, an involvement of symbiotic bacteria, and horizontal gene transfer. Results We describe the identification, annotation and bioinformatic characterisation of the putative paralytic shellfish toxin biosynthesis clusters in an Australian isolate of Anabaena circinalis and an American isolate of Aphanizomenon sp., both members of the Nostocales. These putative PST gene clusters span approximately 28 kb and contain genes coding for the biosynthesis and export of the toxin. A putative insertion/excision site in the Australian Anabaena circinalis AWQC131C was identified, and the organization and evolution of the gene clusters are discussed. A biosynthetic pathway leading to the formation of saxitoxin and its analogues in these organisms is proposed. Conclusion The PST biosynthesis gene cluster presents a mosaic structure, whereby genes have apparently transposed in segments of varying size, resulting in different gene arrangements in all three sxt clusters sequenced so far. The gene cluster organizational structure and sequence similarity seems to reflect the phylogeny of the producer organisms, indicating that the gene clusters have an ancient origin, or that their lateral transfer was also an ancient event. The knowledge we gain from the characterisation of the PST biosynthesis gene clusters, including the identity and sequence of the genes involved in the biosynthesis, may also afford the identification of these gene clusters in dinoflagellates, the cause of human mortalities and significant financial loss to the tourism and shellfish industries. PMID:19331657
Mediator and RNA polymerase II clusters associate in transcription-dependent condensates.
Cho, Won-Ki; Spille, Jan-Hendrik; Hecht, Micca; Lee, Choongman; Li, Charles; Grube, Valentin; Cisse, Ibrahim I
2018-06-21
Models of gene control have emerged from genetic and biochemical studies, with limited consideration of the spatial organization and dynamics of key components in living cells. Here we used live cell super-resolution and light sheet imaging to study the organization and dynamics of the Mediator coactivator and RNA polymerase II (Pol II) directly. Mediator and Pol II each form small transient and large stable clusters in living embryonic stem cells. Mediator and Pol II are colocalized in the stable clusters, which associate with chromatin, have properties of phase-separated condensates, and are sensitive to transcriptional inhibitors. We suggest that large clusters of Mediator, recruited by transcription factors at large or clustered enhancer elements, interact with large Pol II clusters in transcriptional condensates in vivo. Copyright © 2018, American Association for the Advancement of Science.
Throckmorton, Kurt; Wiemann, Philipp; Keller, Nancy P.
2015-01-01
Fungal polyketides are a diverse class of natural products, or secondary metabolites (SMs), with a wide range of bioactivities often associated with toxicity. Here, we focus on a group of non-reducing polyketide synthases (NR-PKSs) in the fungal phylum Ascomycota that lack a thioesterase domain for product release, group V. Although widespread in ascomycete taxa, this group of NR-PKSs is notably absent in the mycotoxigenic genus Fusarium and, surprisingly, found in genera not known for their secondary metabolite production (e.g., the mycorrhizal genus Oidiodendron, the powdery mildew genus Blumeria, and the causative agent of white-nose syndrome in bats, Pseudogymnoascus destructans). This group of NR-PKSs, in association with the other enzymes encoded by their gene clusters, produces a variety of different chemical classes including naphthacenediones, anthraquinones, benzophenones, grisandienes, and diphenyl ethers. We discuss the modification of and transitions between these chemical classes, the requisite enzymes, and the evolution of the SM gene clusters that encode them. Integrating this information, we predict the likely products of related but uncharacterized SM clusters, and we speculate upon the utility of these classes of SMs as virulence factors or chemical defenses to various plant, animal, and insect pathogens, as well as mutualistic fungi. PMID:26378577
Transcriptional organization of the DNA region controlling expression of the K99 gene cluster.
Roosendaal, B; Damoiseaux, J; Jordi, W; de Graaf, F K
1989-01-01
The transcriptional organization of the K99 gene cluster was investigated in two ways. First, the DNA region, containing the transcriptional signals was analyzed using a transcription vector system with Escherichia coli galactokinase (GalK) as assayable marker and second, an in vitro transcription system was employed. A detailed analysis of the transcription signals revealed that a strong promoter PA and a moderate promoter PB are located upstream of fanA and fanB, respectively. No promoter activity was detected in the intercistronic region between fanB and fanC. Factor-dependent terminators of transcription were detected and are probably located in the intercistronic region between fanA and fanB (T1), and between fanB and fanC (T2). A third terminator (T3) was observed between fanC and fanD and has an efficiency of 90%. Analysis of the regulatory region in an in vitro transcription system confirmed the location of the respective transcription signals. A model for the transcriptional organization of the K99 cluster is presented. Indications were obtained that the trans-acting regulatory polypeptides FanA and FanB both function as anti-terminators. A model for the regulation of expression of the K99 gene cluster is postulated.
Mining subspace clusters from DNA microarray data using large itemset techniques.
Chang, Ye-In; Chen, Jiun-Rung; Tsai, Yueh-Chi
2009-05-01
Mining subspace clusters from the DNA microarrays could help researchers identify those genes which commonly contribute to a disease, where a subspace cluster indicates a subset of genes whose expression levels are similar under a subset of conditions. Since in a DNA microarray, the number of genes is far larger than the number of conditions, those previous proposed algorithms which compute the maximum dimension sets (MDSs) for any two genes will take a long time to mine subspace clusters. In this article, we propose the Large Itemset-Based Clustering (LISC) algorithm for mining subspace clusters. Instead of constructing MDSs for any two genes, we construct only MDSs for any two conditions. Then, we transform the task of finding the maximal possible gene sets into the problem of mining large itemsets from the condition-pair MDSs. Since we are only interested in those subspace clusters with gene sets as large as possible, it is desirable to pay attention to those gene sets which have reasonable large support values in the condition-pair MDSs. From our simulation results, we show that the proposed algorithm needs shorter processing time than those previous proposed algorithms which need to construct gene-pair MDSs.
Patel, Vidushi S; Cooper, Steven JB; Deakin, Janine E; Fulton, Bob; Graves, Tina; Warren, Wesley C; Wilson, Richard K; Graves, Jennifer AM
2008-01-01
Background Vertebrate alpha (α)- and beta (β)-globin gene families exemplify the way in which genomes evolve to produce functional complexity. From tandem duplication of a single globin locus, the α- and β-globin clusters expanded, and then were separated onto different chromosomes. The previous finding of a fossil β-globin gene (ω) in the marsupial α-cluster, however, suggested that duplication of the α-β cluster onto two chromosomes, followed by lineage-specific gene loss and duplication, produced paralogous α- and β-globin clusters in birds and mammals. Here we analyse genomic data from an egg-laying monotreme mammal, the platypus (Ornithorhynchus anatinus), to explore haemoglobin evolution at the stem of the mammalian radiation. Results The platypus α-globin cluster (chromosome 21) contains embryonic and adult α- globin genes, a β-like ω-globin gene, and the GBY globin gene with homology to cytoglobin, arranged as 5'-ζ-ζ'-αD-α3-α2-α1-ω-GBY-3'. The platypus β-globin cluster (chromosome 2) contains single embryonic and adult globin genes arranged as 5'-ε-β-3'. Surprisingly, all of these globin genes were expressed in some adult tissues. Comparison of flanking sequences revealed that all jawed vertebrate α-globin clusters are flanked by MPG-C16orf35 and LUC7L, whereas all bird and mammal β-globin clusters are embedded in olfactory genes. Thus, the mammalian α- and β-globin clusters are orthologous to the bird α- and β-globin clusters respectively. Conclusion We propose that α- and β-globin clusters evolved from an ancient MPG-C16orf35-α-β-GBY-LUC7L arrangement 410 million years ago. A copy of the original β (represented by ω in marsupials and monotremes) was inserted into an array of olfactory genes before the amniote radiation (>315 million years ago), then duplicated and diverged to form orthologous clusters of β-globin genes with different expression profiles in different lineages. PMID:18657265
Identification of an Imprinted Gene Cluster in the X-Inactivation Center
Kobayashi, Shin; Totoki, Yasushi; Soma, Miki; Matsumoto, Kazuya; Fujihara, Yoshitaka; Toyoda, Atsushi; Sakaki, Yoshiyuki; Okabe, Masaru; Ishino, Fumitoshi
2013-01-01
Mammalian development is strongly influenced by the epigenetic phenomenon called genomic imprinting, in which either the paternal or the maternal allele of imprinted genes is expressed. Paternally expressed Xist, an imprinted gene, has been considered as a single cis-acting factor to inactivate the paternally inherited X chromosome (Xp) in preimplantation mouse embryos. This means that X-chromosome inactivation also entails gene imprinting at a very early developmental stage. However, the precise mechanism of imprinted X-chromosome inactivation remains unknown and there is little information about imprinted genes on X chromosomes. In this study, we examined whether there are other imprinted genes than Xist expressed from the inactive paternal X chromosome and expressed in female embryos at the preimplantation stage. We focused on small RNAs and compared their expression patterns between sexes by tagging the female X chromosome with green fluorescent protein. As a result, we identified two micro (mi)RNAs–miR-374-5p and miR-421-3p–mapped adjacent to Xist that were predominantly expressed in female blastocysts. Allelic expression analysis revealed that these miRNAs were indeed imprinted and expressed from the Xp. Further analysis of the imprinting status of adjacent locus led to the discovery of a large cluster of imprinted genes expressed from the Xp: Jpx, Ftx and Zcchc13. To our knowledge, this is the first identified cluster of imprinted genes in the cis-acting regulatory region termed the X-inactivation center. This finding may help in understanding the molecular mechanisms regulating imprinted X-chromosome inactivation during early mammalian development. PMID:23940725
Identification of an imprinted gene cluster in the X-inactivation center.
Kobayashi, Shin; Totoki, Yasushi; Soma, Miki; Matsumoto, Kazuya; Fujihara, Yoshitaka; Toyoda, Atsushi; Sakaki, Yoshiyuki; Okabe, Masaru; Ishino, Fumitoshi
2013-01-01
Mammalian development is strongly influenced by the epigenetic phenomenon called genomic imprinting, in which either the paternal or the maternal allele of imprinted genes is expressed. Paternally expressed Xist, an imprinted gene, has been considered as a single cis-acting factor to inactivate the paternally inherited X chromosome (Xp) in preimplantation mouse embryos. This means that X-chromosome inactivation also entails gene imprinting at a very early developmental stage. However, the precise mechanism of imprinted X-chromosome inactivation remains unknown and there is little information about imprinted genes on X chromosomes. In this study, we examined whether there are other imprinted genes than Xist expressed from the inactive paternal X chromosome and expressed in female embryos at the preimplantation stage. We focused on small RNAs and compared their expression patterns between sexes by tagging the female X chromosome with green fluorescent protein. As a result, we identified two micro (mi)RNAs-miR-374-5p and miR-421-3p-mapped adjacent to Xist that were predominantly expressed in female blastocysts. Allelic expression analysis revealed that these miRNAs were indeed imprinted and expressed from the Xp. Further analysis of the imprinting status of adjacent locus led to the discovery of a large cluster of imprinted genes expressed from the Xp: Jpx, Ftx and Zcchc13. To our knowledge, this is the first identified cluster of imprinted genes in the cis-acting regulatory region termed the X-inactivation center. This finding may help in understanding the molecular mechanisms regulating imprinted X-chromosome inactivation during early mammalian development.
Addlesee, Hugh A.; Fiedor, Leszek; Hunter, C. Neil
2000-01-01
The purple photosynthetic bacterium Rhodobacter sphaeroides has within its genome a cluster of photosynthesis-related genes approximately 41 kb in length. In an attempt to identify genes involved in the terminal esterification stage of bacteriochlorophyll biosynthesis, a previously uncharacterized 5-kb region of this cluster was sequenced. Four open reading frames (ORFs) were identified, and each was analyzed by transposon mutagenesis. The product of one of these ORFs, bchG, shows close homologies with (bacterio)chlorophyll synthetases, and mutants in this gene were found to accumulate bacteriopheophorbide, the metal-free derivative of the bacteriochlorophyll precursor bacteriochlorophyllide, suggesting that bchG is responsible for the esterification of bacteriochlorophyllide with an alcohol moiety. This assignment of function to bchG was verified by the performance of assays demonstrating the ability of BchG protein, heterologously synthesized in Escherichia coli, to esterify bacteriochlorophyllide with geranylgeranyl pyrophosphate in vitro, thereby generating bacteriochlorophyll. This step is pivotal to the assembly of a functional photosystem in R. sphaeroides, a model organism for the study of structure-function relationships in photosynthesis. A second gene, orf177, is a member of a large family of isopentenyl diphosphate isomerases, while sequence homologies suggest that a third gene, orf427, may encode an assembly factor for photosynthetic complexes. The function of the remaining ORF, bchP, is the subject of a separate paper (H. Addlesee and C. N. Hunter, J. Bacteriol. 181:7248–7255, 1999). An operonal arrangement of the genes is proposed. PMID:10809697
Båge, Tove; Lagervall, Maria; Jansson, Leif; Lundeberg, Joakim; Yucel-Lindberg, Tülay
2012-01-01
Periodontitis is a chronic inflammatory disease affecting the soft tissue and bone that surrounds the teeth. Despite extensive research, distinctive genes responsible for the disease have not been identified. The objective of this study was to elucidate transcriptome changes in periodontitis, by investigating gene expression profiles in gingival tissue obtained from periodontitis-affected and healthy gingiva from the same patient, using RNA-sequencing. Gingival biopsies were obtained from a disease-affected and a healthy site from each of 10 individuals diagnosed with periodontitis. Enrichment analysis performed among uniquely expressed genes for the periodontitis-affected and healthy tissues revealed several regulated pathways indicative of inflammation for the periodontitis-affected condition. Hierarchical clustering of the sequenced biopsies demonstrated clustering according to the degree of inflammation, as observed histologically in the biopsies, rather than clustering at the individual level. Among the top 50 upregulated genes in periodontitis-affected tissues, we investigated two genes which have not previously been demonstrated to be involved in periodontitis. These included interferon regulatory factor 4 and chemokine (C-C motif) ligand 18, which were also expressed at the protein level in gingival biopsies from patients with periodontitis. In conclusion, this study provides a first step towards a quantitative comprehensive insight into the transcriptome changes in periodontitis. We demonstrate for the first time site-specific local variation in gene expression profiles of periodontitis-affected and healthy tissues obtained from patients with periodontitis, using RNA-seq. Further, we have identified novel genes expressed in periodontitis tissues, which may constitute potential therapeutic targets for future treatment strategies of periodontitis. PMID:23029519
The transcriptome of a complete episode of acute otitis media.
Hernandez, Michelle; Leichtle, Anke; Pak, Kwang; Webster, Nicholas J; Wasserman, Stephen I; Ryan, Allen F
2015-04-03
Otitis media is the most common disease of childhood, and represents an important health challenge to the 10-15% of children who experience chronic/recurrent middle ear infections. The middle ear undergoes extensive modifications during otitis media, potentially involving changes in the expression of many genes. Expression profiling offers an opportunity to discover novel genes and pathways involved in this common childhood disease. The middle ears of 320 WBxB6 F1 hybrid mice were inoculated with non-typeable Haemophilus influenzae (NTHi) or PBS (sham control). Two independent samples were generated for each time point and condition, from initiation of infection to resolution. RNA was profiled on Affymetrix mouse 430 2.0 whole-genome microarrays. Approximately 8% of the sampled transcripts defined the signature of acute NTHi-induced otitis media across time. Hierarchical clustering of signal intensities revealed several temporal gene clusters. Network and pathway enrichment analysis of these clusters identified sets of genes involved in activation of the innate immune response, negative regulation of immune response, changes in epithelial and stromal cell markers, and the recruitment/function of neutrophils and macrophages. We also identified key transcriptional regulators related to events in otitis media, which likely determine the expression of these gene clusters. A list of otitis media susceptibility genes, derived from genome-wide association and candidate gene studies, was significantly enriched during the early induction phase and the middle re-modeling phase of otitis but not in the resolution phase. Our results further indicate that positive versus negative regulation of inflammatory processes occur with highly similar kinetics during otitis media, underscoring the importance of anti-inflammatory responses in controlling pathogenesis. The results characterize the global gene response during otitis media and identify key signaling and transcription factor networks that control the defense of the middle ear against infection. These networks deserve further attention, as dysregulated immune defense and inflammatory responses may contribute to recurrent or chronic otitis in children.
Environmental adaptability and stress tolerance of Laribacter hongkongensis: a genome-wide analysis
2011-01-01
Background Laribacter hongkongensis is associated with community-acquired gastroenteritis and traveler's diarrhea and it can reside in human, fish, frogs and water. In this study, we performed an in-depth annotation of the genes in its genome related to adaptation to the various environmental niches. Results L. hongkongensis possessed genes for DNA repair and recombination, basal transcription, alternative σ-factors and 109 putative transcription factors, allowing DNA repair and global changes in gene expression in response to different environmental stresses. For acid stress, it possessed a urease gene cassette and two arc gene clusters. For alkaline stress, it possessed six CDSs for transporters of the monovalent cation/proton antiporter-2 and NhaC Na+:H+ antiporter families. For heavy metals acquisition and tolerance, it possessed CDSs for iron and nickel transport and efflux pumps for other metals. For temperature stress, it possessed genes related to chaperones and chaperonins, heat shock proteins and cold shock proteins. For osmotic stress, 25 CDSs were observed, mostly related to regulators for potassium ion, proline and glutamate transport. For oxidative and UV light stress, genes for oxidant-resistant dehydratase, superoxide scavenging, hydrogen peroxide scavenging, exclusion and export of redox-cycling antibiotics, redox balancing, DNA repair, reduction of disulfide bonds, limitation of iron availability and reduction of iron-sulfur clusters are present. For starvation, it possessed phosphorus and, despite being asaccharolytic, carbon starvation-related CDSs. Conclusions The L. hongkongensis genome possessed a high variety of genes for adaptation to acid, alkaline, temperature, osmotic, oxidative, UV light and starvation stresses and acquisition of and tolerance to heavy metals. PMID:21711489
Adamek, Martina; Alanjary, Mohammad; Sales-Ortells, Helena; Goodfellow, Michael; Bull, Alan T; Winkler, Anika; Wibberg, Daniel; Kalinowski, Jörn; Ziemert, Nadine
2018-06-01
Genome mining tools have enabled us to predict biosynthetic gene clusters that might encode compounds with valuable functions for industrial and medical applications. With the continuously increasing number of genomes sequenced, we are confronted with an overwhelming number of predicted clusters. In order to guide the effective prioritization of biosynthetic gene clusters towards finding the most promising compounds, knowledge about diversity, phylogenetic relationships and distribution patterns of biosynthetic gene clusters is necessary. Here, we provide a comprehensive analysis of the model actinobacterial genus Amycolatopsis and its potential for the production of secondary metabolites. A phylogenetic characterization, together with a pan-genome analysis showed that within this highly diverse genus, four major lineages could be distinguished which differed in their potential to produce secondary metabolites. Furthermore, we were able to distinguish gene cluster families whose distribution correlated with phylogeny, indicating that vertical gene transfer plays a major role in the evolution of secondary metabolite gene clusters. Still, the vast majority of the diverse biosynthetic gene clusters were derived from clusters unique to the genus, and also unique in comparison to a database of known compounds. Our study on the locations of biosynthetic gene clusters in the genomes of Amycolatopsis' strains showed that clusters acquired by horizontal gene transfer tend to be incorporated into non-conserved regions of the genome thereby allowing us to distinguish core and hypervariable regions in Amycolatopsis genomes. Using a comparative genomics approach, it was possible to determine the potential of the genus Amycolatopsis to produce a huge diversity of secondary metabolites. Furthermore, the analysis demonstrates that horizontal and vertical gene transfer play an important role in the acquisition and maintenance of valuable secondary metabolites. Our results cast light on the interconnections between secondary metabolite gene clusters and provide a way to prioritize biosynthetic pathways in the search and discovery of novel compounds.
Jadhav, Rohit R; Ye, Zhenqing; Huang, Rui-Lan; Liu, Joseph; Hsu, Pei-Yin; Huang, Yi-Wen; Rangel, Leticia B; Lai, Hung-Cheng; Roa, Juan Carlos; Kirma, Nameer B; Huang, Tim Hui-Ming; Jin, Victor X
2015-01-01
Recent genome-wide analysis has shown that DNA methylation spans long stretches of chromosome regions consisting of clusters of contiguous CpG islands or gene families. Hypermethylation of various gene clusters has been reported in many types of cancer. In this study, we conducted methyl-binding domain capture (MBDCap) sequencing (MBD-seq) analysis on a breast cancer cohort consisting of 77 patients and 10 normal controls, as well as a panel of 38 breast cancer cell lines. Bioinformatics analysis determined seven gene clusters with a significant difference in overall survival (OS) and further revealed a distinct feature that the conservation of a large gene cluster (approximately 70 kb) metallothionein-1 (MT1) among 45 species is much lower than the average of all RefSeq genes. Furthermore, we found that DNA methylation is an important epigenetic regulator contributing to gene repression of MT1 gene cluster in both ERα positive (ERα+) and ERα negative (ERα-) breast tumors. In silico analysis revealed much lower gene expression of this cluster in The Cancer Genome Atlas (TCGA) cohort for ERα + tumors. To further investigate the role of estrogen, we conducted 17β-estradiol (E2) and demethylating agent 5-aza-2'-deoxycytidine (DAC) treatment in various breast cancer cell types. Cell proliferation and invasion assays suggested MT1F and MT1M may play an anti-oncogenic role in breast cancer. Our data suggests that DNA methylation in large contiguous gene clusters can be potential prognostic markers of breast cancer. Further investigation of these clusters revealed that estrogen mediates epigenetic repression of MT1 cluster in ERα + breast cancer cell lines. In all, our studies identify thousands of breast tumor hypermethylated regions for the first time, in particular, discovering seven large contiguous hypermethylated gene clusters.
Long noncoding RNA HOTTIP cooperates with CCCTC-binding factor to coordinate HOXA gene expression.
Wang, Feng; Tang, Zhongqiong; Shao, Honglian; Guo, Jun; Tan, Tao; Dong, Yang; Lin, Lianbing
2018-06-12
The spatiotemporal control of HOX gene expression is dependent on positional identity and often correlated to their genomic location within each loci. Maintenance of HOX expression patterns is under complex transcriptional and epigenetic regulation, which is not well understood. Here we demonstrate that HOTTIP, a lincRNA transcribed from the 5' edge of the HOXA locus, physically associates with the CCCTC-binding factor (CTCF) that serves as an insulator by organizing HOXA cluster into disjoint domains, to cooperatively maintain the chromatin modifications of HOXA genes and thus coordinate the transcriptional activation of distal HOXA genes in human foreskin fibroblasts. Our results reveal the functional connection of HOTTIP and CTCF, and shed light on lincRNAs in gene activation and CTCF mediated chromatin organization. Copyright © 2018 Elsevier Inc. All rights reserved.
WordCluster: detecting clusters of DNA words and genomic elements
2011-01-01
Background Many k-mers (or DNA words) and genomic elements are known to be spatially clustered in the genome. Well established examples are the genes, TFBSs, CpG dinucleotides, microRNA genes and ultra-conserved non-coding regions. Currently, no algorithm exists to find these clusters in a statistically comprehensible way. The detection of clustering often relies on densities and sliding-window approaches or arbitrarily chosen distance thresholds. Results We introduce here an algorithm to detect clusters of DNA words (k-mers), or any other genomic element, based on the distance between consecutive copies and an assigned statistical significance. We implemented the method into a web server connected to a MySQL backend, which also determines the co-localization with gene annotations. We demonstrate the usefulness of this approach by detecting the clusters of CAG/CTG (cytosine contexts that can be methylated in undifferentiated cells), showing that the degree of methylation vary drastically between inside and outside of the clusters. As another example, we used WordCluster to search for statistically significant clusters of olfactory receptor (OR) genes in the human genome. Conclusions WordCluster seems to predict biological meaningful clusters of DNA words (k-mers) and genomic entities. The implementation of the method into a web server is available at http://bioinfo2.ugr.es/wordCluster/wordCluster.php including additional features like the detection of co-localization with gene regions or the annotation enrichment tool for functional analysis of overlapped genes. PMID:21261981
Williams, Bronwyn W; Scribner, Kim T
2010-01-01
Reintroductions and translocations are increasingly used to repatriate or increase probabilities of persistence for animal and plant species. Genetic and demographic characteristics of founding individuals and suitability of habitat at release sites are commonly believed to affect the success of these conservation programs. Genetic divergence among multiple source populations of American martens (Martes americana) and well documented introduction histories permitted analyses of post-introduction dispersion from release sites and development of genetic clusters in the Upper Peninsula (UP) of Michigan <50 years following release. Location and size of spatial genetic clusters and measures of individual-based autocorrelation were inferred using 11 microsatellite loci. We identified three genetic clusters in geographic proximity to original release locations. Estimated distances of effective gene flow based on spatial autocorrelation varied greatly among genetic clusters (30-90 km). Spatial contiguity of genetic clusters has been largely maintained with evidence for admixture primarily in localized regions, suggesting recent contact or locally retarded rates of gene flow. Data provide guidance for future studies of the effects of permeabilities of different land-cover and land-use features to dispersal and of other biotic and environmental factors that may contribute to the colonization process and development of spatial genetic associations.
Koren, Omry; Knights, Dan; Gonzalez, Antonio; Waldron, Levi; Segata, Nicola; Knight, Rob; Huttenhower, Curtis; Ley, Ruth E
2013-01-01
Recent analyses of human-associated bacterial diversity have categorized individuals into 'enterotypes' or clusters based on the abundances of key bacterial genera in the gut microbiota. There is a lack of consensus, however, on the analytical basis for enterotypes and on the interpretation of these results. We tested how the following factors influenced the detection of enterotypes: clustering methodology, distance metrics, OTU-picking approaches, sequencing depth, data type (whole genome shotgun (WGS) vs.16S rRNA gene sequence data), and 16S rRNA region. We included 16S rRNA gene sequences from the Human Microbiome Project (HMP) and from 16 additional studies and WGS sequences from the HMP and MetaHIT. In most body sites, we observed smooth abundance gradients of key genera without discrete clustering of samples. Some body habitats displayed bimodal (e.g., gut) or multimodal (e.g., vagina) distributions of sample abundances, but not all clustering methods and workflows accurately highlight such clusters. Because identifying enterotypes in datasets depends not only on the structure of the data but is also sensitive to the methods applied to identifying clustering strength, we recommend that multiple approaches be used and compared when testing for enterotypes.
Waldron, Levi; Segata, Nicola; Knight, Rob; Huttenhower, Curtis; Ley, Ruth E.
2013-01-01
Recent analyses of human-associated bacterial diversity have categorized individuals into ‘enterotypes’ or clusters based on the abundances of key bacterial genera in the gut microbiota. There is a lack of consensus, however, on the analytical basis for enterotypes and on the interpretation of these results. We tested how the following factors influenced the detection of enterotypes: clustering methodology, distance metrics, OTU-picking approaches, sequencing depth, data type (whole genome shotgun (WGS) vs.16S rRNA gene sequence data), and 16S rRNA region. We included 16S rRNA gene sequences from the Human Microbiome Project (HMP) and from 16 additional studies and WGS sequences from the HMP and MetaHIT. In most body sites, we observed smooth abundance gradients of key genera without discrete clustering of samples. Some body habitats displayed bimodal (e.g., gut) or multimodal (e.g., vagina) distributions of sample abundances, but not all clustering methods and workflows accurately highlight such clusters. Because identifying enterotypes in datasets depends not only on the structure of the data but is also sensitive to the methods applied to identifying clustering strength, we recommend that multiple approaches be used and compared when testing for enterotypes. PMID:23326225
Large clusters of co-expressed genes in the Drosophila genome.
Boutanaev, Alexander M; Kalmykova, Alla I; Shevelyov, Yuri Y; Nurminsky, Dmitry I
2002-12-12
Clustering of co-expressed, non-homologous genes on chromosomes implies their co-regulation. In lower eukaryotes, co-expressed genes are often found in pairs. Clustering of genes that share aspects of transcriptional regulation has also been reported in higher eukaryotes. To advance our understanding of the mode of coordinated gene regulation in multicellular organisms, we performed a genome-wide analysis of the chromosomal distribution of co-expressed genes in Drosophila. We identified a total of 1,661 testes-specific genes, one-third of which are clustered on chromosomes. The number of clusters of three or more genes is much higher than expected by chance. We observed a similar trend for genes upregulated in the embryo and in the adult head, although the expression pattern of individual genes cannot be predicted on the basis of chromosomal position alone. Our data suggest that the prevalent mechanism of transcriptional co-regulation in higher eukaryotes operates with extensive chromatin domains that comprise multiple genes.
Unusual Gene Order and Organization of the Sea Urchin Hox Cluster
DOE Office of Scientific and Technical Information (OSTI.GOV)
Cameron, R A; Rowen, L; Nesbitt, R
2005-10-11
The highly consistent gene order and axial colinear expression patterns found in vertebrate hox gene clusters are less well conserved across the rest of bilaterians. We report the first deuterostome instance of an intact hox cluster with a unique gene order where the paralog groups are not expressed in a sequential manner. The finished sequence from BAC clones from the genome of the sea urchin, Strongylocentrotus purpuratus, reveals a gene order wherein the anterior genes (Hox1, Hox2 and Hox3) lie nearest the posterior genes in the cluster such that the most 3 gene is Hox5. (The gene order is :more » 5-Hox1, 2, 3, 11/13c, 11/13b, 11/13a, 9/10, 8, 7, 6, 5 - 3). The finished sequence result is corroborated by restriction mapping evidence and BAC-end scaffold analyses. Comparisons with a putative ancestral deuterostome Hox gene cluster suggest that the rearrangements leading to the sea urchin gene order were many and complex.« less
Unusual Gene Order and Organization of the Sea Urchin HoxCluster
DOE Office of Scientific and Technical Information (OSTI.GOV)
Richardson, Paul M.; Lucas, Susan; Cameron, R. Andrew
2005-05-10
The highly consistent gene order and axial colinear expression patterns found in vertebrate hox gene clusters are less well conserved across the rest of bilaterians. We report the first deuterostome instance of an intact hox cluster with a unique gene order where the paralog groups are not expressed in a sequential manner. The finished sequence from BAC clones from the genome of the sea urchin, Strongylocentrotus purpuratus, reveals a gene order wherein the anterior genes (Hox1, Hox2 and Hox3) lie nearest the posterior genes in the cluster such that the most 3' gene is Hox5. (The gene order is :more » 5'-Hox1,2, 3, 11/13c, 11/13b, '11/13a, 9/10, 8, 7, 6, 5 - 3)'. The finished sequence result is corroborated by restriction mapping evidence and BAC-end scaffold analyses. Comparisons with a putative ancestral deuterostome Hox gene cluster suggest that the rearrangements leading to the sea urchin gene order were many and complex.« less
TFIIIC Bound DNA Elements in Nuclear Organization and Insulation
Kirkland, Jacob G.; Raab, Jesse R.
2012-01-01
tRNA genes (tDNAs) have been known to have barrier insulator function in budding yeast, Saccharomyces cerevisiae, for over a decade. tDNAs also play a role in genome organization by clustering at sites in the nucleus and both of these functions are dependent on the transcription factor TFIIIC. More recently TFIIIC bound sites devoid of pol III, termed Extra-TFIIIC sites (ETC) have been identified in budding yeast and these sites also function as insulators and affect genome organization. Subsequent studies in Schizosaccharomyces pombe showed that TFIIIC bound sites were insulators and also functioned as Chromosome Organization Clamps (COC); tethering the sites to the nuclear periphery. Very recently studies have moved to mammalian systems where pol III genes and their associated factors have been investigated in both mouse and human cells. Short Interspersed Nuclear Elements (SINEs) that bind TFIIIC, function as insulator elements and tDNAs can also function as both enhancer -blocking and barrier insulators in these organisms. It was also recently shown that tDNAs cluster with other tDNAs and with ETCs but not with pol II transcribed genes. Intriguingly, TFIIIC is often found near pol II transcription start sites and it remains unclear what the consequences of TFIIIC based genomic organization are and what influence pol III factors have on pol II transcribed genes and vise versa. In this review we provide a comprehensive overview of the known data on pol III factors in insulation and genome organization and identify the many open questions that require further investigation. \\ PMID:23000638
Dass, J Febin Prabhu; Sudandiradoss, C
2012-07-15
5-HT (5-Hydroxy-tryptamine) or serotonin receptors are found both in central and peripheral nervous system as well as in non-neuronal tissues. In the animal and human nervous system, serotonin produces various functional effects through a variety of membrane bound receptors. In this study, we focus on 5-HT receptor family from different mammals and examined the factors that account for codon and nucleotide usage variation. A total of 110 homologous coding sequences from 11 different mammalian species were analyzed using relative synonymous codon usage (RSCU), correspondence analysis (COA) and hierarchical cluster analysis together with nucleotide base usage frequency of chemically similar amino acid codons. The mean effective number of codon (ENc) value of 37.06 for 5-HT(6) shows very high codon bias within the family and may be due to high selective translational efficiency. The COA and Spearman's rank correlation reveals that the nucleotide compositional mutation bias as the major factors influencing the codon usage in serotonin receptor genes. The hierarchical cluster analysis suggests that gene function is another dominant factor that affects the codon usage bias, while species is a minor factor. Nucleotide base usage was reported using Goldman, Engelman, Stietz (GES) scale reveals the presence of high uracil (>45%) content at functionally important hydrophobic regions. Our in silico approach will certainly help for further investigations on critical inference on evolution, structure, function and gene expression aspects of 5-HT receptors family which are potential antipsychotic drug targets. Copyright © 2012 Elsevier B.V. All rights reserved.
Genome-Wide Prediction of Metabolic Enzymes, Pathways, and Gene Clusters in Plants
Schläpfer, Pascal; Zhang, Peifen; Wang, Chuan; ...
2017-04-01
Plant metabolism underpins many traits of ecological and agronomic importance. Plants produce numerous compounds to cope with their environments but the biosynthetic pathways for most of these compounds have not yet been elucidated. To engineer and improve metabolic traits, we will need comprehensive and accurate knowledge of the organization and regulation of plant metabolism at the genome scale. Here, we present a computational pipeline to identify metabolic enzymes, pathways, and gene clusters from a sequenced genome. Using this pipeline, we generated metabolic pathway databases for 22 species and identified metabolic gene clusters from 18 species. This unified resource can bemore » used to conduct a wide array of comparative studies of plant metabolism. Using the resource, we discovered a widespread occurrence of metabolic gene clusters in plants: 11,969 clusters from 18 species. The prevalence of metabolic gene clusters offers an intriguing possibility of an untapped source for uncovering new metabolite biosynthesis pathways. For example, more than 1,700 clusters contain enzymes that could generate a specialized metabolite scaffold (signature enzymes) and enzymes that modify the scaffold (tailoring enzymes). In four species with sufficient gene expression data, we identified 43 highly coexpressed clusters that contain signature and tailoring enzymes, of which eight were characterized previously to be functional pathways. Finally, we identified patterns of genome organization that implicate local gene duplication and, to a lesser extent, single gene transposition as having played roles in the evolution of plant metabolic gene clusters.« less
Genome-Wide Prediction of Metabolic Enzymes, Pathways, and Gene Clusters in Plants1[OPEN
Zhang, Peifen; Kim, Taehyong; Banf, Michael; Chavali, Arvind K.; Nilo-Poyanco, Ricardo; Bernard, Thomas
2017-01-01
Plant metabolism underpins many traits of ecological and agronomic importance. Plants produce numerous compounds to cope with their environments but the biosynthetic pathways for most of these compounds have not yet been elucidated. To engineer and improve metabolic traits, we need comprehensive and accurate knowledge of the organization and regulation of plant metabolism at the genome scale. Here, we present a computational pipeline to identify metabolic enzymes, pathways, and gene clusters from a sequenced genome. Using this pipeline, we generated metabolic pathway databases for 22 species and identified metabolic gene clusters from 18 species. This unified resource can be used to conduct a wide array of comparative studies of plant metabolism. Using the resource, we discovered a widespread occurrence of metabolic gene clusters in plants: 11,969 clusters from 18 species. The prevalence of metabolic gene clusters offers an intriguing possibility of an untapped source for uncovering new metabolite biosynthesis pathways. For example, more than 1,700 clusters contain enzymes that could generate a specialized metabolite scaffold (signature enzymes) and enzymes that modify the scaffold (tailoring enzymes). In four species with sufficient gene expression data, we identified 43 highly coexpressed clusters that contain signature and tailoring enzymes, of which eight were characterized previously to be functional pathways. Finally, we identified patterns of genome organization that implicate local gene duplication and, to a lesser extent, single gene transposition as having played roles in the evolution of plant metabolic gene clusters. PMID:28228535
Genome-Wide Prediction of Metabolic Enzymes, Pathways, and Gene Clusters in Plants
DOE Office of Scientific and Technical Information (OSTI.GOV)
Schläpfer, Pascal; Zhang, Peifen; Wang, Chuan
Plant metabolism underpins many traits of ecological and agronomic importance. Plants produce numerous compounds to cope with their environments but the biosynthetic pathways for most of these compounds have not yet been elucidated. To engineer and improve metabolic traits, we will need comprehensive and accurate knowledge of the organization and regulation of plant metabolism at the genome scale. Here, we present a computational pipeline to identify metabolic enzymes, pathways, and gene clusters from a sequenced genome. Using this pipeline, we generated metabolic pathway databases for 22 species and identified metabolic gene clusters from 18 species. This unified resource can bemore » used to conduct a wide array of comparative studies of plant metabolism. Using the resource, we discovered a widespread occurrence of metabolic gene clusters in plants: 11,969 clusters from 18 species. The prevalence of metabolic gene clusters offers an intriguing possibility of an untapped source for uncovering new metabolite biosynthesis pathways. For example, more than 1,700 clusters contain enzymes that could generate a specialized metabolite scaffold (signature enzymes) and enzymes that modify the scaffold (tailoring enzymes). In four species with sufficient gene expression data, we identified 43 highly coexpressed clusters that contain signature and tailoring enzymes, of which eight were characterized previously to be functional pathways. Finally, we identified patterns of genome organization that implicate local gene duplication and, to a lesser extent, single gene transposition as having played roles in the evolution of plant metabolic gene clusters.« less
Genome-Wide Prediction of Metabolic Enzymes, Pathways, and Gene Clusters in Plants.
Schläpfer, Pascal; Zhang, Peifen; Wang, Chuan; Kim, Taehyong; Banf, Michael; Chae, Lee; Dreher, Kate; Chavali, Arvind K; Nilo-Poyanco, Ricardo; Bernard, Thomas; Kahn, Daniel; Rhee, Seung Y
2017-04-01
Plant metabolism underpins many traits of ecological and agronomic importance. Plants produce numerous compounds to cope with their environments but the biosynthetic pathways for most of these compounds have not yet been elucidated. To engineer and improve metabolic traits, we need comprehensive and accurate knowledge of the organization and regulation of plant metabolism at the genome scale. Here, we present a computational pipeline to identify metabolic enzymes, pathways, and gene clusters from a sequenced genome. Using this pipeline, we generated metabolic pathway databases for 22 species and identified metabolic gene clusters from 18 species. This unified resource can be used to conduct a wide array of comparative studies of plant metabolism. Using the resource, we discovered a widespread occurrence of metabolic gene clusters in plants: 11,969 clusters from 18 species. The prevalence of metabolic gene clusters offers an intriguing possibility of an untapped source for uncovering new metabolite biosynthesis pathways. For example, more than 1,700 clusters contain enzymes that could generate a specialized metabolite scaffold (signature enzymes) and enzymes that modify the scaffold (tailoring enzymes). In four species with sufficient gene expression data, we identified 43 highly coexpressed clusters that contain signature and tailoring enzymes, of which eight were characterized previously to be functional pathways. Finally, we identified patterns of genome organization that implicate local gene duplication and, to a lesser extent, single gene transposition as having played roles in the evolution of plant metabolic gene clusters. © 2017 American Society of Plant Biologists. All Rights Reserved.
Liu, L L; Liu, M J; Ma, M
2015-09-28
The central task of this study was to mine the gene-to-medium relationship. Adequate knowledge of this relationship could potentially improve the accuracy of differentially expressed gene mining. One of the approaches to differentially expressed gene mining uses conventional clustering algorithms to identify the gene-to-medium relationship. Compared to conventional clustering algorithms, self-organization maps (SOMs) identify the nonlinear aspects of the gene-to-medium relationships by mapping the input space into another higher dimensional feature space. However, SOMs are not suitable for huge datasets consisting of millions of samples. Therefore, a new computational model, the Function Clustering Self-Organization Maps (FCSOMs), was developed. FCSOMs take advantage of the theory of granular computing as well as advanced statistical learning methodologies, and are built specifically for each information granule (a function cluster of genes), which are intelligently partitioned by the clustering algorithm provided by the DAVID_6.7 software platform. However, only the gene functions, and not their expression values, are considered in the fuzzy clustering algorithm of DAVID. Compared to the clustering algorithm of DAVID, these experimental results show a marked improvement in the accuracy of classification with the application of FCSOMs. FCSOMs can handle huge datasets and their complex classification problems, as each FCSOM (modeled for each function cluster) can be easily parallelized.
Bao, Yun-Juan; Liang, Zhong; Mayfield, Jeffrey A.; McShan, William M.; Lee, Shaun W.; Ploplis, Victoria A.; Castellino, Francis J.
2016-01-01
Symmetric genomic rearrangements around replication axes in genomes are commonly observed in prokaryotic genomes, including Group A Streptococcus (GAS). However, asymmetric rearrangements are rare. Our previous studies showed that the hypervirulent invasive GAS strain, M23ND, containing an inactivated transcriptional regulator system, covRS, exhibits unique extensive asymmetric rearrangements, which reconstructed a genomic structure distinct from other GAS genomes. In the current investigation, we identified the rearrangement events and examined the genetic consequences and evolutionary implications underlying the rearrangements. By comparison with a close phylogenetic relative, M18-MGAS8232, we propose a molecular model wherein a series of asymmetric rearrangements have occurred in M23ND, involving translocations, inversions and integrations mediated by multiple factors, viz., rRNA-comX (factor for late competence), transposons and phage-encoded gene segments. Assessments of the cumulative gene orientations and GC skews reveal that the asymmetric genomic rearrangements did not affect the general genomic integrity of the organism. However, functional distributions reveal re-clustering of a broad set of CovRS-regulated actively transcribed genes, including virulence factors and metabolic genes, to the same leading strand, with high confidence (p-value ~10−10). The re-clustering of the genes suggests a potential selection advantage for the spatial proximity to the transcription complexes, which may contain the global transcriptional regulator, CovRS, and other RNA polymerases. Their proximities allow for efficient transcription of the genes required for growth, virulence and persistence. A new paradigm of survival strategies of GAS strains is provided through multiple genomic rearrangements, while, at the same time, maintaining genomic integrity. PMID:27329479
A model for evolution and regulation of nicotine biosynthesis regulon in tobacco.
Kajikawa, Masataka; Sierro, Nicolas; Hashimoto, Takashi; Shoji, Tsubasa
2017-06-03
In tobacco, the defense alkaloid nicotine is produced in roots and accumulates mainly in leaves. Signaling mediated by jasmonates (JAs) induces the formation of nicotine via a series of structural genes that constitute a regulon and are coordinated by JA-responsive transcription factors of the ethylene response factor (ERF) family. Early steps in the pyrrolidine and pyridine biosynthesis pathways likely arose through duplication of the polyamine and nicotinamide adenine dinucleotide (NAD) biosynthetic pathways, respectively, followed by recruitment of duplicated primary metabolic genes into the nicotine biosynthesis regulon. Transcriptional regulation of nicotine biosynthesis by ERF and cooperatively-acting MYC2 transcription factors is implied by the frequency of cognate cis-regulatory elements for these factors in the promoter regions of the downstream structural genes. Indeed, a mutant tobacco with low nicotine content was found to have a large chromosomal deletion in a cluster of closely related ERF genes at the nicotine-controlling NICOTINE2 (NIC2) locus.
2013-01-01
Background The adipose tissue is an endocrine regulator and a risk factor for atherosclerosis and cardiovascular disease when by excessive accumulation induces obesity. Although the adipose tissue is also a reservoir for stem cells (ASC) their function and “stemcellness” has been questioned. Our aim was to investigate the mechanisms by which obesity affects subcutaneous white adipose tissue (WAT) stem cells. Results Transcriptomics, in silico analysis, real-time polymerase chain reaction (PCR) and western blots were performed on isolated stem cells from subcutaneous abdominal WAT of morbidly obese patients (ASCmo) and of non-obese individuals (ASCn). ASCmo and ASCn gene expression clustered separately from each other. ASCmo showed downregulation of “stemness” genes and upregulation of adipogenic and inflammatory genes with respect to ASCn. Moreover, the application of bioinformatics and Ingenuity Pathway Analysis (IPA) showed that the transcription factor Smad3 was tentatively affected in obese ASCmo. Validation of this target confirmed a significantly reduced Smad3 nuclear translocation in the isolated ASCmo. Conclusions The transcriptomic profile of the stem cells reservoir in obese subcutaneous WAT is highly modified with significant changes in genes regulating stemcellness, lineage commitment and inflammation. In addition to body mass index, cardiovascular risk factor clustering further affect the ASC transcriptomic profile inducing loss of multipotency and, hence, capacity for tissue repair. In summary, the stem cells in the subcutaneous WAT niche of obese patients are already committed to adipocyte differentiation and show an upregulated inflammatory gene expression associated to their loss of stemcellness. PMID:24040759
Tanaka-Tsuno, Fumiko; Mizukami-Murata, Satomi; Murata, Yoshinori; Nakamura, Toshihide; Ando, Akira; Takagi, Hiroshi; Shima, Jun
2007-10-01
In the modern baking industry, high-sucrose-tolerant (HS) and maltose-utilizing (LS) yeast were developed using breeding techniques and are now used commercially. Sugar utilization and high-sucrose tolerance differ significantly between HS and LS yeasts. We analysed the gene expression profiles of HS and LS yeasts under different sucrose conditions in order to determine their basic physiology. Two-way hierarchical clustering was performed to obtain the overall patterns of gene expression. The clustering clearly showed that the gene expression patterns of LS yeast differed from those of HS yeast. Quality threshold clustering was used to identify the gene clusters containing upregulated genes (cluster 1) and downregulated genes (cluster 2) under high-sucrose conditions. Clusters 1 and 2 contained numerous genes involved in carbon and nitrogen metabolism, respectively. The expression level of the genes involved in the metabolism of glycerol and trehalose, which are known to be osmoprotectants, in LS yeast was higher than that in HS yeast under sucrose concentrations of 5-40%. No clear correlation was found between the expression level of the genes involved in the biosynthesis of the osmoprotectants and the intracellular contents of the osmoprotectants. The present gene expression data were compared with data previously reported in a comprehensive analysis of a gene deletion strain collection. Welch's t-test for this comparison showed that the relative growth rates of the deletion strains whose deletion occurred in genes belonging to cluster 1 were significantly higher than the average growth rates of all deletion strains. Copyright 2007 John Wiley & Sons, Ltd.
Differential Retention of Gene Functions in a Secondary Metabolite Cluster.
Reynolds, Hannah T; Slot, Jason C; Divon, Hege H; Lysøe, Erik; Proctor, Robert H; Brown, Daren W
2017-08-01
In fungi, distribution of secondary metabolite (SM) gene clusters is often associated with host- or environment-specific benefits provided by SMs. In the plant pathogen Alternaria brassicicola (Dothideomycetes), the DEP cluster confers an ability to synthesize the SM depudecin, a histone deacetylase inhibitor that contributes weakly to virulence. The DEP cluster includes genes encoding enzymes, a transporter, and a transcription regulator. We investigated the distribution and evolution of the DEP cluster in 585 fungal genomes and found a wide but sporadic distribution among Dothideomycetes, Sordariomycetes, and Eurotiomycetes. We confirmed DEP gene expression and depudecin production in one fungus, Fusarium langsethiae. Phylogenetic analyses suggested 6-10 horizontal gene transfers (HGTs) of the cluster, including a transfer that led to the presence of closely related cluster homologs in Alternaria and Fusarium. The analyses also indicated that HGTs were frequently followed by loss/pseudogenization of one or more DEP genes. Independent cluster inactivation was inferred in at least four fungal classes. Analyses of transitions among functional, pseudogenized, and absent states of DEP genes among Fusarium species suggest enzyme-encoding genes are lost at higher rates than the transporter (DEP3) and regulatory (DEP6) genes. The phenotype of an experimentally-induced DEP3 mutant of Fusarium did not support the hypothesis that selective retention of DEP3 and DEP6 protects fungi from exogenous depudecin. Together, the results suggest that HGT and gene loss have contributed significantly to DEP cluster distribution, and that some DEP genes provide a greater fitness benefit possibly due to a differential tendency to form network connections. Published by Oxford University Press on behalf of the Society for Molecular Biology and Evolution 2017. This work is written by US Government employees and is in the public domain in the US.
Circadian Enhancers Coordinate Multiple Phases of Rhythmic Gene Transcription In Vivo
Fang, Bin; Everett, Logan J.; Jager, Jennifer; Briggs, Erika; Armour, Sean M.; Feng, Dan; Roy, Ankur; Gerhart-Hines, Zachary; Sun, Zheng; Lazar, Mitchell A.
2014-01-01
SUMMARY Mammalian transcriptomes display complex circadian rhythms with multiple phases of gene expression that cannot be accounted for by current models of the molecular clock. We have determined the underlying mechanisms by measuring nascent RNA transcription around the clock in mouse liver. Unbiased examination of eRNAs that cluster in specific circadian phases identified functional enhancers driven by distinct transcription factors (TFs). We further identify on a global scale the components of the TF cistromes that function to orchestrate circadian gene expression. Integrated genomic analyses also revealed novel mechanisms by which a single circadian factor controls opposing transcriptional phases. These findings shed new light on the diversity and specificity of TF function in the generation of multiple phases of circadian gene transcription in a mammalian organ. PMID:25416951
Circadian enhancers coordinate multiple phases of rhythmic gene transcription in vivo.
Fang, Bin; Everett, Logan J; Jager, Jennifer; Briggs, Erika; Armour, Sean M; Feng, Dan; Roy, Ankur; Gerhart-Hines, Zachary; Sun, Zheng; Lazar, Mitchell A
2014-11-20
Mammalian transcriptomes display complex circadian rhythms with multiple phases of gene expression that cannot be accounted for by current models of the molecular clock. We have determined the underlying mechanisms by measuring nascent RNA transcription around the clock in mouse liver. Unbiased examination of enhancer RNAs (eRNAs) that cluster in specific circadian phases identified functional enhancers driven by distinct transcription factors (TFs). We further identify on a global scale the components of the TF cistromes that function to orchestrate circadian gene expression. Integrated genomic analyses also revealed mechanisms by which a single circadian factor controls opposing transcriptional phases. These findings shed light on the diversity and specificity of TF function in the generation of multiple phases of circadian gene transcription in a mammalian organ.
Pimentel, Harold; Parra, Marilyn; Gee, Sherry L.; ...
2015-11-03
Differentiating erythroblasts execute a dynamic alternative splicing program shown here to include extensive and diverse intron retention (IR) events. Cluster analysis revealed hundreds of developmentallydynamic introns that exhibit increased IR in mature erythroblasts, and are enriched in functions related to RNA processing such as SF3B1 spliceosomal factor. Distinct, developmentally-stable IR clusters are enriched in metal-ion binding functions and include mitoferrin genes SLC25A37 and SLC25A28 that are critical for iron homeostasis. Some IR transcripts are abundant, e.g. comprising ~50% of highly-expressed SLC25A37 and SF3B1 transcripts in late erythroblasts, and thereby limiting functional mRNA levels. IR transcripts tested were predominantly nuclearlocalized. Splicemore » site strength correlated with IR among stable but not dynamic intron clusters, indicating distinct regulation of dynamically-increased IR in late erythroblasts. Retained introns were preferentially associated with alternative exons with premature termination codons (PTCs). High IR was observed in disease-causing genes including SF3B1 and the RNA binding protein FUS. Comparative studies demonstrated that the intron retention program in erythroblasts shares features with other tissues but ultimately is unique to erythropoiesis. Finally, we conclude that IR is a multi-dimensional set of processes that post-transcriptionally regulate diverse gene groups during normal erythropoiesis, misregulation of which could be responsible for human disease.« less
DOE Office of Scientific and Technical Information (OSTI.GOV)
Pimentel, Harold; Parra, Marilyn; Gee, Sherry L.
Differentiating erythroblasts execute a dynamic alternative splicing program shown here to include extensive and diverse intron retention (IR) events. Cluster analysis revealed hundreds of developmentallydynamic introns that exhibit increased IR in mature erythroblasts, and are enriched in functions related to RNA processing such as SF3B1 spliceosomal factor. Distinct, developmentally-stable IR clusters are enriched in metal-ion binding functions and include mitoferrin genes SLC25A37 and SLC25A28 that are critical for iron homeostasis. Some IR transcripts are abundant, e.g. comprising ~50% of highly-expressed SLC25A37 and SF3B1 transcripts in late erythroblasts, and thereby limiting functional mRNA levels. IR transcripts tested were predominantly nuclearlocalized. Splicemore » site strength correlated with IR among stable but not dynamic intron clusters, indicating distinct regulation of dynamically-increased IR in late erythroblasts. Retained introns were preferentially associated with alternative exons with premature termination codons (PTCs). High IR was observed in disease-causing genes including SF3B1 and the RNA binding protein FUS. Comparative studies demonstrated that the intron retention program in erythroblasts shares features with other tissues but ultimately is unique to erythropoiesis. Finally, we conclude that IR is a multi-dimensional set of processes that post-transcriptionally regulate diverse gene groups during normal erythropoiesis, misregulation of which could be responsible for human disease.« less
Shang, Haihong; Li, Wei; Zou, Changsong; Yuan, Youlu
2013-07-01
NAC domain proteins are plant-specific transcription factors known to play diverse roles in various plant developmental processes. In the present study, we performed the first comprehensive study of the NAC gene family in Gossypium raimondii Ulbr., incorporating phylogenetic, chromosomal location, gene structure, conserved motif, and expression profiling analyses. We identified 145 NAC transcription factor (NAC-TF) genes that were phylogenetically clustered into 18 distinct subfamilies. Of these, 127 NAC-TF genes were distributed across the 13 chromosomes, 80 (55%) were preferentially retained duplicates located in both duplicated regions and six were located in triplicated chromosomal regions. The majority of NAC-TF genes showed temporal-, spatial-, and tissue-specific expression patterns based on transcriptomic and qRT-PCR analyses. However, the expression patterns of several duplicate genes were partially redundant, suggesting the occurrence of sub-functionalization during their evolution. Based on their genomic organization, we concluded that genomic duplications contributed significantly to the expansion of the NAC-TF gene family in G. raimondii. Comprehensive analysis of their expression profiles could provide novel insights into the functional divergence among members of the NAC gene family in G. raimondii. © 2013 Institute of Botany, Chinese Academy of Sciences.
Stevenson, G; Andrianopoulos, K; Hobbs, M; Reeves, P R
1996-01-01
Colanic acid (CA) is an extracellular polysaccharide produced by most Escherichia coli strains as well as by other species of the family Enterobacteriaceae. We have determined the sequence of a 23-kb segment of the E. coli K-12 chromosome which includes the cluster of genes necessary for production of CA. The CA cluster comprises 19 genes. Two other sequenced genes (orf1.3 and galF), which are situated between the CA cluster and the O-antigen cluster, were shown to be unnecessary for CA production. The CA cluster includes genes for synthesis of GDP-L-fucose, one of the precursors of CA, and the gene for one of the enzymes in this pathway (GDP-D-mannose 4,6-dehydratase) was identified by biochemical assay. Six of the inferred proteins show sequence similarity to glycosyl transferases, and two others have sequence similarity to acetyl transferases. Another gene (wzx) is predicted to encode a protein with multiple transmembrane segments and may function in export of the CA repeat unit from the cytoplasm into the periplasm in a process analogous to O-unit export. The first three genes of the cluster are predicted to encode an outer membrane lipoprotein, a phosphatase, and an inner membrane protein with an ATP-binding domain. Since homologs of these genes are found in other extracellular polysaccharide gene clusters, they may have a common function, such as export of polysaccharide from the cell. PMID:8759852
Ficklin, Stephen P.; Luo, Feng; Feltus, F. Alex
2010-01-01
Discovering gene sets underlying the expression of a given phenotype is of great importance, as many phenotypes are the result of complex gene-gene interactions. Gene coexpression networks, built using a set of microarray samples as input, can help elucidate tightly coexpressed gene sets (modules) that are mixed with genes of known and unknown function. Functional enrichment analysis of modules further subdivides the coexpressed gene set into cofunctional gene clusters that may coexist in the module with other functionally related gene clusters. In this study, 45 coexpressed gene modules and 76 cofunctional gene clusters were discovered for rice (Oryza sativa) using a global, knowledge-independent paradigm and the combination of two network construction methodologies. Some clusters were enriched for previously characterized mutant phenotypes, providing evidence for specific gene sets (and their annotated molecular functions) that underlie specific phenotypes. PMID:20668062
Ficklin, Stephen P; Luo, Feng; Feltus, F Alex
2010-09-01
Discovering gene sets underlying the expression of a given phenotype is of great importance, as many phenotypes are the result of complex gene-gene interactions. Gene coexpression networks, built using a set of microarray samples as input, can help elucidate tightly coexpressed gene sets (modules) that are mixed with genes of known and unknown function. Functional enrichment analysis of modules further subdivides the coexpressed gene set into cofunctional gene clusters that may coexist in the module with other functionally related gene clusters. In this study, 45 coexpressed gene modules and 76 cofunctional gene clusters were discovered for rice (Oryza sativa) using a global, knowledge-independent paradigm and the combination of two network construction methodologies. Some clusters were enriched for previously characterized mutant phenotypes, providing evidence for specific gene sets (and their annotated molecular functions) that underlie specific phenotypes.
Bender, Carol L.; Alarcón-Chaidez, Francisco; Gross, Dennis C.
1999-01-01
Coronatine, syringomycin, syringopeptin, tabtoxin, and phaseolotoxin are the most intensively studied phytotoxins of Pseudomonas syringae, and each contributes significantly to bacterial virulence in plants. Coronatine functions partly as a mimic of methyl jasmonate, a hormone synthesized by plants undergoing biological stress. Syringomycin and syringopeptin form pores in plasma membranes, a process that leads to electrolyte leakage. Tabtoxin and phaseolotoxin are strongly antimicrobial and function by inhibiting glutamine synthetase and ornithine carbamoyltransferase, respectively. Genetic analysis has revealed the mechanisms responsible for toxin biosynthesis. Coronatine biosynthesis requires the cooperation of polyketide and peptide synthetases for the assembly of the coronafacic and coronamic acid moieties, respectively. Tabtoxin is derived from the lysine biosynthetic pathway, whereas syringomycin, syringopeptin, and phaseolotoxin biosynthesis requires peptide synthetases. Activation of phytotoxin synthesis is controlled by diverse environmental factors including plant signal molecules and temperature. Genes involved in the regulation of phytotoxin synthesis have been located within the coronatine and syringomycin gene clusters; however, additional regulatory genes are required for the synthesis of these and other phytotoxins. Global regulatory genes such as gacS modulate phytotoxin production in certain pathovars, indicating the complexity of the regulatory circuits controlling phytotoxin synthesis. The coronatine and syringomycin gene clusters have been intensively characterized and show potential for constructing modified polyketides and peptides. Genetic reprogramming of peptide and polyketide synthetases has been successful, and portions of the coronatine and syringomycin gene clusters could be valuable resources in developing new antimicrobial agents. PMID:10357851
Dong, Ying; Matigian, Nick; Harvey, Tracey J; Samaratunga, Hemamali; Hooper, John D; Clements, Judith A
2008-02-01
Abstract Tissue kallikrein (kallikrein 1) was first identified in pancreas and is the namesake of the kallikrein-related peptidase (KLK) family. KLK1 and the other 14 members of the human KLK family are encoded by 15 serine protease genes clustered at chromosome 19q13.4. Our Northern blot analysis of 19 normal human tissues for expression of KLK4 to KLK15 identified pancreas as a common expression site for the gene cluster spanning KLK5 to KLK13, as well as for KLK15 which is located adjacent to KLK1. Consistent with previous reports detailing the ability of KLK genes to generate organ- and disease-specific transcripts, detailed molecular and in silico analyses indicated that KLK5 and KLK7 generate transcripts in pancreas variant from those in skin or ovary. Consistently, we identified in the promoters of these KLK genes motifs which conform with consensus binding sites for transcription factors conferring pancreatic expression. In addition, immunohistochemical analysis revealed predominant localisation of KLK5 and KLK7 in acinar cells of the exocrine pancreas, suggesting roles for these enzymes in digestion. Our data also support expression patterns derived from gene duplication events in the human KLK cluster. These findings suggest that, in addition to KLK1, other related KLK enzymes will function in the exocrine pancreas.
HOX genes in human lung: altered expression in primary pulmonary hypertension and emphysema.
Golpon, H A; Geraci, M W; Moore, M D; Miller, H L; Miller, G J; Tuder, R M; Voelkel, N F
2001-03-01
HOX genes belong to the large family of homeodomain genes that function as transcription factors. Animal studies indicate that they play an essential role in lung development. We investigated the expression pattern of HOX genes in human lung tissue by using microarray and degenerate reverse transcriptase-polymerase chain reaction survey techniques. HOX genes predominantly from the 3' end of clusters A and B were expressed in normal human adult lung and among them HOXA5 was the most abundant, followed by HOXB2 and HOXB6. In fetal (12 weeks old) and diseased lung specimens (emphysema, primary pulmonary hypertension) additional HOX genes from clusters C and D were expressed. Using in situ hybridization, transcripts for HOXA5 were predominantly found in alveolar septal and epithelial cells, both in normal and diseased lungs. A 2.5-fold increase in HOXA5 mRNA expression was demonstrated by quantitative reverse transcriptase-polymerase chain reaction in primary pulmonary hypertension lung specimens when compared to normal lung tissue. In conclusion, we demonstrate that HOX genes are selectively expressed in the human lung. Differences in the pattern of HOX gene expression exist among fetal, adult, and diseased lung specimens. The altered pattern of HOX gene expression may contribute to the development of pulmonary diseases.
Golpon, Heiko A.; Geraci, Mark W.; Moore, Mark D.; Miller, Heidi L.; Miller, Gary J.; Tuder, Rubin M.; Voelkel, Norbert F.
2001-01-01
HOX genes belong to the large family of homeodomain genes that function as transcription factors. Animal studies indicate that they play an essential role in lung development. We investigated the expression pattern of HOX genes in human lung tissue by using microarray and degenerate reverse transcriptase-polymerase chain reaction survey techniques. HOX genes predominantly from the 3′ end of clusters A and B were expressed in normal human adult lung and among them HOXA5 was the most abundant, followed by HOXB2 and HOXB6. In fetal (12 weeks old) and diseased lung specimens (emphysema, primary pulmonary hypertension) additional HOX genes from clusters C and D were expressed. Using in situ hybridization, transcripts for HOXA5 were predominantly found in alveolar septal and epithelial cells, both in normal and diseased lungs. A 2.5-fold increase in HOXA5 mRNA expression was demonstrated by quantitative reverse transcriptase-polymerase chain reaction in primary pulmonary hypertension lung specimens when compared to normal lung tissue. In conclusion, we demonstrate that HOX genes are selectively expressed in the human lung. Differences in the pattern of HOX gene expression exist among fetal, adult, and diseased lung specimens. The altered pattern of HOX gene expression may contribute to the development of pulmonary diseases. PMID:11238043
The sirodesmin biosynthetic gene cluster of the plant pathogenic fungus Leptosphaeria maculans.
Gardiner, Donald M; Cozijnsen, Anton J; Wilson, Leanne M; Pedras, M Soledade C; Howlett, Barbara J
2004-09-01
Sirodesmin PL is a phytotoxin produced by the fungus Leptosphaeria maculans, which causes blackleg disease of canola (Brassica napus). This phytotoxin belongs to the epipolythiodioxopiperazine (ETP) class of toxins produced by fungi including mammalian and plant pathogens. We report the cloning of a cluster of genes with predicted roles in the biosynthesis of sirodesmin PL and show via gene disruption that one of these genes (encoding a two-module non-ribosomal peptide synthetase) is essential for sirodesmin PL biosynthesis. Of the nine genes in the cluster tested, all are co-regulated with the production of sirodesmin PL in culture. A similar cluster is present in the genome of the opportunistic human pathogen Aspergillus fumigatus and is most likely responsible for the production of gliotoxin, which is also an ETP. Homologues of the genes in the cluster were also identified in expressed sequence tags of the ETP producing fungus Chaetomium globosum. Two other fungi with publicly available genome sequences, Magnaporthe grisea and Fusarium graminearum, had similar gene clusters. A comparative analysis of all four clusters is presented. This is the first report of the genes responsible for the biosynthesis of an ETP. Copyright 2004 Blackwell Publishing Ltd
Many nonuniversal archaeal ribosomal proteins are found in conserved gene clusters
WANG, JIACHEN; DASGUPTA, INDRANI; FOX, GEORGE E.
2009-01-01
The genomic associations of the archaeal ribosomal proteins, (r-proteins), were examined in detail. The archaeal versions of the universal r-protein genes are typically in clusters similar or identical and to those found in bacteria. Of the 35 nonuniversal archaeal r-protein genes examined, the gene encoding L18e was found to be associated with the conserved L13 cluster, whereas the genes for S4e, L32e and L19e were found in the archaeal version of the spc operon. Eleven nonuniversal protein genes were not associated with any common genomic context. Of the remaining 19 protein genes, 17 were convincingly assigned to one of 10 previously unrecognized gene clusters. Examination of the gene content of these clusters revealed multiple associations with genes involved in the initiation of protein synthesis, transcription or other cellular processes. The lack of such associations in the universal clusters suggests that initially the ribosome evolved largely independently of other processes. More recently it likely has evolved in concert with other cellular systems. It was also verified that a second copy of the gene encoding L7ae found in some bacteria is actually a homolog of the gene encoding L30e and should be annotated as such. PMID:19478915
Association with AflR in Endosomes Reveals New Functions for AflJ in Aflatoxin Biosynthesis
Ehrlich, Kenneth C.; Mack, Brian M.; Wei, Qijian; Li, Ping; Roze, Ludmila V.; Dazzo, Frank; Cary, Jeffrey W.; Bhatnagar, Deepak; Linz, John E.
2012-01-01
Aflatoxins are the most potent naturally occurring carcinogens of fungal origin. Biosynthesis of aflatoxin involves the coordinated expression of more than 25 genes. The function of one gene in the aflatoxin gene cluster, aflJ, is not entirely understood but, because previous studies demonstrated a physical interaction between the Zn2Cys6 transcription factor AflR and AflJ, AflJ was proposed to act as a transcriptional co-activator. Image analysis revealed that, in the absence of aflJ in A. parasiticus, endosomes cluster within cells and near septa. AflJ fused to yellow fluorescent protein complemented the mutation in A. parasiticus ΔaflJ and localized mainly in endosomes. We found that AflJ co-localizes with AflR both in endosomes and in nuclei. Chromatin immunoprecipitation did not detect AflJ binding at known AflR DNA recognition sites suggesting that AflJ either does not bind to these sites or binds to them transiently. Based on these data, we hypothesize that AflJ assists in AflR transport to or from the nucleus, thus controlling the availability of AflR for transcriptional activation of aflatoxin biosynthesis cluster genes. AflJ may also assist in directing endosomes to the cytoplasmic membrane for aflatoxin export. PMID:23342682
Chai, Xiaoqiang; Han, Yanan; Yang, Jian; Zhao, Xianxian; Liu, Yewang; Hou, Xugang; Tang, Yiheng; Zhao, Shirong; Li, Xiao
2016-02-01
The molecular pathogenesis of infection by hepatitis B virus with human is extremely complex and heterogeneous. To date the molecular information is not clearly defined despite intensive research efforts. Thus, studies aimed at transcription and regulation during virus infection or combined researches of those already known to be beneficial are needed. With the purpose of identifying the transcriptional regulators related to infection of hepatitis B virus in gene level, the gene expression profiles from some normal individuals and hepatitis B patients were analyzed in our study. In this work, the differential expressed genes were selected primarily. The several genes among those were validated in an independent set by qRT-PCR. Then the differentially co-expression analysis was conducted to identify differentially co-expressed links and differential co-expressed genes. Next, the analysis of the regulatory impact factors was performed through mapping the links and regulatory data. In order to give a further insight to these regulators, the co-expression gene modules were identified using a threshold-based hierarchical clustering method. Incidentally, the construction of the regulatory network was generated using the computer software. A total of 137,284 differentially co-expressed links and 780 differential co-expressed genes were identified. These co-expressed genes were significantly enriched inflammatory response. The results of regulatory impact factors revealed several crucial regulators related to hepatocellular carcinoma and other high-rank regulators. Meanwhile, more than one hundred co-expression gene modules were identified using clustering method. In our study, some important transcriptional regulators were identified using a computational method, which may enhance the understanding of disease mechanisms and lead to an improved treatment of hepatitis B. However, further experimental studies are required to confirm these findings. Copyright © 2015 Elsevier Masson SAS. All rights reserved.
Allcock, Richard J N; Barrow, Alexander D; Forbes, Simon; Beck, Stephan; Trowsdale, John
2003-02-01
We have characterized a cluster of single immunoglobulin variable (IgV) domain receptors centromeric of the major histocompatibility complex (MHC) on human chromosome 6. In addition to triggering receptor expressed on myeloid cells (TREM)-1 and TREM2, the cluster contains NKp44, a triggering receptor whose expression is limited to NK cells. We identified three new related genes and two gene fragments within a cluster of approximately 200 kb. Two of the three new genes lack charged residues in their transmembrane domain tails. Further, one of the genes contains two potential immunotyrosine Inhibitory motifs in its cytoplasmic tail, suggesting that it delivers inhibitory signals. The human and mouse TREM clusters appear to have diverged such that there are unique sequences in each species. Finally, each gene in the TREM cluster was expressed in a different range of cell types.
Susca, Antonia; Proctor, Robert H; Butchko, Robert A E; Haidukowski, Miriam; Stea, Gaetano; Logrieco, Antonio; Moretti, Antonio
2014-12-01
The ability to produce fumonisin mycotoxins varies among members of the black aspergilli. Previously, analyses of selected genes in the fumonisin biosynthetic gene (fum) cluster in black aspergilli from California grapes indicated that fumonisin-nonproducing isolates of Aspergillus welwitschiae lack six fum genes, but nonproducing isolates of Aspergillus niger do not. In the current study, analyses of black aspergilli from grapes from the Mediterranean Basin indicate that the genomic context of the fum cluster is the same in isolates of A. niger and A. welwitschiae regardless of fumonisin-production ability and that full-length clusters occur in producing isolates of both species and nonproducing isolates of A. niger. In contrast, the cluster has undergone an eight-gene deletion in fumonisin-nonproducing isolates of A. welwitschiae. Phylogenetic analyses suggest each species consists of a mixed population of fumonisin-producing and nonproducing individuals, and that existence of both production phenotypes may provide a selective advantage to these species. Differences in gene content of fum cluster homologues and phylogenetic relationships of fum genes suggest that the mutation(s) responsible for the nonproduction phenotype differs, and therefore arose independently, in the two species. Partial fum cluster homologues were also identified in genome sequences of four other black Aspergillus species. Gene content of these partial clusters and phylogenetic relationships of fum sequences indicate that non-random partial deletion of the cluster has occurred multiple times among the species. This in turn suggests that an intact cluster and fumonisin production were once more widespread among black aspergilli. Copyright © 2014 Elsevier Inc. All rights reserved.
Lependina, I N; Churnosov, M I; Artamentova, L A; Ishchuk, M A; Tegako, O V; Balanovskaia, E V
2008-04-01
The characteristics of the gene pools of indigenous populations of Ukraine and Belarus have been studied using 28 alleles of 10 loci of biochemical gene markers (HP, GC, TF, PI, C'3, ACP1, GLO1, PGM1, ESD, and 6-PGD). The gene pools of the Russian and Ukrainian indigenous populations of Belgorod oblast (Russia) and the indigenous populations of Ukraine and Belarus have been compared. Cluster analysis, multidimensional scaling, and factor analysis of the obtained data have been used to determine the position of the Belgorod population gene pool in the Eastern Slavic gene pool system.
The Dynamics of Transcript Abundance during Cellularization of Developing Barley Endosperm1[OPEN
Zhang, Runxuan; Burton, Rachel A; Shirley, Neil J.; Little, Alan; Morris, Jenny; Milne, Linda
2016-01-01
Within the cereal grain, the endosperm and its nutrient reserves are critical for successful germination and in the context of grain utilization. The identification of molecular determinants of early endosperm development, particularly regulators of cell division and cell wall deposition, would help predict end-use properties such as yield, quality, and nutritional value. Custom microarray data have been generated using RNA isolated from developing barley grain endosperm 3 d to 8 d after pollination (DAP). Comparisons of transcript abundance over time revealed 47 gene expression modules that can be clustered into 10 broad groups. Superimposing these modules upon cytological data allowed patterns of transcript abundance to be linked with key stages of early grain development. Here, attention was focused on how the datasets could be mined to explore and define the processes of cell wall biosynthesis, remodeling, and degradation. Using a combination of spatial molecular network and gene ontology enrichment analyses, it is shown that genes involved in cell wall metabolism are found in multiple modules, but cluster into two main groups that exhibit peak expression at 3 DAP to 4 DAP and 5 DAP to 8 DAP. The presence of transcription factor genes in these modules allowed candidate genes for the control of wall metabolism during early barley grain development to be identified. The data are publicly available through a dedicated web interface (https://ics.hutton.ac.uk/barseed/), where they can be used to interrogate co- and differential expression for any other genes, groups of genes, or transcription factors expressed during early endosperm development. PMID:26754666
Miyamoto, Kiyoko T.; Komatsu, Mamoru
2014-01-01
Mycosporines and mycosporine-like amino acids (MAAs), including shinorine (mycosporine-glycine-serine) and porphyra-334 (mycosporine-glycine-threonine), are UV-absorbing compounds produced by cyanobacteria, fungi, and marine micro- and macroalgae. These MAAs have the ability to protect these organisms from damage by environmental UV radiation. Although no reports have described the production of MAAs and the corresponding genes involved in MAA biosynthesis from Gram-positive bacteria to date, genome mining of the Gram-positive bacterial database revealed that two microorganisms belonging to the order Actinomycetales, Actinosynnema mirum DSM 43827 and Pseudonocardia sp. strain P1, possess a gene cluster homologous to the biosynthetic gene clusters identified from cyanobacteria. When the two strains were grown in liquid culture, Pseudonocardia sp. accumulated a very small amount of MAA-like compound in a medium-dependent manner, whereas A. mirum did not produce MAAs under any culture conditions, indicating that the biosynthetic gene cluster of A. mirum was in a cryptic state in this microorganism. In order to characterize these biosynthetic gene clusters, each biosynthetic gene cluster was heterologously expressed in an engineered host, Streptomyces avermitilis SUKA22. Since the resultant transformants carrying the entire biosynthetic gene cluster controlled by an alternative promoter produced mainly shinorine, this is the first confirmation of a biosynthetic gene cluster for MAA from Gram-positive bacteria. Furthermore, S. avermitilis SUKA22 transformants carrying the biosynthetic gene cluster for MAA of A. mirum accumulated not only shinorine and porphyra-334 but also a novel MAA. Structure elucidation revealed that the novel MAA is mycosporine-glycine-alanine, which substitutes l-alanine for the l-serine of shinorine. PMID:24907338
Miyamoto, Kiyoko T; Komatsu, Mamoru; Ikeda, Haruo
2014-08-01
Mycosporines and mycosporine-like amino acids (MAAs), including shinorine (mycosporine-glycine-serine) and porphyra-334 (mycosporine-glycine-threonine), are UV-absorbing compounds produced by cyanobacteria, fungi, and marine micro- and macroalgae. These MAAs have the ability to protect these organisms from damage by environmental UV radiation. Although no reports have described the production of MAAs and the corresponding genes involved in MAA biosynthesis from Gram-positive bacteria to date, genome mining of the Gram-positive bacterial database revealed that two microorganisms belonging to the order Actinomycetales, Actinosynnema mirum DSM 43827 and Pseudonocardia sp. strain P1, possess a gene cluster homologous to the biosynthetic gene clusters identified from cyanobacteria. When the two strains were grown in liquid culture, Pseudonocardia sp. accumulated a very small amount of MAA-like compound in a medium-dependent manner, whereas A. mirum did not produce MAAs under any culture conditions, indicating that the biosynthetic gene cluster of A. mirum was in a cryptic state in this microorganism. In order to characterize these biosynthetic gene clusters, each biosynthetic gene cluster was heterologously expressed in an engineered host, Streptomyces avermitilis SUKA22. Since the resultant transformants carrying the entire biosynthetic gene cluster controlled by an alternative promoter produced mainly shinorine, this is the first confirmation of a biosynthetic gene cluster for MAA from Gram-positive bacteria. Furthermore, S. avermitilis SUKA22 transformants carrying the biosynthetic gene cluster for MAA of A. mirum accumulated not only shinorine and porphyra-334 but also a novel MAA. Structure elucidation revealed that the novel MAA is mycosporine-glycine-alanine, which substitutes l-alanine for the l-serine of shinorine. Copyright © 2014, American Society for Microbiology. All Rights Reserved.
Cardoza, R. E.; Malmierca, M. G.; Hermosa, M. R.; Alexander, N. J.; McCormick, S. P.; Proctor, R. H.; Tijerino, A. M.; Rumbero, A.; Monte, E.; Gutiérrez, S.
2011-01-01
Trichothecenes are mycotoxins produced by Trichoderma, Fusarium, and at least four other genera in the fungal order Hypocreales. Fusarium has a trichothecene biosynthetic gene (TRI) cluster that encodes transport and regulatory proteins as well as most enzymes required for the formation of the mycotoxins. However, little is known about trichothecene biosynthesis in the other genera. Here, we identify and characterize TRI gene orthologues (tri) in Trichoderma arundinaceum and Trichoderma brevicompactum. Our results indicate that both Trichoderma species have a tri cluster that consists of orthologues of seven genes present in the Fusarium TRI cluster. Organization of genes in the cluster is the same in the two Trichoderma species but differs from the organization in Fusarium. Sequence and functional analysis revealed that the gene (tri5) responsible for the first committed step in trichothecene biosynthesis is located outside the cluster in both Trichoderma species rather than inside the cluster as it is in Fusarium. Heterologous expression analysis revealed that two T. arundinaceum cluster genes (tri4 and tri11) differ in function from their Fusarium orthologues. The Tatri4-encoded enzyme catalyzes only three of the four oxygenation reactions catalyzed by the orthologous enzyme in Fusarium. The Tatri11-encoded enzyme catalyzes a completely different reaction (trichothecene C-4 hydroxylation) than the Fusarium orthologue (trichothecene C-15 hydroxylation). The results of this study indicate that although some characteristics of the tri/TRI cluster have been conserved during evolution of Trichoderma and Fusarium, the cluster has undergone marked changes, including gene loss and/or gain, gene rearrangement, and divergence of gene function. PMID:21642405
USDA-ARS?s Scientific Manuscript database
Cyclopiazonic acid (CPA), an indole-tetramic acid toxin, is produced by many species of Aspergillus and Penicillium. In addition to CPA Aspergillus flavus produces polyketide-derived carcinogenic aflatoxins (AFs). AF biosynthesis genes form a gene cluster in a subtelomeric region. Isolates of A. fla...
Dai, Zhimin; Guo, Xue; Yin, Huaqun; Liang, Yili; Cong, Jing; Liu, Xueduan
2014-01-01
Biological nitrogen fixation is an essential function of acid mine drainage (AMD) microbial communities. However, most acidophiles in AMD environments are uncultured microorganisms and little is known about the diversity of nitrogen-fixing genes and structure of nif gene cluster in AMD microbial communities. In this study, we used metagenomic sequencing to isolate nif genes in the AMD microbial community from Dexing Copper Mine, China. Meanwhile, a metagenome microarray containing 7,776 large-insertion fosmids was constructed to screen novel nif gene clusters. Metagenomic analyses revealed that 742 sequences were identified as nif genes including structural subunit genes nifH, nifD, nifK and various additional genes. The AMD community is massively dominated by the genus Acidithiobacillus. However, the phylogenetic diversity of nitrogen-fixing microorganisms is much higher than previously thought in the AMD community. Furthermore, a 32.5-kb genomic sequence harboring nif, fix and associated genes was screened by metagenome microarray. Comparative genome analysis indicated that most nif genes in this cluster are most similar to those of Herbaspirillum seropedicae, but the organization of the nif gene cluster had significant differences from H. seropedicae. Sequence analysis and reverse transcription PCR also suggested that distinct transcription units of nif genes exist in this gene cluster. nifQ gene falls into the same transcription unit with fixABCX genes, which have not been reported in other diazotrophs before. All of these results indicated that more novel diazotrophs survive in the AMD community.
Yin, Huaqun; Liang, Yili; Cong, Jing; Liu, Xueduan
2014-01-01
Biological nitrogen fixation is an essential function of acid mine drainage (AMD) microbial communities. However, most acidophiles in AMD environments are uncultured microorganisms and little is known about the diversity of nitrogen-fixing genes and structure of nif gene cluster in AMD microbial communities. In this study, we used metagenomic sequencing to isolate nif genes in the AMD microbial community from Dexing Copper Mine, China. Meanwhile, a metagenome microarray containing 7,776 large-insertion fosmids was constructed to screen novel nif gene clusters. Metagenomic analyses revealed that 742 sequences were identified as nif genes including structural subunit genes nifH, nifD, nifK and various additional genes. The AMD community is massively dominated by the genus Acidithiobacillus. However, the phylogenetic diversity of nitrogen-fixing microorganisms is much higher than previously thought in the AMD community. Furthermore, a 32.5-kb genomic sequence harboring nif, fix and associated genes was screened by metagenome microarray. Comparative genome analysis indicated that most nif genes in this cluster are most similar to those of Herbaspirillum seropedicae, but the organization of the nif gene cluster had significant differences from H. seropedicae. Sequence analysis and reverse transcription PCR also suggested that distinct transcription units of nif genes exist in this gene cluster. nifQ gene falls into the same transcription unit with fixABCX genes, which have not been reported in other diazotrophs before. All of these results indicated that more novel diazotrophs survive in the AMD community. PMID:24498417
Qin, Jin-Hong; Zhang, Qing; Zhang, Zhi-Ming; Zhong, Yi; Yang, Yang; Hu, Bao-Yu; Zhao, Guo-Ping; Guo, Xiao-Kui
2008-06-01
DNA microarray analysis was used to compare the differential gene expression profiles between Leptospira interrogans serovar Lai type strain 56601 and its corresponding attenuated strain IPAV. A 22-kb genomic island covering a cluster of 34 genes (i.e., genes LA0186 to LA0219) was actively expressed in both strains but concomitantly upregulated in strain 56601 in contrast to that of IPAV. Reverse transcription-PCR assays proved that the gene cluster comprised five transcripts. Gene annotation of this cluster revealed characteristics of a putative prophage-like remnant with at least 8 of 34 sequences encoding prophage-like proteins, of which the LA0195 protein is probably a putative prophage CI-like regulator. The transcription initiation activities of putative promoter-regulatory sequences of transcripts I, II, and III, all proximal to the LA0195 gene, were further analyzed in the Escherichia coli promoter probe vector pKK232-8 by assaying the reporter chloramphenicol acetyltransferase (CAT) activities. The strong promoter activities of both transcripts I and II indicated by the E. coli CAT assay were well correlated with the in vitro sequence-specific binding of the recombinant LA0195 protein to the corresponding promoter probes detected by the electrophoresis mobility shift assay. On the other hand, the promoter activity of transcript III was very low in E. coli and failed to show active binding to the LA0195 protein in vitro. These results suggested that the LA0195 protein is likely involved in the transcription of transcripts I and II. However, the identical complete DNA sequences of this prophage remnant from these two strains strongly suggests that possible regulatory factors or signal transduction systems residing outside of this region within the genome may be responsible for the differential expression profiling in these two strains.
Panesso, Diana; Reyes, Jinnethe; Rincón, Sandra; Díaz, Lorena; Galloway-Peña, Jessica; Zurita, Jeannete; Carrillo, Carlos; Merentes, Altagracia; Guzmán, Manuel; Adachi, Javier A.; Murray, Barbara E.; Arias, Cesar A.
2010-01-01
Enterococcus faecium has emerged as an important nosocomial pathogen worldwide, and this trend has been associated with the dissemination of a genetic lineage designated clonal cluster 17 (CC17). Enterococcal isolates were collected prospectively (2006 to 2008) from 32 hospitals in Colombia, Ecuador, Perú, and Venezuela and subjected to antimicrobial susceptibility testing. Genotyping was performed with all vancomycin-resistant E. faecium (VREfm) isolates by pulsed-field gel electrophoresis (PFGE) and multilocus sequence typing. All VREfm isolates were evaluated for the presence of 16 putative virulence genes (14 fms genes, the esp gene of E. faecium [espEfm], and the hyl gene of E. faecium [hylEfm]) and plasmids carrying the fms20-fms21 (pilA), hylEfm, and vanA genes. Of 723 enterococcal isolates recovered, E. faecalis was the most common (78%). Vancomycin resistance was detected in 6% of the isolates (74% of which were E. faecium). Eleven distinct PFGE types were found among the VREfm isolates, with most belonging to sequence types 412 and 18. The ebpAEfm-ebpBEfm-ebpCEfm (pilB) and fms11-fms19-fms16 clusters were detected in all VREfm isolates from the region, whereas espEfm and hylEfm were detected in 69% and 23% of the isolates, respectively. The fms20-fms21 (pilA) cluster, which encodes a putative pilus-like protein, was found on plasmids from almost all VREfm isolates and was sometimes found to coexist with hylEfm and the vanA gene cluster. The population genetics of VREfm in South America appear to resemble those of such strains in the United States in the early years of the CC17 epidemic. The overwhelming presence of plasmids encoding putative virulence factors and vanA genes suggests that E. faecium from the CC17 genogroup may disseminate in the region in the coming years. PMID:20220167
Duronio, Robert J.; Marzluff, William F.
2017-01-01
ABSTRACT Metazoan replication-dependent (RD) histone genes encode the only known cellular mRNAs that are not polyadenylated. These mRNAs end instead in a conserved stem-loop, which is formed by an endonucleolytic cleavage of the pre-mRNA. The genes for all 5 histone proteins are clustered in all metazoans and coordinately regulated with high levels of expression during S phase. Production of histone mRNAs occurs in a nuclear body called the Histone Locus Body (HLB), a subdomain of the nucleus defined by a concentration of factors necessary for histone gene transcription and pre-mRNA processing. These factors include the scaffolding protein NPAT, essential for histone gene transcription, and FLASH and U7 snRNP, both essential for histone pre-mRNA processing. Histone gene expression is activated by Cyclin E/Cdk2-mediated phosphorylation of NPAT at the G1-S transition. The concentration of factors within the HLB couples transcription with pre-mRNA processing, enhancing the efficiency of histone mRNA biosynthesis. PMID:28059623
BFDCA: A Comprehensive Tool of Using Bayes Factor for Differential Co-Expression Analysis.
Wang, Duolin; Wang, Juexin; Jiang, Yuexu; Liang, Yanchun; Xu, Dong
2017-02-03
Comparing the gene-expression profiles between biological conditions is useful for understanding gene regulation underlying complex phenotypes. Along this line, analysis of differential co-expression (DC) has gained attention in the recent years, where genes under one condition have different co-expression patterns compared with another. We developed an R package Bayes Factor approach for Differential Co-expression Analysis (BFDCA) for DC analysis. BFDCA is unique in integrating various aspects of DC patterns (including Shift, Cross, and Re-wiring) into one uniform Bayes factor. We tested BFDCA using simulation data and experimental data. Simulation results indicate that BFDCA outperforms existing methods in accuracy and robustness of detecting DC pairs and DC modules. Results of using experimental data suggest that BFDCA can cluster disease-related genes into functional DC subunits and estimate the regulatory impact of disease-related genes well. BFDCA also achieves high accuracy in predicting case-control phenotypes by using significant DC gene pairs as markers. BFDCA is publicly available at http://dx.doi.org/10.17632/jdz4vtvnm3.1. Copyright © 2016 Elsevier Ltd. All rights reserved.
Sood, Archit; Jaiswal, Varun; Chanumolu, Sree Krishna; Malhotra, Nikhil; Pal, Tarun; Chauhan, Rajinder Singh
2014-11-01
Jatropha (Jatropha curcas L.) and Castor bean (Ricinus communis) are oilseed crops of family Euphorbiaceae with the potential of producing high quality biodiesel and having industrial value. Both the bioenergy plants are becoming susceptible to various biotic stresses directly affecting the oil quality and content. No report exists as of today on analysis of Nucleotide Binding Site-Leucine Rich Repeat (NBS-LRR) gene repertoire and defense response transcription factors in both the plant species. In silico analysis of whole genomes and transcriptomes identified 47 new NBS-LRR genes in both the species and 122 and 318 defense response related transcription factors in Jatropha and Castor bean, respectively. The identified NBS-LRR genes and defense response transcription factors were mapped onto the respective genomes. Common and unique NBS-LRR genes and defense related transcription factors were identified in both the plant species. All NBS-LRR genes in both the species were characterized into Toll/interleukin-1 receptor NBS-LRRs (TNLs) and coiled-coil NBS-LRRs (CNLs), position on contigs, gene clusters and motifs and domains distribution. Transcript abundance or expression values were measured for all NBS-LRR genes and defense response transcription factors, suggesting their functional role. The current study provides a repertoire of NBS-LRR genes and transcription factors which can be used in not only dissecting the molecular basis of disease resistance phenotype but also in developing disease resistant genotypes in Jatropha and Castor bean through transgenic or molecular breeding approaches.
Sekigami, Yuka; Kobayashi, Takuya; Omi, Ai; Nishitsuji, Koki; Ikuta, Tetsuro; Fujiyama, Asao; Satoh, Noriyuki; Saiga, Hidetoshi
2017-01-01
Hox gene clusters with at least 13 paralog group (PG) members are common in vertebrate genomes and in that of amphioxus. Ascidians, which belong to the subphylum Tunicata (Urochordata), are phylogenetically positioned between vertebrates and amphioxus, and traditionally divided into two groups: the Pleurogona and the Enterogona. An enterogonan ascidian, Ciona intestinalis ( Ci ), possesses nine Hox genes localized on two chromosomes; thus, the Hox gene cluster is disintegrated. We investigated the Hox gene cluster of a pleurogonan ascidian, Halocynthia roretzi ( Hr ) to investigate whether Hox gene cluster disintegration is common among ascidians, and if so, how such disintegration occurred during ascidian or tunicate evolution. Our phylogenetic analysis reveals that the Hr Hox gene complement comprises nine members, including one with a relatively divergent Hox homeodomain sequence. Eight of nine Hr Hox genes were orthologous to Ci-Hox1 , 2, 3, 4, 5, 10, 12 and 13. Following the phylogenetic classification into 13 PGs, we designated Hr Hox genes as Hox1, 2, 3, 4, 5, 10, 11/12/13.a , 11/12/13.b and HoxX . To address the chromosomal arrangement of the nine Hox genes, we performed two-color chromosomal fluorescent in situ hybridization, which revealed that the nine Hox genes are localized on a single chromosome in Hr , distinct from their arrangement in Ci . We further examined the order of the nine Hox genes on the chromosome by chromosome/scaffold walking. This analysis suggested a gene order of Hox1 , 11/12/13.b, 11/12/13.a, 10, 5, X, followed by either Hox4, 3, 2 or Hox2, 3, 4 on the chromosome. Based on the present results and those previously reported in Ci , we discuss the establishment of the Hox gene complement and disintegration of Hox gene clusters during the course of ascidian or tunicate evolution. The Hox gene cluster and the genome must have experienced extensive reorganization during the course of evolution from the ancestral tunicate to Hr and Ci . Nevertheless, some features are shared in Hox gene components and gene arrangement on the chromosomes, suggesting that Hox gene cluster disintegration in ascidians involved early events common to tunicates as well as later ascidian lineage-specific events.
Ehrlich, Kenneth C; Mack, Brian M
2014-06-23
Fifty six secondary metabolite biosynthesis gene clusters are predicted to be in the Aspergillus flavus genome. In spite of this, the biosyntheses of only seven metabolites, including the aflatoxins, kojic acid, cyclopiazonic acid and aflatrem, have been assigned to a particular gene cluster. We used RNA-seq to compare expression of secondary metabolite genes in gene clusters for the closely related fungi A. parasiticus, A. oryzae, and A. flavus S and L sclerotial morphotypes. The data help to refine the identification of probable functional gene clusters within these species. Our results suggest that A. flavus, a prevalent contaminant of maize, cottonseed, peanuts and tree nuts, is capable of producing metabolites which, besides aflatoxin, could be an underappreciated contributor to its toxicity.
Ehrlich, Kenneth C.; Mack, Brian M.
2014-01-01
Fifty six secondary metabolite biosynthesis gene clusters are predicted to be in the Aspergillus flavus genome. In spite of this, the biosyntheses of only seven metabolites, including the aflatoxins, kojic acid, cyclopiazonic acid and aflatrem, have been assigned to a particular gene cluster. We used RNA-seq to compare expression of secondary metabolite genes in gene clusters for the closely related fungi A. parasiticus, A. oryzae, and A. flavus S and L sclerotial morphotypes. The data help to refine the identification of probable functional gene clusters within these species. Our results suggest that A. flavus, a prevalent contaminant of maize, cottonseed, peanuts and tree nuts, is capable of producing metabolites which, besides aflatoxin, could be an underappreciated contributor to its toxicity. PMID:24960201
Gene Cluster Encoding Cholate Catabolism in Rhodococcus spp.
Wilbrink, Maarten H.; Casabon, Israël; Stewart, Gordon R.; Liu, Jie; van der Geize, Robert; Eltis, Lindsay D.
2012-01-01
Bile acids are highly abundant steroids with important functions in vertebrate digestion. Their catabolism by bacteria is an important component of the carbon cycle, contributes to gut ecology, and has potential commercial applications. We found that Rhodococcus jostii RHA1 grows well on cholate, as well as on its conjugates, taurocholate and glycocholate. The transcriptome of RHA1 growing on cholate revealed 39 genes upregulated on cholate, occurring in a single gene cluster. Reverse transcriptase quantitative PCR confirmed that selected genes in the cluster were upregulated 10-fold on cholate versus on cholesterol. One of these genes, kshA3, encoding a putative 3-ketosteroid-9α-hydroxylase, was deleted and found essential for growth on cholate. Two coenzyme A (CoA) synthetases encoded in the cluster, CasG and CasI, were heterologously expressed. CasG was shown to transform cholate to cholyl-CoA, thus initiating side chain degradation. CasI was shown to form CoA derivatives of steroids with isopropanoyl side chains, likely occurring as degradation intermediates. Orthologous gene clusters were identified in all available Rhodococcus genomes, as well as that of Thermomonospora curvata. Moreover, Rhodococcus equi 103S, Rhodococcus ruber Chol-4 and Rhodococcus erythropolis SQ1 each grew on cholate. In contrast, several mycolic acid bacteria lacking the gene cluster were unable to grow on cholate. Our results demonstrate that the above-mentioned gene cluster encodes cholate catabolism and is distinct from a more widely occurring gene cluster encoding cholesterol catabolism. PMID:23024343
Aguiar, Bruno; Vieira, Jorge; Cunha, Ana E.; Fonseca, Nuno A.; Reboiro-Jato, David; Reboiro-Jato, Miguel; Fdez-Riverola, Florentino; Raspé, Olivier; Vieira, Cristina P.
2013-01-01
S-RNase-based gametophytic self-incompatibility evolved once before the split of the Asteridae and Rosidae. In Prunus (tribe Amygdaloideae of Rosaceae), the self-incompatibility S-pollen is a single F-box gene that presents the expected evolutionary signatures. In Malus and Pyrus (subtribe Pyrinae of Rosaceae), however, clusters of F-box genes (called SFBBs) have been described that are expressed in pollen only and are linked to the S-RNase gene. Although polymorphic, SFBB genes present levels of diversity lower than those of the S-RNase gene. They have been suggested as putative S-pollen genes, in a system of non-self recognition by multiple factors. Subsets of allelic products of the different SFBB genes interact with non-self S-RNases, marking them for degradation, and allowing compatible pollinations. This study performed a detailed characterization of SFBB genes in Sorbus aucuparia (Pyrinae) to address three predictions of the non-self recognition by multiple factors model. As predicted, the number of SFBB genes was large to account for the many S-RNase specificities. Secondly, like the S-RNase gene, the SFBB genes were old. Thirdly, amino acids under positive selection—those that could be involved in specificity determination—were identified when intra-haplotype SFBB genes were analysed using codon models. Overall, the findings reported here support the non-self recognition by multiple factors model. PMID:23606363
Lee, Sunhee; Reth, Alexander; Meletzus, Dietmar; Sevilla, Myrna; Kennedy, Christina
2000-01-01
A major 30.5-kb cluster of nif and associated genes of Acetobacter diazotrophicus (syn. Gluconacetobacter diazotrophicus), a nitrogen-fixing endophyte of sugarcane, was sequenced and analyzed. This cluster represents the largest assembly of contiguous nif-fix and associated genes so far characterized in any diazotrophic bacterial species. Northern blots and promoter sequence analysis indicated that the genes are organized into eight transcriptional units. The overall arrangement of genes is most like that of the nif-fix cluster in Azospirillum brasilense, while the individual gene products are more similar to those in species of Rhizobiaceae or in Rhodobacter capsulatus. PMID:11092875
Malviya, N; Gupta, S; Singh, V K; Yadav, M K; Bisht, N C; Sarangi, B K; Yadav, D
2015-02-01
The DNA binding with One Finger (Dof) protein is a plant specific transcription factor involved in the regulation of wide range of processes. The analysis of whole genome sequence of pigeonpea has identified 38 putative Dof genes (CcDof) distributed on 8 chromosomes. A total of 17 out of 38 CcDof genes were found to be intronless. A comprehensive in silico characterization of CcDof gene family including the gene structure, chromosome location, protein motif, phylogeny, gene duplication and functional divergence has been attempted. The phylogenetic analysis resulted in 3 major clusters with closely related members in phylogenetic tree revealed common motif distribution. The in silico cis-regulatory element analysis revealed functional diversity with predominance of light responsive and stress responsive elements indicating the possibility of these CcDof genes to be associated with photoperiodic control and biotic and abiotic stress. The duplication pattern showed that tandem duplication is predominant over segmental duplication events. The comparative phylogenetic analysis of these Dof proteins along with 78 soybean, 36 Arabidopsis and 30 rice Dof proteins revealed 7 major clusters. Several groups of orthologs and paralogs were identified based on phylogenetic tree constructed. Our study provides useful information for functional characterization of CcDof genes.
Athey, Taryn B T; Vaillancourt, Katy; Frenette, Michel; Fittipaldi, Nahuel; Gottschalk, Marcelo; Grenier, Daniel
2016-01-01
Recently, we reported the purification and characterization of three distinct lantibiotics (named suicin 90-1330, suicin 3908, and suicin 65) produced by Streptococcus suis . In this study, we investigated the distribution of the three suicin lantibiotic gene clusters among serotype 2 S. suis strains belonging to sequence type (ST) 25 and ST28, the two dominant STs identified in North America. The genomes of 102 strains were interrogated for the presence of suicin gene clusters encoding suicins 90-1330, 3908, and 65. The gene cluster encoding suicin 65 was the most prevalent and mainly found among ST25 strains. In contrast, none of the genes related to suicin 90-1330 production were identified in 51 ST25 strains nor in 35/51 ST28 strains. However, the complete suicin 90-1330 gene cluster was found in ten ST28 strains, although some genes in the cluster were truncated in three of these isolates. The vast majority (101/102) of S. suis strains did not possess any of the genes encoding suicin 3908. In conclusion, this study indicates heterogeneous distribution of suicin genes in S. suis .
Heart morphogenesis gene regulatory networks revealed by temporal expression analysis.
Hill, Jonathon T; Demarest, Bradley; Gorsi, Bushra; Smith, Megan; Yost, H Joseph
2017-10-01
During embryogenesis the heart forms as a linear tube that then undergoes multiple simultaneous morphogenetic events to obtain its mature shape. To understand the gene regulatory networks (GRNs) driving this phase of heart development, during which many congenital heart disease malformations likely arise, we conducted an RNA-seq timecourse in zebrafish from 30 hpf to 72 hpf and identified 5861 genes with altered expression. We clustered the genes by temporal expression pattern, identified transcription factor binding motifs enriched in each cluster, and generated a model GRN for the major gene batteries in heart morphogenesis. This approach predicted hundreds of regulatory interactions and found batteries enriched in specific cell and tissue types, indicating that the approach can be used to narrow the search for novel genetic markers and regulatory interactions. Subsequent analyses confirmed the GRN using two mutants, Tbx5 and nkx2-5 , and identified sets of duplicated zebrafish genes that do not show temporal subfunctionalization. This dataset provides an essential resource for future studies on the genetic/epigenetic pathways implicated in congenital heart defects and the mechanisms of cardiac transcriptional regulation. © 2017. Published by The Company of Biologists Ltd.
Bacterium induces cryptic meroterpenoid pathway in the pathogenic fungus Aspergillus fumigatus.
König, Claudia C; Scherlach, Kirstin; Schroeckh, Volker; Horn, Fabian; Nietzsche, Sandor; Brakhage, Axel A; Hertweck, Christian
2013-05-27
Stimulating encounter: The intimate, physical interaction between the soil-derived bacterium Streptomyces rapamycinicus and the human pathogenic fungus Aspergillus fumigatus led to the activation of an otherwise silent polyketide synthase (PKS) gene cluster coding for an unusual prenylated polyphenol (fumicycline A). The meroterpenoid pathway is regulated by a pathway-specific activator gene as well as by epigenetic factors. Copyright © 2013 WILEY-VCH Verlag GmbH & Co. KGaA, Weinheim.
Mihali, Troco K.; Carmichael, Wayne W.; Neilan, Brett A.
2011-01-01
Saxitoxin and its analogs cause the paralytic shellfish-poisoning syndrome, adversely affecting human health and coastal shellfish industries worldwide. Here we report the isolation, sequencing, annotation, and predicted pathway of the saxitoxin biosynthetic gene cluster in the cyanobacterium Lyngbya wollei. The gene cluster spans 36 kb and encodes enzymes for the biosynthesis and export of the toxins. The Lyngbya wollei saxitoxin gene cluster differs from previously identified saxitoxin clusters as it contains genes that are unique to this cluster, whereby the carbamoyltransferase is truncated and replaced by an acyltransferase, explaining the unique toxin profile presented by Lyngbya wollei. These findings will enable the creation of toxin probes, for water monitoring purposes, as well as proof-of-concept for the combinatorial biosynthesis of these natural occurring alkaloids for the production of novel, biologically active compounds. PMID:21347365
Merino-Puerto, Victoria; Herrero, Antonia
2013-01-01
The filamentous, heterocyst-forming cyanobacteria perform oxygenic photosynthesis in vegetative cells and nitrogen fixation in heterocysts, and their filaments can be hundreds of cells long. In the model heterocyst-forming cyanobacterium Anabaena sp. strain PCC 7120, the genes in the fraC-fraD-fraE operon are required for filament integrity mainly under conditions of nitrogen deprivation. The fraC operon transcript partially overlaps gene all2395, which lies in the opposite DNA strand and ends 1 bp beyond fraE. Gene all2395 produces transcripts of 1.35 kb (major transcript) and 2.2 kb (minor transcript) that overlap fraE and whose expression is dependent on the N-control transcription factor NtcA. Insertion of a gene cassette containing transcriptional terminators between fraE and all2395 prevented production of the antisense RNAs and resulted in an increased length of the cyanobacterial filaments. Deletion of all2395 resulted in a larger increase of filament length and in impaired growth, mainly under N2-fixing conditions and specifically on solid medium. We denote all2395 the fraF gene, which encodes a protein restricting filament length. A FraF-green fluorescent protein (GFP) fusion protein accumulated significantly in heterocysts. Similar to some heterocyst differentiation-related proteins such as HglK, HetL, and PatL, FraF is a pentapeptide repeat protein. We conclude that the fraC-fraD-fraE←fraF gene cluster (where the arrow indicates a change in orientation), in which cis antisense RNAs are produced, regulates morphology by encoding proteins that influence positively (FraC, FraD, FraE) or negatively (FraF) the length of the filament mainly under conditions of nitrogen deprivation. This gene cluster is often conserved in heterocyst-forming cyanobacteria. PMID:23813733
Vital, Marius; Gao, Jiarong; Rizzo, Mike; Harrison, Tara; Tiedje, James M
2015-01-01
Butyrate-producing bacteria have an important role in maintaining host health. They are well studied in human and medically associated animal models; however, much less is known for other Vertebrata. We investigated the butyrate-producing community in hindgut-fermenting Mammalia (n=38), Aves (n=8) and Reptilia (n=8) using a gene-targeted pyrosequencing approach of the terminal genes of the main butyrate-synthesis pathways, namely butyryl-CoA:acetate CoA-transferase (but) and butyrate kinase (buk). Most animals exhibit high gene abundances, and clear diet-specific signatures were detected with but genes significantly enriched in omnivores and herbivores compared with carnivores. But dominated the butyrate-producing community in these two groups, whereas buk was more abundant in many carnivorous animals. Clustering of protein sequences (5% cutoff) of the combined communities (but and buk) placed carnivores apart from other diet groups, except for noncarnivorous Carnivora, which clustered together with carnivores. The majority of clusters (but: 5141 and buk: 2924) did not show close relation to any reference sequences from public databases (identity <90%) demonstrating a large ‘unknown diversity'. Each diet group had abundant signature taxa, where buk genes linked to Clostridium perfringens dominated in carnivores and but genes associated with Ruminococcaceae bacterium D16 were specific for herbivores and omnivores. Whereas 16S rRNA gene analysis showed similar overall patterns, it was unable to reveal communities at the same depth and resolution as the functional gene-targeted approach. This study demonstrates that butyrate producers are abundant across vertebrates exhibiting great functional redundancy and that diet is the primary determinant governing the composition of the butyrate-producing guild. PMID:25343515
Vital, Marius; Gao, Jiarong; Rizzo, Mike; Harrison, Tara; Tiedje, James M
2015-03-17
Butyrate-producing bacteria have an important role in maintaining host health. They are well studied in human and medically associated animal models; however, much less is known for other Vertebrata. We investigated the butyrate-producing community in hindgut-fermenting Mammalia (n = 38), Aves (n = 8) and Reptilia (n = 8) using a gene-targeted pyrosequencing approach of the terminal genes of the main butyrate-synthesis pathways, namely butyryl-CoA:acetate CoA-transferase (but) and butyrate kinase (buk). Most animals exhibit high gene abundances, and clear diet-specific signatures were detected with but genes significantly enriched in omnivores and herbivores compared with carnivores. But dominated the butyrate-producing community in these two groups, whereas buk was more abundant in many carnivorous animals. Clustering of protein sequences (5% cutoff) of the combined communities (but and buk) placed carnivores apart from other diet groups, except for noncarnivorous Carnivora, which clustered together with carnivores. The majority of clusters (but: 5141 and buk: 2924) did not show close relation to any reference sequences from public databases (identity <90%) demonstrating a large 'unknown diversity'. Each diet group had abundant signature taxa, where buk genes linked to Clostridium perfringens dominated in carnivores and but genes associated with Ruminococcaceae bacterium D16 were specific for herbivores and omnivores. Whereas 16S rRNA gene analysis showed similar overall patterns, it was unable to reveal communities at the same depth and resolution as the functional gene-targeted approach. This study demonstrates that butyrate producers are abundant across vertebrates exhibiting great functional redundancy and that diet is the primary determinant governing the composition of the butyrate-producing guild.
Wu, Mengmeng; Huang, Haidong; Li, Guoqiang; Ren, Yi; Shi, Zhong; Li, Xiaoyan; Dai, Xiaohui; Gao, Ge; Ren, Mengnan; Ma, Ting
2017-04-21
Although clustering of genes from the same metabolic pathway is a widespread phenomenon, the evolution of the polysaccharide biosynthetic gene cluster remains poorly understood. To determine the evolution of this pathway, we identified a scattered production pathway of the polysaccharide sanxan by Sphingomonas sanxanigenens NX02, and compared the distribution of genes between sphingan-producing and other Sphingomonadaceae strains. This allowed us to determine how the scattered sanxan pathway developed, and how the polysaccharide gene cluster evolved. Our findings suggested that the evolution of microbial polysaccharide biosynthesis gene clusters is a lengthy cyclic process comprising cluster 1 → scatter → cluster 2. The sanxan biosynthetic pathway proved the existence of a dispersive process. We also report the complete genome sequence of NX02, in which we identified many unstable genetic elements and powerful secretion systems. Furthermore, nine enzymes for the formation of activated precursors, four glycosyltransferases, four acyltransferases, and four polymerization and export proteins were identified. These genes were scattered in the NX02 genome, and the positive regulator SpnA of sphingans synthesis could not regulate sanxan production. Finally, we concluded that the evolution of the sanxan pathway was independent. NX02 evolved naturally as a polysaccharide producing strain over a long-time evolution involving gene acquisitions and adaptive mutations.
Dattenböck, Christoph; Tisch, Doris; Schuster, Andre; Monroy, Alberto Alonso; Hinterdobler, Wolfgang; Schmoll, Monika
2018-01-01
Trichoderma reesei is one of the most frequently used filamentous fungi in industry for production of homologous and heterologous proteins. The ability to use sexual crossing in this fungus was discovered several years ago and opens up new perspectives for industrial strain improvement and investigation of gene regulation. Here we investigated the female sterile strain QM6a in comparison to the fertile isolate CBS999.97 and backcrossed derivatives of QM6a, which have regained fertility (FF1 and FF2 strains) in both mating types under conditions of sexual development. We found considerable differences in gene regulation between strains with the CBS999.97 genetic background and the QM6a background. Regulation patterns of QM6a largely clustered with the backcrossed FF1 and FF2 strains. Differential regulation between QM6a and FF1/FF2 as well as clustering of QM6a patterns with those of CBS999.97 strains was also observed. Consistent mating type dependent regulation was limited to mating type genes and those involved in pheromone response, but included also nta1 encoding a putative N-terminal amidase previously not associated with development. Comparison of female sterile QM6a with female fertile strains showed differential expression in genes encoding several transcription factors, metabolic genes and genes involved in secondary metabolism. Evaluation of the functions of genes specifically regulated under conditions of sexual development and of genes with highest levels of transcripts under these conditions indicated a relevance of secondary metabolism for sexual development in T. reesei . Among others, the biosynthetic genes of the recently characterized SOR cluster are in this gene group. However, these genes are not essential for sexual development, but rather have a function in protection and defence against competitors during reproduction.
Roelen, Bernard A J; de Graaff, Wim; Forlani, Sylvie; Deschamps, Jacqueline
2002-11-01
The molecular mechanism underlying the 3' to 5' polarity of induction of mouse Hox genes is still elusive. While relief from a cluster-encompassing repression was shown to lead to all Hoxd genes being expressed like the 3'most of them, Hoxd1 (Kondo and Duboule, 1999), the molecular basis of initial activation of this 3'most gene, is not understood yet. We show that, already before primitive streak formation, prior to initial expression of the first Hox gene, a dramatic transcriptional stimulation of the 3'most genes, Hoxb1 and Hoxb2, is observed upon a short pulse of exogenous retinoic acid (RA), whereas it is not in the case for more 5', cluster-internal, RA-responsive Hoxb genes. In contrast, the RA-responding Hoxb1lacZ transgene that faithfully mimics the endogenous gene (Marshall et al., 1994) did not exhibit the sensitivity of Hoxb1 to precocious activation. We conclude that polarity in initial activation of Hoxb genes reflects a greater availability of 3'Hox genes for transcription, suggesting a pre-existing (susceptibility to) opening of the chromatin structure at the 3' extremity of the cluster. We discuss the data in the context of prevailing models involving differential chromatin opening in the directionality of clustered Hox gene transcription, and regarding the importance of the cluster context for correct timing of initial Hox gene expression.Interestingly, Cdx1 manifested the same early transcriptional availability as Hoxb1. Copyright 2002 Elsevier Science Ireland Ltd.
Gasmi, Najla; Jacques, Pierre-Etienne; Klimova, Natalia; Guo, Xiao; Ricciardi, Alessandra; Robert, François; Turcotte, Bernard
2014-10-01
In the yeast Saccharomyces cerevisiae, fermentation is the major pathway for energy production, even under aerobic conditions. However, when glucose becomes scarce, ethanol produced during fermentation is used as a carbon source, requiring a shift to respiration. This adaptation results in massive reprogramming of gene expression. Increased expression of genes for gluconeogenesis and the glyoxylate cycle is observed upon a shift to ethanol and, conversely, expression of some fermentation genes is reduced. The zinc cluster proteins Cat8, Sip4, and Rds2, as well as Adr1, have been shown to mediate this reprogramming of gene expression. In this study, we have characterized the gene YBR239C encoding a putative zinc cluster protein and it was named ERT1 (ethanol regulated transcription factor 1). ChIP-chip analysis showed that Ert1 binds to a limited number of targets in the presence of glucose. The strongest enrichment was observed at the promoter of PCK1 encoding an important gluconeogenic enzyme. With ethanol as the carbon source, enrichment was observed with many additional genes involved in gluconeogenesis and mitochondrial function. Use of lacZ reporters and quantitative RT-PCR analyses demonstrated that Ert1 regulates expression of its target genes in a manner that is highly redundant with other regulators of gluconeogenesis. Interestingly, in the presence of ethanol, Ert1 is a repressor of PDC1 encoding an important enzyme for fermentation. We also show that Ert1 binds directly to the PCK1 and PDC1 promoters. In summary, Ert1 is a novel factor involved in the regulation of gluconeogenesis as well as a key fermentation gene. Copyright © 2014 by the Genetics Society of America.
Coral comparative genomics reveal expanded Hox cluster in the cnidarian-bilaterian ancestor.
DuBuc, Timothy Q; Ryan, Joseph F; Shinzato, Chuya; Satoh, Nori; Martindale, Mark Q
2012-12-01
The key developmental role of the Hox cluster of genes was established prior to the last common ancestor of protostomes and deuterostomes and the subsequent evolution of this cluster has played a major role in the morphological diversity exhibited in extant bilaterians. Despite 20 years of research into cnidarian Hox genes, the nature of the cnidarian-bilaterian ancestral Hox cluster remains unclear. In an attempt to further elucidate this critical phylogenetic node, we have characterized the Hox cluster of the recently sequenced Acropora digitifera genome. The A. digitifera genome contains two anterior Hox genes (PG1 and PG2) linked to an Eve homeobox gene and an Anthox1A gene, which is thought to be either a posterior or posterior/central Hox gene. These data show that the Hox cluster of the cnidarian-bilaterian ancestor was more extensive than previously thought. The results are congruent with the existence of an ancient set of constraints on the Hox cluster and reinforce the importance of incorporating a wide range of animal species to reconstruct critical ancestral nodes.
Broad spectrum antibiotic compounds and use thereof
Koglin, Alexander; Strieker, Matthias
2016-07-05
The discovery of a non-ribosomal peptide synthetase (NRPS) gene cluster in the genome of Clostridium thermocellum that produces a secondary metabolite that is assembled outside of the host membrane is described. Also described is the identification of homologous NRPS gene clusters from several additional microorganisms. The secondary metabolites produced by the NRPS gene clusters exhibit broad spectrum antibiotic activity. Thus, antibiotic compounds produced by the NRPS gene clusters, and analogs thereof, their use for inhibiting bacterial growth, and methods of making the antibiotic compounds are described.
A scan statistic to extract causal gene clusters from case-control genome-wide rare CNV data.
Nishiyama, Takeshi; Takahashi, Kunihiko; Tango, Toshiro; Pinto, Dalila; Scherer, Stephen W; Takami, Satoshi; Kishino, Hirohisa
2011-05-26
Several statistical tests have been developed for analyzing genome-wide association data by incorporating gene pathway information in terms of gene sets. Using these methods, hundreds of gene sets are typically tested, and the tested gene sets often overlap. This overlapping greatly increases the probability of generating false positives, and the results obtained are difficult to interpret, particularly when many gene sets show statistical significance. We propose a flexible statistical framework to circumvent these problems. Inspired by spatial scan statistics for detecting clustering of disease occurrence in the field of epidemiology, we developed a scan statistic to extract disease-associated gene clusters from a whole gene pathway. Extracting one or a few significant gene clusters from a global pathway limits the overall false positive probability, which results in increased statistical power, and facilitates the interpretation of test results. In the present study, we applied our method to genome-wide association data for rare copy-number variations, which have been strongly implicated in common diseases. Application of our method to a simulated dataset demonstrated the high accuracy of this method in detecting disease-associated gene clusters in a whole gene pathway. The scan statistic approach proposed here shows a high level of accuracy in detecting gene clusters in a whole gene pathway. This study has provided a sound statistical framework for analyzing genome-wide rare CNV data by incorporating topological information on the gene pathway.
Haakensen, Vilde D; Lingjaerde, Ole Christian; Lüders, Torben; Riis, Margit; Prat, Aleix; Troester, Melissa A; Holmen, Marit M; Frantzen, Jan Ole; Romundstad, Linda; Navjord, Dina; Bukholm, Ida K; Johannesen, Tom B; Perou, Charles M; Ursin, Giske; Kristensen, Vessela N; Børresen-Dale, Anne-Lise; Helland, Aslaug
2011-11-01
Increased understanding of the variability in normal breast biology will enable us to identify mechanisms of breast cancer initiation and the origin of different subtypes, and to better predict breast cancer risk. Gene expression patterns in breast biopsies from 79 healthy women referred to breast diagnostic centers in Norway were explored by unsupervised hierarchical clustering and supervised analyses, such as gene set enrichment analysis and gene ontology analysis and comparison with previously published genelists and independent datasets. Unsupervised hierarchical clustering identified two separate clusters of normal breast tissue based on gene-expression profiling, regardless of clustering algorithm and gene filtering used. Comparison of the expression profile of the two clusters with several published gene lists describing breast cells revealed that the samples in cluster 1 share characteristics with stromal cells and stem cells, and to a certain degree with mesenchymal cells and myoepithelial cells. The samples in cluster 1 also share many features with the newly identified claudin-low breast cancer intrinsic subtype, which also shows characteristics of stromal and stem cells. More women belonging to cluster 1 have a family history of breast cancer and there is a slight overrepresentation of nulliparous women in cluster 1. Similar findings were seen in a separate dataset consisting of histologically normal tissue from both breasts harboring breast cancer and from mammoplasty reductions. This is the first study to explore the variability of gene expression patterns in whole biopsies from normal breasts and identified distinct subtypes of normal breast tissue. Further studies are needed to determine the specific cell contribution to the variation in the biology of normal breasts, how the clusters identified relate to breast cancer risk and their possible link to the origin of the different molecular subtypes of breast cancer.
[Genome-wide identification and analysis of WRKY transcription factors in Medicago truncatula].
Song, Hui; Nan, Zhibiao
2014-02-01
WRKY gene family plays important roles in plant by involving in transcriptional regulations during various physiologically processes such as development, metabolism and responses to biotic and abiotic stresses. WRKY genes have been identified in various plants. However, only few WRKY genes in Medicago truncatula have been identified with systematic analysis and comparison. In this study, we identified 93 WRKY genes through analyses of M. truncatula genome. These genes include 19 type-I genes, 49 type II genes and 13 type-III genes, and 12 non-regular type genes. All of these genes were characterized through analyses of gene duplication, chromosomal locations, structural diversity, conserved protein motifs and phylogenetic relations. The results showed that 11 times of gene duplication event occurred in WRKY gene family involving 24 genes. WRKY genes, containing 6 gene clusters, are unevenly distributed into chromosome 1 to 6, and there is the purifying selection pressure in WRKY group III genes.
Clustering Algorithms: Their Application to Gene Expression Data
Oyelade, Jelili; Isewon, Itunuoluwa; Oladipupo, Funke; Aromolaran, Olufemi; Uwoghiren, Efosa; Ameh, Faridah; Achas, Moses; Adebiyi, Ezekiel
2016-01-01
Gene expression data hide vital information required to understand the biological process that takes place in a particular organism in relation to its environment. Deciphering the hidden patterns in gene expression data proffers a prodigious preference to strengthen the understanding of functional genomics. The complexity of biological networks and the volume of genes present increase the challenges of comprehending and interpretation of the resulting mass of data, which consists of millions of measurements; these data also inhibit vagueness, imprecision, and noise. Therefore, the use of clustering techniques is a first step toward addressing these challenges, which is essential in the data mining process to reveal natural structures and identify interesting patterns in the underlying data. The clustering of gene expression data has been proven to be useful in making known the natural structure inherent in gene expression data, understanding gene functions, cellular processes, and subtypes of cells, mining useful information from noisy data, and understanding gene regulation. The other benefit of clustering gene expression data is the identification of homology, which is very important in vaccine design. This review examines the various clustering algorithms applicable to the gene expression data in order to discover and provide useful knowledge of the appropriate clustering technique that will guarantee stability and high degree of accuracy in its analysis procedure. PMID:27932867
Gene Expression Profiles of Sporadic Canine Hemangiosarcoma Are Uniquely Associated with Breed
Tamburini, Beth A.; Trapp, Susan; Phang, Tzu Lip; Schappa, Jill T.; Hunter, Lawrence E.; Modiano, Jaime F.
2009-01-01
The role an individual's genetic background plays on phenotype and biological behavior of sporadic tumors remains incompletely understood. We showed previously that lymphomas from Golden Retrievers harbor defined, recurrent chromosomal aberrations that occur less frequently in lymphomas from other dog breeds, suggesting spontaneous canine tumors provide suitable models to define how heritable traits influence cancer genotypes. Here, we report a complementary approach using gene expression profiling in a naturally occurring endothelial sarcoma of dogs (hemangiosarcoma). Naturally occurring hemangiosarcomas of Golden Retrievers clustered separately from those of non-Golden Retrievers, with contributions from transcription factors, survival factors, and from pro-inflammatory and angiogenic genes, and which were exclusively present in hemangiosarcoma and not in other tumors or normal cells (i.e., they were not due simply to variation in these genes among breeds). Vascular Endothelial Growth Factor Receptor 1 (VEGFR1) was among genes preferentially enriched within known pathways derived from gene set enrichment analysis when characterizing tumors from Golden Retrievers versus other breeds. Heightened VEGFR1 expression in these tumors also was apparent at the protein level and targeted inhibition of VEGFR1 increased proliferation of hemangiosarcoma cells derived from tumors of Golden Retrievers, but not from other breeds. Our results suggest heritable factors mold gene expression phenotypes, and consequently biological behavior in sporadic, naturally occurring tumors. PMID:19461996
Genomic analyses of bacterial porin-cytochrome gene clusters
Shi, Liang; Fredrickson, James K.; Zachara, John M.
2014-11-26
In this study, the porin-cytochrome (Pcc) protein complex is responsible for trans-outer membrane electron transfer during extracellular reduction of Fe(III) by the dissimilatory metal-reducing bacterium Geobacter sulfurreducens PCA. The identified and characterized Pcc complex of G. sulfurreducens PCA consists of a porin-like outer-membrane protein, a periplasmic 8-heme c type cytochrome (c-Cyt) and an outer-membrane 12-heme c-Cyt, and the genes encoding the Pcc proteins are clustered in the same regions of genome (i.e., the pcc gene clusters) of G. sulfurreducens PCA. A survey of additionally microbial genomes has identified the pcc gene clusters in all sequenced Geobacter spp. and other bacteriamore » from six different phyla, including Anaeromyxobacter dehalogenans 2CP-1, A. dehalogenans 2CP-C, Anaeromyxobacter sp. K, Candidatus Kuenenia stuttgartiensis, Denitrovibrio acetiphilus DSM 12809, Desulfurispirillum indicum S5, Desulfurivibrio alkaliphilus AHT2, Desulfurobacterium thermolithotrophum DSM 11699, Desulfuromonas acetoxidans DSM 684, Ignavibacterium album JCM 16511, and Thermovibrio ammonificans HB-1. The numbers of genes in the pcc gene clusters vary, ranging from two to nine. Similar to the metal-reducing (Mtr) gene clusters of other Fe(III)-reducing bacteria, such as Shewanella spp., additional genes that encode putative c-Cyts with predicted cellular localizations at the cytoplasmic membrane, periplasm and outer membrane often associate with the pcc gene clusters. This suggests that the Pcc-associated c-Cyts may be part of the pathways for extracellular electron transfer reactions. The presence of pcc gene clusters in the microorganisms that do not reduce solid-phase Fe(III) and Mn(IV) oxides, such as D. alkaliphilus AHT2 and I. album JCM 16511, also suggests that some of the pcc gene clusters may be involved in extracellular electron transfer reactions with the substrates other than Fe(III) and Mn(IV) oxides.« less
Silva, Bruno; Nunes, Alexandra; Vale, Filipa F; Rocha, Raquel; Gomes, João Paulo; Dias, Ricardo; Oleastro, Mónica
2017-08-01
Helicobacter pylori virulence is associated with different clinical outcomes. The existence of an intact dupA gene from tfs4b cluster has been suggested as a predictor for duodenal ulcer development. However, the role of tfs plasticity zone clusters in the development of ulcers remains unclear. We studied several H. pylori strains to characterize the gene arrangement of tfs3 and tfs4 clusters and their impact in the inflammatory response by infected gastric cells. The genome of 14 H. pylori strains isolated from Western patients, pediatric (n=10) and adult (n=4), was fully sequenced using the Illumina platform MiSeq, in addition to eight pediatric strains previously sequenced. These strains were used to infect human gastric cells, and the secreted interleukin-8 (IL-8) was quantified by ELISA. The expression of virB2, dupA, virB8, virB10, and virB6 was assessed by quantitative PCR in adherent and nonadherent fractions of H. pylori during in vitro co-infection, at different pH values. We have found that cagA-positive H. pylori strains harboring a complete tfs plasticity zone cluster significantly induce increased production of IL-8 from gastric cells. We have also found that the region spanning from virB2 to virB10 genes constitutes an operon, whose expression is increased in the adherent fraction of bacteria during infection, as well as in both adherent and nonadherent fractions at acidic conditions. A complete tfs plasticity zone cluster is a virulence factor that may be important for the colonization of H. pylori and to the development of severe outcomes of the infection with cagA-positive strains. © 2017 John Wiley & Sons Ltd.
Dhar, Alok; Polev, Dmitrii E.; Masharsky, Alexey E.; Rogozin, Igor B.; Pavlov, Youri I.
2015-01-01
Mutations in genomes of species are frequently distributed non-randomly, resulting in mutation clusters, including recently discovered kataegis in tumors. DNA editing deaminases play the prominent role in the etiology of these mutations. To gain insight into the enigmatic mechanisms of localized hypermutagenesis that lead to cluster formation, we analyzed the mutational single nucleotide variations (SNV) data obtained by whole-genome sequencing of drug-resistant mutants induced in yeast diploids by AID/APOBEC deaminase and base analog 6-HAP. Deaminase from sea lamprey, PmCDA1, induced robust clusters, while 6-HAP induced a few weak ones. We found that PmCDA1, AID, and APOBEC1 deaminases preferentially mutate the beginning of the actively transcribed genes. Inactivation of transcription initiation factor Sub1 strongly reduced deaminase-induced can1 mutation frequency, but, surprisingly, did not decrease the total SNV load in genomes. However, the SNVs in the genomes of the sub1 clones were re-distributed, and the effect of mutation clustering in the regions of transcription initiation was even more pronounced. At the same time, the mutation density in the protein-coding regions was reduced, resulting in the decrease of phenotypically detected mutants. We propose that the induction of clustered mutations by deaminases involves: a) the exposure of ssDNA strands during transcription and loss of protection of ssDNA due to the depletion of ssDNA-binding proteins, such as Sub1, and b) attainment of conditions favorable for APOBEC action in subpopulation of cells, leading to enzymatic deamination within the currently expressed genes. This model is applicable to both the initial and the later stages of oncogenic transformation and explains variations in the distribution of mutations and kataegis events in different tumor cells. PMID:25941824
Uhong Lü, Yuhong; Liu, Xiaoli; Wang, Miao; Li, Yuanyuan; Liu, Ning; Bao, Yuxin; Liu, Minghao; Li, Xiaoqian; Wang, Yinyin; Qian, Shenyan; Yue, Changwu; Huang, Ying
2016-09-01
In order to obtain the natural products synthesized by the three putative xiamycin biosynthesis gene clusters which were predicted via antiSMASH during the genome mining of marine Streptomyces sp. FXJ 7.388, Streptomyces sp. FXJ 8.012, and Streptomyces olivaceus FXJ 7.023. Sixteen genes involved in xiamycin assembly, modification, and regulation with higher identity than the newest reported xiamycin biosynthetic gene cluster from marine Streptomyces sp. SCSIO 02999, Streptomyces sp. HKI0576, and Streptomyces sp. FXJ 7.388 were discovered via gene cluster comparative analysis. A ribosome engineering strategy was adopted to activate such cryptic gene clusters with different final concentrations antibiotics that act on the ribosome, and two indolosesquiterpenes were isolated from idlethaldose streptomycin-resistant Streptomyces sp. FXJ 7.388 strains. However, no such product was detected in Streptomyces sp. FXJ 8.012 and Streptomyces olivaceus FXJ 7.023 under the same treatment. This result suggested that these genes might hold the least gene content for xiamycin biosynthesis.
DOE Office of Scientific and Technical Information (OSTI.GOV)
Hamilton, A T; Huntley, S; Tran-Gyamfi, M
Although most genes are conserved as one-to-one orthologs in different mammalian orders, certain gene families have evolved to comprise different numbers and types of protein-coding genes through independent series of gene duplications, divergence and gene loss in each evolutionary lineage. One such family encodes KRAB-zinc finger (KRAB-ZNF) genes, which are likely to function as transcriptional repressors. One KRAB-ZNF subfamily, the ZNF91 clade, has expanded specifically in primates to comprise more than 110 loci in the human genome, yielding large gene clusters in human chromosomes 19 and 7 and smaller clusters or isolated copies at other chromosomal locations. Although phylogenetic analysismore » indicates that many of these genes arose before the split between old world monkeys and new world monkeys, the ZNF91 subfamily has continued to expand and diversify throughout the evolution of apes and humans. The paralogous loci are distinguished by sequence divergence within their zinc finger arrays indicating a selection for proteins with different DNA binding specificities. RT-PCR and in situ hybridization data show that some of these ZNF genes can have tissue-specific expression patterns, however many KRAB-ZNFs that are near-ubiquitous could also be playing very specific roles in halting target pathways in all tissues except for a few, where the target is released by the absence of its repressor. The number of variant KRAB-ZNF proteins is increased not only because of the large number of loci, but also because many loci can produce multiple splice variants, which because of the modular structure of these genes may have separate and perhaps even conflicting regulatory roles. The lineage-specific duplication and rapid divergence of this family of transcription factor genes suggests a role in determining species-specific biological differences and the evolution of novel primate traits.« less
Genes encoding cuticular proteins are components of the Nimrod gene cluster in Drosophila.
Cinege, Gyöngyi; Zsámboki, János; Vidal-Quadras, Maite; Uv, Anne; Csordás, Gábor; Honti, Viktor; Gábor, Erika; Hegedűs, Zoltán; Varga, Gergely I B; Kovács, Attila L; Juhász, Gábor; Williams, Michael J; Andó, István; Kurucz, Éva
2017-08-01
The Nimrod gene cluster, located on the second chromosome of Drosophila melanogaster, is the largest synthenic unit of the Drosophila genome. Nimrod genes show blood cell specific expression and code for phagocytosis receptors that play a major role in fruit fly innate immune functions. We previously identified three homologous genes (vajk-1, vajk-2 and vajk-3) located within the Nimrod cluster, which are unrelated to the Nimrod genes, but are homologous to a fourth gene (vajk-4) located outside the cluster. Here we show that, unlike the Nimrod candidates, the Vajk proteins are expressed in cuticular structures of the late embryo and the late pupa, indicating that they contribute to cuticular barrier functions. Copyright © 2017 Elsevier Ltd. All rights reserved.
Noninvasive analysis of the sputum transcriptome discriminates clinical phenotypes of asthma.
Yan, Xiting; Chu, Jen-Hwa; Gomez, Jose; Koenigs, Maria; Holm, Carole; He, Xiaoxuan; Perez, Mario F; Zhao, Hongyu; Mane, Shrikant; Martinez, Fernando D; Ober, Carole; Nicolae, Dan L; Barnes, Kathleen C; London, Stephanie J; Gilliland, Frank; Weiss, Scott T; Raby, Benjamin A; Cohn, Lauren; Chupp, Geoffrey L
2015-05-15
The airway transcriptome includes genes that contribute to the pathophysiologic heterogeneity seen in individuals with asthma. We analyzed sputum gene expression for transcriptomic endotypes of asthma (TEA), gene signatures that discriminate phenotypes of disease. Gene expression in the sputum and blood of patients with asthma was measured using Affymetrix microarrays. Unsupervised clustering analysis based on pathways from the Kyoto Encyclopedia of Genes and Genomes was used to identify TEA clusters. Logistic regression analysis of matched blood samples defined an expression profile in the circulation to determine the TEA cluster assignment in a cohort of children with asthma to replicate clinical phenotypes. Three TEA clusters were identified. TEA cluster 1 had the most subjects with a history of intubation (P = 0.05), a lower prebronchodilator FEV1 (P = 0.006), a higher bronchodilator response (P = 0.03), and higher exhaled nitric oxide levels (P = 0.04) compared with the other TEA clusters. TEA cluster 2, the smallest cluster, had the most subjects that were hospitalized for asthma (P = 0.04). TEA cluster 3, the largest cluster, had normal lung function, low exhaled nitric oxide levels, and lower inhaled steroid requirements. Evaluation of TEA clusters in children confirmed that TEA clusters 1 and 2 are associated with a history of intubation (P = 5.58 × 10(-6)) and hospitalization (P = 0.01), respectively. There are common patterns of gene expression in the sputum and blood of children and adults that are associated with near-fatal, severe, and milder asthma.
Noninvasive Analysis of the Sputum Transcriptome Discriminates Clinical Phenotypes of Asthma
Yan, Xiting; Chu, Jen-Hwa; Gomez, Jose; Koenigs, Maria; Holm, Carole; He, Xiaoxuan; Perez, Mario F.; Zhao, Hongyu; Mane, Shrikant; Martinez, Fernando D.; Ober, Carole; Nicolae, Dan L.; Barnes, Kathleen C.; London, Stephanie J.; Gilliland, Frank; Weiss, Scott T.; Raby, Benjamin A.; Cohn, Lauren
2015-01-01
Rationale: The airway transcriptome includes genes that contribute to the pathophysiologic heterogeneity seen in individuals with asthma. Objectives: We analyzed sputum gene expression for transcriptomic endotypes of asthma (TEA), gene signatures that discriminate phenotypes of disease. Methods: Gene expression in the sputum and blood of patients with asthma was measured using Affymetrix microarrays. Unsupervised clustering analysis based on pathways from the Kyoto Encyclopedia of Genes and Genomes was used to identify TEA clusters. Logistic regression analysis of matched blood samples defined an expression profile in the circulation to determine the TEA cluster assignment in a cohort of children with asthma to replicate clinical phenotypes. Measurements and Main Results: Three TEA clusters were identified. TEA cluster 1 had the most subjects with a history of intubation (P = 0.05), a lower prebronchodilator FEV1 (P = 0.006), a higher bronchodilator response (P = 0.03), and higher exhaled nitric oxide levels (P = 0.04) compared with the other TEA clusters. TEA cluster 2, the smallest cluster, had the most subjects that were hospitalized for asthma (P = 0.04). TEA cluster 3, the largest cluster, had normal lung function, low exhaled nitric oxide levels, and lower inhaled steroid requirements. Evaluation of TEA clusters in children confirmed that TEA clusters 1 and 2 are associated with a history of intubation (P = 5.58 × 10−6) and hospitalization (P = 0.01), respectively. Conclusions: There are common patterns of gene expression in the sputum and blood of children and adults that are associated with near-fatal, severe, and milder asthma. PMID:25763605
Analysis of genetic association using hierarchical clustering and cluster validation indices.
Pagnuco, Inti A; Pastore, Juan I; Abras, Guillermo; Brun, Marcel; Ballarin, Virginia L
2017-10-01
It is usually assumed that co-expressed genes suggest co-regulation in the underlying regulatory network. Determining sets of co-expressed genes is an important task, based on some criteria of similarity. This task is usually performed by clustering algorithms, where the genes are clustered into meaningful groups based on their expression values in a set of experiment. In this work, we propose a method to find sets of co-expressed genes, based on cluster validation indices as a measure of similarity for individual gene groups, and a combination of variants of hierarchical clustering to generate the candidate groups. We evaluated its ability to retrieve significant sets on simulated correlated and real genomics data, where the performance is measured based on its detection ability of co-regulated sets against a full search. Additionally, we analyzed the quality of the best ranked groups using an online bioinformatics tool that provides network information for the selected genes. Copyright © 2017 Elsevier Inc. All rights reserved.
TFIIIC bound DNA elements in nuclear organization and insulation.
Kirkland, Jacob G; Raab, Jesse R; Kamakaka, Rohinton T
2013-01-01
tRNA genes (tDNAs) have been known to have barrier insulator function in budding yeast, Saccharomyces cerevisiae, for over a decade. tDNAs also play a role in genome organization by clustering at sites in the nucleus and both of these functions are dependent on the transcription factor TFIIIC. More recently TFIIIC bound sites devoid of pol III, termed Extra-TFIIIC sites (ETC) have been identified in budding yeast and these sites also function as insulators and affect genome organization. Subsequent studies in Schizosaccharomyces pombe showed that TFIIIC bound sites were insulators and also functioned as Chromosome Organization Clamps (COC); tethering the sites to the nuclear periphery. Very recently studies have moved to mammalian systems where pol III genes and their associated factors have been investigated in both mouse and human cells. Short interspersed nuclear elements (SINEs) that bind TFIIIC, function as insulator elements and tDNAs can also function as both enhancer - blocking and barrier insulators in these organisms. It was also recently shown that tDNAs cluster with other tDNAs and with ETCs but not with pol II transcribed genes. Intriguingly, TFIIIC is often found near pol II transcription start sites and it remains unclear what the consequences of TFIIIC based genomic organization are and what influence pol III factors have on pol II transcribed genes and vice versa. In this review we provide a comprehensive overview of the known data on pol III factors in insulation and genome organization and identify the many open questions that require further investigation. This article is part of a Special Issue entitled: Transcription by Odd Pols. Copyright © 2012 Elsevier B.V. All rights reserved.
Morata, Jordi; Puigdomènech, Pere
2017-02-08
Cucurbitaceae species contain a significantly lower number of genes coding for proteins with similarity to plant resistance genes belonging to the NBS-LRR family than other plant species of similar genome size. A large proportion of these genes are organized in clusters that appear to be hotspots of variability. The genomes of the Cucurbitaceae species measured until now are intermediate in size (between 350 and 450 Mb) and they apparently have not undergone any genome duplications beside those at the origin of eudicots. The cluster containing the largest number of NBS-LRR genes has previously been analyzed in melon and related species and showed a high degree of interspecific and intraspecific variability. It was of interest to study whether similar behavior occurred in other cluster of the same family of genes. The cluster of NBS-LRR genes located in melon chromosome 9 was analyzed and compared with the syntenic regions in other cucurbit genomes. This is the second cluster in number within this species and it contains nine sequences with a NBS-LRR annotation including two genes, Fom1 and Prv, providing resistance against Fusarium and Ppapaya ring-spot virus (PRSV). The variability within the melon species appears to consist essentially of single nucleotide polymorphisms. Clusters of similar genes are present in the syntenic regions of the two species of Cucurbitaceae that were sequenced, cucumber and watermelon. Most of the genes in the syntenic clusters can be aligned between species and a hypothesis of generation of the cluster is proposed. The number of genes in the watermelon cluster is similar to that in melon while a higher number of genes (12) is present in cucumber, a species with a smaller genome than melon. After comparing genome resequencing data of 115 cucumber varieties, deletion of a group of genes is observed in a group of varieties of Indian origin. Clusters of genes coding for NBS-LRR proteins in cucurbits appear to have specific variability in different regions of the genome and between different species. This observation is in favour of considering that the adaptation of plant species to changing environments is based upon the variability that may occur at any location in the genome and that has been produced by specific mechanisms of sequence variation acting on plant genomes. This information could be useful both to understand the evolution of species and for plant breeding.
Jothi, R; Mohanty, Sraban Kumar; Ojha, Aparajita
2016-04-01
Gene expression data clustering is an important biological process in DNA microarray analysis. Although there have been many clustering algorithms for gene expression analysis, finding a suitable and effective clustering algorithm is always a challenging problem due to the heterogeneous nature of gene profiles. Minimum Spanning Tree (MST) based clustering algorithms have been successfully employed to detect clusters of varying shapes and sizes. This paper proposes a novel clustering algorithm using Eigenanalysis on Minimum Spanning Tree based neighborhood graph (E-MST). As MST of a set of points reflects the similarity of the points with their neighborhood, the proposed algorithm employs a similarity graph obtained from k(') rounds of MST (k(')-MST neighborhood graph). By studying the spectral properties of the similarity matrix obtained from k(')-MST graph, the proposed algorithm achieves improved clustering results. We demonstrate the efficacy of the proposed algorithm on 12 gene expression datasets. Experimental results show that the proposed algorithm performs better than the standard clustering algorithms. Copyright © 2016 Elsevier Ltd. All rights reserved.
Transcriptome analysis of resistant soybean roots infected by Meloidogyne javanica
de Sá, Maria Eugênia Lisei; Conceição Lopes, Marcus José; de Araújo Campos, Magnólia; Paiva, Luciano Vilela; dos Santos, Regina Maria Amorim; Beneventi, Magda Aparecida; Firmino, Alexandre Augusto Pereira; de Sá, Maria Fátima Grossi
2012-01-01
Soybean is an important crop for Brazilian agribusiness. However, many factors can limit its production, especially root-knot nematode infection. Studies on the mechanisms employed by the resistant soybean genotypes to prevent infection by these nematodes are of great interest for breeders. For these reasons, the aim of this work is to characterize the transcriptome of soybean line PI 595099-Meloidogyne javanica interaction through expression analysis. Two cDNA libraries were obtained using a pool of RNA from PI 595099 uninfected and M. javanica (J2) infected roots, collected at 6, 12, 24, 48, 96, 144 and 192 h after inoculation. Around 800 ESTs (Expressed Sequence Tags) were sequenced and clustered into 195 clusters. In silico subtraction analysis identified eleven differentially expressed genes encoding putative proteins sharing amino acid sequence similarities by using BlastX: metallothionein, SLAH4 (SLAC1 Homologue 4), SLAH1 (SLAC1 Homologue 1), zinc-finger proteins, AN1-type proteins, auxin-repressed proteins, thioredoxin and nuclear transport factor 2 (NTF-2). Other genes were also found exclusively in nematode stressed soybean roots, such as NAC domain-containing proteins, MADS-box proteins, SOC1 (suppressor of overexpression of constans 1) proteins, thioredoxin-like protein 4-Coumarate-CoA ligase and the transcription factor (TF) MYBZ2. Among the genes identified in non-stressed roots only were Ser/Thr protein kinases, wound-induced basic protein, ethylene-responsive family protein, metallothionein-like protein cysteine proteinase inhibitor (cystatin) and Putative Kunitz trypsin protease inhibitor. An understanding of the roles of these differentially expressed genes will provide insights into the resistance mechanisms and candidate genes involved in soybean-M. javanica interaction and contribute to more effective control of this pathogen. PMID:22802712
Wang, Ruoxin; Su, Chao; Wang, Xinting; Fu, Qiang; Gao, Xingjie; Zhang, Chunyan; Yang, Jie; Yang, Xi; Wei, Minxin
2018-01-01
Mammalian cardiomyocytes may permanently lose their ability to proliferate after birth. Therefore, studying the proliferation and growth arrest of cardiomyocytes during the postnatal period may enhance the current understanding regarding this molecular mechanism. The present study identified the differentially expressed genes in hearts obtained from 24 h‑old mice, which contain proliferative cardiomyocytes; 7‑day‑old mice, in which the cardiomyocytes are undergoing a proliferative burst; and 10‑week‑old mice, which contain growth‑arrested cardiomyocytes, using global gene expression analysis. Furthermore, myocardial proliferation and growth arrest were analyzed from numerous perspectives, including Gene Ontology annotation, cluster analysis, pathway enrichment and network construction. The results of a Gene Ontology analysis indicated that, with increasing age, enriched gene function was not only associated with cell cycle, cell division and mitosis, but was also associated with metabolic processes and protein synthesis. In the pathway analysis, 'cell cycle', proliferation pathways, such as the 'PI3K‑AKT signaling pathway', and 'metabolic pathways' were well represented. Notably, the cluster analysis revealed that bone morphogenetic protein (BMP)1, BMP10, cyclin E2, E2F transcription factor 1 and insulin like growth factor 1 exhibited increased expression in hearts obtained from 7‑day‑old mice. In addition, the signal transduction pathway associated with the cell cycle was identified. The present study primarily focused on genes with altered expression, including downregulated anaphase promoting complex subunit 1, cell division cycle (CDC20), cyclin dependent kinase 1, MYC proto-oncogene, bHLH transcription factor and CDC25C, and upregulated growth arrest and DNA damage inducible α in 10-week group, which may serve important roles in postnatal myocardial cell cycle arrest. In conclusion, these data may provide important information regarding myocardial proliferation and development.
Roche, Béatrice; Huguenot, Allison; Barras, Frédéric; Py, Béatrice
2015-02-01
In eukaryotes, frataxin deficiency (FXN) causes severe phenotypes including loss of iron-sulfur (Fe-S) cluster protein activity, accumulation of mitochondrial iron and leads to the neurodegenerative disease Friedreich's ataxia. In contrast, in prokaryotes, deficiency in the FXN homolog, CyaY, was reported not to cause any significant phenotype, questioning both its importance and its actual contribution to Fe-S cluster biogenesis. Because FXN is conserved between eukaryotes and prokaryotes, this surprising discrepancy prompted us to reinvestigate the role of CyaY in Escherichia coli. We report that CyaY (i) potentiates E. coli fitness, (ii) belongs to the ISC pathway catalyzing the maturation of Fe-S cluster-containing proteins and (iii) requires iron-rich conditions for its contribution to be significant. A genetic interaction was discovered between cyaY and iscX, the last gene of the isc operon. Deletion of both genes showed an additive effect on Fe-S cluster protein maturation, which led, among others, to increased resistance to aminoglycosides and increased sensitivity to lambda phage infection. Together, these in vivo results establish the importance of CyaY as a member of the ISC-mediated Fe-S cluster biogenesis pathway in E. coli, like it does in eukaryotes, and validate IscX as a new bona fide Fe-S cluster biogenesis factor. © 2014 John Wiley & Sons Ltd.
High-throughput platform for the discovery of elicitors of silent bacterial gene clusters.
Seyedsayamdost, Mohammad R
2014-05-20
Over the past decade, bacterial genome sequences have revealed an immense reservoir of biosynthetic gene clusters, sets of contiguous genes that have the potential to produce drugs or drug-like molecules. However, the majority of these gene clusters appear to be inactive for unknown reasons prompting terms such as "cryptic" or "silent" to describe them. Because natural products have been a major source of therapeutic molecules, methods that rationally activate these silent clusters would have a profound impact on drug discovery. Herein, a new strategy is outlined for awakening silent gene clusters using small molecule elicitors. In this method, a genetic reporter construct affords a facile read-out for activation of the silent cluster of interest, while high-throughput screening of small molecule libraries provides potential inducers. This approach was applied to two cryptic gene clusters in the pathogenic model Burkholderia thailandensis. The results not only demonstrate a prominent activation of these two clusters, but also reveal that the majority of elicitors are themselves antibiotics, most in common clinical use. Antibiotics, which kill B. thailandensis at high concentrations, act as inducers of secondary metabolism at low concentrations. One of these antibiotics, trimethoprim, served as a global activator of secondary metabolism by inducing at least five biosynthetic pathways. Further application of this strategy promises to uncover the regulatory networks that activate silent gene clusters while at the same time providing access to the vast array of cryptic molecules found in bacteria.
Liu, Yimo; Wang, Zheming; Liu, Juan; ...
2014-09-24
The multiheme, outer membrane c-type cytochrome (c-Cyt) OmcB of Geobacter sulfurreducens was previously proposed to mediate electron transfer across the outer membrane. However, the underlying mechanism has remained uncharacterized. In G. sulfurreducens, the omcB gene is part of two tandem four-gene clusters, each is predicted to encode a transcriptional factor (OrfR/OrfS), a porin-like outer membrane protein (OmbB/OmbC), a periplasmic c-type cytochrome (OmaB/OmaC), and an outer membrane c-Cyt (OmcB/OmcC), respectively. Here we showed that OmbB/OmbC, OmaB/OmaC and OmcB/OmcC of G. sulfurreducens PCA formed the porin-cytochrome (Pcc) protein complexes, which were involved in transferring electrons across the outer membrane. The isolated Pccmore » protein complexes reconstituted in proteoliposomes transferred electrons from reduced methyl viologen across the lipid bilayer of liposomes to Fe(III)-citrate and ferrihydrite. The pcc clusters were found in all eight sequenced Geobacter and 11 other bacterial genomes from six different phyla, demonstrating a widespread distribution of Pcc protein complexes in phylogenetically diverse bacteria. Deletion of ombB-omaB-omcB-orfS-ombC-omaC-omcC gene clusters had no impact on the growth of G. sulfurreducens PCA with fumarate, but diminished the ability of G. sulfurreducens PCA to reduce Fe(III)-citrate and ferrihydrite. Finally, complementation with the ombB-omaB-omcB gene cluster restored the ability of G. sulfurreducens PCA to reduce Fe(III)-citrate and ferrihydrite.« less
A conserved gene cluster as a putative functional unit in insect innate immunity.
Somogyi, Kálmán; Sipos, Botond; Pénzes, Zsolt; Andó, István
2010-11-05
The Nimrod gene superfamily is an important component of the innate immune response. The majority of its member genes are located in close proximity within the Drosophila melanogaster genome and they lie in a larger conserved cluster ("Nimrod cluster"), made up of non-related groups (families, superfamilies) of genes. This cluster has been a part of the Arthropod genomes for about 300-350 million years. The available data suggest that the Nimrod cluster is a functional module of the insect innate immune response. Copyright © 2010 Federation of European Biochemical Societies. Published by Elsevier B.V. All rights reserved.
Zhang, Xiujun; Parry, Ronald J.
2007-01-01
The pyrrolomycins are a family of polyketide antibiotics, some of which contain a nitro group. To gain insight into the nitration mechanism associated with the formation of these antibiotics, the pyrrolomycin biosynthetic gene cluster from Actinosporangium vitaminophilum was cloned. Sequencing of ca. 56 kb of A. vitaminophilum DNA revealed 35 open reading frames (ORFs). Sequence analysis revealed a clear relationship between some of these ORFs and the biosynthetic gene cluster for pyoluteorin, a structurally related antibiotic. Since a gene transfer system could not be devised for A. vitaminophilum, additional proof for the identity of the cloned gene cluster was sought by cloning the pyrrolomycin gene cluster from Streptomyces sp. strain UC 11065, a transformable pyrrolomycin producer. Sequencing of ca. 26 kb of UC 11065 DNA revealed the presence of 17 ORFs, 15 of which exhibit strong similarity to ORFs in the A. vitaminophilum cluster as well as a nearly identical organization. Single-crossover disruption of two genes in the UC 11065 cluster abolished pyrrolomycin production in both cases. These results confirm that the genetic locus cloned from UC 11065 is essential for pyrrolomycin production, and they also confirm that the highly similar locus in A. vitaminophilum encodes pyrrolomycin biosynthetic genes. Sequence analysis revealed that both clusters contain genes encoding the two components of an assimilatory nitrate reductase. This finding suggests that nitrite is required for the formation of the nitrated pyrrolomycins. However, sequence analysis did not provide additional insights into the nitration process, suggesting the operation of a novel nitration mechanism. PMID:17158935
Functional characterization of the role of rpfA in Xylella fastidiosa
USDA-ARS?s Scientific Manuscript database
Xylella fastidiosa coordinates virulence in grapevines via quorum sensing signal molecules that are regulated and synthesized by the rpf gene cluster (regulation of pathogenicity factors). rpfA encodes aconitate hydratase and could play a regulator role involved in virulence. To elucidate the role o...
Barczak, Amy K; Avraham, Roi; Singh, Shantanu; Luo, Samantha S; Zhang, Wei Ran; Bray, Mark-Anthony; Hinman, Amelia E; Thompson, Matthew; Nietupski, Raymond M; Golas, Aaron; Montgomery, Paul; Fitzgerald, Michael; Smith, Roger S; White, Dylan W; Tischler, Anna D; Carpenter, Anne E; Hung, Deborah T
2017-05-01
A key to the pathogenic success of Mycobacterium tuberculosis (Mtb), the causative agent of tuberculosis, is the capacity to survive within host macrophages. Although several factors required for this survival have been identified, a comprehensive knowledge of such factors and how they work together to manipulate the host environment to benefit bacterial survival are not well understood. To systematically identify Mtb factors required for intracellular growth, we screened an arrayed, non-redundant Mtb transposon mutant library by high-content imaging to characterize the mutant-macrophage interaction. Based on a combination of imaging features, we identified mutants impaired for intracellular survival. We then characterized the phenotype of infection with each mutant by profiling the induced macrophage cytokine response. Taking a systems-level approach to understanding the biology of identified mutants, we performed a multiparametric analysis combining pathogen and host phenotypes to predict functional relationships between mutants based on clustering. Strikingly, mutants defective in two well-known virulence factors, the ESX-1 protein secretion system and the virulence lipid phthiocerol dimycocerosate (PDIM), clustered together. Building upon the shared phenotype of loss of the macrophage type I interferon (IFN) response to infection, we found that PDIM production and export are required for coordinated secretion of ESX-1-substrates, for phagosomal permeabilization, and for downstream induction of the type I IFN response. Multiparametric clustering also identified two novel genes that are required for PDIM production and induction of the type I IFN response. Thus, multiparametric analysis combining host and pathogen infection phenotypes can be used to identify novel functional relationships between genes that play a role in infection.
Cao, Huojun; Amendt, Brad A
2016-11-01
Developmental dental anomalies are common forms of congenital defects. The molecular mechanisms of dental anomalies are poorly understood. Systematic approaches such as clustering genes based on similar expression patterns could identify novel genes involved in dental anomalies and provide a framework for understanding molecular regulatory mechanisms of these genes during tooth development (odontogenesis). A python package (pySAPC) of sparse affinity propagation clustering algorithm for large datasets was developed. Whole genome pair-wise similarity was calculated based on expression pattern similarity based on 45 microarrays of several stages during odontogenesis. pySAPC identified 743 gene clusters based on expression pattern similarity during mouse tooth development. Three clusters are significantly enriched for genes associated with dental anomalies (with FDR <0.1). The three clusters of genes have distinct expression patterns during odontogenesis. Clustering genes based on similar expression profiles recovered several known regulatory relationships for genes involved in odontogenesis, as well as many novel genes that may be involved with the same genetic pathways as genes that have already been shown to contribute to dental defects. By using sparse similarity matrix, pySAPC use much less memory and CPU time compared with the original affinity propagation program that uses a full similarity matrix. This python package will be useful for many applications where dataset(s) are too large to use full similarity matrix. This article is part of a Special Issue entitled "System Genetics" Guest Editor: Dr. Yudong Cai and Dr. Tao Huang. Copyright © 2016. Published by Elsevier B.V.
Evolution of Chemical Diversity in Echinocandin Lipopeptide Antifungal Metabolites
Yue, Qun; Chen, Li; Zhang, Xiaoling; Li, Kuan; Sun, Jingzu; Liu, Xingzhong
2015-01-01
The echinocandins are a class of antifungal drugs that includes caspofungin, micafungin, and anidulafungin. Gene clusters encoding most of the structural complexity of the echinocandins provided a framework for hypotheses about the evolutionary history and chemical logic of echinocandin biosynthesis. Gene orthologs among echinocandin-producing fungi were identified. Pathway genes, including the nonribosomal peptide synthetases (NRPSs), were analyzed phylogenetically to address the hypothesis that these pathways represent descent from a common ancestor. The clusters share cooperative gene contents and linkages among the different strains. Individual pathway genes analyzed in the context of similar genes formed unique echinocandin-exclusive phylogenetic lineages. The echinocandin NRPSs, along with the NRPS from the inp gene cluster in Aspergillus nidulans and its orthologs, comprise a novel lineage among fungal NRPSs. NRPS adenylation domains from different species exhibited a one-to-one correspondence between modules and amino acid specificity that is consistent with models of tandem duplication and subfunctionalization. Pathway gene trees and Ascomycota phylogenies are congruent and consistent with the hypothesis that the echinocandin gene clusters have a common origin. The disjunct Eurotiomycete-Leotiomycete distribution appears to be consistent with a scenario of vertical descent accompanied by incomplete lineage sorting and loss of the clusters from most lineages of the Ascomycota. We present evidence for a single evolutionary origin of the echinocandin family of gene clusters and a progression of structural diversification in two fungal classes that diverged approximately 290 to 390 million years ago. Lineage-specific gene cluster evolution driven by selection of new chemotypes contributed to diversification of the molecular functionalities. PMID:26024901
Clusters of antibiotic resistance genes enriched together stay together in swine agriculture
Johnson, Timothy A.; Stedtfeld, Robert D.; Wang, Qiong; ...
2016-04-12
Antibiotic resistance is a worldwide health risk, but the influence of animal agriculture on the genetic context and enrichment of individual antibiotic resistance alleles remains unclear. Using quantitative PCR followed by amplicon sequencing, we quantified and sequenced 44 genes related to antibiotic resistance, mobile genetic elements, and bacterial phylogeny in microbiomes from U.S. laboratory swine and from swine farms from three Chinese regions. We identified highly abundant resistance clusters: groups of resistance and mobile genetic element alleles that cooccur. For example, the abundance of genes conferring resistance to six classes of antibiotics together with class 1 integrase and the abundancemore » of IS6100-type transposons in three Chinese regions are directly correlated. These resistance cluster genes likely colocalize in microbial genomes in the farms. Resistance cluster alleles were dramatically enriched (up to 1 to 10% as abundant as 16S rRNA) and indicate that multidrug-resistant bacteria are likely the norm rather than an exception in these communities. This enrichment largely occurred independently of phylogenetic composition; thus, resistance clusters are likely present in many bacterial taxa. Furthermore, resistance clusters contain resistance genes that confer resistance to antibiotics independently of their particular use on the farms. Selection for these clusters is likely due to the use of only a subset of the broad range of chemicals to which the clusters confer resistance. The scale of animal agriculture and its wastes, the enrichment and horizontal gene transfer potential of the clusters, and the vicinity of large human populations suggest that managing this resistance reservoir is important for minimizing human risk.Agricultural antibiotic use results in clusters of cooccurring resistance genes that together confer resistance to multiple antibiotics. The use of a single antibiotic could select for an entire suite of resistance genes if they are genetically linked. No links to bacterial membership were observed for these clusters of resistance genes. These findings urge deeper understanding of colocalization of resistance genes and mobile genetic elements in resistance islands and their distribution throughout antibiotic-exposed microbiomes. In addition, as governments seek to combat the rise in antibiotic resistance, a balance is sought between ensuring proper animal health and welfare and preserving medically important antibiotics for therapeutic use. Metagenomic and genomic monitoring will be critical to determine if resistance genes can be reduced in animal microbiomes, or if these gene clusters will continue to be coselected by antibiotics not deemed medically important for human health but used for growth promotion or by medically important antibiotics used therapeutically.« less
Clusters of antibiotic resistance genes enriched together stay together in swine agriculture
DOE Office of Scientific and Technical Information (OSTI.GOV)
Johnson, Timothy A.; Stedtfeld, Robert D.; Wang, Qiong
Antibiotic resistance is a worldwide health risk, but the influence of animal agriculture on the genetic context and enrichment of individual antibiotic resistance alleles remains unclear. Using quantitative PCR followed by amplicon sequencing, we quantified and sequenced 44 genes related to antibiotic resistance, mobile genetic elements, and bacterial phylogeny in microbiomes from U.S. laboratory swine and from swine farms from three Chinese regions. We identified highly abundant resistance clusters: groups of resistance and mobile genetic element alleles that cooccur. For example, the abundance of genes conferring resistance to six classes of antibiotics together with class 1 integrase and the abundancemore » of IS6100-type transposons in three Chinese regions are directly correlated. These resistance cluster genes likely colocalize in microbial genomes in the farms. Resistance cluster alleles were dramatically enriched (up to 1 to 10% as abundant as 16S rRNA) and indicate that multidrug-resistant bacteria are likely the norm rather than an exception in these communities. This enrichment largely occurred independently of phylogenetic composition; thus, resistance clusters are likely present in many bacterial taxa. Furthermore, resistance clusters contain resistance genes that confer resistance to antibiotics independently of their particular use on the farms. Selection for these clusters is likely due to the use of only a subset of the broad range of chemicals to which the clusters confer resistance. The scale of animal agriculture and its wastes, the enrichment and horizontal gene transfer potential of the clusters, and the vicinity of large human populations suggest that managing this resistance reservoir is important for minimizing human risk.Agricultural antibiotic use results in clusters of cooccurring resistance genes that together confer resistance to multiple antibiotics. The use of a single antibiotic could select for an entire suite of resistance genes if they are genetically linked. No links to bacterial membership were observed for these clusters of resistance genes. These findings urge deeper understanding of colocalization of resistance genes and mobile genetic elements in resistance islands and their distribution throughout antibiotic-exposed microbiomes. In addition, as governments seek to combat the rise in antibiotic resistance, a balance is sought between ensuring proper animal health and welfare and preserving medically important antibiotics for therapeutic use. Metagenomic and genomic monitoring will be critical to determine if resistance genes can be reduced in animal microbiomes, or if these gene clusters will continue to be coselected by antibiotics not deemed medically important for human health but used for growth promotion or by medically important antibiotics used therapeutically.« less
Clusters of Antibiotic Resistance Genes Enriched Together Stay Together in Swine Agriculture.
Johnson, Timothy A; Stedtfeld, Robert D; Wang, Qiong; Cole, James R; Hashsham, Syed A; Looft, Torey; Zhu, Yong-Guan; Tiedje, James M
2016-04-12
Antibiotic resistance is a worldwide health risk, but the influence of animal agriculture on the genetic context and enrichment of individual antibiotic resistance alleles remains unclear. Using quantitative PCR followed by amplicon sequencing, we quantified and sequenced 44 genes related to antibiotic resistance, mobile genetic elements, and bacterial phylogeny in microbiomes from U.S. laboratory swine and from swine farms from three Chinese regions. We identified highly abundant resistance clusters: groups of resistance and mobile genetic element alleles that cooccur. For example, the abundance of genes conferring resistance to six classes of antibiotics together with class 1 integrase and the abundance of IS6100-type transposons in three Chinese regions are directly correlated. These resistance cluster genes likely colocalize in microbial genomes in the farms. Resistance cluster alleles were dramatically enriched (up to 1 to 10% as abundant as 16S rRNA) and indicate that multidrug-resistant bacteria are likely the norm rather than an exception in these communities. This enrichment largely occurred independently of phylogenetic composition; thus, resistance clusters are likely present in many bacterial taxa. Furthermore, resistance clusters contain resistance genes that confer resistance to antibiotics independently of their particular use on the farms. Selection for these clusters is likely due to the use of only a subset of the broad range of chemicals to which the clusters confer resistance. The scale of animal agriculture and its wastes, the enrichment and horizontal gene transfer potential of the clusters, and the vicinity of large human populations suggest that managing this resistance reservoir is important for minimizing human risk. Agricultural antibiotic use results in clusters of cooccurring resistance genes that together confer resistance to multiple antibiotics. The use of a single antibiotic could select for an entire suite of resistance genes if they are genetically linked. No links to bacterial membership were observed for these clusters of resistance genes. These findings urge deeper understanding of colocalization of resistance genes and mobile genetic elements in resistance islands and their distribution throughout antibiotic-exposed microbiomes. As governments seek to combat the rise in antibiotic resistance, a balance is sought between ensuring proper animal health and welfare and preserving medically important antibiotics for therapeutic use. Metagenomic and genomic monitoring will be critical to determine if resistance genes can be reduced in animal microbiomes, or if these gene clusters will continue to be coselected by antibiotics not deemed medically important for human health but used for growth promotion or by medically important antibiotics used therapeutically. Copyright © 2016 Johnson et al.
Jung, Inuk; Jo, Kyuri; Kang, Hyejin; Ahn, Hongryul; Yu, Youngjae; Kim, Sun
2017-12-01
Identifying biologically meaningful gene expression patterns from time series gene expression data is important to understand the underlying biological mechanisms. To identify significantly perturbed gene sets between different phenotypes, analysis of time series transcriptome data requires consideration of time and sample dimensions. Thus, the analysis of such time series data seeks to search gene sets that exhibit similar or different expression patterns between two or more sample conditions, constituting the three-dimensional data, i.e. gene-time-condition. Computational complexity for analyzing such data is very high, compared to the already difficult NP-hard two dimensional biclustering algorithms. Because of this challenge, traditional time series clustering algorithms are designed to capture co-expressed genes with similar expression pattern in two sample conditions. We present a triclustering algorithm, TimesVector, specifically designed for clustering three-dimensional time series data to capture distinctively similar or different gene expression patterns between two or more sample conditions. TimesVector identifies clusters with distinctive expression patterns in three steps: (i) dimension reduction and clustering of time-condition concatenated vectors, (ii) post-processing clusters for detecting similar and distinct expression patterns and (iii) rescuing genes from unclassified clusters. Using four sets of time series gene expression data, generated by both microarray and high throughput sequencing platforms, we demonstrated that TimesVector successfully detected biologically meaningful clusters of high quality. TimesVector improved the clustering quality compared to existing triclustering tools and only TimesVector detected clusters with differential expression patterns across conditions successfully. The TimesVector software is available at http://biohealth.snu.ac.kr/software/TimesVector/. sunkim.bioinfo@snu.ac.kr. Supplementary data are available at Bioinformatics online. © The Author 2017. Published by Oxford University Press. All rights reserved. For Permissions, please e-mail: journals.permissions@oup.com
The Human Paraoxonase Gene Cluster As a Target in the Treatment of Atherosclerosis
She, Zhi-Gang; Chen, Hou-Zao; Yan, Yunfei; Li, Hongliang
2012-01-01
Abstract The paraoxonase (PON) gene cluster contains three adjacent gene members, PON1, PON2, and PON3. Originating from the same fungus lactonase precursor, all of the three PON genes share high sequence identity and a similar β propeller protein structure. PON1 and PON3 are primarily expressed in the liver and secreted into the serum upon expression, whereas PON2 is ubiquitously expressed and remains inside the cell. Each PON member has high catalytic activity toward corresponding artificial organophosphate, and all exhibit activities to lactones. Therefore, all three members of the family are regarded as lactonases. Under physiological conditions, they act to degrade metabolites of polyunsaturated fatty acids and homocysteine (Hcy) thiolactone, among other compounds. By detoxifying both oxidized low-density lipoprotein and Hcy thiolactone, PONs protect against atherosclerosis and coronary artery diseases, as has been illustrated by many types of in vitro and in vivo experimental evidence. Clinical observations focusing on gene polymorphisms also indicate that PON1, PON2, and PON3 are protective against coronary artery disease. Many other conditions, such as diabetes, metabolic syndrome, and aging, have been shown to relate to PONs. The abundance and/or activity of PONs can be regulated by lipoproteins and their metabolites, biological macromolecules, pharmacological treatments, dietary factors, and lifestyle. In conclusion, both previous results and ongoing studies provide evidence, making the PON cluster a prospective target for the treatment of atherosclerosis. Antioxid. Redox Signal. 16, 597–632. PMID:21867409
ORGANIZATION OF THE nif GENES OF THE NONHETEROCYSTOUS CYANOBACTERIUM TRICHODESMIUM SP. IMS101.
Dominic, Benny; Zani, Sabino; Chen, Yi-Bu; Mellon, Mark T; Zehr, Jonathan P
2000-08-26
An approximately 16-kb fragment of the Trichodesmium sp. IMS101 (a nonheterocystous filamentous cyanobacterium) "conventional"nif gene cluster was cloned and sequenced. The gene organization of the Trichodesmium and Anabaena variabilis vegetative (nif 2) nitrogenase gene clusters spanning the region from nif B to nif W are similar except for the absence of two open reading frames (ORF3 and ORF1) in Trichodesmium. The Trichodesmium nif EN genes encode a fused Nif EN polypeptide that does not appear to be processed into individual Nif E and Nif N polypeptides. Fused nif EN genes were previously found in the A. variabilis nif 2 genes, but we have found that fused nif EN genes are widespread in the nonheterocystous cyanobacteria. Although the gene organization of the nonheterocystous filamentous Trichodesmium nif gene cluster is very similar to that of the A. variabilis vegetative nif 2 gene cluster, phylogenetic analysis of nif sequences do not support close relatedness of Trichodesmium and A. variabilis vegetative (nif 2) nitrogenase genes.
A Critical Role for CRM1 in Regulating HOXA Gene Transcription in CALM-AF10 Leukemias
Conway, Amanda E.; Haldeman, Jonathan M.; Wechsler, Daniel S.; Lavau, Catherine P.
2014-01-01
The leukemogenic CALM-AF10 fusion protein is found in patients with immature acute myeloid and T-lymphoid malignancies. CALM-AF10 leukemias display abnormal H3K79 methylation and increased HOXA cluster gene transcription. Elevated expression of HOXA genes is critical for leukemia maintenance and progression; however, the precise mechanism by which CALM-AF10 alters HOXA gene expression is unclear. We previously determined that CALM contains a CRM1-dependent nuclear export signal (NES), which is both necessary and sufficient for CALM-AF10-mediated leukemogenesis. Here, we find that interaction of CALM-AF10 with the nuclear export receptor CRM1 is necessary for activating HOXA gene expression. We show that CRM1 localizes to HOXA loci where it recruits CALM-AF10, leading to transcriptional and epigenetic activation of HOXA genes. Genetic and pharmacological inhibition of the CALM-CRM1 interaction prevents CALM-AF10 enrichment at HOXA chromatin, resulting in immediate loss of transcription. These results provide a comprehensive mechanism by which the CALM-AF10 translocation activates the critical HOXA cluster genes. Furthermore, this report identifies a novel function of CRM1: the ability to bind chromatin and recruit the NES-containing CALM-AF10 transcription factor. PMID:25027513
Yang, Seoyeon; Lee, Ji-Yeon; Hur, Ho; Oh, Ji Hoon; Kim, Myoung Hee
2018-05-28
Tamoxifen (TAM) is commonly used to treat estrogen receptor (ER)-positive breast cancer. Despite the remarkable benefits, resistance to TAM presents a serious therapeutic challenge. Since several HOX transcription factors have been proposed as strong candidates in the development of resistance to TAM therapy in breast cancer, we generated an in vitro model of acquired TAM resistance using ER-positive MCF7 breast cancer cells (MCF7-TAMR), and analyzed the expression pattern and epigenetic states of HOX genes. HOXB cluster genes were uniquely up-regulated in MCF7-TAMR cells. Survival analysis of in slico data showed the correlation of high expression of HOXB genes with poor response to TAM in ER-positive breast cancer patients treated with TAM. Gain- and loss-of-function experiments showed that the overexpression of multi HOXB genes in MCF7 renders cancer cells more resistant to TAM, whereas the knockdown restores TAM sensitivity. Furthermore, activation of HOXB genes in MCF7-TAMR was associated with histone modifications, particularly the gain of H3K9ac. These findings imply that the activation of HOXB genes mediate the development of TAM resistance, and represent a target for development of new strategies to prevent or reverse TAM resistance.
Lampreys, the jawless vertebrates, contain only two ParaHox gene clusters.
Zhang, Huixian; Ravi, Vydianathan; Tay, Boon-Hui; Tohari, Sumanty; Pillai, Nisha E; Prasad, Aravind; Lin, Qiang; Brenner, Sydney; Venkatesh, Byrappa
2017-08-22
ParaHox genes ( Gsx , Pdx , and Cdx ) are an ancient family of developmental genes closely related to the Hox genes. They play critical roles in the patterning of brain and gut. The basal chordate, amphioxus, contains a single ParaHox cluster comprising one member of each family, whereas nonteleost jawed vertebrates contain four ParaHox genomic loci with six or seven ParaHox genes. Teleosts, which have experienced an additional whole-genome duplication, contain six ParaHox genomic loci with six ParaHox genes. Jawless vertebrates, represented by lampreys and hagfish, are the most ancient group of vertebrates and are crucial for understanding the origin and evolution of vertebrate gene families. We have previously shown that lampreys contain six Hox gene loci. Here we report that lampreys contain only two ParaHox gene clusters (designated as α- and β-clusters) bearing five ParaHox genes ( Gsxα , Pdxα , Cdxα , Gsxβ , and Cdxβ ). The order and orientation of the three genes in the α-cluster are identical to that of the single cluster in amphioxus. However, the orientation of Gsxβ in the β-cluster is inverted. Interestingly, Gsxβ is expressed in the eye, unlike its homologs in jawed vertebrates, which are expressed mainly in the brain. The lamprey Pdxα is expressed in the pancreas similar to jawed vertebrate Pdx genes, indicating that the pancreatic expression of Pdx was acquired before the divergence of jawless and jawed vertebrate lineages. It is likely that the lamprey Pdxα plays a crucial role in pancreas specification and insulin production similar to the Pdx of jawed vertebrates.
Genome-based exploration of the specialized metabolic capacities of the genus Rhodococcus.
Ceniceros, Ana; Dijkhuizen, Lubbert; Petrusma, Mirjan; Medema, Marnix H
2017-08-09
Bacteria of the genus Rhodococcus are well known for their ability to degrade a large range of organic compounds. Some rhodococci are free-living, saprophytic bacteria; others are animal and plant pathogens. Recently, several studies have shown that their genomes encode putative pathways for the synthesis of a large number of specialized metabolites that are likely to be involved in microbe-microbe and host-microbe interactions. To systematically explore the specialized metabolic potential of this genus, we here performed a comprehensive analysis of the biosynthetic coding capacity across publicly available rhododoccal genomes, and compared these with those of several Mycobacterium strains as well as that of their mutual close relative Amycolicicoccus subflavus. Comparative genomic analysis shows that most predicted biosynthetic gene cluster families in these strains are clade-specific and lack any homology with gene clusters encoding the production of known natural products. Interestingly, many of these clusters appear to encode the biosynthesis of lipopeptides, which may play key roles in the diverse environments were rhodococci thrive, by acting as biosurfactants, pathogenicity factors or antimicrobials. We also identified several gene cluster families that are universally shared among all three genera, which therefore may have a more 'primary' role in their physiology. Inactivation of these clusters by mutagenesis might help to generate weaker strains that can be used as live vaccines. The genus Rhodococcus thus provides an interesting target for natural product discovery, in view of its large and mostly uncharacterized biosynthetic repertoire, its relatively fast growth and the availability of effective genetic tools for its genomic modification.
Sartor, Maureen A.; Schnekenburger, Michael; Marlowe, Jennifer L.; Reichard, John F.; Wang, Ying; Fan, Yunxia; Ma, Ci; Karyala, Saikumar; Halbleib, Danielle; Liu, Xiangdong; Medvedovic, Mario; Puga, Alvaro
2009-01-01
Background The vertebrate aryl hydrocarbon receptor (AHR) is a ligand-activated transcription factor that regulates cellular responses to environmental polycyclic and halogenated compounds. The naive receptor is believed to reside in an inactive cytosolic complex that translocates to the nucleus and induces transcription of xenobiotic detoxification genes after activation by ligand. Objectives We conducted an integrative genomewide analysis of AHR gene targets in mouse hepatoma cells and determined whether AHR regulatory functions may take place in the absence of an exogenous ligand. Methods The network of AHR-binding targets in the mouse genome was mapped through a multipronged approach involving chromatin immunoprecipitation/chip and global gene expression signatures. The findings were integrated into a prior functional knowledge base from Gene Ontology, interaction networks, Kyoto Encyclopedia of Genes and Genomes pathways, sequence motif analysis, and literature molecular concepts. Results We found the naive receptor in unstimulated cells bound to an extensive array of gene clusters with functions in regulation of gene expression, differentiation, and pattern specification, connecting multiple morphogenetic and developmental programs. Activation by the ligand displaced the receptor from some of these targets toward sites in the promoters of xenobiotic metabolism genes. Conclusions The vertebrate AHR appears to possess unsuspected regulatory functions that may be potential targets of environmental injury. PMID:19654925
Genomics of Dementia: APOE- and CYP2D6-Related Pharmacogenetics
Cacabelos, Ramón; Martínez, Rocío; Fernández-Novoa, Lucía; Carril, Juan C.; Lombardi, Valter; Carrera, Iván; Corzo, Lola; Tellado, Iván; Leszek, Jerzy; McKay, Adam; Takeda, Masatoshi
2012-01-01
Dementia is a major problem of health in developed societies. Alzheimer's disease (AD), vascular dementia, and mixed dementia account for over 90% of the most prevalent forms of dementia. Both genetic and environmental factors are determinant for the phenotypic expression of dementia. AD is a complex disorder in which many different gene clusters may be involved. Most genes screened to date belong to different proteomic and metabolomic pathways potentially affecting AD pathogenesis. The ε4 variant of the APOE gene seems to be a major risk factor for both degenerative and vascular dementia. Metabolic factors, cerebrovascular disorders, and epigenetic phenomena also contribute to neurodegeneration. Five categories of genes are mainly involved in pharmacogenomics: genes associated with disease pathogenesis, genes associated with the mechanism of action of a particular drug, genes associated with phase I and phase II metabolic reactions, genes associated with transporters, and pleiotropic genes and/or genes associated with concomitant pathologies. The APOE and CYP2D6 genes have been extensively studied in AD. The therapeutic response to conventional drugs in patients with AD is genotype specific, with CYP2D6-PMs, CYP2D6-UMs, and APOE-4/4 carriers acting as the worst responders. APOE and CYP2D6 may cooperate, as pleiotropic genes, in the metabolism of drugs and hepatic function. The introduction of pharmacogenetic procedures into AD pharmacological treatment may help to optimize therapeutics. PMID:22482072
Li, Jun; Tai, Cui; Deng, Zixin; Zhong, Weihong; He, Yongqun; Ou, Hong-Yu
2017-01-10
VRprofile is a Web server that facilitates rapid investigation of virulence and antibiotic resistance genes, as well as extends these trait transfer-related genetic contexts, in newly sequenced pathogenic bacterial genomes. The used backend database MobilomeDB was firstly built on sets of known gene cluster loci of bacterial type III/IV/VI/VII secretion systems and mobile genetic elements, including integrative and conjugative elements, prophages, class I integrons, IS elements and pathogenicity/antibiotic resistance islands. VRprofile is thus able to co-localize the homologs of these conserved gene clusters using HMMer or BLASTp searches. With the integration of the homologous gene cluster search module with a sequence composition module, VRprofile has exhibited better performance for island-like region predictions than the other widely used methods. In addition, VRprofile also provides an integrated Web interface for aligning and visualizing identified gene clusters with MobilomeDB-archived gene clusters, or a variety set of bacterial genomes. VRprofile might contribute to meet the increasing demands of re-annotations of bacterial variable regions, and aid in the real-time definitions of disease-relevant gene clusters in pathogenic bacteria of interest. VRprofile is freely available at http://bioinfo-mml.sjtu.edu.cn/VRprofile. © The Author 2017. Published by Oxford University Press. All rights reserved. For Permissions, please email: journals.permissions@oup.com.
Zhou, Zhenxing; Xu, Qingqing; Bu, Qingting; Guo, Yuanyang; Liu, Shuiping; Liu, Yu; Du, Yiling; Li, Yongquan
2015-02-09
Genomic sequencing of actinomycetes has revealed the presence of numerous gene clusters seemingly capable of natural product biosynthesis, yet most clusters are cryptic under laboratory conditions. Bioinformatics analysis of the completely sequenced genome of Streptomyces chattanoogensis L10 (CGMCC 2644) revealed a silent angucycline biosynthetic gene cluster. The overexpression of a pathway-specific activator gene under the constitutive ermE* promoter successfully triggered the expression of the angucycline biosynthetic genes. Two novel members of the angucycline antibiotic family, chattamycins A and B, were further isolated and elucidated. Biological activity assays demonstrated that chattamycin B possesses good antitumor activities against human cancer cell lines and moderate antibacterial activities. The results presented here provide a feasible method to activate silent angucycline biosynthetic gene clusters to discover potential new drug leads. © 2015 WILEY-VCH Verlag GmbH & Co. KGaA, Weinheim.
The ergot alkaloid gene cluster: functional analyses and evolutionary aspects.
Lorenz, Nicole; Haarmann, Thomas; Pazoutová, Sylvie; Jung, Manfred; Tudzynski, Paul
2009-01-01
Ergot alkaloids and their derivatives have been traditionally used as therapeutic agents in migraine, blood pressure regulation and help in childbirth and abortion. Their production in submerse culture is a long established biotechnological process. Ergot alkaloids are produced mainly by members of the genus Claviceps, with Claviceps purpurea as best investigated species concerning the biochemistry of ergot alkaloid synthesis (EAS). Genes encoding enzymes involved in EAS have been shown to be clustered; functional analyses of EAS cluster genes have allowed to assign specific functions to several gene products. Various Claviceps species differ with respect to their host specificity and their alkaloid content; comparison of the ergot alkaloid clusters in these species (and of clavine alkaloid clusters in other genera) yields interesting insights into the evolution of cluster structure. This review focuses on recently published and also yet unpublished data on the structure and evolution of the EAS gene cluster and on the function and regulation of cluster genes. These analyses have also significant biotechnological implications: the characterization of non-ribosomal peptide synthetases (NRPS) involved in the synthesis of the peptide moiety of ergopeptines opened interesting perspectives for the synthesis of ergot alkaloids; on the other hand, defined mutants could be generated producing interesting intermediates or only single peptide alkaloids (instead of the alkaloid mixtures usually produced by industrial strains).
Schorn, Michelle A; Alanjary, Mohammad M; Aguinaldo, Kristen; Korobeynikov, Anton; Podell, Sheila; Patin, Nastassia; Lincecum, Tommie; Jensen, Paul R; Ziemert, Nadine; Moore, Bradley S
2016-12-01
Traditional natural product discovery methods have nearly exhausted the accessible diversity of microbial chemicals, making new sources and techniques paramount in the search for new molecules. Marine actinomycete bacteria have recently come into the spotlight as fruitful producers of structurally diverse secondary metabolites, and remain relatively untapped. In this study, we sequenced 21 marine-derived actinomycete strains, rarely studied for their secondary metabolite potential and under-represented in current genomic databases. We found that genome size and phylogeny were good predictors of biosynthetic gene cluster diversity, with larger genomes rivalling the well-known marine producers in the Streptomyces and Salinispora genera. Genomes in the Micrococcineae suborder, however, had consistently the lowest number of biosynthetic gene clusters. By networking individual gene clusters into gene cluster families, we were able to computationally estimate the degree of novelty each genus contributed to the current sequence databases. Based on the similarity measures between all actinobacteria in the Joint Genome Institute's Atlas of Biosynthetic gene Clusters database, rare marine genera show a high degree of novelty and diversity, with Corynebacterium, Gordonia, Nocardiopsis, Saccharomonospora and Pseudonocardia genera representing the highest gene cluster diversity. This research validates that rare marine actinomycetes are important candidates for exploration, as they are relatively unstudied, and their relatives are historically rich in secondary metabolites.
Schorn, Michelle A.; Alanjary, Mohammad M.; Aguinaldo, Kristen; Korobeynikov, Anton; Podell, Sheila; Patin, Nastassia; Lincecum, Tommie; Jensen, Paul R.; Ziemert, Nadine
2016-01-01
Traditional natural product discovery methods have nearly exhausted the accessible diversity of microbial chemicals, making new sources and techniques paramount in the search for new molecules. Marine actinomycete bacteria have recently come into the spotlight as fruitful producers of structurally diverse secondary metabolites, and remain relatively untapped. In this study, we sequenced 21 marine-derived actinomycete strains, rarely studied for their secondary metabolite potential and under-represented in current genomic databases. We found that genome size and phylogeny were good predictors of biosynthetic gene cluster diversity, with larger genomes rivalling the well-known marine producers in the Streptomyces and Salinispora genera. Genomes in the Micrococcineae suborder, however, had consistently the lowest number of biosynthetic gene clusters. By networking individual gene clusters into gene cluster families, we were able to computationally estimate the degree of novelty each genus contributed to the current sequence databases. Based on the similarity measures between all actinobacteria in the Joint Genome Institute's Atlas of Biosynthetic gene Clusters database, rare marine genera show a high degree of novelty and diversity, with Corynebacterium, Gordonia, Nocardiopsis, Saccharomonospora and Pseudonocardia genera representing the highest gene cluster diversity. This research validates that rare marine actinomycetes are important candidates for exploration, as they are relatively unstudied, and their relatives are historically rich in secondary metabolites. PMID:27902408
Methionine sulphoxide reductases protect iron-sulphur clusters from oxidative inactivation in yeast
Sideri, Theodora C.; Willetts, Sylvia A.; Avery, Simon V.
2008-01-01
Methionine residues and iron-sulphur (FeS) clusters are primary targets of reactive oxygen species in the proteins of microorganisms. Here we show that methionine redox-modifications help to preserve essential FeS cluster activities in yeast. Mutants defective for the highly conserved methionine sulphoxide reductases (MSRs; which re-reduce oxidized methionines) are sensitive to many pro-oxidants, but here exhibited an unexpected copper resistance. This phenotype was mimicked by methionine sulphoxide supplementation. Microarray analyses highlighted several Cu and Fe homeostasis genes that were upregulated in the mxrΔ double mutant, which lacks both of the yeast MSRs. Of the upregulated genes, the Cu-binding Fe-transporter Fet3p proved to be required for the Cu-resistance phenotype. FET3 is known to be regulated by the Aft1 transcription factor, which responds to low mitochondrial FeS-cluster status. Here, constitutive Aft1p expression in the wild type reproduced the Cu-resistance phenotype, and FeS cluster functions were found to be defective in the mxrΔ mutant. Genetic perturbation of FeS activity also mimicked FET3-dependent Cu resistance. 55Fe-labeling studies showed that FeS clusters are turned over more rapidly in the mxrΔ mutant than the wild type, consistent with elevated oxidative targeting of the clusters in MSR-deficient cells. The potential underlying molecular mechanisms of this targeting are discussed. Moreover, the results indicate an important new role for cellular MSR enzymes, in helping to protect the essential function of FeS clusters in aerobic settings. PMID:19202110
Scoring clustering solutions by their biological relevance.
Gat-Viks, I; Sharan, R; Shamir, R
2003-12-12
A central step in the analysis of gene expression data is the identification of groups of genes that exhibit similar expression patterns. Clustering gene expression data into homogeneous groups was shown to be instrumental in functional annotation, tissue classification, regulatory motif identification, and other applications. Although there is a rich literature on clustering algorithms for gene expression analysis, very few works addressed the systematic comparison and evaluation of clustering results. Typically, different clustering algorithms yield different clustering solutions on the same data, and there is no agreed upon guideline for choosing among them. We developed a novel statistically based method for assessing a clustering solution according to prior biological knowledge. Our method can be used to compare different clustering solutions or to optimize the parameters of a clustering algorithm. The method is based on projecting vectors of biological attributes of the clustered elements onto the real line, such that the ratio of between-groups and within-group variance estimators is maximized. The projected data are then scored using a non-parametric analysis of variance test, and the score's confidence is evaluated. We validate our approach using simulated data and show that our scoring method outperforms several extant methods, including the separation to homogeneity ratio and the silhouette measure. We apply our method to evaluate results of several clustering methods on yeast cell-cycle gene expression data. The software is available from the authors upon request.
Bobe, Julien; Montfort, Jerôme; Nguyen, Thaovi; Fostier, Alexis
2006-01-01
Background The hormonal control of oocyte maturation and ovulation as well as the molecular mechanisms of nuclear maturation have been thoroughly studied in fish. In contrast, the other molecular events occurring in the ovary during post-vitellogenesis have received far less attention. Methods Nylon microarrays displaying 9152 rainbow trout cDNAs were hybridized using RNA samples originating from ovarian tissue collected during late vitellogenesis, post-vitellogenesis and oocyte maturation. Differentially expressed genes were identified using a statistical analysis. A supervised clustering analysis was performed using only differentially expressed genes in order to identify gene clusters exhibiting similar expression profiles. In addition, specific genes were selected and their preovulatory ovarian expression was analyzed using real-time PCR. Results From the statistical analysis, 310 differentially expressed genes were identified. Among those genes, 90 were up-regulated at the time of oocyte maturation while 220 exhibited an opposite pattern. After clustering analysis, 90 clones belonging to 3 gene clusters exhibiting the most remarkable expression patterns were kept for further analysis. Using real-time PCR analysis, we observed a strong up-regulation of ion and water transport genes such as aquaporin 4 (aqp4) and pendrin (slc26). In addition, a dramatic up-regulation of vasotocin (avt) gene was observed. Furthermore, angiotensin-converting-enzyme 2 (ace2), coagulation factor V (cf5), adam 22, and the chemokine cxcl14 genes exhibited a sharp up-regulation at the time of oocyte maturation. Finally, ovarian aromatase (cyp19a1) exhibited a dramatic down-regulation over the post-vitellogenic period while a down-regulation of Cytidine monophosphate-N-acetylneuraminic acid hydroxylase (cmah) was observed at the time of oocyte maturation. Conclusion We showed the over or under expression of more that 300 genes, most of them being previously unstudied or unknown in the fish preovulatory ovary. Our data confirmed the down-regulation of estrogen synthesis genes during the preovulatory period. In addition, the strong up-regulation of aqp4 and slc26 genes prior to ovulation suggests their participation in the oocyte hydration process occurring at that time. Furthermore, among the most up-regulated clones, several genes such as cxcl14, ace2, adam22, cf5 have pro-inflammatory, vasodilatory, proteolytics and coagulatory functions. The identity and expression patterns of those genes support the theory comparing ovulation to an inflammatory-like reaction. PMID:16872517
DOE Office of Scientific and Technical Information (OSTI.GOV)
Liebhaber, S.A.; Weiss, I.; Cash, F.E.
Synthesis of normal human hemoglobin A, {alpha}{sub 2}{beta}{sub 2}, is based upon balanced expression of genes in the {alpha}-globin gene cluster on chromosome 15 and the {beta}-globin gene cluster on chromosome 11. Full levels of erythroid-specific activation of the {beta}-globin cluster depend on sequences located at a considerable distance 5{prime} to the {beta}-globin gene, referred to as the locus-activating or dominant control region. The existence of an analogous element(s) upstream of the {alpha}-globin cluster has been suggested from observations on naturally occurring deletions and experimental studies. The authors have identified an individual with {alpha}-thalassemia in whom structurally normal {alpha}-globin genesmore » have been inactivated in cis by a discrete de novo 35-kilobase deletion located {approximately}30 kilobases 5{prime} from the {alpha}-globin gene cluster. They conclude that this deletion inactivates expression of the {alpha}-globin genes by removing one or more of the previously identified upstream regulatory sequences that are critical to expression of the {alpha}-globin genes.« less
Discovery of a Phosphonoacetic Acid Derived Natural Product by Pathway Refactoring.
Freestone, Todd S; Ju, Kou-San; Wang, Bin; Zhao, Huimin
2017-02-17
The activation of silent natural product gene clusters is a synthetic biology problem of great interest. As the rate at which gene clusters are identified outpaces the discovery rate of new molecules, this unknown chemical space is rapidly growing, as too are the rewards for developing technologies to exploit it. One class of natural products that has been underrepresented is phosphonic acids, which have important medical and agricultural uses. Hundreds of phosphonic acid biosynthetic gene clusters have been identified encoding for unknown molecules. Although methods exist to elicit secondary metabolite gene clusters in native hosts, they require the strain to be amenable to genetic manipulation. One method to circumvent this is pathway refactoring, which we implemented in an effort to discover new phosphonic acids from a gene cluster from Streptomyces sp. strain NRRL F-525. By reengineering this cluster for expression in the production host Streptomyces lividans, utility of refactoring is demonstrated with the isolation of a novel phosphonic acid, O-phosphonoacetic acid serine, and the characterization of its biosynthesis. In addition, a new biosynthetic branch point is identified with a phosphonoacetaldehyde dehydrogenase, which was used to identify additional phosphonic acid gene clusters that share phosphonoacetic acid as an intermediate.
The intact dupA cluster is a more reliable Helicobacter pylori virulence marker than dupA alone.
Jung, Sung Woo; Sugimoto, Mitsushige; Shiota, Seiji; Graham, David Y; Yamaoka, Yoshio
2012-01-01
The duodenal ulcer promoting (dupA) gene, located in the plasticity region of Helicobacter pylori, is associated with duodenal ulcer development. dupA was predicted to form a type IV secretory system (T4SS) with vir genes around dupA (dupA cluster). We investigated the prevalence of dupA and dupA clusters and clarified associations between the dupA cluster status and clinical outcomes in the U.S. population. In all, 245 H. pylori strains were examined using PCR to evaluate the status of dupA and the adjacent vir genes predicted to form T4SS, in addition to the status of cag pathogenicity island (PAI). The associations between dupA cluster status and interleukin-8 (IL-8) and IL-12 production were also examined. The presence of dupA and all adjacent vir genes were defined as a complete dupA cluster. Many variations related to the status of dupA and dupA cluster genes were identified. Concurrent H. pylori infection and the presence of a complete dupA cluster increases duodenal ulcer risk compared to H. pylori infection with incomplete dupA cluster or without the dupA gene independent on the cag PAI status (adjusted odds ratio, 2.13; 95% confidence interval, 1.13 to 4.03). Gastric mucosal IL-8 levels were also significantly higher in the complete dupA cluster group than in other groups (P=0.01). In conclusion, although the causal relationship between the dupA cluster and duodenal ulcer development is not proved, the presence of a complete dupA cluster but not dupA alone, is associated with duodenal ulcer development.
The Intact dupA Cluster Is a More Reliable Helicobacter pylori Virulence Marker than dupA Alone
Jung, Sung Woo; Sugimoto, Mitsushige; Shiota, Seiji; Graham, David Y.
2012-01-01
The duodenal ulcer promoting (dupA) gene, located in the plasticity region of Helicobacter pylori, is associated with duodenal ulcer development. dupA was predicted to form a type IV secretory system (T4SS) with vir genes around dupA (dupA cluster). We investigated the prevalence of dupA and dupA clusters and clarified associations between the dupA cluster status and clinical outcomes in the U.S. population. In all, 245 H. pylori strains were examined using PCR to evaluate the status of dupA and the adjacent vir genes predicted to form T4SS, in addition to the status of cag pathogenicity island (PAI). The associations between dupA cluster status and interleukin-8 (IL-8) and IL-12 production were also examined. The presence of dupA and all adjacent vir genes were defined as a complete dupA cluster. Many variations related to the status of dupA and dupA cluster genes were identified. Concurrent H. pylori infection and the presence of a complete dupA cluster increases duodenal ulcer risk compared to H. pylori infection with incomplete dupA cluster or without the dupA gene independent on the cag PAI status (adjusted odds ratio, 2.13; 95% confidence interval, 1.13 to 4.03). Gastric mucosal IL-8 levels were also significantly higher in the complete dupA cluster group than in other groups (P = 0.01). In conclusion, although the causal relationship between the dupA cluster and duodenal ulcer development is not proved, the presence of a complete dupA cluster but not dupA alone, is associated with duodenal ulcer development. PMID:22038914
Fast gene ontology based clustering for microarray experiments.
Ovaska, Kristian; Laakso, Marko; Hautaniemi, Sampsa
2008-11-21
Analysis of a microarray experiment often results in a list of hundreds of disease-associated genes. In order to suggest common biological processes and functions for these genes, Gene Ontology annotations with statistical testing are widely used. However, these analyses can produce a very large number of significantly altered biological processes. Thus, it is often challenging to interpret GO results and identify novel testable biological hypotheses. We present fast software for advanced gene annotation using semantic similarity for Gene Ontology terms combined with clustering and heat map visualisation. The methodology allows rapid identification of genes sharing the same Gene Ontology cluster. Our R based semantic similarity open-source package has a speed advantage of over 2000-fold compared to existing implementations. From the resulting hierarchical clustering dendrogram genes sharing a GO term can be identified, and their differences in the gene expression patterns can be seen from the heat map. These methods facilitate advanced annotation of genes resulting from data analysis.
Cormier, Hubert; Rudkowska, Iwona; Paradis, Ann-Marie; Thifault, Elisabeth; Garneau, Véronique; Lemieux, Simone; Couture, Patrick; Vohl, Marie-Claude
2012-08-01
Eicosapentaenoic and docosahexaenoic acids have been reported to have a variety of beneficial effects on cardiovascular disease risk factors. However, a large inter-individual variability in the plasma lipid response to an omega-3 (n-3) polyunsaturated fatty acid (PUFA) supplementation is observed in different studies. Genetic variations may influence plasma lipid responsiveness. The aim of the present study was to examine the effects of a supplementation with n-3 PUFA on the plasma lipid profile in relation to the presence of single-nucleotide polymorphisms (SNPs) in the fatty acid desaturase (FADS) gene cluster. A total of 208 subjects from Quebec City area were supplemented with 3 g/day of n-3 PUFA, during six weeks. In a statistical model including the effect of the genotype, the supplementation and the genotype by supplementation interaction, SNP rs174546 was significantly associated (p = 0.02) with plasma triglyceride (TG) levels, pre- and post-supplementation. The n-3 supplementation had an independent effect on plasma TG levels and no significant genotype by supplementation interaction effects were observed. In summary, our data support the notion that the FADS gene cluster is a major determinant of plasma TG levels. SNP rs174546 may be an important SNP associated with plasma TG levels and FADS1 gene expression independently of a nutritional intervention with n-3 PUFA.
Cormier, Hubert; Rudkowska, Iwona; Paradis, Ann-Marie; Thifault, Elisabeth; Garneau, Véronique; Lemieux, Simone; Couture, Patrick; Vohl, Marie-Claude
2012-01-01
Eicosapentaenoic and docosahexaenoic acids have been reported to have a variety of beneficial effects on cardiovascular disease risk factors. However, a large inter-individual variability in the plasma lipid response to an omega-3 (n-3) polyunsaturated fatty acid (PUFA) supplementation is observed in different studies. Genetic variations may influence plasma lipid responsiveness. The aim of the present study was to examine the effects of a supplementation with n-3 PUFA on the plasma lipid profile in relation to the presence of single-nucleotide polymorphisms (SNPs) in the fatty acid desaturase (FADS) gene cluster. A total of 208 subjects from Quebec City area were supplemented with 3 g/day of n-3 PUFA, during six weeks. In a statistical model including the effect of the genotype, the supplementation and the genotype by supplementation interaction, SNP rs174546 was significantly associated (p = 0.02) with plasma triglyceride (TG) levels, pre- and post-supplementation. The n-3 supplementation had an independent effect on plasma TG levels and no significant genotype by supplementation interaction effects were observed. In summary, our data support the notion that the FADS gene cluster is a major determinant of plasma TG levels. SNP rs174546 may be an important SNP associated with plasma TG levels and FADS1 gene expression independently of a nutritional intervention with n-3 PUFA. PMID:23016130
Campbell, Elsie L; Hagen, Kari D; Chen, Rui; Risser, Douglas D; Ferreira, Daniela P; Meeks, John C
2015-02-15
In cyanobacterial Nostoc species, substratum-dependent gliding motility is confined to specialized nongrowing filaments called hormogonia, which differentiate from vegetative filaments as part of a conditional life cycle and function as dispersal units. Here we confirm that Nostoc punctiforme hormogonia are positively phototactic to white light over a wide range of intensities. N. punctiforme contains two gene clusters (clusters 2 and 2i), each of which encodes modular cyanobacteriochrome-methyl-accepting chemotaxis proteins (MCPs) and other proteins that putatively constitute a basic chemotaxis-like signal transduction complex. Transcriptional analysis established that all genes in clusters 2 and 2i, plus two additional clusters (clusters 1 and 3) with genes encoding MCPs lacking cyanobacteriochrome sensory domains, are upregulated during the differentiation of hormogonia. Mutational analysis determined that only genes in cluster 2i are essential for positive phototaxis in N. punctiforme hormogonia; here these genes are designated ptx (for phototaxis) genes. The cluster is unusual in containing complete or partial duplicates of genes encoding proteins homologous to the well-described chemotaxis elements CheY, CheW, MCP, and CheA. The cyanobacteriochrome-MCP gene (ptxD) lacks transmembrane domains and has 7 potential binding sites for bilins. The transcriptional start site of the ptx genes does not resemble a sigma 70 consensus recognition sequence; moreover, it is upstream of two genes encoding gas vesicle proteins (gvpA and gvpC), which also are expressed only in the hormogonium filaments of N. punctiforme. Copyright © 2015, American Society for Microbiology. All Rights Reserved.
Molecular Analysis of SCARECROW Genes Expressed in White Lupin Cluster Roots
USDA-ARS?s Scientific Manuscript database
The Scarecrow (SCR) transcription factor plays a crucial role in root cell radial patterning and is required for maintenance of the quiescent center and differentiation of the endodermis. In response to phosphorus (P) deficiency, white lupin (Lupinus albus L.) root surface area increases some 50- to...
Identifying a gene expression signature of cluster headache in blood
Eising, Else; Pelzer, Nadine; Vijfhuizen, Lisanne S.; Vries, Boukje de; Ferrari, Michel D.; ‘t Hoen, Peter A. C.; Terwindt, Gisela M.; van den Maagdenberg, Arn M. J. M.
2017-01-01
Cluster headache is a relatively rare headache disorder, typically characterized by multiple daily, short-lasting attacks of excruciating, unilateral (peri-)orbital or temporal pain associated with autonomic symptoms and restlessness. To better understand the pathophysiology of cluster headache, we used RNA sequencing to identify differentially expressed genes and pathways in whole blood of patients with episodic (n = 19) or chronic (n = 20) cluster headache in comparison with headache-free controls (n = 20). Gene expression data were analysed by gene and by module of co-expressed genes with particular attention to previously implicated disease pathways including hypocretin dysregulation. Only moderate gene expression differences were identified and no associations were found with previously reported pathogenic mechanisms. At the level of functional gene sets, associations were observed for genes involved in several brain-related mechanisms such as GABA receptor function and voltage-gated channels. In addition, genes and modules of co-expressed genes showed a role for intracellular signalling cascades, mitochondria and inflammation. Although larger study samples may be required to identify the full range of involved pathways, these results indicate a role for mitochondria, intracellular signalling and inflammation in cluster headache. PMID:28074859
Huang, Shi-Ming; Zhao, Xia; Zhao, Xue-Mei; Wang, Xiao-Ying; Li, Shan-Shan; Zhu, Yu-Hui
2014-01-01
Renal transplantation is the preferred method for most patients with end-stage renal disease, however, acute renal allograft rejection is still a major risk factor for recipients leading to renal injury. To improve the early diagnosis and treatment of acute rejection, study on the molecular mechanism of it is urgent. MicroRNA (miRNA) expression profile and mRNA expression profile of acute renal allograft rejection and well-functioning allograft downloaded from ArrayExpress database were applied to identify differentially expressed (DE) miRNAs and DE mRNAs. DE miRNAs targets were predicted by combining five algorithm. By overlapping the DE mRNAs and DE miRNAs targets, common genes were obtained. Differentially co-expressed genes (DCGs) were identified by differential co-expression profile (DCp) and differential co-expression enrichment (DCe) methods in Differentially Co-expressed Genes and Links (DCGL) package. Then, co-expression network of DCGs and the cluster analysis were performed. Functional enrichment analysis for DCGs was undergone. A total of 1270 miRNA targets were predicted and 698 DE mRNAs were obtained. While overlapping miRNA targets and DE mRNAs, 59 common genes were gained. We obtained 103 DCGs and 5 transcription factors (TFs) based on regulatory impact factors (RIF), then built the regulation network of miRNA targets and DE mRNAs. By clustering the co-expression network, 5 modules were obtained. Thereinto, module 1 had the highest degree and module 2 showed the most number of DCGs and common genes. TF CEBPB and several common genes, such as RXRA, BASP1 and AKAP10, were mapped on the co-expression network. C1R showed the highest degree in the network. These genes might be associated with human acute renal allograft rejection. We conducted biological analysis on integration of DE mRNA and DE miRNA in acute renal allograft rejection, displayed gene expression patterns and screened out genes and TFs that may be related to acute renal allograft rejection.
Huang, Shi-Ming; Zhao, Xia; Zhao, Xue-Mei; Wang, Xiao-Ying; Li, Shan-Shan; Zhu, Yu-Hui
2014-01-01
Objectives: Renal transplantation is the preferred method for most patients with end-stage renal disease, however, acute renal allograft rejection is still a major risk factor for recipients leading to renal injury. To improve the early diagnosis and treatment of acute rejection, study on the molecular mechanism of it is urgent. Methods: MicroRNA (miRNA) expression profile and mRNA expression profile of acute renal allograft rejection and well-functioning allograft downloaded from ArrayExpress database were applied to identify differentially expressed (DE) miRNAs and DE mRNAs. DE miRNAs targets were predicted by combining five algorithm. By overlapping the DE mRNAs and DE miRNAs targets, common genes were obtained. Differentially co-expressed genes (DCGs) were identified by differential co-expression profile (DCp) and differential co-expression enrichment (DCe) methods in Differentially Co-expressed Genes and Links (DCGL) package. Then, co-expression network of DCGs and the cluster analysis were performed. Functional enrichment analysis for DCGs was undergone. Results: A total of 1270 miRNA targets were predicted and 698 DE mRNAs were obtained. While overlapping miRNA targets and DE mRNAs, 59 common genes were gained. We obtained 103 DCGs and 5 transcription factors (TFs) based on regulatory impact factors (RIF), then built the regulation network of miRNA targets and DE mRNAs. By clustering the co-expression network, 5 modules were obtained. Thereinto, module 1 had the highest degree and module 2 showed the most number of DCGs and common genes. TF CEBPB and several common genes, such as RXRA, BASP1 and AKAP10, were mapped on the co-expression network. C1R showed the highest degree in the network. These genes might be associated with human acute renal allograft rejection. Conclusions: We conducted biological analysis on integration of DE mRNA and DE miRNA in acute renal allograft rejection, displayed gene expression patterns and screened out genes and TFs that may be related to acute renal allograft rejection. PMID:25664019
2016-01-01
Color variation provides the opportunity to investigate the genetic basis of evolution and selection. Reptiles are less studied than mammals. Comparative genomics approaches allow for knowledge gained in one species to be leveraged for use in another species. We describe a comparative vertebrate analysis of conserved regulatory modules in pythons aimed at assessing bioinformatics evidence that transcription factors important in mammalian pigmentation phenotypes may also be important in python pigmentation phenotypes. We identified 23 python orthologs of mammalian genes associated with variation in coat color phenotypes for which we assessed the extent of pairwise protein sequence identity between pythons and mouse, dog, horse, cow, chicken, anole lizard, and garter snake. We next identified a set of melanocyte/pigment associated transcription factors (CREB, FOXD3, LEF-1, MITF, POU3F2, and USF-1) that exhibit relatively conserved sequence similarity within their DNA binding regions across species based on orthologous alignments across multiple species. Finally, we identified 27 evolutionarily conserved clusters of transcription factor binding sites within ~200-nucleotide intervals of the 1500-nucleotide upstream regions of AIM1, DCT, MC1R, MITF, MLANA, OA1, PMEL, RAB27A, and TYR from Python bivittatus. Our results provide insight into pigment phenotypes in pythons. PMID:27698666
Irizarry, Kristopher J L; Bryden, Randall L
2016-01-01
Color variation provides the opportunity to investigate the genetic basis of evolution and selection. Reptiles are less studied than mammals. Comparative genomics approaches allow for knowledge gained in one species to be leveraged for use in another species. We describe a comparative vertebrate analysis of conserved regulatory modules in pythons aimed at assessing bioinformatics evidence that transcription factors important in mammalian pigmentation phenotypes may also be important in python pigmentation phenotypes. We identified 23 python orthologs of mammalian genes associated with variation in coat color phenotypes for which we assessed the extent of pairwise protein sequence identity between pythons and mouse, dog, horse, cow, chicken, anole lizard, and garter snake. We next identified a set of melanocyte/pigment associated transcription factors (CREB, FOXD3, LEF-1, MITF, POU3F2, and USF-1) that exhibit relatively conserved sequence similarity within their DNA binding regions across species based on orthologous alignments across multiple species. Finally, we identified 27 evolutionarily conserved clusters of transcription factor binding sites within ~200-nucleotide intervals of the 1500-nucleotide upstream regions of AIM1, DCT, MC1R, MITF, MLANA, OA1, PMEL, RAB27A, and TYR from Python bivittatus . Our results provide insight into pigment phenotypes in pythons.
2012-01-01
Background Time-course gene expression data such as yeast cell cycle data may be periodically expressed. To cluster such data, currently used Fourier series approximations of periodic gene expressions have been found not to be sufficiently adequate to model the complexity of the time-course data, partly due to their ignoring the dependence between the expression measurements over time and the correlation among gene expression profiles. We further investigate the advantages and limitations of available models in the literature and propose a new mixture model with autoregressive random effects of the first order for the clustering of time-course gene-expression profiles. Some simulations and real examples are given to demonstrate the usefulness of the proposed models. Results We illustrate the applicability of our new model using synthetic and real time-course datasets. We show that our model outperforms existing models to provide more reliable and robust clustering of time-course data. Our model provides superior results when genetic profiles are correlated. It also gives comparable results when the correlation between the gene profiles is weak. In the applications to real time-course data, relevant clusters of coregulated genes are obtained, which are supported by gene-function annotation databases. Conclusions Our new model under our extension of the EMMIX-WIRE procedure is more reliable and robust for clustering time-course data because it adopts a random effects model that allows for the correlation among observations at different time points. It postulates gene-specific random effects with an autocorrelation variance structure that models coregulation within the clusters. The developed R package is flexible in its specification of the random effects through user-input parameters that enables improved modelling and consequent clustering of time-course data. PMID:23151154
Kato, Hiroki; Tsunematsu, Yuta; Yamamoto, Tsuyoshi; Namiki, Takuya; Kishimoto, Shinji; Noguchi, Hiroshi; Watanabe, Kenji
2016-07-01
To rapidly identify novel natural products and their associated biosynthetic genes from underutilized and genetically difficult-to-manipulate microbes, we developed a method that uses (1) chemical screening to isolate novel microbial secondary metabolites, (2) bioinformatic analyses to identify a potential biosynthetic gene cluster and (3) heterologous expression of the genes in a convenient host to confirm the identity of the gene cluster and the proposed biosynthetic mechanism. The chemical screen was achieved by searching known natural product databases with data from liquid chromatographic and high-resolution mass spectrometric analyses collected on the extract from a target microbe culture. Using this method, we were able to isolate two new meroterpenes, subglutinols C (1) and D (2), from an entomopathogenic filamentous fungus Metarhizium robertsii ARSEF 23. Bioinformatics analysis of the genome allowed us to identify a gene cluster likely to be responsible for the formation of subglutinols. Heterologous expression of three genes from the gene cluster encoding a polyketide synthase, a prenyltransferase and a geranylgeranyl pyrophosphate synthase in Aspergillus nidulans A1145 afforded an α-pyrone-fused uncyclized diterpene, the expected intermediate of the subglutinol biosynthesis, thereby confirming the gene cluster to be responsible for the subglutinol biosynthesis. These results indicate the usefulness of our methodology in isolating new natural products and identifying their associated biosynthetic gene cluster from microbes that are not amenable to genetic manipulation. Our method should facilitate the natural product discovery efforts by expediting the identification of new secondary metabolites and their associated biosynthetic genes from a wider source of microbes.
Bushley, Kathryn E.; Raja, Rajani; Jaiswal, Pankaj; Cumbie, Jason S.; Nonogaki, Mariko; Boyd, Alexander E.; Owensby, C. Alisha; Knaus, Brian J.; Elser, Justin; Miller, Daniel; Di, Yanming; McPhail, Kerry L.; Spatafora, Joseph W.
2013-01-01
The ascomycete fungus Tolypocladium inflatum, a pathogen of beetle larvae, is best known as the producer of the immunosuppressant drug cyclosporin. The draft genome of T. inflatum strain NRRL 8044 (ATCC 34921), the isolate from which cyclosporin was first isolated, is presented along with comparative analyses of the biosynthesis of cyclosporin and other secondary metabolites in T. inflatum and related taxa. Phylogenomic analyses reveal previously undetected and complex patterns of homology between the nonribosomal peptide synthetase (NRPS) that encodes for cyclosporin synthetase (simA) and those of other secondary metabolites with activities against insects (e.g., beauvericin, destruxins, etc.), and demonstrate the roles of module duplication and gene fusion in diversification of NRPSs. The secondary metabolite gene cluster responsible for cyclosporin biosynthesis is described. In addition to genes necessary for cyclosporin biosynthesis, it harbors a gene for a cyclophilin, which is a member of a family of immunophilins known to bind cyclosporin. Comparative analyses support a lineage specific origin of the cyclosporin gene cluster rather than horizontal gene transfer from bacteria or other fungi. RNA-Seq transcriptome analyses in a cyclosporin-inducing medium delineate the boundaries of the cyclosporin cluster and reveal high levels of expression of the gene cluster cyclophilin. In medium containing insect hemolymph, weaker but significant upregulation of several genes within the cyclosporin cluster, including the highly expressed cyclophilin gene, was observed. T. inflatum also represents the first reference draft genome of Ophiocordycipitaceae, a third family of insect pathogenic fungi within the fungal order Hypocreales, and supports parallel and qualitatively distinct radiations of insect pathogens. The T. inflatum genome provides additional insight into the evolution and biosynthesis of cyclosporin and lays a foundation for further investigations of the role of secondary metabolite gene clusters and their metabolites in fungal biology. PMID:23818858
Outcome-Driven Cluster Analysis with Application to Microarray Data.
Hsu, Jessie J; Finkelstein, Dianne M; Schoenfeld, David A
2015-01-01
One goal of cluster analysis is to sort characteristics into groups (clusters) so that those in the same group are more highly correlated to each other than they are to those in other groups. An example is the search for groups of genes whose expression of RNA is correlated in a population of patients. These genes would be of greater interest if their common level of RNA expression were additionally predictive of the clinical outcome. This issue arose in the context of a study of trauma patients on whom RNA samples were available. The question of interest was whether there were groups of genes that were behaving similarly, and whether each gene in the cluster would have a similar effect on who would recover. For this, we develop an algorithm to simultaneously assign characteristics (genes) into groups of highly correlated genes that have the same effect on the outcome (recovery). We propose a random effects model where the genes within each group (cluster) equal the sum of a random effect, specific to the observation and cluster, and an independent error term. The outcome variable is a linear combination of the random effects of each cluster. To fit the model, we implement a Markov chain Monte Carlo algorithm based on the likelihood of the observed data. We evaluate the effect of including outcome in the model through simulation studies and describe a strategy for prediction. These methods are applied to trauma data from the Inflammation and Host Response to Injury research program, revealing a clustering of the genes that are informed by the recovery outcome.
Transcriptome Analysis of a Premature Leaf Senescence Mutant of Common Wheat (Triticum aestivum L.)
Xia, Chuan; Zhang, Lichao; Dong, Chunhao; Liu, Xu; Kong, Xiuying
2018-01-01
Leaf senescence is an important agronomic trait that affects both crop yield and quality. In this study, we characterized a premature leaf senescence mutant of wheat (Triticum aestivum L.) obtained by ethylmethane sulfonate (EMS) mutagenesis, named m68. Genetic analysis showed that the leaf senescence phenotype of m68 is controlled by a single recessive nuclear gene. We compared the transcriptome of wheat leaves between the wild type (WT) and the m68 mutant at four time points. Differentially expressed gene (DEG) analysis revealed many genes that were closely related to senescence genes. Gene Ontology (GO) enrichment analysis suggested that transcription factors and protein transport genes might function in the beginning of leaf senescence, while genes that were associated with chlorophyll and carbon metabolism might function in the later stage. Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analysis showed that the genes that are involved in plant hormone signal transduction were significantly enriched. Through expression pattern clustering of DEGs, we identified 1012 genes that were induced during senescence, and we found that the WRKY family and zinc finger transcription factors might be more important than other transcription factors in the early stage of leaf senescence. These results will not only support further gene cloning and functional analysis of m68, but also facilitate the study of leaf senescence in wheat. PMID:29534430
da Costa, L B; Rajala-Schultz, P J; Hoet, A; Seo, K S; Fogt, K; Moon, B S
2014-11-01
The objective of this study was to assess the role of teat skin colonization in Staphylococcus aureus intramammary infections (IMI) by evaluating genetic relatedness of Staph. aureus isolates from milk and teat skin of dairy cows using pulsed-field gel electrophoresis and characterizing the isolates based on the carriage of virulence genes. Cows in 4 known Staph. aureus-positive herds were sampled and Staph. aureus was detected in 43 quarters of 20 cows, with 10 quarters positive in both milk and skin (20 isolates), 18 positive only in milk, and 15 only on teat skin. Quarters with teat skin colonized with Staph. aureus were 4.5 times more likely to be diagnosed with Staph. aureus IMI than quarters not colonized on teat skin. Three main clusters were identified by pulsed-field gel electrophoresis using a cutoff of 80% similarity. All 3 clusters included both milk and skin isolates. The majority of isolates (72%) belonged to one predominant cluster (B), with 60% of isolates in the cluster originating from milk and 40% from teat skin. Genotypic variability was observed within 10 pairs (formed by isolates originating from milk and teat skin of the same quarter), where isolates in 5 out of the 10 pairs belonged to the same cluster. Forty-two virulence factors were screened using PCR. Some virulence factors were carried more frequently by teat skin isolates than by milk isolates or isolates from quarters with high somatic cell counts. Isolates in the predominant cluster B carried virulence factors clfA and clfB significantly more often than isolates in the minor clusters, which may have assisted them in becoming predominant in the herds. The present findings suggest that teat skin colonization with Staph. aureus can be an important factor involved in Staph. aureus IMI. Copyright © 2014 American Dairy Science Association. Published by Elsevier Inc. All rights reserved.
González-Pedrajo, Bertha; de la Mora, Javier; Ballado, Teresa; Camarena, Laura; Dreyfus, Georges
2002-11-13
In this work, we show evidence regarding the functionality of a large cluster of flagellar genes in Rhodobacter sphaeroides. The genes of this cluster, flgGHIJKL and orf-1, are mainly involved in the formation of the basal body, and flgK and flgL encode the hook-associated proteins HAP1 and HAP3. In general, these genes showed a good similarity as compared with those reported for Salmonella enterica. However, flgJ and flgK showed particular features that make them unique among the flagellar sequences already reported. flgJ is only a third of the size reported for flgJ from Salmonella; whereas flgK is about three times larger than any other flgK sequence previously known. Our results indicate that both genes are functional, and their products are essential for flagellar assembly. In contrast, the interruption of orf-1, did not affect motility suggesting that this sequence, if functional, is not indispensable for flagellar assembly. Finally, we present genetic evidence suggesting that the flgGHIJKL genes are expressed as a single transcriptional unit depending on the sigma-54 factor.
Engert, Christoph G; Droste, Rita; van Oudenaarden, Alexander; Horvitz, H Robert
2018-04-01
To better understand the tissue-specific regulation of chromatin state in cell-fate determination and animal development, we defined the tissue-specific expression of all 36 C. elegans presumptive lysine methyltransferase (KMT) genes using single-molecule fluorescence in situ hybridization (smFISH). Most KMTs were expressed in only one or two tissues. The germline was the tissue with the broadest KMT expression. We found that the germline-expressed C. elegans protein SET-17, which has a SET domain similar to that of the PRDM9 and PRDM7 SET-domain proteins, promotes fertility by regulating gene expression in primary spermatocytes. SET-17 drives the transcription of spermatocyte-specific genes from four genomic clusters to promote spermatid development. SET-17 is concentrated in stable chromatin-associated nuclear foci at actively transcribed msp (major sperm protein) gene clusters, which we term msp locus bodies. Our results reveal the function of a PRDM9/7-family SET-domain protein in spermatocyte transcription. We propose that the spatial intranuclear organization of chromatin factors might be a conserved mechanism in tissue-specific control of transcription.
Cancer Detection in Microarray Data Using a Modified Cat Swarm Optimization Clustering Approach
M, Pandi; R, Balamurugan; N, Sadhasivam
2017-12-29
Objective: A better understanding of functional genomics can be obtained by extracting patterns hidden in gene expression data. This could have paramount implications for cancer diagnosis, gene treatments and other domains. Clustering may reveal natural structures and identify interesting patterns in underlying data. The main objective of this research was to derive a heuristic approach to detection of highly co-expressed genes related to cancer from gene expression data with minimum Mean Squared Error (MSE). Methods: A modified CSO algorithm using Harmony Search (MCSO-HS) for clustering cancer gene expression data was applied. Experiment results are analyzed using two cancer gene expression benchmark datasets, namely for leukaemia and for breast cancer. Result: The results indicated MCSO-HS to be better than HS and CSO, 13% and 9% with the leukaemia dataset. For breast cancer dataset improvement was by 22% and 17%, respectively, in terms of MSE. Conclusion: The results showed MCSO-HS to outperform HS and CSO with both benchmark datasets. To validate the clustering results, this work was tested with internal and external cluster validation indices. Also this work points to biological validation of clusters with gene ontology in terms of function, process and component. Creative Commons Attribution License
Wan, B; Yarbrough, J W; Schultz, T W
2008-01-01
This study was undertaken to test the hypothesis that structurally similar PAHs induce similar gene expression profiles. THP-1 cells were exposed to a series of 12 selected PAHs at 50 microM for 24 hours and gene expressions profiles were analyzed using both unsupervised and supervised methods. Clustering analysis of gene expression profiles revealed that the 12 tested chemicals were grouped into five clusters. Within each cluster, the gene expression profiles are more similar to each other than to the ones outside the cluster. One-methylanthracene and 1-methylfluorene were found to have the most similar profiles; dibenzothiophene and dibenzofuran were found to share common profiles with fluorine. As expression pattern comparisons were expanded, similarity in genomic fingerprint dropped off dramatically. Prediction analysis of microarrays (PAM) based on the clustering pattern generated 49 predictor genes that can be used for sample discrimination. Moreover, a significant analysis of Microarrays (SAM) identified 598 genes being modulated by tested chemicals with a variety of biological processes, such as cell cycle, metabolism, and protein binding and KEGG pathways being significantly (p < 0.05) affected. It is feasible to distinguish structurally different PAHs based on their genomic fingerprints, which are mechanism based.
Fox, Ellen M.; Gardiner, Donald M.; Keller, Nancy P.; Howlett, Barbara J.
2008-01-01
A gene, sirZ, encoding a Zn(II)2Cys6 DNA binding protein is present in a cluster of genes responsible for the biosynthesis of the epipolythiodioxopiperazine (ETP) toxin, sirodesmin PL in the ascomycete plant pathogen, Leptosphaeria maculans. RNA-mediated silencing of sirZ gives rise to transformants that produce only residual amounts of sirodesmin PL and display a decrease in the transcription of several sirodesmin PL biosynthetic genes. This indicates that SirZ is a major regulator of this gene cluster. Proteins similar to SirZ are encoded in the gliotoxin biosynthetic gene cluster of Aspergillus fumigatus (gliZ) and in an ETP-like cluster in Penicillium lilacinoechinulatum (PlgliZ). Despite its high level of sequence similarity to gliZ, PlgliZ is unable to complement the gliotoxin-deficiency of a mutant of gliZ in A. fumigatus. Putative binding sites for these regulatory proteins in the promoters of genes in these clusters were predicted using bioinformatic analysis. These sites are similar to those commonly bound by other proteins with Zn(II)2Cys6 DNA binding domains. PMID:18023597
Evidence against the selfish operon theory.
Pál, Csaba; Hurst, Laurence D
2004-06-01
According to the selfish operon hypothesis, the clustering of genes and their subsequent organization into operons is beneficial for the constituent genes because it enables the horizontal gene transfer of weakly selected, functionally coupled genes. The majority of these are expected to be non-essential genes. From our analysis of the Escherichia coli genome, we conclude that the selfish operon hypothesis is unlikely to provide a general explanation for clustering nor can it account for the gene composition of operons. Contrary to expectations, essential genes with related functions have an especially strong tendency to cluster, even if they are not in operons. Moreover, essential genes are particularly abundant in operons.
Crack, Jason C.; Munnoch, John; Dodd, Erin L.; Knowles, Felicity; Al Bassam, Mahmoud M.; Kamali, Saeed; Holland, Ashley A.; Cramer, Stephen P.; Hamilton, Chris J.; Johnson, Michael K.; Thomson, Andrew J.; Hutchings, Matthew I.; Le Brun, Nick E.
2015-01-01
The Rrf2 family transcription factor NsrR controls expression of genes in a wide range of bacteria in response to nitric oxide (NO). The precise form of the NO-sensing module of NsrR is the subject of controversy because NsrR proteins containing either [2Fe-2S] or [4Fe-4S] clusters have been observed previously. Optical, Mössbauer, resonance Raman spectroscopies and native mass spectrometry demonstrate that Streptomyces coelicolor NsrR (ScNsrR), previously reported to contain a [2Fe-2S] cluster, can be isolated containing a [4Fe-4S] cluster. ChIP-seq experiments indicated that the ScNsrR regulon is small, consisting of only hmpA1, hmpA2, and nsrR itself. The hmpA genes encode NO-detoxifying flavohemoglobins, indicating that ScNsrR has a specialized regulatory function focused on NO detoxification and is not a global regulator like some NsrR orthologues. EMSAs and DNase I footprinting showed that the [4Fe-4S] form of ScNsrR binds specifically and tightly to an 11-bp inverted repeat sequence in the promoter regions of the identified target genes and that DNA binding is abolished following reaction with NO. Resonance Raman data were consistent with cluster coordination by three Cys residues and one oxygen-containing residue, and analysis of ScNsrR variants suggested that highly conserved Glu-85 may be the fourth ligand. Finally, we demonstrate that some low molecular weight thiols, but importantly not physiologically relevant thiols, such as cysteine and an analogue of mycothiol, bind weakly to the [4Fe-4S] cluster, and exposure of this bound form to O2 results in cluster conversion to the [2Fe-2S] form, which does not bind to DNA. These data help to account for the observation of [2Fe-2S] forms of NsrR. PMID:25771538
Missing link in the evolution of Hox clusters.
Ogishima, Soichi; Tanaka, Hiroshi
2007-01-31
Hox cluster has key roles in regulating the patterning of the antero-posterior axis in a metazoan embryo. It consists of the anterior, central and posterior genes; the central genes have been identified only in bilaterians, but not in cnidarians, and are responsible for archiving morphological complexity in bilaterian development. However, their evolutionary history has not been revealed, that is, there has been a "missing link". Here we show the evolutionary history of Hox clusters of 18 bilaterians and 2 cnidarians by using a new method, "motif-based reconstruction", examining the gain/loss processes of evolutionarily conserved sequences, "motifs", outside the homeodomain. We successfully identified the missing link in the evolution of Hox clusters between the cnidarian-bilaterian ancestor and the bilaterians as the ancestor of the central genes, which we call the proto-central gene. Exploring the correspondent gene with the proto-central gene, we found that one of the acoela Hox genes has the same motif repertory as that of the proto-central gene. This interesting finding suggests that the acoela Hox cluster corresponds with the missing link in the evolution of the Hox cluster between the cnidarian-bilaterian ancestor and the bilaterians. Our findings suggested that motif gains/diversifications led to the explosive diversity of the bilaterian body plan.
Reyes-Dominguez, Yazmid; Boedi, Stefan; Sulyok, Michael; Wiesenberger, Gerlinde; Stoppacher, Norbert; Krska, Rudolf; Strauss, Joseph
2012-01-01
Chromatin modifications and heterochromatic marks have been shown to be involved in the regulation of secondary metabolism gene clusters in the fungal model system Aspergillus nidulans. We examine here the role of HEP1, the heterochromatin protein homolog of Fusarium graminearum, for the production of secondary metabolites. Deletion of Hep1 in a PH-1 background strongly influences expression of genes required for the production of aurofusarin and the main tricothecene metabolite DON. In the Hep1 deletion strains AUR genes are highly up-regulated and aurofusarin production is greatly enhanced suggesting a repressive role for heterochromatin on gene expression of this cluster. Unexpectedly, gene expression and metabolites are lower for the trichothecene cluster suggesting a positive function of Hep1 for DON biosynthesis. However, analysis of histone modifications in chromatin of AUR and DON gene promoters reveals that in both gene clusters the H3K9me3 heterochromatic mark is strongly reduced in the Hep1 deletion strain. This, and the finding that a DON-cluster flanking gene is up-regulated, suggests that the DON biosynthetic cluster is repressed by HEP1 directly and indirectly. Results from this study point to a conserved mode of secondary metabolite (SM) biosynthesis regulation in fungi by chromatin modifications and the formation of facultative heterochromatin. PMID:22100541
Kibinge, Nelson; Ono, Naoaki; Horie, Masafumi; Sato, Tetsuo; Sugiura, Tadao; Altaf-Ul-Amin, Md; Saito, Akira; Kanaya, Shigehiko
2016-06-01
Conventionally, workflows examining transcription regulation networks from gene expression data involve distinct analytical steps. There is a need for pipelines that unify data mining and inference deduction into a singular framework to enhance interpretation and hypotheses generation. We propose a workflow that merges network construction with gene expression data mining focusing on regulation processes in the context of transcription factor driven gene regulation. The pipeline implements pathway-based modularization of expression profiles into functional units to improve biological interpretation. The integrated workflow was implemented as a web application software (TransReguloNet) with functions that enable pathway visualization and comparison of transcription factor activity between sample conditions defined in the experimental design. The pipeline merges differential expression, network construction, pathway-based abstraction, clustering and visualization. The framework was applied in analysis of actual expression datasets related to lung, breast and prostrate cancer. Copyright © 2016 Elsevier Inc. All rights reserved.
2011-01-01
Background Increased understanding of the variability in normal breast biology will enable us to identify mechanisms of breast cancer initiation and the origin of different subtypes, and to better predict breast cancer risk. Methods Gene expression patterns in breast biopsies from 79 healthy women referred to breast diagnostic centers in Norway were explored by unsupervised hierarchical clustering and supervised analyses, such as gene set enrichment analysis and gene ontology analysis and comparison with previously published genelists and independent datasets. Results Unsupervised hierarchical clustering identified two separate clusters of normal breast tissue based on gene-expression profiling, regardless of clustering algorithm and gene filtering used. Comparison of the expression profile of the two clusters with several published gene lists describing breast cells revealed that the samples in cluster 1 share characteristics with stromal cells and stem cells, and to a certain degree with mesenchymal cells and myoepithelial cells. The samples in cluster 1 also share many features with the newly identified claudin-low breast cancer intrinsic subtype, which also shows characteristics of stromal and stem cells. More women belonging to cluster 1 have a family history of breast cancer and there is a slight overrepresentation of nulliparous women in cluster 1. Similar findings were seen in a separate dataset consisting of histologically normal tissue from both breasts harboring breast cancer and from mammoplasty reductions. Conclusion This is the first study to explore the variability of gene expression patterns in whole biopsies from normal breasts and identified distinct subtypes of normal breast tissue. Further studies are needed to determine the specific cell contribution to the variation in the biology of normal breasts, how the clusters identified relate to breast cancer risk and their possible link to the origin of the different molecular subtypes of breast cancer. PMID:22044755
Liu, Shi-Huo; Li, Hong-Fei; Yang, Yang; Yang, Rui-Lin; Yang, Wen-Jia; Jiang, Hong-Bo; Dou, Wei; Smagghe, Guy; Wang, Jin-Jun
2018-05-01
Chitinases (Chts) and chitin deacetylases (CDAs) are important enzymes required for chitin metabolism in insects. In this study, 12 Cht-related genes (including seven Cht genes and five imaginal disc growth factor genes) and 6 CDA genes (encoding seven proteins) were identified in Bactrocera dorsalis using genome-wide searching and transcript profiling. Based on the conserved sequences and phylogenetic relationships, 12 Cht-related proteins were clustered into eight groups (group I-V and VII-IX). Further domain architecture analysis showed that all contained at least one chitinase catalytic domain, however, only four (BdCht5, BdCht7, BdCht8 and BdCht10) possessed chitin-binding domains. The subsequent phylogenetic analysis revealed that seven CDAs were clustered into five groups (group I-V), and all had one chitin deacetylase catalytic domain. However, only six exhibited chitin-binding domains. Finally, the development- and tissue-specific expression profiling showed that transcript levels of the 12 Cht-related genes and 6 CDA genes varied considerably among eggs, larvae, pupae and adults, as well as among different tissues of larvae and adults. Our findings illustrate the structural differences and expression patterns of Cht and CDA genes in B. dorsalis, and provide important information for the development of new pest control strategies based on these vital enzymes. Copyright © 2018. Published by Elsevier Inc.
Patel, Vidushi S; Ezaz, Tariq; Deakin, Janine E; Graves, Jennifer A Marshall
2010-12-01
The haemoglobin protein, required for oxygen transportation in the body, is encoded by α- and β-globin genes that are arranged in clusters. The transpositional model for the evolution of distinct α-globin and β-globin clusters in amniotes is much simpler than the previously proposed whole genome duplication model. According to this model, all jawed vertebrates share one ancient region containing α- and β-globin genes and several flanking genes in the order MPG-C16orf35-(α-β)-GBY-LUC7L that has been conserved for more than 410 million years, whereas amniotes evolved a distinct β-globin cluster by insertion of a transposed β-globin gene from this ancient region into a cluster of olfactory receptors flanked by CCKBR and RRM1. It could not be determined whether this organisation is conserved in all amniotes because of the paucity of information from non-avian reptiles. To fill in this gap, we examined globin gene organisation in a squamate reptile, the Australian bearded dragon lizard, Pogona vitticeps (Agamidae). We report here that the α-globin cluster (HBK, HBA) is flanked by C16orf35 and GBY and is located on a pair of microchromosomes, whereas the β-globin cluster is flanked by RRM1 on the 3' end and is located on the long arm of chromosome 3. However, the CCKBR gene that flanks the β-globin cluster on the 5' end in other amniotes is located on the short arm of chromosome 5 in P. vitticeps, indicating that a chromosomal break between the β-globin cluster and CCKBR occurred at least in the agamid lineage. Our data from a reptile species provide further evidence to support the transpositional model for the evolution of β-globin gene cluster in amniotes.
Origin and Evolution of the Sponge Aggregation Factor Gene Family
Grice, Laura F.; Gauthier, Marie E.A.; Roper, Kathrein E.; Fernàndez-Busquets, Xavier; Degnan, Sandie M.
2017-01-01
Although discriminating self from nonself is a cardinal animal trait, metazoan allorecognition genes do not appear to be homologous. Here, we characterize the Aggregation Factor (AF) gene family, which encodes putative allorecognition factors in the demosponge Amphimedon queenslandica, and trace its evolution across 24 sponge (Porifera) species. The AF locus in Amphimedon is comprised of a cluster of five similar genes that encode Calx-beta and Von Willebrand domains and a newly defined Wreath domain, and are highly polymorphic. Further AF variance appears to be generated through individualistic patterns of RNA editing. The AF gene family varies between poriferans, with protein sequences and domains diagnostic of the AF family being present in Amphimedon and other demosponges, but absent from other sponge classes. Within the demosponges, AFs vary widely with no two species having the same AF repertoire or domain organization. The evolution of AFs suggests that their diversification occurs via high allelism, and the continual and rapid gain, loss and shuffling of domains over evolutionary time. Given the marked differences in metazoan allorecognition genes, we propose the rapid evolution of AFs in sponges provides a model for understanding the extensive diversification of self–nonself recognition systems in the animal kingdom. PMID:28104746
2010-01-01
Background Clock family genes encode transcription factors that regulate clock-controlled genes and thus regulate many physiological mechanisms/processes in a circadian fashion. Clock1 duplicates and copies of Clock3 and NPAS2-like genes were partially characterized (genomic sequencing) and mapped using family-based indels/SNPs in rainbow trout (RT)(Oncorhynchus mykiss), Arctic charr (AC)(Salvelinus alpinus), and Atlantic salmon (AS)(Salmo salar) mapping panels. Results Clock1 duplicates mapped to linkage groups RT-8/-24, AC-16/-13 and AS-2/-18. Clock3/NPAS2-like genes mapped to RT-9/-20, AC-20/-43, and AS-5. Most of these linkage group regions containing the Clock gene duplicates were derived from the most recent 4R whole genome duplication event specific to the salmonids. These linkage groups contain quantitative trait loci (QTL) for life history and growth traits (i.e., reproduction and cell cycling). Comparative synteny analyses with other model teleost species reveal a high degree of conservation for genes in these chromosomal regions suggesting that functionally related or co-regulated genes are clustered in syntenic blocks. For example, anti-müllerian hormone (amh), regulating sexual maturation, and ornithine decarboxylase antizymes (oaz1 and oaz2), regulating cell cycling, are contained within these syntenic blocks. Conclusions Synteny analyses indicate that regions homologous to major life-history QTL regions in salmonids contain many candidate genes that are likely to influence reproduction and cell cycling. The order of these genes is highly conserved across the vertebrate species examined, and as such, these genes may make up a functional cluster of genes that are likely co-regulated. CLOCK, as a transcription factor, is found within this block and therefore has the potential to cis-regulate the processes influenced by these genes. Additionally, clock-controlled genes (CCGs) are located in other life-history QTL regions within salmonids suggesting that at least in part, trans-regulation of these QTL regions may also occur via Clock expression. PMID:20670436
Darbani, Behrooz; Motawia, Mohammed Saddik; Olsen, Carl Erik; Nour-Eldin, Hussam H.; Møller, Birger Lindberg; Rook, Fred
2016-01-01
Genomic gene clusters for the biosynthesis of chemical defence compounds are increasingly identified in plant genomes. We previously reported the independent evolution of biosynthetic gene clusters for cyanogenic glucoside biosynthesis in three plant lineages. Here we report that the gene cluster for the cyanogenic glucoside dhurrin in Sorghum bicolor additionally contains a gene, SbMATE2, encoding a transporter of the multidrug and toxic compound extrusion (MATE) family, which is co-expressed with the biosynthetic genes. The predicted localisation of SbMATE2 to the vacuolar membrane was demonstrated experimentally by transient expression of a SbMATE2-YFP fusion protein and confocal microscopy. Transport studies in Xenopus laevis oocytes demonstrate that SbMATE2 is able to transport dhurrin. In addition, SbMATE2 was able to transport non-endogenous cyanogenic glucosides, but not the anthocyanin cyanidin 3-O-glucoside or the glucosinolate indol-3-yl-methyl glucosinolate. The genomic co-localisation of a transporter gene with the biosynthetic genes producing the transported compound is discussed in relation to the role self-toxicity of chemical defence compounds may play in the formation of gene clusters. PMID:27841372
Spohn, Marius; Kirchner, Norbert; Kulik, Andreas; Jochim, Angelika; Wolf, Felix; Muenzer, Patrick; Borst, Oliver; Gross, Harald; Wohlleben, Wolfgang
2014-01-01
The emergence of antibiotic-resistant pathogenic bacteria within the last decades is one reason for the urgent need for new antibacterial agents. A strategy to discover new anti-infective compounds is the evaluation of the genetic capacity of secondary metabolite producers and the activation of cryptic gene clusters (genome mining). One genus known for its potential to synthesize medically important products is Amycolatopsis. However, Amycolatopsis japonicum does not produce an antibiotic under standard laboratory conditions. In contrast to most Amycolatopsis strains, A. japonicum is genetically tractable with different methods. In order to activate a possible silent glycopeptide cluster, we introduced a gene encoding the transcriptional activator of balhimycin biosynthesis, the bbr gene from Amycolatopsis balhimycina (bbrAba), into A. japonicum. This resulted in the production of an antibiotically active compound. Following whole-genome sequencing of A. japonicum, 29 cryptic gene clusters were identified by genome mining. One of these gene clusters is a putative glycopeptide biosynthesis gene cluster. Using bioinformatic tools, ristomycin (syn. ristocetin), a type III glycopeptide, which has antibacterial activity and which is used for the diagnosis of von Willebrand disease and Bernard-Soulier syndrome, was deduced as a possible product of the gene cluster. Chemical analyses by high-performance liquid chromatography and mass spectrometry (HPLC-MS), tandem mass spectrometry (MS/MS), and nuclear magnetic resonance (NMR) spectroscopy confirmed the in silico prediction that the recombinant A. japonicum/pRM4-bbrAba synthesizes ristomycin A. PMID:25114137
Elmore, M Holly; McGary, Kriston L; Wisecaver, Jennifer H; Slot, Jason C; Geiser, David M; Sink, Stacy; O'Donnell, Kerry; Rokas, Antonis
2015-02-06
Fungi that have the enzymes cyanase and carbonic anhydrase show a limited capacity to detoxify cyanate, a fungicide employed by both plants and humans. Here, we describe a novel two-gene cluster that comprises duplicated cyanase and carbonic anhydrase copies, which we name the CCA gene cluster, trace its evolution across Ascomycetes, and examine the evolutionary dynamics of its spread among lineages of the Fusarium oxysporum species complex (hereafter referred to as the FOSC), a cosmopolitan clade of purportedly clonal vascular wilt plant pathogens. Phylogenetic analysis of fungal cyanase and carbonic anhydrase genes reveals that the CCA gene cluster arose independently at least twice and is now present in three lineages, namely Cochliobolus lunatus, Oidiodendron maius, and the FOSC. Genome-wide surveys within the FOSC indicate that the CCA gene cluster varies in copy number across isolates, is always located on accessory chromosomes, and is absent in FOSC's closest relatives. Phylogenetic reconstruction of the CCA gene cluster in 163 FOSC strains from a wide variety of hosts suggests a recent history of rampant transfers between isolates. We hypothesize that the independent formation of the CCA gene cluster in different fungal lineages and its spread across FOSC strains may be associated with resistance to plant-produced cyanates or to use of cyanate fungicides in agriculture. © The Author(s) 2015. Published by Oxford University Press on behalf of the Society for Molecular Biology and Evolution.
Identifying potential maternal genes of Bombyx mori using digital gene expression profiling
Xu, Pingzhen
2018-01-01
Maternal genes present in mature oocytes play a crucial role in the early development of silkworm. Although maternal genes have been widely studied in many other species, there has been limited research in Bombyx mori. High-throughput next generation sequencing provides a practical method for gene discovery on a genome-wide level. Herein, a transcriptome study was used to identify maternal-related genes from silkworm eggs. Unfertilized eggs from five different stages of early development were used to detect the changing situation of gene expression. The expressed genes showed different patterns over time. Seventy-six maternal genes were annotated according to homology analysis with Drosophila melanogaster. More than half of the differentially expressed maternal genes fell into four expression patterns, while the expression patterns showed a downward trend over time. The functional annotation of these material genes was mainly related to transcription factor activity, growth factor activity, nucleic acid binding, RNA binding, ATP binding, and ion binding. Additionally, twenty-two gene clusters including maternal genes were identified from 18 scaffolds. Altogether, we plotted a profile for the maternal genes of Bombyx mori using a digital gene expression profiling method. This will provide the basis for maternal-specific signature research and improve the understanding of the early development of silkworm. PMID:29462160
de Jonge, Ronnie; Ebert, Malaika K; Huitt-Roehl, Callie R; Pal, Paramita; Suttle, Jeffrey C; Spanner, Rebecca E; Neubauer, Jonathan D; Jurick, Wayne M; Stott, Karina A; Secor, Gary A; Thomma, Bart P H J; Van de Peer, Yves; Townsend, Craig A; Bolton, Melvin D
2018-06-12
Species in the genus Cercospora cause economically devastating diseases in sugar beet, maize, rice, soy bean, and other major food crops. Here, we sequenced the genome of the sugar beet pathogen Cercospora beticola and found it encodes 63 putative secondary metabolite gene clusters, including the cercosporin toxin biosynthesis ( CTB ) cluster. We show that the CTB gene cluster has experienced multiple duplications and horizontal transfers across a spectrum of plant pathogenic fungi, including the wide-host range Colletotrichum genus as well as the rice pathogen Magnaporthe oryzae Although cercosporin biosynthesis has been thought to rely on an eight-gene CTB cluster, our phylogenomic analysis revealed gene collinearity adjacent to the established cluster in all CTB cluster-harboring species. We demonstrate that the CTB cluster is larger than previously recognized and includes cercosporin facilitator protein, previously shown to be involved with cercosporin autoresistance, and four additional genes required for cercosporin biosynthesis, including the final pathway enzymes that install the unusual cercosporin methylenedioxy bridge. Lastly, we demonstrate production of cercosporin by Colletotrichum fioriniae , the first known cercosporin producer within this agriculturally important genus. Thus, our results provide insight into the intricate evolution and biology of a toxin critical to agriculture and broaden the production of cercosporin to another fungal genus containing many plant pathogens of important crops worldwide. Copyright © 2018 the Author(s). Published by PNAS.
Tambat, Subodh; Vasudevan, Madavan
2016-01-01
Although salt tolerance is a feature representative of halophytes, most studies on this topic in plants have been conducted on glycophytes. Transcriptome profiles are also available for only a limited number of halophytes. Hence, the present study was conducted to understand the molecular basis of salt tolerance through the transcriptome profiling of the halophyte Suaeda maritima, which is an emerging plant model for research on salt tolerance. Illumina sequencing revealed 72,588 clustered transcripts, including 27,434 that were annotated using BLASTX. Salt application resulted in the 2-fold or greater upregulation of 647 genes and downregulation of 735 genes. Of these, 391 proteins were homologous to proteins in the COGs (cluster of orthologous groups) database, and the majorities were grouped into the poorly characterized category. Approximately 50% of the genes assigned to MapMan pathways showed homology to S. maritima. The majority of such genes represented transcription factors. Several genes also contributed to cell wall and carbohydrate metabolism, ion relation, redox responses and G protein, phosphoinositide and hormone signaling. Real-time PCR was used to validate the results of the deep sequencing for the most of the genes. This study demonstrates the expression of protein kinase C, the target of diacylglycerol in phosphoinositide signaling, for the first time in plants. This study further reveals that the biochemical and molecular responses occurring at several levels are associated with salt tolerance in S. maritima. At the structural level, adaptations to high salinity levels include the remodeling of cell walls and the modification of membrane lipids. At the cellular level, the accumulation of glycinebetaine and the sequestration and exclusion of Na+ appear to be important. Moreover, this study also shows that the processes related to salt tolerance might be highly complex, as reflected by the salt-induced enhancement of transcription factor expression, including hormone-responsive factors, and that this process might be initially triggered by G protein and phosphoinositide signaling. PMID:27682829
Zhu, Y B; Xie, X Q; Li, Z Y; Bai, H; Dong, L; Dong, Z P; Dong, J G
2014-08-28
The nucleotide-binding site (NBS) disease-resistance genes are the largest category of plant disease-resistance gene analogs. The complete set of disease-resistant candidate genes, which encode the NBS sequence, was filtered in the genomes of two varieties of foxtail millet (Yugu1 and 'Zhang gu'). This study investigated a number of characteristics of the putative NBS genes, such as structural diversity and phylogenetic relationships. A total of 269 and 281 NBS-coding sequences were identified in Yugu1 and 'Zhang gu', respectively. When the two databases were compared, 72 genes were found to be identical and 164 genes showed more than 90% similarity. Physical positioning and gene family analysis of the NBS disease-resistance genes in the genome revealed that the number of genes on each chromosome was similar in both varieties. The eighth chromosome contained the largest number of genes and the ninth chromosome contained the lowest number of genes. Exactly 34 gene clusters containing the 161 genes were found in the Yugu1 genome, with each cluster containing 4.7 genes on average. In comparison, the 'Zhang gu' genome possessed 28 gene clusters, which had 151 genes, with an average of 5.4 genes in each cluster. The largest gene cluster, located on the eighth chromosome, contained 12 genes in the Yugu1 database, whereas it contained 16 genes in the 'Zhang gu' database. The classification results showed that the CC-NBS-LRR gene made up the largest part of each chromosome in the two databases. Two TIR-NBS genes were also found in the Yugu1 genome.
[Chromosomal large fragment deletion induced by CRISPR/Cas9 gene editing system].
Cheng, L H; Liu, Y; Niu, T
2017-05-14
Objective: Using CRISPR-Cas9 gene editing technology to achieve a number of genes co-deletion on the same chromosome. Methods: CRISPR-Cas9 lentiviral plasmid that could induce deletion of Aloxe3-Alox12b-Alox8 cluster genes located on mouse 11B3 chromosome was constructed via molecular clone. HEK293T cells were transfected to package lentivirus of CRISPR or Cas9 cDNA, then mouse NIH3T3 cells were infected by lentivirus and genomic DNA of these cells was extracted. The deleted fragment was amplified by PCR, TA clone, Sanger sequencing and other techniques were used to confirm the deletion of Aloxe3-Alox12b-Alox8 cluster genes. Results: The CRISPR-Cas9 lentiviral plasmid, which could induce deletion of Aloxe3-Alox12b-Alox8 cluster genes, was successfully constructed. Deletion of target chromosome fragment (Aloxe3-Alox12b-Alox8 cluster genes) was verified by PCR. The deletion of Aloxe3-Alox12b-Alox8 cluster genes was affirmed by TA clone, Sanger sequencing, and the breakpoint junctions of the CRISPR-Cas9 system mediate cutting events were accurately recombined, insertion mutation did not occur between two cleavage sites at all. Conclusion: Large fragment deletion of Aloxe3-Alox12b-Alox8 cluster genes located on mouse chromosome 11B3 was successfully induced by CRISPR-Cas9 gene editing system.
2013-01-01
Background The antifungal therapy caspofungin is a semi-synthetic derivative of pneumocandin B0, a lipohexapeptide produced by the fungus Glarea lozoyensis, and was the first member of the echinocandin class approved for human therapy. The nonribosomal peptide synthetase (NRPS)-polyketide synthases (PKS) gene cluster responsible for pneumocandin biosynthesis from G. lozoyensis has not been elucidated to date. In this study, we report the elucidation of the pneumocandin biosynthetic gene cluster by whole genome sequencing of the G. lozoyensis wild-type strain ATCC 20868. Results The pneumocandin biosynthetic gene cluster contains a NRPS (GLNRPS4) and a PKS (GLPKS4) arranged in tandem, two cytochrome P450 monooxygenases, seven other modifying enzymes, and genes for L-homotyrosine biosynthesis, a component of the peptide core. Thus, the pneumocandin biosynthetic gene cluster is significantly more autonomous and organized than that of the recently characterized echinocandin B gene cluster. Disruption mutants of GLNRPS4 and GLPKS4 no longer produced the pneumocandins (A0 and B0), and the Δglnrps4 and Δglpks4 mutants lost antifungal activity against the human pathogenic fungus Candida albicans. In addition to pneumocandins, the G. lozoyensis genome encodes a rich repertoire of natural product-encoding genes including 24 PKSs, six NRPSs, five PKS-NRPS hybrids, two dimethylallyl tryptophan synthases, and 14 terpene synthases. Conclusions Characterization of the gene cluster provides a blueprint for engineering new pneumocandin derivatives with improved pharmacological properties. Whole genome estimation of the secondary metabolite-encoding genes from G. lozoyensis provides yet another example of the huge potential for drug discovery from natural products from the fungal kingdom. PMID:23688303
USDA-ARS?s Scientific Manuscript database
Fungi that have the enzymes cyanase and carbonic anhydrase show a limited capacity to detoxify cyanate, a fungicide employed by both plants and humans. Here, we describe a novel two-gene cluster that comprises duplicated cyanase and carbonic anhydrase copies, which we name the CCA gene cluster, trac...
The impact of polyploidy on the evolution of a complex NB-LRR resistance gene cluster in soybean
USDA-ARS?s Scientific Manuscript database
A comparative genomics approach was used to investigate the evolution of a complex NB-LRR gene cluster found in soybean (Glycine max), common bean (Phaseolus vulgaris), and other legumes. In soybean, the cluster is associated with several disease resistance (R) genes of known function including Rpg1...
Mäder, Lisa; Blank, Anna E; Capper, David; Jansong, Janina; Baumgarten, Peter; Wirsik, Naita M; Zachskorn, Cornelia; Ehlers, Jakob; Seifert, Michael; Klink, Barbara; Liebner, Stefan; Niclou, Simone; Naumann, Ulrike; Harter, Patrick N; Mittelbronn, Michel
2018-05-08
Epithelial-to-mesenchymal transition (EMT) is supposed to be responsible for increased invasion and metastases in epithelial cancer cells. The activation of EMT genes has further been proposed to be important in the process of malignant transformation of primary CNS tumors. Since the cellular source and clinical impact of EMT factors in primary CNS tumors still remain unclear, we aimed at deciphering their distribution in vivo and clinico-pathological relevance in human gliomas. We investigated 350 glioma patients for the expression of the key EMT factors SLUG and TWIST by immunohistochemistry and immunofluorescence related to morpho-genetic alterations such as EGFR -amplification, IDH-1 (R132H) mutation and 1p/19q LOH. Furthermore, transcriptional cluster and survival analyses were performed. Our data illustrate that SLUG and TWIST are overexpressed in gliomas showing vascular proliferation such as pilocytic astrocytomas and glioblastomas. EMT factors are exclusively expressed by non-neoplastic pericytes/vessel-associated mural cells (VAMCs). They are not associated with patient survival but correlate with pericytic/VAMC genes in glioblastoma cluster analysis. In summary, the upregulation of EMT genes in pilocytic astrocytomas and glioblastomas reflects the level of activation of pericytes/VAMCs in newly formed blood vessels. Our results underscore that the negative prognostic potential of the EMT signature in the group of diffuse gliomas of WHO grade II-IV does most likely not derive from glioma cells but rather reflects the degree of proliferating mural cells thereby constituting a potential target for future alternative treatment approaches.
[Estimation of individual breast cancer risk: relevance and limits of risk estimation models].
De Pauw, A; Stoppa-Lyonnet, D; Andrieu, N; Asselain, B
2009-10-01
Several risk estimation models for breast or ovarian cancers have been developed these last decades. All these models take into account the family history, with different levels of sophistication. Gail model was developed in 1989 taking into account the family history (0, 1 or > or = 2 affected relatives) and several environmental factors. In 1990, Claus model was the first to integrate explicit assumptions about genetic effects, assuming a single gene dominantly inherited occurring with a low frequency in the population. BRCAPRO model, posterior to the identification of BRCA1 and BRCA2, assumes a restricted transmission with only these two dominantly inherited genes. BOADICEA model adds the effect of a polygenic component to the effect of BRCA1 and BRCA2 to explain the residual clustering of breast cancer. At last, IBIS model assumes a third dominantly inherited gene to explain this residual clustering. Moreover, this model incorporates environmental factors. We applied the Claus, BRCAPRO, BOADICEA and IBIS models to four clinical situations, corresponding to more or less heavy family histories, in order to study the consistency of the risk estimates. The three more recent models (BRCAPRO, BOADICEA and IBIS) gave the closer estimations. These estimates could be useful in clinical practice in front of complex analysis of breast and/or ovarian cancers family history.
Yang, Haixuan; Seoighe, Cathal
2016-01-01
Nonnegative Matrix Factorization (NMF) has proved to be an effective method for unsupervised clustering analysis of gene expression data. By the nonnegativity constraint, NMF provides a decomposition of the data matrix into two matrices that have been used for clustering analysis. However, the decomposition is not unique. This allows different clustering results to be obtained, resulting in different interpretations of the decomposition. To alleviate this problem, some existing methods directly enforce uniqueness to some extent by adding regularization terms in the NMF objective function. Alternatively, various normalization methods have been applied to the factor matrices; however, the effects of the choice of normalization have not been carefully investigated. Here we investigate the performance of NMF for the task of cancer class discovery, under a wide range of normalization choices. After extensive evaluations, we observe that the maximum norm showed the best performance, although the maximum norm has not previously been used for NMF. Matlab codes are freely available from: http://maths.nuigalway.ie/~haixuanyang/pNMF/pNMF.htm.
Wood, Gwendolyn E.; Haydock, Andrew K.; Leigh, John A.
2003-01-01
Methanococcus maripaludis is a mesophilic species of Archaea capable of producing methane from two substrates: hydrogen plus carbon dioxide and formate. To study the latter, we identified the formate dehydrogenase genes of M. maripaludis and found that the genome contains two gene clusters important for formate utilization. Phylogenetic analysis suggested that the two formate dehydrogenase gene sets arose from duplication events within the methanococcal lineage. The first gene cluster encodes homologs of formate dehydrogenase α (FdhA) and β (FdhB) subunits and a putative formate transporter (FdhC) as well as a carbonic anhydrase analog. The second gene cluster encodes only FdhA and FdhB homologs. Mutants lacking either fdhA gene exhibited a partial growth defect on formate, whereas a double mutant was completely unable to grow on formate as a sole methanogenic substrate. Investigation of fdh gene expression revealed that transcription of both gene clusters is controlled by the presence of H2 and not by the presence of formate. PMID:12670979
Liu, Yonghong; Liu, Yuanyuan; Wu, Jiaming; Roizman, Bernard; Zhou, Grace Guoying
2018-04-03
Analyses of the levels of mRNAs encoding IFIT1, IFI16, RIG-1, MDA5, CXCL10, LGP2, PUM1, LSD1, STING, and IFNβ in cell lines from which the gene encoding LGP2, LSD1, PML, HDAC4, IFI16, PUM1, STING, MDA5, IRF3, or HDAC 1 had been knocked out, as well as the ability of these cell lines to support the replication of HSV-1, revealed the following: ( i ) Cell lines lacking the gene encoding LGP2, PML, or HDAC4 (cluster 1) exhibited increased levels of expression of partially overlapping gene networks. Concurrently, these cell lines produced from 5 fold to 12 fold lower yields of HSV-1 than the parental cells. ( ii ) Cell lines lacking the genes encoding STING, LSD1, MDA5, IRF3, or HDAC 1 (cluster 2) exhibited decreased levels of mRNAs of partially overlapping gene networks. Concurrently, these cell lines produced virus yields that did not differ from those produced by the parental cell line. The genes up-regulated in cell lines forming cluster 1, overlapped in part with genes down-regulated in cluster 2. The key conclusions are that gene knockouts and subsequent selection for growth causes changes in expression of multiple genes, and hence the phenotype of the cell lines cannot be ascribed to a single gene; the patterns of gene expression may be shared by multiple knockouts; and the enhanced immunity to viral replication by cluster 1 knockout cell lines but not by cluster 2 cell lines suggests that in parental cells, the expression of innate resistance to infection is specifically repressed.
Chou, A; Burke, J
1999-05-01
DNA sequence clustering has become a valuable method in support of gene discovery and gene expression analysis. Our interest lies in leveraging the sequence diversity within clusters of expressed sequence tags (ESTs) to model gene structure for the study of gene variants that arise from, among other things, alternative mRNA splicing, polymorphism, and divergence after gene duplication, fusion, and translocation events. In previous work, CRAW was developed to discover gene variants from assembled clusters of ESTs. Most importantly, novel gene features (the differing units between gene variants, for example alternative exons, polymorphisms, transposable elements, etc.) that are specialized to tissue, disease, population, or developmental states can be identified when these tools collate DNA source information with gene variant discrimination. While the goal is complete automation of novel feature and gene variant detection, current methods are far from perfect and hence the development of effective tools for visualization and exploratory data analysis are of paramount importance in the process of sifting through candidate genes and validating targets. We present CRAWview, a Java based visualization extension to CRAW. Features that vary between gene forms are displayed using an automatically generated color coded index. The reporting format of CRAWview gives a brief, high level summary report to display overlap and divergence within clusters of sequences as well as the ability to 'drill down' and see detailed information concerning regions of interest. Additionally, the alignment viewing and editing capabilities of CRAWview make it possible to interactively correct frame-shifts and otherwise edit cluster assemblies. We have implemented CRAWview as a Java application across windows NT/95 and UNIX platforms. A beta version of CRAWview will be freely available to academic users from Pangea Systems (http://www.pangeasystems.com). Contact :
Zhai, Ying; Bai, Silei; Liu, Jingjing; Yang, Liyuan; Han, Li; Huang, Xueshi; He, Jing
2016-04-22
Dithiolopyrrolone group antibiotics characterized by an electronically unique dithiolopyrrolone heterobicyclic core are known for their antibacterial, antifungal, insecticidal and antitumor activities. Recently the biosynthetic gene clusters for two dithiolopyrrolone compounds, holomycin and thiomarinol, have been identified respectively in different bacterial species. Here, we report a novel dithiolopyrrolone biosynthetic gene cluster (aut) isolated from Streptomyces thioluteus DSM 40027 which produces two pyrrothine derivatives, aureothricin and thiolutin. By comparison with other characterized dithiolopyrrolone clusters, eight genes in the aut cluster were verified to be responsible for the assembly of dithiolopyrrolone core. The aut cluster was further confirmed by heterologous expression and in-frame gene deletion experiments. Intriguingly, we found that the heterogenetic thioesterase HlmK derived from the holomycin (hlm) gene cluster in Streptomyces clavuligerus significantly improved heterologous biosynthesis of dithiolopyrrolones in Streptomyces albus through coexpression with the aut cluster. In the previous studies, HlmK was considered invalid because it has a Ser to Gly point mutation within the canonical Ser-His-Asp catalytic triad of thioesterases. However, gene inactivation and complementation experiments in our study unequivocally demonstrated that HlmK is an active distinctive type II thioesterase that plays a beneficial role in dithiolopyrrolone biosynthesis. Copyright © 2016 Elsevier Inc. All rights reserved.
Post-genome research on the biosynthesis of ergot alkaloids.
Li, Shu-Ming; Unsöld, Inge A
2006-10-01
Genome sequencing provides new opportunities and challenges for identifying genes for the biosynthesis of secondary metabolites. A putative biosynthetic gene cluster of fumigaclavine C, an ergot alkaloid of the clavine type, was identified in the genome sequence of ASPERGILLUS FUMIGATUS by a bioinformatic approach. This cluster spans 22 kb of genomic DNA and comprises at least 11 open reading frames (ORFs). Seven of them are orthologous to genes from the biosynthetic gene cluster of ergot alkaloids in CLAVICEPS PURPUREA. Experimental evidence of the identified cluster was provided by heterologous expression and biochemical characterization of two ORFs, FgaPT1 and FgaPT2, in the cluster of A. FUMIGATUS, which show remarkable similarities to dimethylallyltryptophan synthase from C. PURPUREA and function as prenyltransferases. FgaPT2 converts L-tryptophan to dimethylallyltryptophan and thereby catalyzes the first step of ergot alkaloid biosynthesis, whilst FgaPT1 catalyzes the last step of the fumigaclavine C biosynthesis, i. e., the prenylation of fumigaclavine A at C-2 position of the indole nucleus. In addition to information obtained from the gene cluster of ergot alkaloids from C. PURPUREA, the identification of the biosynthetic gene cluster of fumigaclavine C in A. FUMIGATUS opens an alternative way to study the biosynthesis of ergot alkaloids in fungi.
Statistical indicators of collective behavior and functional clusters in gene networks of yeast
NASA Astrophysics Data System (ADS)
Živković, J.; Tadić, B.; Wick, N.; Thurner, S.
2006-03-01
We analyze gene expression time-series data of yeast (S. cerevisiae) measured along two full cell-cycles. We quantify these data by using q-exponentials, gene expression ranking and a temporal mean-variance analysis. We construct gene interaction networks based on correlation coefficients and study the formation of the corresponding giant components and minimum spanning trees. By coloring genes according to their cell function we find functional clusters in the correlation networks and functional branches in the associated trees. Our results suggest that a percolation point of functional clusters can be identified on these gene expression correlation networks.
Genome Engineering and Modification Toward Synthetic Biology for the Production of Antibiotics.
Zou, Xuan; Wang, Lianrong; Li, Zhiqiang; Luo, Jie; Wang, Yunfu; Deng, Zixin; Du, Shiming; Chen, Shi
2018-01-01
Antibiotic production is often governed by large gene clusters composed of genes related to antibiotic scaffold synthesis, tailoring, regulation, and resistance. With the expansion of genome sequencing, a considerable number of antibiotic gene clusters has been isolated and characterized. The emerging genome engineering techniques make it possible towards more efficient engineering of antibiotics. In addition to genomic editing, multiple synthetic biology approaches have been developed for the exploration and improvement of antibiotic natural products. Here, we review the progress in the development of these genome editing techniques used to engineer new antibiotics, focusing on three aspects of genome engineering: direct cloning of large genomic fragments, genome engineering of gene clusters, and regulation of gene cluster expression. This review will not only summarize the current uses of genomic engineering techniques for cloning and assembly of antibiotic gene clusters or for altering antibiotic synthetic pathways but will also provide perspectives on the future directions of rebuilding biological systems for the design of novel antibiotics. © 2017 Wiley Periodicals, Inc.
Querying Co-regulated Genes on Diverse Gene Expression Datasets Via Biclustering.
Deveci, Mehmet; Küçüktunç, Onur; Eren, Kemal; Bozdağ, Doruk; Kaya, Kamer; Çatalyürek, Ümit V
2016-01-01
Rapid development and increasing popularity of gene expression microarrays have resulted in a number of studies on the discovery of co-regulated genes. One important way of discovering such co-regulations is the query-based search since gene co-expressions may indicate a shared role in a biological process. Although there exist promising query-driven search methods adapting clustering, they fail to capture many genes that function in the same biological pathway because microarray datasets are fraught with spurious samples or samples of diverse origin, or the pathways might be regulated under only a subset of samples. On the other hand, a class of clustering algorithms known as biclustering algorithms which simultaneously cluster both the items and their features are useful while analyzing gene expression data, or any data in which items are related in only a subset of their samples. This means that genes need not be related in all samples to be clustered together. Because many genes only interact under specific circumstances, biclustering may recover the relationships that traditional clustering algorithms can easily miss. In this chapter, we briefly summarize the literature using biclustering for querying co-regulated genes. Then we present a novel biclustering approach and evaluate its performance by a thorough experimental analysis.
Karbalaei, Reza; Allahyari, Marzieh; Rezaei-Tavirani, Mostafa; Asadzadeh-Aghdaei, Hamid; Zali, Mohammad Reza
2018-01-01
Analysis reconstruction networks from two diseases, NAFLD and Alzheimer`s diseases and their relationship based on systems biology methods. NAFLD and Alzheimer`s diseases are two complex diseases, with progressive prevalence and high cost for countries. There are some reports on relation and same spreading pathways of these two diseases. In addition, they have some similar risk factors, exclusively lifestyle such as feeding, exercises and so on. Therefore, systems biology approach can help to discover their relationship. DisGeNET and STRING databases were sources of disease genes and constructing networks. Three plugins of Cytoscape software, including ClusterONE, ClueGO and CluePedia, were used to analyze and cluster networks and enrichment of pathways. An R package used to define best centrality method. Finally, based on degree and Betweenness, hubs and bottleneck nodes were defined. Common genes between NAFLD and Alzheimer`s disease were 190 genes that used construct a network with STRING database. The resulting network contained 182 nodes and 2591 edges and comprises from four clusters. Enrichment of these clusters separately lead to carbohydrate metabolism, long chain fatty acid and regulation of JAK-STAT and IL-17 signaling pathways, respectively. Also seven genes selected as hub-bottleneck include: IL6, AKT1, TP53, TNF, JUN, VEGFA and PPARG. Enrichment of these proteins and their first neighbors in network by OMIM database lead to diabetes and obesity as ancestors of NAFLD and AD. Systems biology methods, specifically PPI networks, can be useful for analyzing complicated related diseases. Finding Hub and bottleneck proteins should be the goal of drug designing and introducing disease markers.
Novel genomic island modifies DNA with 7-deazaguanine derivatives
Thiaville, Jennifer J.; Kellner, Stefanie M.; Yuan, Yifeng; Hutinet, Geoffrey; Thiaville, Patrick C.; Jumpathong, Watthanachai; Mohapatra, Susovan; Brochier-Armanet, Celine; Letarov, Andrey V.; Hillebrand, Roman; Malik, Chanchal K.; Rizzo, Carmelo J.; Dedon, Peter C.; de Crécy-Lagard, Valérie
2016-01-01
The discovery of ∼20-kb gene clusters containing a family of paralogs of tRNA guanosine transglycosylase genes, called tgtA5, alongside 7-cyano-7-deazaguanine (preQ0) synthesis and DNA metabolism genes, led to the hypothesis that 7-deazaguanine derivatives are inserted in DNA. This was established by detecting 2’-deoxy-preQ0 and 2’-deoxy-7-amido-7-deazaguanosine in enzymatic hydrolysates of DNA extracted from the pathogenic, Gram-negative bacteria Salmonella enterica serovar Montevideo. These modifications were absent in the closely related S. enterica serovar Typhimurium LT2 and from a mutant of S. Montevideo, each lacking the gene cluster. This led us to rename the genes of the S. Montevideo cluster as dpdA-K for 7-deazapurine in DNA. Similar gene clusters were analyzed in ∼150 phylogenetically diverse bacteria, and the modifications were detected in DNA from other organisms containing these clusters, including Kineococcus radiotolerans, Comamonas testosteroni, and Sphingopyxis alaskensis. Comparative genomic analysis shows that, in Enterobacteriaceae, the cluster is a genomic island integrated at the leuX locus, and the phylogenetic analysis of the TgtA5 family is consistent with widespread horizontal gene transfer. Comparison of transformation efficiencies of modified or unmodified plasmids into isogenic S. Montevideo strains containing or lacking the cluster strongly suggests a restriction–modification role for the cluster in Enterobacteriaceae. Another preQ0 derivative, 2’-deoxy-7-formamidino-7-deazaguanosine, was found in the Escherichia coli bacteriophage 9g, as predicted from the presence of homologs of genes involved in the synthesis of the archaeosine tRNA modification. These results illustrate a deep and unexpected evolutionary connection between DNA and tRNA metabolism. PMID:26929322
NASA Astrophysics Data System (ADS)
Pagnuco, Inti A.; Pastore, Juan I.; Abras, Guillermo; Brun, Marcel; Ballarin, Virginia L.
2016-04-01
It is usually assumed that co-expressed genes suggest co-regulation in the underlying regulatory network. Determining sets of co-expressed genes is an important task, where significative groups of genes are defined based on some criteria. This task is usually performed by clustering algorithms, where the whole family of genes, or a subset of them, are clustered into meaningful groups based on their expression values in a set of experiment. In this work we used a methodology based on the Silhouette index as a measure of cluster quality for individual gene groups, and a combination of several variants of hierarchical clustering to generate the candidate groups, to obtain sets of co-expressed genes for two real data examples. We analyzed the quality of the best ranked groups, obtained by the algorithm, using an online bioinformatics tool that provides network information for the selected genes. Moreover, to verify the performance of the algorithm, considering the fact that it doesn’t find all possible subsets, we compared its results against a full search, to determine the amount of good co-regulated sets not detected.
Tiaden, André; Spirig, Thomas; Sahr, Tobias; Wälti, Martin A; Boucke, Karin; Buchrieser, Carmen; Hilbi, Hubert
2010-05-01
The amoebae-resistant opportunistic pathogen Legionella pneumophila employs a biphasic life cycle to replicate in host cells and spread to new niches. Upon entering the stationary growth phase, the bacteria switch to a transmissive (virulent) state, which involves a complex regulatory network including the lqs gene cluster (lqsA-lqsR-hdeD-lqsS). LqsR is a putative response regulator that promotes host-pathogen interactions and represses replication. The autoinducer synthase LqsA catalyses the production of the diffusible signalling molecule 3-hydroxypentadecan-4-one (LAI-1) that is presumably recognized by the sensor kinase LqsS. Here, we analysed L. pneumophila strains lacking lqsA or lqsS. Compared with wild-type L. pneumophila, the DeltalqsS strain was more salt-resistant and impaired for the Icm/Dot type IV secretion system-dependent uptake by phagocytes. Legionella pneumophila strains lacking lqsS, lqsR or the alternative sigma factor rpoS sedimented more slowly and produced extracellular filaments. Deletion of lqsA moderately reduced the uptake of L. pneumophila by phagocytes, and the defect was complemented by expressing lqsA in trans. Unexpectedly, the overexpression of lqsA also restored the virulence defect and reduced filament production of L. pneumophila mutant strains lacking lqsS or lqsR, but not the phenotypes of strains lacking rpoS or icmT. These results suggest that LqsA products also signal through sensors not encoded by the lqs gene cluster. A transcriptome analysis of the DeltalqsA and DeltalqsS mutant strains revealed that under the conditions tested, lqsA regulated only few genes, whereas lqsS upregulated the expression of 93 genes at least twofold. These include 52 genes clustered in a 133 kb high plasticity genomic island, which is flanked by putative DNA-mobilizing genes and encodes multiple metal ion efflux pumps. Upon overexpression of lqsA, a cluster of 19 genes in the genomic island was also upregulated, suggesting that LqsA and LqsS participate in the same regulatory circuit.
Rudolf, Jeffrey D.; Yan, Xiaohui; Shen, Ben
2015-01-01
The enediynes are one of the most fascinating families of bacterial natural products given their unprecedented molecular architecture and extraordinary cytotoxicity. Enediynes are rare with only 11 structurally characterized members and four additional members isolated in their cycloaromatized form. Recent advances in DNA sequencing have resulted in an explosion of microbial genomes. A virtual survey of the GenBank and JGI genome databases revealed 87 enediyne biosynthetic gene clusters from 78 bacteria strains, implying enediynes are more common than previously thought. Here we report the construction and analysis of an enediyne genome neighborhood network (GNN) as a high-throughput approach to analyze secondary metabolite gene clusters. Analysis of the enediyne GNN facilitated rapid gene cluster annotation, revealed genetic trends in enediyne biosynthetic gene clusters resulting in a simple prediction scheme to determine 9- vs 10-membered enediyne gene clusters, and supported a genomic-based strain prioritization method for enediyne discovery. PMID:26318027
Hierarchical Dirichlet process model for gene expression clustering
2013-01-01
Clustering is an important data processing tool for interpreting microarray data and genomic network inference. In this article, we propose a clustering algorithm based on the hierarchical Dirichlet processes (HDP). The HDP clustering introduces a hierarchical structure in the statistical model which captures the hierarchical features prevalent in biological data such as the gene express data. We develop a Gibbs sampling algorithm based on the Chinese restaurant metaphor for the HDP clustering. We apply the proposed HDP algorithm to both regulatory network segmentation and gene expression clustering. The HDP algorithm is shown to outperform several popular clustering algorithms by revealing the underlying hierarchical structure of the data. For the yeast cell cycle data, we compare the HDP result to the standard result and show that the HDP algorithm provides more information and reduces the unnecessary clustering fragments. PMID:23587447
Gautier, Aude; Le Gac, Florence; Lareyre, Jean-Jacques
2011-02-01
The gonadal soma-derived factor (GSDF) belongs to the transforming growth factor-β superfamily and is conserved in teleostean fish species. Gsdf is specifically expressed in the gonads, and gene expression is restricted to the granulosa and Sertoli cells in trout and medaka. The gsdf gene expression is correlated to early testis differentiation in medaka and was shown to stimulate primordial germ cell and spermatogonia proliferation in trout. In the present study, we show that the gsdf gene localizes to a syntenic chromosomal fragment conserved among vertebrates although no gsdf-related gene is detected on the corresponding genomic region in tetrapods. We demonstrate using quantitative RT-PCR that most of the genes localized in the synteny are specifically expressed in medaka gonads. Gsdf is the only gene of the synteny with a much higher expression in the testis compared to the ovary. In contrast, gene expression pattern analysis of the gsdf surrounding genes (nup54, aff1, klhl8, sdad1, and ptpn13) indicates that these genes are preferentially expressed in the female gonads. The tissue distribution of these genes is highly similar in medaka and zebrafish, two teleostean species that have diverged more than 110 million years ago. The cellular localization of these genes was determined in medaka gonads using the whole-mount in situ hybridization technique. We confirm that gsdf gene expression is restricted to Sertoli and granulosa cells in contact with the premeiotic and meiotic cells. The nup54 gene is expressed in spermatocytes and previtellogenic oocytes. Transcripts corresponding to the ovary-specific genes (aff1, klhl8, and sdad1) are detected only in previtellogenic oocytes. No expression was detected in the gonocytes in 10 dpf embryos. In conclusion, we show that the gsdf gene localizes to a syntenic chromosomal fragment harboring evolutionary conserved genes in vertebrates. These genes are preferentially expressed in previtelloogenic oocytes, and thus, they display a different cellular localization compared to that of the gsdf gene indicating that the later gene is not co-regulated. Interestingly, our study identifies new clustered genes that are specifically expressed in previtellogenic oocytes (nup54, aff1, klhl8, sdad1). Copyright © 2010 Elsevier B.V. All rights reserved.
NASA Technical Reports Server (NTRS)
Mjolsness, Eric; Castano, Rebecca; Mann, Tobias; Wold, Barbara
2000-01-01
We provide preliminary evidence that existing algorithms for inferring small-scale gene regulation networks from gene expression data can be adapted to large-scale gene expression data coming from hybridization microarrays. The essential steps are (I) clustering many genes by their expression time-course data into a minimal set of clusters of co-expressed genes, (2) theoretically modeling the various conditions under which the time-courses are measured using a continuous-time analog recurrent neural network for the cluster mean time-courses, (3) fitting such a regulatory model to the cluster mean time courses by simulated annealing with weight decay, and (4) analysing several such fits for commonalities in the circuit parameter sets including the connection matrices. This procedure can be used to assess the adequacy of existing and future gene expression time-course data sets for determining transcriptional regulatory relationships such as coregulation.
Wang, Zhao-Xin; Li, Shu-Ming; Heide, Lutz
2000-01-01
The biosynthetic gene cluster of the aminocoumarin antibiotic coumermycin A1 was cloned by screening of a cosmid library of Streptomyces rishiriensis DSM 40489 with heterologous probes from a dTDP-glucose 4,6-dehydratase gene, involved in deoxysugar biosynthesis, and from the aminocoumarin resistance gyrase gene gyrBr. Sequence analysis of a 30.8-kb region upstream of gyrBr revealed the presence of 28 complete open reading frames (ORFs). Fifteen of the identified ORFs showed, on average, 84% identity to corresponding ORFs in the biosynthetic gene cluster of novobiocin, another aminocoumarin antibiotic. Possible functions of 17 ORFs in the biosynthesis of coumermycin A1 could be assigned by comparison with sequences in GenBank. Experimental proof for the function of the identified gene cluster was provided by an insertional gene inactivation experiment, which resulted in an abolishment of coumermycin A1 production. PMID:11036020
Saavedra, Milene T; Quon, Bradley S; Faino, Anna; Caceres, Silvia M; Poch, Katie R; Sanders, Linda A; Malcolm, Kenneth C; Nichols, David P; Sagel, Scott D; Taylor-Cousar, Jennifer L; Leach, Sonia M; Strand, Matthew; Nick, Jerry A
2018-05-01
Cystic fibrosis pulmonary exacerbations accelerate pulmonary decline and increase mortality. Previously, we identified a 10-gene leukocyte panel measured directly from whole blood, which indicates response to exacerbation treatment. We hypothesized that molecular characteristics of exacerbations could also predict future disease severity. We tested whether a 10-gene panel measured from whole blood could identify patient cohorts at increased risk for severe morbidity and mortality, beyond standard clinical measures. Transcript abundance for the 10-gene panel was measured from whole blood at the beginning of exacerbation treatment (n = 57). A hierarchical cluster analysis of subjects based on their gene expression was performed, yielding four molecular clusters. An analysis of cluster membership and outcomes incorporating an independent cohort (n = 21) was completed to evaluate robustness of cluster partitioning of genes to predict severe morbidity and mortality. The four molecular clusters were analyzed for differences in forced expiratory volume in 1 second, C-reactive protein, return to baseline forced expiratory volume in 1 second after treatment, time to next exacerbation, and time to morbidity or mortality events (defined as lung transplant referral, lung transplant, intensive care unit admission for respiratory insufficiency, or death). Clustering based on gene expression discriminated between patient groups with significant differences in forced expiratory volume in 1 second, admission frequency, and overall morbidity and mortality. At 5 years, all subjects in cluster 1 (very low risk) were alive and well, whereas 90% of subjects in cluster 4 (high risk) had suffered a major event (P = 0.0001). In multivariable analysis, the ability of gene expression to predict clinical outcomes remained significant, despite adjustment for forced expiratory volume in 1 second, sex, and admission frequency. The robustness of gene clustering to categorize patients appropriately in terms of clinical characteristics, and short- and long-term clinical outcomes, remained consistent, even when adding in a secondary population with significantly different clinical outcomes. Whole blood gene expression profiling allows molecular classification of acute pulmonary exacerbations, beyond standard clinical measures, providing a predictive tool for identifying subjects at increased risk for mortality and disease progression.
Hsu, Arthur L; Tang, Sen-Lin; Halgamuge, Saman K
2003-11-01
Current Self-Organizing Maps (SOMs) approaches to gene expression pattern clustering require the user to predefine the number of clusters likely to be expected. Hierarchical clustering methods used in this area do not provide unique partitioning of data. We describe an unsupervised dynamic hierarchical self-organizing approach, which suggests an appropriate number of clusters, to perform class discovery and marker gene identification in microarray data. In the process of class discovery, the proposed algorithm identifies corresponding sets of predictor genes that best distinguish one class from other classes. The approach integrates merits of hierarchical clustering with robustness against noise known from self-organizing approaches. The proposed algorithm applied to DNA microarray data sets of two types of cancers has demonstrated its ability to produce the most suitable number of clusters. Further, the corresponding marker genes identified through the unsupervised algorithm also have a strong biological relationship to the specific cancer class. The algorithm tested on leukemia microarray data, which contains three leukemia types, was able to determine three major and one minor cluster. Prediction models built for the four clusters indicate that the prediction strength for the smaller cluster is generally low, therefore labelled as uncertain cluster. Further analysis shows that the uncertain cluster can be subdivided further, and the subdivisions are related to two of the original clusters. Another test performed using colon cancer microarray data has automatically derived two clusters, which is consistent with the number of classes in data (cancerous and normal). JAVA software of dynamic SOM tree algorithm is available upon request for academic use. A comparison of rectangular and hexagonal topologies for GSOM is available from http://www.mame.mu.oz.au/mechatronics/journalinfo/Hsu2003supp.pdf
Wee, Bryan A; Woolfit, Megan; Beatson, Scott A; Petty, Nicola K
2013-01-01
Legionella encodes multiple classes of Type IV Secretion Systems (T4SSs), including the Dot/Icm protein secretion system that is essential for intracellular multiplication in amoebal and human hosts. Other T4SSs not essential for virulence are thought to facilitate the acquisition of niche-specific adaptation genes including the numerous effector genes that are a hallmark of this genus. Previously, we identified two novel gene clusters in the draft genome of Legionella pneumophila strain 130b that encode homologues of a subtype of T4SS, the genomic island-associated T4SS (GI-T4SS), usually associated with integrative and conjugative elements (ICE). In this study, we performed genomic analyses of 14 homologous GI-T4SS clusters found in eight publicly available Legionella genomes and show that this cluster is unusually well conserved in a region of high plasticity. Phylogenetic analyses show that Legionella GI-T4SSs are substantially divergent from other members of this subtype of T4SS and represent a novel clade of GI-T4SSs only found in this genus. The GI-T4SS was found to be under purifying selection, suggesting it is functional and may play an important role in the evolution and adaptation of Legionella. Like other GI-T4SSs, the Legionella clusters are also associated with ICEs, but lack the typical integration and replication modules of related ICEs. The absence of complete replication and DNA pre-processing modules, together with the presence of Legionella-specific regulatory elements, suggest the Legionella GI-T4SS-associated ICE is unique and may employ novel mechanisms of regulation, maintenance and excision. The Legionella GI-T4SS cluster was found to be associated with several cargo genes, including numerous antibiotic resistance and virulence factors, which may confer a fitness benefit to the organism. The in-silico characterisation of this new T4SS furthers our understanding of the diversity of secretion systems involved in the frequent horizontal gene transfers that allow Legionella to adapt to and exploit diverse environmental niches.
Wee, Bryan A.; Woolfit, Megan; Beatson, Scott A.; Petty, Nicola K.
2013-01-01
Legionella encodes multiple classes of Type IV Secretion Systems (T4SSs), including the Dot/Icm protein secretion system that is essential for intracellular multiplication in amoebal and human hosts. Other T4SSs not essential for virulence are thought to facilitate the acquisition of niche-specific adaptation genes including the numerous effector genes that are a hallmark of this genus. Previously, we identified two novel gene clusters in the draft genome of Legionella pneumophila strain 130b that encode homologues of a subtype of T4SS, the genomic island-associated T4SS (GI-T4SS), usually associated with integrative and conjugative elements (ICE). In this study, we performed genomic analyses of 14 homologous GI-T4SS clusters found in eight publicly available Legionella genomes and show that this cluster is unusually well conserved in a region of high plasticity. Phylogenetic analyses show that Legionella GI-T4SSs are substantially divergent from other members of this subtype of T4SS and represent a novel clade of GI-T4SSs only found in this genus. The GI-T4SS was found to be under purifying selection, suggesting it is functional and may play an important role in the evolution and adaptation of Legionella. Like other GI-T4SSs, the Legionella clusters are also associated with ICEs, but lack the typical integration and replication modules of related ICEs. The absence of complete replication and DNA pre-processing modules, together with the presence of Legionella-specific regulatory elements, suggest the Legionella GI-T4SS-associated ICE is unique and may employ novel mechanisms of regulation, maintenance and excision. The Legionella GI-T4SS cluster was found to be associated with several cargo genes, including numerous antibiotic resistance and virulence factors, which may confer a fitness benefit to the organism. The in-silico characterisation of this new T4SS furthers our understanding of the diversity of secretion systems involved in the frequent horizontal gene transfers that allow Legionella to adapt to and exploit diverse environmental niches. PMID:24358157
Wang, Jen-Chyong; Spiegel, Noah; Bertelsen, Sarah; Le, Nhung; McKenna, Nicholas; Budde, John P.; Harari, Oscar; Kapoor, Manav; Brooks, Andrew; Hancock, Dana; Tischfield, Jay; Foroud, Tatiana; Bierut, Laura J.; Steinbach, Joe Henry; Edenberg, Howard J.; Traynor, Bryan J.; Goate, Alison M.
2013-01-01
Variants within the gene cluster encoding α3, α5, and β4 nicotinic receptor subunits are major risk factors for substance dependence. The strongest impact on risk is associated with variation in the CHRNA5 gene, where at least two mechanisms are at work: amino acid variation and altered mRNA expression levels. The risk allele of the non-synonymous variant (rs16969968; D398N) primarily occurs on the haplotype containing the low mRNA expression allele. In populations of European ancestry, there are approximately 50 highly correlated variants in the CHRNA5-CHRNA3-CHRNB4 gene cluster and the adjacent PSMA4 gene region that are associated with CHRNA5 mRNA levels. It is not clear which of these variants contribute to the changes in CHRNA5 transcript level. Because populations of African ancestry have reduced linkage disequilibrium among variants spanning this gene cluster, eQTL mapping in subjects of African ancestry could potentially aid in defining the functional variants that affect CHRNA5 mRNA levels. We performed quantitative allele specific gene expression using frontal cortices derived from 49 subjects of African ancestry and 111 subjects of European ancestry. This method measures allele-specific transcript levels in the same individual, which eliminates other biological variation that occurs when comparing expression levels between different samples. This analysis confirmed that substance dependence associated variants have a direct cis-regulatory effect on CHRNA5 transcript levels in human frontal cortices of African and European ancestry and identified 10 highly correlated variants, located in a 9 kb region, that are potential functional variants modifying CHRNA5 mRNA expression levels. PMID:24303001
Tripathi, G.; Rangaswamy, D.; Borkar, M.; Prasad, N.; Sharma, R. K.; Sankhwar, S. N.; Agrawal, S.
2015-01-01
We evaluated whether polymorphisms in interleukin (IL-1) gene cluster (IL-1 alpha [IL-1A], IL-1 beta [IL-1B], and IL-1 receptor antagonist [IL-1RN]) are associated with end stage renal disease (ESRD). A total of 258 ESRD patients and 569 ethnicity matched controls were examined for IL-1 gene cluster. These were genotyped for five single-nucleotide gene polymorphisms in the IL-1A, IL-1B and IL-1RN genes and a variable number of tandem repeats (VNTR) in the IL-1RN. The IL-1B − 3953 and IL-1RN + 8006 polymorphism frequencies were significantly different between the two groups. At IL-1B, the T allele of − 3953C/T was increased among ESRD (P = 0.0001). A logistic regression model demonstrated that two repeat (240 base pair [bp]) of the IL-1Ra VNTR polymorphism was associated with ESRD (P = 0.0001). The C/C/C/C/C/1 haplotype was more prevalent in ESRD = 0.007). No linkage disequilibrium (LD) was observed between six loci of IL-1 gene. We further conducted a meta-analysis of existing studies and found that there is a strong association of IL-1 RN VNTR 86 bp repeat polymorphism with susceptibility to ESRD (odds ratio = 2.04, 95% confidence interval = 1.48-2.82; P = 0.000). IL-1B − 5887, +8006 and the IL-1RN VNTR polymorphisms have been implicated as potential risk factors for ESRD. The meta-analysis showed a strong association of IL-1RN 86 bp VNTR polymorphism with susceptibility to ESRD. PMID:25684870
USDA-ARS?s Scientific Manuscript database
The zinc finger transcription factor nsdC is required for both sexual development and aflatoxin production in the saprophytic fungus Aspergillus flavus. While previous work with an nsdC knockout mutant was conducted in Aspergillus nidulans and A. flavus strain 3357, here we demonstrate perturbations...
CRISPR-Cas Targeting of Host Genes as an Antiviral Strategy.
Chen, Shuliang; Yu, Xiao; Guo, Deyin
2018-01-16
Currently, a new gene editing tool-the Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) associated (Cas) system-is becoming a promising approach for genetic manipulation at the genomic level. This simple method, originating from the adaptive immune defense system in prokaryotes, has been developed and applied to antiviral research in humans. Based on the characteristics of virus-host interactions and the basic rules of nucleic acid cleavage or gene activation of the CRISPR-Cas system, it can be used to target both the virus genome and host factors to clear viral reservoirs and prohibit virus infection or replication. Here, we summarize recent progress of the CRISPR-Cas technology in editing host genes as an antiviral strategy.
Engineering Synthetic Gene Circuits in Living Cells with CRISPR Technology.
Jusiak, Barbara; Cleto, Sara; Perez-Piñera, Pablo; Lu, Timothy K
2016-07-01
One of the goals of synthetic biology is to build regulatory circuits that control cell behavior, for both basic research purposes and biomedical applications. The ability to build transcriptional regulatory devices depends on the availability of programmable, sequence-specific, and effective synthetic transcription factors (TFs). The prokaryotic clustered regularly interspaced short palindromic repeat (CRISPR) system, recently harnessed for transcriptional regulation in various heterologous host cells, offers unprecedented ease in designing synthetic TFs. We review how CRISPR can be used to build synthetic gene circuits and discuss recent advances in CRISPR-mediated gene regulation that offer the potential to build increasingly complex, programmable, and efficient gene circuits in the future. Copyright © 2016. Published by Elsevier Ltd.
Stathakis, D. G.; Pentz, E. S.; Freeman, M. E.; Kullman, J.; Hankins, G. R.; Pearlson, N. J.; Wright, TRF.
1995-01-01
We report the complete molecular organization of the Dopa decarboxylase gene cluster. Mutagenesis screens recovered 77 new Df(2L)TW130 recessive lethal mutations. These new alleles combined with 263 previously isolated mutations in the cluster to define 18 essential genes. In addition, seven new deficiencies were isolated and characterized. Deficiency mapping, restriction fragment length polymorphism (RFLP) analysis and P-element-mediated germline transformation experiments determined the gene order for all 18 loci. Genomic and cDNA restriction endonuclease mapping, Northern blot analysis and DNA sequencing provided information on exact gene location, mRNA size and transcriptional direction for most of these loci. In addition, this analysis identified two transcription units that had not previously been identified by extensive mutagenesis screening. Most of the loci are contained within two dense subclusters. We discuss the effectiveness of mutagens and strategies used in our screens, the variable mutability of loci within the genome of Drosophila melanogaster, the cytological and molecular organization of the Ddc gene cluster, the validity of the one band-one gene hypothesis and a possible purpose for the clustering of genes in the Ddc region. PMID:8647399
Canales, Javier; Moyano, Tomás C.; Villarroel, Eva; Gutiérrez, Rodrigo A.
2014-01-01
Nitrogen (N) is an essential macronutrient for plant growth and development. Plants adapt to changes in N availability partly by changes in global gene expression. We integrated publicly available root microarray data under contrasting nitrate conditions to identify new genes and functions important for adaptive nitrate responses in Arabidopsis thaliana roots. Overall, more than 2000 genes exhibited changes in expression in response to nitrate treatments in Arabidopsis thaliana root organs. Global regulation of gene expression by nitrate depends largely on the experimental context. However, despite significant differences from experiment to experiment in the identity of regulated genes, there is a robust nitrate response of specific biological functions. Integrative gene network analysis uncovered relationships between nitrate-responsive genes and 11 highly co-expressed gene clusters (modules). Four of these gene network modules have robust nitrate responsive functions such as transport, signaling, and metabolism. Network analysis hypothesized G2-like transcription factors are key regulatory factors controlling transport and signaling functions. Our meta-analysis highlights the role of biological processes not studied before in the context of the nitrate response such as root hair development and provides testable hypothesis to advance our understanding of nitrate responses in plants. PMID:24570678
Aguirre von Wobeser, Eneas; Ibelings, Bas W.; Bok, Jasper; Krasikov, Vladimir; Huisman, Jef; Matthijs, Hans C.P.
2011-01-01
Physiological adaptation and genome-wide expression profiles of the cyanobacterium Synechocystis sp. strain PCC 6803 in response to gradual transitions between nitrogen-limited and light-limited growth conditions were measured in continuous cultures. Transitions induced changes in pigment composition, light absorption coefficient, photosynthetic electron transport, and specific growth rate. Physiological changes were accompanied by reproducible changes in the expression of several hundred open reading frames, genes with functions in photosynthesis and respiration, carbon and nitrogen assimilation, protein synthesis, phosphorus metabolism, and overall regulation of cell function and proliferation. Cluster analysis of the nearly 1,600 regulated open reading frames identified eight clusters, each showing a different temporal response during the transitions. Two large clusters mirrored each other. One cluster included genes involved in photosynthesis, which were up-regulated during light-limited growth but down-regulated during nitrogen-limited growth. Conversely, genes in the other cluster were down-regulated during light-limited growth but up-regulated during nitrogen-limited growth; this cluster included several genes involved in nitrogen uptake and assimilation. These results demonstrate complementary regulation of gene expression for two major metabolic activities of cyanobacteria. Comparison with batch-culture experiments revealed interesting differences in gene expression between batch and continuous culture and illustrates that continuous-culture experiments can pick up subtle changes in cell physiology and gene expression. PMID:21205618
Glenn, Anthony E.; Davis, C. Britton; Gao, Minglu; Gold, Scott E.; Mitchell, Trevor R.; Proctor, Robert H.; Stewart, Jane E.; Snook, Maurice E.
2016-01-01
Microbes encounter a broad spectrum of antimicrobial compounds in their environments and often possess metabolic strategies to detoxify such xenobiotics. We have previously shown that Fusarium verticillioides, a fungal pathogen of maize known for its production of fumonisin mycotoxins, possesses two unlinked loci, FDB1 and FDB2, necessary for detoxification of antimicrobial compounds produced by maize, including the γ-lactam 2-benzoxazolinone (BOA). In support of these earlier studies, microarray analysis of F. verticillioides exposed to BOA identified the induction of multiple genes at FDB1 and FDB2, indicating the loci consist of gene clusters. One of the FDB1 cluster genes encoded a protein having domain homology to the metallo-β-lactamase (MBL) superfamily. Deletion of this gene (MBL1) rendered F. verticillioides incapable of metabolizing BOA and thus unable to grow on BOA-amended media. Deletion of other FDB1 cluster genes, in particular AMD1 and DLH1, did not affect BOA degradation. Phylogenetic analyses and topology testing of the FDB1 and FDB2 cluster genes suggested two horizontal transfer events among fungi, one being transfer of FDB1 from Fusarium to Colletotrichum, and the second being transfer of the FDB2 cluster from Fusarium to Aspergillus. Together, the results suggest that plant-derived xenobiotics have exerted evolutionary pressure on these fungi, leading to horizontal transfer of genes that enhance fitness or virulence. PMID:26808652
Genome-Wide Significant Association between Alcohol Dependence and a Variant in the ADH Gene Cluster
Frank, Josef; Cichon, Sven; Treutlein, Jens; Ridinger, Monika; Mattheisen, Manuel; Hoffmann, Per; Herms, Stefan; Wodarz, Norbert; Soyka, Michael; Zill, Peter; Maier, Wolfgang; Mössner, Rainald; Gaebel, Wolfgang; Dahmen, Norbert; Scherbaum, Norbert; Schmäl, Christine; Steffens, Michael; Lucae, Susanne; Ising, Marcus; Müller-Myhsok, Bertram; Nöthen, Markus M; Mann, Karl; Kiefer, Falk; Rietschel, Marcella
2011-01-01
Alcohol dependence (AD) is an important contributory factor to the global burden of disease. The etiology of AD involves both environmental and genetic factors, and the disorder has a heritability of around 50%. The aim of the present study was to identify susceptibility genes for AD by performing a genome-wide association study (GWAS). The sample comprised 1,333 male in-patients with severe DSM-IV AD and 2,168 controls. These included 487 patients and 1,358 controls from a previous GWAS study by our group. All individuals were of German descent. Single marker tests and a polygenic score based analysis to assess the combined contribution of multiple markers with small effects were performed. The SNP rs1789891, which is located between the ADH1B and ADH1C genes, achieved genome-wide significance (p=1.27E–8; OR=1.46). Other markers from this region were also associated with AD, and conditional analyses indicated that these made a partially independent contribution. The SNP rs1789891 is in complete linkage disequilibrium with the functional Arg272Gln variant (p=1.24E–7, OR=1.31) of the ADH1C gene, which has been reported to modify the rate of ethanol oxidation to acetaldehyde in vitro. A polygenic score based approach produced a significant result (p=9.66E–9). This is the first GWAS of AD to provide genome-wide significant support for the role of the ADH gene cluster and to suggest a polygenic component to the etiology of AD. The latter result suggests that many more AD susceptibility genes still await identification. PMID:22004471
Stefan, Mihaela; Simmons, Rebecca A; Bertera, Suzanne; Trucco, Massimo; Esni, Farzad; Drain, Peter; Nicholls, Robert D
2011-05-01
Prader-Willi syndrome (PWS) is a multisystem disorder caused by genetic loss of function of a cluster of imprinted, paternally expressed genes. Neonatal failure to thrive in PWS is followed by childhood-onset hyperphagia and obesity among other endocrine and behavioral abnormalities. PWS is typically assumed to be caused by an unknown hypothalamic-pituitary dysfunction, but the underlying pathogenesis remains unknown. A transgenic deletion mouse model (TgPWS) has severe failure to thrive, with very low levels of plasma insulin and glucagon in fetal and neonatal life prior to and following onset of progressive hypoglycemia. In this study, we tested the hypothesis that primary deficits in pancreatic islet development or function may play a fundamental role in the TgPWS neonatal phenotype. Major pancreatic islet hormones (insulin, glucagon) were decreased in TgPWS mice, consistent with plasma levels. Immunohistochemical analysis of the pancreas demonstrated disrupted morphology of TgPWS islets, with reduced α- and β-cell mass arising from an increase in apoptosis. Furthermore, in vivo and in vitro studies show that the rate of insulin secretion is significantly impaired in TgPWS β-cells. In TgPWS pancreas, mRNA levels for genes encoding all pancreatic hormones, other secretory factors, and the ISL1 transcription factor are upregulated by either a compensatory response to plasma hormone deficiencies or a primary effect of a deleted gene. Our findings identify a cluster of imprinted genes required for the development, survival, coordinate regulation of genes encoding hormones, and secretory function of pancreatic endocrine cells, which may underlie the neonatal phenotype of the TgPWS mouse model.
antiSMASH 3.0—a comprehensive resource for the genome mining of biosynthetic gene clusters
Blin, Kai; Duddela, Srikanth; Krug, Daniel; Kim, Hyun Uk; Bruccoleri, Robert; Lee, Sang Yup; Fischbach, Michael A; Müller, Rolf; Wohlleben, Wolfgang; Breitling, Rainer; Takano, Eriko
2015-01-01
Abstract Microbial secondary metabolism constitutes a rich source of antibiotics, chemotherapeutics, insecticides and other high-value chemicals. Genome mining of gene clusters that encode the biosynthetic pathways for these metabolites has become a key methodology for novel compound discovery. In 2011, we introduced antiSMASH, a web server and stand-alone tool for the automatic genomic identification and analysis of biosynthetic gene clusters, available at http://antismash.secondarymetabolites.org. Here, we present version 3.0 of antiSMASH, which has undergone major improvements. A full integration of the recently published ClusterFinder algorithm now allows using this probabilistic algorithm to detect putative gene clusters of unknown types. Also, a new dereplication variant of the ClusterBlast module now identifies similarities of identified clusters to any of 1172 clusters with known end products. At the enzyme level, active sites of key biosynthetic enzymes are now pinpointed through a curated pattern-matching procedure and Enzyme Commission numbers are assigned to functionally classify all enzyme-coding genes. Additionally, chemical structure prediction has been improved by incorporating polyketide reduction states. Finally, in order for users to be able to organize and analyze multiple antiSMASH outputs in a private setting, a new XML output module allows offline editing of antiSMASH annotations within the Geneious software. PMID:25948579
β-globin gene cluster haplotypes in ethnic minority populations of southwest China
Sun, Hao; Liu, Hongxian; Huang, Kai; Lin, Keqin; Huang, Xiaoqin; Chu, Jiayou; Ma, Shaohui; Yang, Zhaoqing
2017-01-01
The genetic diversity and relationships among ethnic minority populations of southwest China were investigated using seven polymorphic restriction enzyme sites in the β-globin gene cluster. The haplotypes of 1392 chromosomes from ten ethnic populations living in southwest China were determined. Linkage equilibrium and recombination hotspot were found between the 5′ sites and 3′ sites of the β-globin gene cluster. 5′ haplotypes 2 (+−−−), 6 (−++−+), 9 (−++++) and 3′ haplotype FW3 (−+) were the predominant haplotypes. Notably, haplotype 9 frequency was significantly high in the southwest populations, indicating their difference with other Chinese. The interpopulation differentiation of southwest Chinese minority populations is less than those in populations of northern China and other continents. Phylogenetic analysis shows that populations sharing same ethnic origin or language clustered to each other, indicating current β-globin cluster diversity in the Chinese populations reflects their ethnic origin and linguistic affiliations to a great extent. This study characterizes β-globin gene cluster haplotypes in southwest Chinese minorities for the first time, and reveals the genetic variability and affinity of these populations using β-globin cluster haplotype frequencies. The results suggest that ethnic origin plays an important role in shaping variations of the β-globin gene cluster in the southwestern ethnic populations of China. PMID:28205625
Schertel, Claus; Albarca, Monica; Rockel-Bauer, Claudia; Kelley, Nicholas W.; Bischof, Johannes; Hens, Korneel
2015-01-01
Transcription factors (TFs) are key regulators of cell fate. The estimated 755 genes that encode DNA binding domain-containing proteins comprise ∼5% of all Drosophila genes. However, the majority has remained uncharacterized so far due to the lack of proper genetic tools. We generated 594 site-directed transgenic Drosophila lines that contain integrations of individual UAS-TF constructs to facilitate spatiotemporally controlled misexpression in vivo. All transgenes were expressed in the developing wing, and two-thirds induced specific phenotypic defects. In vivo knockdown of the same genes yielded a phenotype for 50%, with both methods indicating a great potential for misexpression to characterize novel functions in wing growth, patterning, and development. Thus, our UAS-TF library provides an important addition to the genetic toolbox of Drosophila research, enabling the identification of several novel wing development-related TFs. In parallel, we established the chromatin landscape of wing imaginal discs by ChIP-seq analyses of five chromatin marks and RNA Pol II. Subsequent clustering revealed six distinct chromatin states, with two clusters showing enrichment for both active and repressive marks. TFs that carry such “bivalent” chromatin are highly enriched for causing misexpression phenotypes in the wing, and analysis of existing expression data shows that these TFs tend to be differentially expressed across the wing disc. Thus, bivalently marked chromatin can be used as a marker for spatially regulated TFs that are functionally relevant in a developing tissue. PMID:25568052
Expressed sequence tag analysis of guinea pig (Cavia porcellus) eye tissues for NEIBank
Simpanya, Mukoma F.; Wistow, Graeme; Gao, James; David, Larry L.; Giblin, Frank J.
2008-01-01
Purpose To characterize gene expression patterns in guinea pig ocular tissues and identify orthologs of human genes from NEIBank expressed sequence tags. Methods RNA was extracted from dissected eye tissues of 2.5-month-old guinea pigs to make three unamplified and unnormalized cDNA libraries in the pCMVSport-6 vector for the lens, retina, and eye minus lens and retina. Over 4,000 clones were sequenced from each library and were analyzed using GRIST for clustering and gene identification. Lens crystallin EST data were validated using two-dimensional electrophoresis (2-DE), matrix assisted laser desorption (MALDI), and electrospray ionization mass spectrometry (ESIMS). Results Combined data from the three libraries generated a total of 6,694 distinctive gene clusters, with each library having between 1,000 and 3,000 clusters. Approximately 60% of the total gene clusters were novel cDNA sequences and had significant homologies to other mammalian sequences in GenBank. Complete cDNA sequences were obtained for many guinea pig lens proteins, including αA/αAinsert-, γN-, and γS-crystallins, lengsin and GRIFIN. The ratio of αA- to αB-crystallin on 2-DE gels was 8: 1 in the lens nucleus and 6.5: 1 in the cortex. Analysis of ESTs, genome sequence, and proteins (by MALDI), did not reveal any evidence for the presence of γD-, γE-, and γF-crystallin in the guinea pig. Predicted masses of many guinea pig lens crystallins were confirmed by ESIMS analysis. For the retina, orthologs of human phototransduction genes were found, such as Rhodopsin, S-antigen (Sag, Arrestin), and Transducin. The guinea-pig ortholog of NRL, a key rod photoreceptor-specific transcription factor, was also represented in EST data. In the ‘rest-of-eye’ library, the most abundant transcripts included decorin and keratin 12, representative of the cornea. Conclusions Genomic analysis of guinea pig eye tissues provides sequence-verified clones for future studies. Guinea pig orthologs of many human eye specific genes were identified. Guinea pig gene structures were similar to their human and rodent gene counterparts. Surprisingly, no orthologs of γD-, γE-, and γF-crystallin were found in EST, proteomic, or the current guinea pig genome data. PMID:19104676
Hidalgo, Pedro I; Ullán, Ricardo V; Albillos, Silvia M; Montero, Olimpio; Fernández-Bodega, María Ángeles; García-Estrada, Carlos; Fernández-Aguado, Marta; Martín, Juan-Francisco
2014-01-01
The PR-toxin is a potent mycotoxin produced by Penicillium roqueforti in moulded grains and grass silages and may contaminate blue-veined cheese. The PR-toxin derives from the 15 carbon atoms sesquiterpene aristolochene formed by the aristolochene synthase (encoded by ari1). We have cloned and sequenced a four gene cluster that includes the ari1 gene from P. roqueforti. Gene silencing of each of the four genes (named prx1 to prx4) resulted in a reduction of 65-75% in the production of PR-toxin indicating that the four genes encode enzymes involved in PR-toxin biosynthesis. Interestingly the four silenced mutants overproduce large amounts of mycophenolic acid, an antitumor compound formed by an unrelated pathway suggesting a cross-talk of PR-toxin and mycophenolic acid production. An eleven gene cluster that includes the above mentioned four prx genes and a 14-TMS drug/H(+) antiporter was found in the genome of Penicillium chrysogenum. This eleven gene cluster has been reported to be very poorly expressed in a transcriptomic study of P. chrysogenum genes under conditions of penicillin production (strongly aerated cultures). We found that this apparently silent gene cluster is able to produce PR-toxin in P. chrysogenum under static culture conditions on hydrated rice medium. Noteworthily, the production of PR-toxin was 2.6-fold higher in P. chrysogenum npe10, a strain deleted in the 56.8kb amplifiable region containing the pen gene cluster, than in the parental strain Wisconsin 54-1255 providing another example of cross-talk between secondary metabolite pathways in this fungus. A detailed PR-toxin biosynthesis pathway is proposed based on all available evidence. Copyright © 2013 Elsevier Inc. All rights reserved.
The nif Gene Operon of the Methanogenic Archaeon Methanococcus maripaludis
Kessler, Peter S.; Blank, Carrine; Leigh, John A.
1998-01-01
Nitrogen fixation occurs in two domains, Archaea and Bacteria. We have characterized a nif (nitrogen fixation) gene cluster in the methanogenic archaeon Methanococcus maripaludis. Sequence analysis revealed eight genes, six with sequence similarity to known nif genes and two with sequence similarity to glnB. The gene order, nifH, ORF105 (similar to glnB), ORF121 (similar to glnB), nifD, nifK, nifE, nifN, and nifX, was the same as that found in part in other diazotrophic methanogens and except for the presence of the glnB-like genes, also resembled the order found in many members of the Bacteria. Using transposon insertion mutagenesis, we determined that an 8-kb region required for nitrogen fixation corresponded to the nif gene cluster. Northern analysis revealed the presence of either a single 7.6-kb nif mRNA transcript or 10 smaller mRNA species containing portions of the large transcript. Polar effects of transposon insertions demonstrated that all of these mRNAs arose from a single promoter region, where transcription initiated 80 bp 5′ to nifH. Distinctive features of the nif gene cluster include the presence of the six primary nif genes in a single operon, the placement of the two glnB-like genes within the cluster, the apparent physical separation of the cluster from any other nif genes that might be in the genome, the fragmentation pattern of the mRNA, and the regulation of expression by a repression mechanism described previously. Our study and others with methanogenic archaea reporting multiple mRNAs arising from gene clusters with only a single putative promoter sequence suggest that mRNA processing following transcription may be a common occurrence in methanogens. PMID:9515920
A genomics based discovery of secondary metabolite biosynthetic gene clusters in Aspergillus ustus.
Pi, Borui; Yu, Dongliang; Dai, Fangwei; Song, Xiaoming; Zhu, Congyi; Li, Hongye; Yu, Yunsong
2015-01-01
Secondary metabolites (SMs) produced by Aspergillus have been extensively studied for their crucial roles in human health, medicine and industrial production. However, the resulting information is almost exclusively derived from a few model organisms, including A. nidulans and A. fumigatus, but little is known about rare pathogens. In this study, we performed a genomics based discovery of SM biosynthetic gene clusters in Aspergillus ustus, a rare human pathogen. A total of 52 gene clusters were identified in the draft genome of A. ustus 3.3904, such as the sterigmatocystin biosynthesis pathway that was commonly found in Aspergillus species. In addition, several SM biosynthetic gene clusters were firstly identified in Aspergillus that were possibly acquired by horizontal gene transfer, including the vrt cluster that is responsible for viridicatumtoxin production. Comparative genomics revealed that A. ustus shared the largest number of SM biosynthetic gene clusters with A. nidulans, but much fewer with other Aspergilli like A. niger and A. oryzae. These findings would help to understand the diversity and evolution of SM biosynthesis pathways in genus Aspergillus, and we hope they will also promote the development of fungal identification methodology in clinic.
A Genomics Based Discovery of Secondary Metabolite Biosynthetic Gene Clusters in Aspergillus ustus
Pi, Borui; Yu, Dongliang; Dai, Fangwei; Song, Xiaoming; Zhu, Congyi; Li, Hongye; Yu, Yunsong
2015-01-01
Secondary metabolites (SMs) produced by Aspergillus have been extensively studied for their crucial roles in human health, medicine and industrial production. However, the resulting information is almost exclusively derived from a few model organisms, including A. nidulans and A. fumigatus, but little is known about rare pathogens. In this study, we performed a genomics based discovery of SM biosynthetic gene clusters in Aspergillus ustus, a rare human pathogen. A total of 52 gene clusters were identified in the draft genome of A. ustus 3.3904, such as the sterigmatocystin biosynthesis pathway that was commonly found in Aspergillus species. In addition, several SM biosynthetic gene clusters were firstly identified in Aspergillus that were possibly acquired by horizontal gene transfer, including the vrt cluster that is responsible for viridicatumtoxin production. Comparative genomics revealed that A. ustus shared the largest number of SM biosynthetic gene clusters with A. nidulans, but much fewer with other Aspergilli like A. niger and A. oryzae. These findings would help to understand the diversity and evolution of SM biosynthesis pathways in genus Aspergillus, and we hope they will also promote the development of fungal identification methodology in clinic. PMID:25706180
Use of keyword hierarchies to interpret gene expression patterns.
Masys, D R; Welsh, J B; Lynn Fink, J; Gribskov, M; Klacansky, I; Corbeil, J
2001-04-01
High-density microarray technology permits the quantitative and simultaneous monitoring of thousands of genes. The interpretation challenge is to extract relevant information from this large amount of data. A growing variety of statistical analysis approaches are available to identify clusters of genes that share common expression characteristics, but provide no information regarding the biological similarities of genes within clusters. The published literature provides a potential source of information to assist in interpretation of clustering results. We describe a data mining method that uses indexing terms ('keywords') from the published literature linked to specific genes to present a view of the conceptual similarity of genes within a cluster or group of interest. The method takes advantage of the hierarchical nature of Medical Subject Headings used to index citations in the MEDLINE database, and the registry numbers applied to enzymes.
Clusters of Antibiotic Resistance Genes Enriched Together Stay Together in Swine Agriculture
Johnson, Timothy A.; Stedtfeld, Robert D.; Wang, Qiong; Cole, James R.; Hashsham, Syed A.; Looft, Torey; Zhu, Yong-Guan
2016-01-01
ABSTRACT Antibiotic resistance is a worldwide health risk, but the influence of animal agriculture on the genetic context and enrichment of individual antibiotic resistance alleles remains unclear. Using quantitative PCR followed by amplicon sequencing, we quantified and sequenced 44 genes related to antibiotic resistance, mobile genetic elements, and bacterial phylogeny in microbiomes from U.S. laboratory swine and from swine farms from three Chinese regions. We identified highly abundant resistance clusters: groups of resistance and mobile genetic element alleles that cooccur. For example, the abundance of genes conferring resistance to six classes of antibiotics together with class 1 integrase and the abundance of IS6100-type transposons in three Chinese regions are directly correlated. These resistance cluster genes likely colocalize in microbial genomes in the farms. Resistance cluster alleles were dramatically enriched (up to 1 to 10% as abundant as 16S rRNA) and indicate that multidrug-resistant bacteria are likely the norm rather than an exception in these communities. This enrichment largely occurred independently of phylogenetic composition; thus, resistance clusters are likely present in many bacterial taxa. Furthermore, resistance clusters contain resistance genes that confer resistance to antibiotics independently of their particular use on the farms. Selection for these clusters is likely due to the use of only a subset of the broad range of chemicals to which the clusters confer resistance. The scale of animal agriculture and its wastes, the enrichment and horizontal gene transfer potential of the clusters, and the vicinity of large human populations suggest that managing this resistance reservoir is important for minimizing human risk. PMID:27073098
Kihara, Takahiro; Hiroe, Ayaka; Ishii-Hyakutake, Manami; Mizuno, Kouhei; Tsuge, Takeharu
2017-08-01
Bacillus cereus and Bacillus megaterium both accumulate polyhydroxyalkanoate (PHA) but their PHA biosynthetic gene (pha) clusters that code for proteins involved in PHA biosynthesis are different. Namely, a gene encoding MaoC-like protein exists in the B. cereus-type pha cluster but not in the B. megaterium-type pha cluster. MaoC-like protein has an R-specific enoyl-CoA hydratase (R-hydratase) activity and is referred to as PhaJ when involved in PHA metabolism. In this study, the pha cluster of B. cereus YB-4 was characterized in terms of PhaJ's function. In an in vitro assay, PhaJ from B. cereus YB-4 (PhaJ YB4 ) exhibited hydration activity toward crotonyl-CoA. In an in vivo assay using Escherichia coli as a host for PHA accumulation, the recombinant strain expressing PhaJ YB4 and PHA synthase led to increased PHA accumulation, suggesting that PhaJ YB4 functioned as a monomer supplier. The monomer composition of the accumulated PHA reflected the substrate specificity of PhaJ YB4 , which appeared to prefer short chain-length substrates. The pha cluster from B. cereus YB-4 functioned to accumulate PHA in E. coli; however, it did not function when the phaJ YB4 gene was deleted. The B. cereus-type pha cluster represents a new example of a pha cluster that contains the gene encoding PhaJ.
Wotton, Karl R; Shimeld, Sebastian M
2011-12-01
In the human genome, members of the FoxC, FoxF, FoxL1, and FoxQ1 gene families are found in two paralagous clusters. One cluster contains the genes FOXQ1, FOXF2, FOXC1 and the second consists of FOXF1, FOXC2, and FOXL1. In jawed vertebrates these genes are known to be expressed in different pharyngeal tissues and all, except FoxQ1, are involved in patterning the early embryonic mesoderm. We have previously traced the evolution of this cluster in the bony vertebrates, and the gene content is identical in the dogfish, a member of the most basally branching lineage of the jawed vertebrates. Here we extend these analyses to jawless vertebrates. Using genomic searches and molecular approaches we have identified homologues of these genes from lampreys. We identify two FoxC genes, two FoxF genes, two FoxQ1 genes and single FoxL1 gene. We examine the embryonic expression of one predominantly mesodermally expressed gene family, FoxC, and the endodermally expressed member of the cluster, FoxQ1. We identified FoxQ1 transcripts in the pharyngeal endoderm, while the two FoxC genes are differentially expressed in the pharyngeal mesenchyme and ectoderm. Furthermore we identify conserved expression of lamprey FoxC genes in the paraxial and intermediate mesoderms. We interpret our results through a chordate-wide comparison of expression patterns and discuss gene content in the context of theories on the evolution of the vertebrate genome. 2011 Elsevier B.V. All rights reserved.
Wang, Hao; Fewer, David P; Holm, Liisa; Rouhiainen, Leo; Sivonen, Kaarina
2014-06-24
Nonribosomal peptides and polyketides are a diverse group of natural products with complex chemical structures and enormous pharmaceutical potential. They are synthesized on modular nonribosomal peptide synthetase (NRPS) and polyketide synthase (PKS) enzyme complexes by a conserved thiotemplate mechanism. Here, we report the widespread occurrence of NRPS and PKS genetic machinery across the three domains of life with the discovery of 3,339 gene clusters from 991 organisms, by examining a total of 2,699 genomes. These gene clusters display extraordinarily diverse organizations, and a total of 1,147 hybrid NRPS/PKS clusters were found. Surprisingly, 10% of bacterial gene clusters lacked modular organization, and instead catalytic domains were mostly encoded as separate proteins. The finding of common occurrence of nonmodular NRPS differs substantially from the current classification. Sequence analysis indicates that the evolution of NRPS machineries was driven by a combination of common descent and horizontal gene transfer. We identified related siderophore NRPS gene clusters that encoded modular and nonmodular NRPS enzymes organized in a gradient. A higher frequency of the NRPS and PKS gene clusters was detected from bacteria compared with archaea or eukarya. They commonly occurred in the phyla of Proteobacteria, Actinobacteria, Firmicutes, and Cyanobacteria in bacteria and the phylum of Ascomycota in fungi. The majority of these NRPS and PKS gene clusters have unknown end products highlighting the power of genome mining in identifying novel genetic machinery for the biosynthesis of secondary metabolites.
González, Víctor M; Aventín, Núria; Centeno, Emilio; Puigdomènech, Pere
2014-12-17
Plant NBS-LRR -resistance genes tend to be found in clusters, which have been shown to be hot spots of genome variability. In melon, half of the 81 predicted NBS-LRR genes group in nine clusters, and a 1 Mb region on linkage group V contains the highest density of R-genes and presence/absence gene polymorphisms found in the melon genome. This region is known to contain the locus of Vat, an agronomically important gene that confers resistance to aphids. However, the presence of duplications makes the sequencing and annotation of R-gene clusters difficult, usually resulting in multi-gapped sequences with higher than average errors. A 1-Mb sequence that contains the largest NBS-LRR gene cluster found in melon was improved using a strategy that combines Illumina paired-end mapping and PCR-based gap closing. Unknown sequence was decreased by 70% while about 3,000 SNPs and small indels were corrected. As a result, the annotations of 18 of a total of 23 NBS-LRR genes found in this region were modified, including additional coding sequences, amino acid changes, correction of splicing boundaries, or fussion of ORFs in common transcription units. A phylogeny analysis of the R-genes and their comparison with syntenic sequences in other cucurbits point to a pattern of local gene amplifications since the diversification of cucurbits from other families, and through speciation within the family. A candidate Vat gene is proposed based on the sequence similarity between a reported Vat gene from a Korean melon cultivar and a sequence fragment previously absent in the unrefined sequence. A sequence refinement strategy allowed substantial improvement of a 1 Mb fragment of the melon genome and the re-annotation of the largest cluster of NBS-LRR gene homologues found in melon. Analysis of the cluster revealed that resistance genes have been produced by sequence duplication in adjacent genome locations since the divergence of cucurbits from other close families, and through the process of speciation within the family a candidate Vat gene was also identified using sequence previously unavailable, which demonstrates the advantages of genome assembly refinements when analyzing complex regions such as those containing clusters of highly similar genes.
Li, Jing-Jing; Hu, Zi-Min; Sun, Zhong-Min; Yao, Jian-Ting; Liu, Fu-Li; Fresia, Pablo; Duan, De-Lin
2017-12-07
Long-term survival in isolated marginal seas of the China coast during the late Pleistocene ice ages is widely believed to be an important historical factor contributing to population genetic structure in coastal marine species. Whether or not contemporary factors (e.g. long-distance dispersal via coastal currents) continue to shape diversity gradients in marine organisms with high dispersal capability remains poorly understood. Our aim was to explore how historical and contemporary factors influenced the genetic diversity and distribution of the brown alga Sargassum thunbergii, which can drift on surface water, leading to long-distance dispersal. We used 11 microsatellites and the plastid RuBisCo spacer to evaluate the genetic diversity of 22 Sargassum thunbergii populations sampled along the China coast. Population structure and differentiation was inferred based on genotype clustering and pairwise F ST and allele-frequency analyses. Integrated genetic analyses revealed two genetic clusters in S. thunbergii that dominated in the Yellow-Bohai Sea (YBS) and East China Sea (ECS) respectively. Higher levels of genetic diversity and variation were detected among populations in the YBS than in the ECS. Bayesian coalescent theory was used to estimate contemporary and historical gene flow. High levels of contemporary gene flow were detected from the YBS (north) to the ECS (south), whereas low levels of historical gene flow occurred between the two regions. Our results suggest that the deep genetic divergence in S. thunbergii along the China coast may result from long-term geographic isolation during glacial periods. The dispersal of S. thunbergii driven by coastal currents may facilitate the admixture between southern and northern regimes. Our findings exemplify how both historical and contemporary forces are needed to understand phylogeographical patterns in coastal marine species with long-distance dispersal.
Hox gene expression during postlarval development of the polychaete Alitta virens.
Bakalenko, Nadezhda I; Novikova, Elena L; Nesterenko, Alexander Y; Kulakova, Milana A
2013-05-01
Hox genes are the family of transcription factors that play a key role in the patterning of the anterior-posterior axis of all bilaterian animals. These genes display clustered organization and colinear expression. Expression boundaries of individual Hox genes usually correspond with morphological boundaries of the body. Previously, we studied Hox gene expression during larval development of the polychaete Alitta virens (formerly Nereis virens) and discovered that Hox genes are expressed in nereid larva according to the spatial colinearity principle. Adult Alitta virens consist of multiple morphologically similar segments, which are formed sequentially in the growth zone. Since the worm grows for most of its life, postlarval segments constantly change their position along the anterior-posterior axis. We studied the expression dynamics of the Hox cluster during postlarval development of the nereid Alitta virens and found that 8 out of 11 Hox genes are transcribed as wide gene-specific gradients in the ventral nerve cord, ectoderm, and mesoderm. The expression domains constantly shift in accordance with the changing proportions of the growing worm, so expression domains of most Hox genes do not have stable anterior or/and posterior boundaries.In the course of our study, we revealed long antisense RNA (asRNA) for some Hox genes. Expression patterns of two of these genes were analyzed using whole-mount in-situ hybridization. This is the first discovery of antisense RNA for Hox genes in Lophotrochozoa. Hox gene expression in juvenile A. virens differs significantly from Hox gene expression patterns both in A. virens larva and in other Bilateria.We suppose that the postlarval function of the Hox genes in this polychaete is to establish and maintain positional coordinates in a constantly growing body, as opposed to creating morphological difference between segments.
Wide distribution of O157-antigen biosynthesis gene clusters in Escherichia coli.
Iguchi, Atsushi; Shirai, Hiroki; Seto, Kazuko; Ooka, Tadasuke; Ogura, Yoshitoshi; Hayashi, Tetsuya; Osawa, Kayo; Osawa, Ro
2011-01-01
Most Escherichia coli O157-serogroup strains are classified as enterohemorrhagic E. coli (EHEC), which is known as an important food-borne pathogen for humans. They usually produce Shiga toxin (Stx) 1 and/or Stx2, and express H7-flagella antigen (or nonmotile). However, O157 strains that do not produce Stxs and express H antigens different from H7 are sometimes isolated from clinical and other sources. Multilocus sequence analysis revealed that these 21 O157:non-H7 strains tested in this study belong to multiple evolutionary lineages different from that of EHEC O157:H7 strains, suggesting a wide distribution of the gene set encoding the O157-antigen biosynthesis in multiple lineages. To gain insight into the gene organization and the sequence similarity of the O157-antigen biosynthesis gene clusters, we conducted genomic comparisons of the chromosomal regions (about 59 kb in each strain) covering the O-antigen gene cluster and its flanking regions between six O157:H7/non-H7 strains. Gene organization of the O157-antigen gene cluster was identical among O157:H7/non-H7 strains, but was divided into two distinct types at the nucleotide sequence level. Interestingly, distribution of the two types did not clearly follow the evolutionary lineages of the strains, suggesting that horizontal gene transfer of both types of O157-antigen gene clusters has occurred independently among E. coli strains. Additionally, detailed sequence comparison revealed that some positions of the repetitive extragenic palindromic (REP) sequences in the regions flanking the O-antigen gene clusters were coincident with possible recombination points. From these results, we conclude that the horizontal transfer of the O157-antigen gene clusters induced the emergence of multiple O157 lineages within E. coli and speculate that REP sequences may involve one of the driving forces for exchange and evolution of O-antigen loci.
Ziemons, Sandra; Koutsantas, Katerina; Becker, Kordula; Dahlmann, Tim; Kück, Ulrich
2017-02-16
Multi-copy gene integration into microbial genomes is a conventional tool for obtaining improved gene expression. For Penicillium chrysogenum, the fungal producer of the beta-lactam antibiotic penicillin, many production strains carry multiple copies of the penicillin biosynthesis gene cluster. This discovery led to the generally accepted view that high penicillin titers are the result of multiple copies of penicillin genes. Here we investigated strain P2niaD18, a production line that carries only two copies of the penicillin gene cluster. We performed pulsed-field gel electrophoresis (PFGE), quantitative qRT-PCR, and penicillin bioassays to investigate production, deletion and overexpression strains generated in the P. chrysogenum P2niaD18 background, in order to determine the copy number of the penicillin biosynthesis gene cluster, and study the expression of one penicillin biosynthesis gene, and the penicillin titer. Analysis of production and recombinant strain showed that the enhanced penicillin titer did not depend on the copy number of the penicillin gene cluster. Our assumption was strengthened by results with a penicillin null strain lacking pcbC encoding isopenicillin N synthase. Reintroduction of one or two copies of the cluster into the pcbC deletion strain restored transcriptional high expression of the pcbC gene, but recombinant strains showed no significantly different penicillin titer compared to parental strains. Here we present a molecular genetic analysis of production and recombinant strains in the P2niaD18 background carrying different copy numbers of the penicillin biosynthesis gene cluster. Our analysis shows that the enhanced penicillin titer does not strictly depend on the copy number of the cluster. Based on these overall findings, we hypothesize that instead, complex regulatory mechanisms are prominently implicated in increased penicillin biosynthesis in production strains.
Nishida, Yuichiro; Adati, Naoki; Ozawa, Ritsuko; Maeda, Aasami; Sakaki, Yoshiyuki; Takeda, Tadayuki
2008-10-28
SH-SY5Y cells exhibit a neuronal phenotype when treated with all-trans retinoic acid (RA), but the molecular mechanism of activation in the signalling pathway mediated by phosphatidylinositol 3-kinase (PI3K) is unclear. To investigate this mechanism, we compared the gene expression profiles in SK-N-SH cells and two subtypes of SH-SY5Y cells (SH-SY5Y-A and SH-SY5Y-E), each of which show a different phenotype during RA-mediated differentiation. SH-SY5Y-A cells differentiated in the presence of RA, whereas RA-treated SH-SY5Y-E cells required additional treatment with brain-derived neurotrophic factor (BDNF) for full differentiation. After exposing cells to a PI3K inhibitor, LY294002, we identified 386 genes and categorised these genes into two clusters dependent on the PI3K signalling pathway during RA-mediated differentiation in SH-SY5Y-A cells. Transcriptional regulation of the gene cluster, including 158 neural genes, was greatly reduced in SK-N-SH cells and partially impaired in SH-SY5Y-E cells, which is consistent with a defect in the neuronal phenotype of these cells. Additional stimulation with BDNF induced a set of neural genes that were down-regulated in RA-treated SH-SY5Y-E cells but were abundant in differentiated SH-SY5Y-A cells. We identified gene clusters controlled by PI3K- and TRKB-mediated signalling pathways during the differentiation of two subtypes of SH-SY5Y cells. The TRKB-mediated bypass pathway compensates for impaired neural function generated by defects in several signalling pathways, including PI3K in SH-SY5Y-E cells. Our expression profiling data will be useful for further elucidation of the signal transduction-transcriptional network involving PI3K or TRKB.
Conversion events in gene clusters
2011-01-01
Background Gene clusters containing multiple similar genomic regions in close proximity are of great interest for biomedical studies because of their associations with inherited diseases. However, such regions are difficult to analyze due to their structural complexity and their complicated evolutionary histories, reflecting a variety of large-scale mutational events. In particular, conversion events can mislead inferences about the relationships among these regions, as traced by traditional methods such as construction of phylogenetic trees or multi-species alignments. Results To correct the distorted information generated by such methods, we have developed an automated pipeline called CHAP (Cluster History Analysis Package) for detecting conversion events. We used this pipeline to analyze the conversion events that affected two well-studied gene clusters (α-globin and β-globin) and three gene clusters for which comparative sequence data were generated from seven primate species: CCL (chemokine ligand), IFN (interferon), and CYP2abf (part of cytochrome P450 family 2). CHAP is freely available at http://www.bx.psu.edu/miller_lab. Conclusions These studies reveal the value of characterizing conversion events in the context of studying gene clusters in complex genomes. PMID:21798034
DOE Office of Scientific and Technical Information (OSTI.GOV)
Gallagher, Kelley A.; Jensen, Paul R.
Background: Considerable advances have been made in our understanding of the molecular genetics of secondary metabolite biosynthesis. Coupled with increased access to genome sequence data, new insight can be gained into the diversity and distributions of secondary metabolite biosynthetic gene clusters and the evolutionary processes that generate them. Here we examine the distribution of gene clusters predicted to encode the biosynthesis of a structurally diverse class of molecules called hybrid isoprenoids (HIs) in the genus Streptomyces. These compounds are derived from a mixed biosynthetic origin that is characterized by the incorporation of a terpene moiety onto a variety of chemicalmore » scaffolds and include many potent antibiotic and cytotoxic agents. Results: One hundred and twenty Streptomyces genomes were searched for HI biosynthetic gene clusters using ABBA prenyltransferases (PTases) as queries. These enzymes are responsible for a key step in HI biosynthesis. The strains included 12 that belong to the ‘MAR4’ clade, a largely marine-derived lineage linked to the production of diverse HI secondary metabolites. We found ABBA PTase homologs in all of the MAR4 genomes, which averaged five copies per strain, compared with 21 % of the non-MAR4 genomes, which averaged one copy per strain. Phylogenetic analyses suggest that MAR4 PTase diversity has arisen by a combination of horizontal gene transfer and gene duplication. Furthermore, there is evidence that HI gene cluster diversity is generated by the horizontal exchange of orthologous PTases among clusters. Many putative HI gene clusters have not been linked to their secondary metabolic products, suggesting that MAR4 strains will yield additional new compounds in this structure class. Finally, we confirm that the mevalonate pathway is not always present in genomes that contain HI gene clusters and thus is not a reliable query for identifying strains with the potential to produce HI secondary metabolites. In conclusion: We found that marine-derived MAR4 streptomycetes possess a relatively high genetic potential for HI biosynthesis. The combination of horizontal gene transfer, duplication, and rearrangement indicate that complex evolutionary processes account for the high level of HI gene cluster diversity in these bacteria, the products of which may provide a yet to be defined adaptation to the marine environment.« less
Gallagher, Kelley A.; Jensen, Paul R.
2015-11-17
Background: Considerable advances have been made in our understanding of the molecular genetics of secondary metabolite biosynthesis. Coupled with increased access to genome sequence data, new insight can be gained into the diversity and distributions of secondary metabolite biosynthetic gene clusters and the evolutionary processes that generate them. Here we examine the distribution of gene clusters predicted to encode the biosynthesis of a structurally diverse class of molecules called hybrid isoprenoids (HIs) in the genus Streptomyces. These compounds are derived from a mixed biosynthetic origin that is characterized by the incorporation of a terpene moiety onto a variety of chemicalmore » scaffolds and include many potent antibiotic and cytotoxic agents. Results: One hundred and twenty Streptomyces genomes were searched for HI biosynthetic gene clusters using ABBA prenyltransferases (PTases) as queries. These enzymes are responsible for a key step in HI biosynthesis. The strains included 12 that belong to the ‘MAR4’ clade, a largely marine-derived lineage linked to the production of diverse HI secondary metabolites. We found ABBA PTase homologs in all of the MAR4 genomes, which averaged five copies per strain, compared with 21 % of the non-MAR4 genomes, which averaged one copy per strain. Phylogenetic analyses suggest that MAR4 PTase diversity has arisen by a combination of horizontal gene transfer and gene duplication. Furthermore, there is evidence that HI gene cluster diversity is generated by the horizontal exchange of orthologous PTases among clusters. Many putative HI gene clusters have not been linked to their secondary metabolic products, suggesting that MAR4 strains will yield additional new compounds in this structure class. Finally, we confirm that the mevalonate pathway is not always present in genomes that contain HI gene clusters and thus is not a reliable query for identifying strains with the potential to produce HI secondary metabolites. In conclusion: We found that marine-derived MAR4 streptomycetes possess a relatively high genetic potential for HI biosynthesis. The combination of horizontal gene transfer, duplication, and rearrangement indicate that complex evolutionary processes account for the high level of HI gene cluster diversity in these bacteria, the products of which may provide a yet to be defined adaptation to the marine environment.« less
2017-06-30
Clustered Regularly Interspaced Short Palindromic Repeat/ CRISPR -associated protein 9 ( CRISPR /Cas9)-based Gene Drives En vi ro nm en ta l L ab or at...Management on Military Lands Clustered Regularly Interspaced Short Palindromic Repeat/ CRISPR -associated protein 9 ( CRISPR /Cas9)-based Gene Drives Ping... CRISPR /Cas9-based Gene Drives for Invasive Species Management on Military Lands” ERDC/EL SR-17-2 ii Abstract Applications of genetic engineering
Boyd, David A.; Willey, Barbara M.; Fawcett, Darlene; Gillani, Nazira; Mulvey, Michael R.
2008-01-01
Enterococcus faecalis N06-0364, exhibiting a vancomycin MIC of 8 μg/ml, was found to harbor a novel d-Ala-d-Ser gene cluster, designated vanL. The vanL gene cluster was similar in organization to the vanC operon, but the VanT serine racemase was encoded by two separate genes, vanTmL (membrane binding) and vanTrL (racemase). PMID:18458129
Sinha, Somya; Raxwal, Vivek K.; Joshi, Bharat; Jagannath, Arun; Katiyar-Agarwal, Surekha; Goel, Shailendra; Kumar, Amar; Agarwal, Manu
2015-01-01
Low temperature is a major abiotic stress that impedes plant growth and development. Brassica juncea is an economically important oil seed crop and is sensitive to freezing stress during pod filling subsequently leading to abortion of seeds. To understand the cold stress mediated global perturbations in gene expression, whole transcriptome of B. juncea siliques that were exposed to sub-optimal temperature was sequenced. Manually self-pollinated siliques at different stages of development were subjected to either short (6 h) or long (12 h) durations of chilling stress followed by construction of RNA-seq libraries and deep sequencing using Illumina's NGS platform. De-novo assembly of B. juncea transcriptome resulted in 133,641 transcripts, whose combined length was 117 Mb and N50 value was 1428 bp. We identified 13,342 differentially regulated transcripts by pair-wise comparison of 18 transcriptome libraries. Hierarchical clustering along with Spearman correlation analysis identified that the differentially expressed genes segregated in two major clusters representing early (5–15 DAP) and late stages (20–30 DAP) of silique development. Further analysis led to the discovery of sub-clusters having similar patterns of gene expression. Two of the sub-clusters (one each from the early and late stages) comprised of genes that were inducible by both the durations of cold stress. Comparison of transcripts from these clusters led to identification of 283 transcripts that were commonly induced by cold stress, and were referred to as “core cold-inducible” transcripts. Additionally, we found that 689 and 100 transcripts were specifically up-regulated by cold stress in early and late stages, respectively. We further explored the expression patterns of gene families encoding for transcription factors (TFs), transcription regulators (TRs) and kinases, and found that cold stress induced protein kinases only during early silique development. We validated the digital gene expression profiles of selected transcripts by qPCR and found a high degree of concordance between the two analyses. To our knowledge this is the first report of transcriptome sequencing of cold-stressed B. juncea siliques. The data generated in this study would be a valuable resource for not only understanding the cold stress signaling pathway but also for introducing cold hardiness in B. juncea. PMID:26579175
The alpha1-fetoprotein locus is activated by a nuclear receptor of the Drosophila FTZ-F1 family.
Galarneau, L; Paré, J F; Allard, D; Hamel, D; Levesque, L; Tugwood, J D; Green, S; Bélanger, L
1996-07-01
The alpha1-fetoprotein (AFP) gene is located between the albumin and alpha-albumin genes and is activated by transcription factor FTF (fetoprotein transcription factor), presumed to transduce early developmental signals to the albumin gene cluster. We have identified FTF as an orphan nuclear receptor of the Drosophila FTZ-F1 family. FTF recognizes the DNA sequence 5'-TCAAGGTCA-3', the canonical recognition motif for FTZ-F1 receptors. cDNA sequence homologies indicate that rat FTF is the ortholog of mouse LRH-1 and Xenopus xFF1rA. Rodent FTF is encoded by a single-copy gene, related to the gene encoding steroidogenic factor 1 (SF-1). The 5.2-kb FTF transcript is translated from several in-frame initiator codons into FTF isoforms (54 to 64 kDa) which appear to bind DNA as monomers, with no need for a specific ligand, similar KdS (approximately equal 3 x 10(-10) M), and similar transcriptional effects. FTF activates the AFP promoter without the use of an amino-terminal activation domain; carboxy-terminus-truncated FTF exerts strong dominant negative effects. In the AFP promoter, FTF recruits an accessory trans-activator which imparts glucocorticoid reactivity upon the AFP gene. FTF binding sites are found in the promoters of other liver-expressed genes, some encoding liver transcription factors; FTF, liver alpha1-antitrypsin promoter factor LFB2, and HNF-3beta promoter factor UF2-H3beta are probably the same factor. FTF is also abundantly expressed in the pancreas and may exert differentiation functions in endodermal sublineages, similar to SF-1 in steroidogenic tissues. HepG2 hepatoma cells seem to express a mutated form of FTF.
A transcriptional dynamic network during Arabidopsis thaliana pollen development.
Wang, Jigang; Qiu, Xiaojie; Li, Yuhua; Deng, Youping; Shi, Tieliu
2011-01-01
To understand transcriptional regulatory networks (TRNs), especially the coordinated dynamic regulation between transcription factors (TFs) and their corresponding target genes during development, computational approaches would represent significant advances in the genome-wide expression analysis. The major challenges for the experiments include monitoring the time-specific TFs' activities and identifying the dynamic regulatory relationships between TFs and their target genes, both of which are currently not yet available at the large scale. However, various methods have been proposed to computationally estimate those activities and regulations. During the past decade, significant progresses have been made towards understanding pollen development at each development stage under the molecular level, yet the regulatory mechanisms that control the dynamic pollen development processes remain largely unknown. Here, we adopt Networks Component Analysis (NCA) to identify TF activities over time course, and infer their regulatory relationships based on the coexpression of TFs and their target genes during pollen development. We carried out meta-analysis by integrating several sets of gene expression data related to Arabidopsis thaliana pollen development (stages range from UNM, BCP, TCP, HP to 0.5 hr pollen tube and 4 hr pollen tube). We constructed a regulatory network, including 19 TFs, 101 target genes and 319 regulatory interactions. The computationally estimated TF activities were well correlated to their coordinated genes' expressions during the development process. We clustered the expression of their target genes in the context of regulatory influences, and inferred new regulatory relationships between those TFs and their target genes, such as transcription factor WRKY34, which was identified that specifically expressed in pollen, and regulated several new target genes. Our finding facilitates the interpretation of the expression patterns with more biological relevancy, since the clusters corresponding to the activity of specific TF or the combination of TFs suggest the coordinated regulation of TFs to their target genes. Through integrating different resources, we constructed a dynamic regulatory network of Arabidopsis thaliana during pollen development with gene coexpression and NCA. The network illustrated the relationships between the TFs' activities and their target genes' expression, as well as the interactions between TFs, which provide new insight into the molecular mechanisms that control the pollen development.
Epstein-Barr Virus oncoprotein super-enhancers control B cell growth
Zhou, Hufeng; Schmidt, Stefanie CS; Jiang, Sizun; Willox, Bradford; Bernhardt, Katharina; Liang, Jun; Johannsen, Eric C; Kharchenko, Peter; Gewurz, Benjamin E; Kieff, Elliott; Zhao, Bo
2015-01-01
Summary Super-enhancers are clusters of gene-regulatory sites bound by multiple transcription factors that govern cell transcription, development, phenotype, and oncogenesis. By examining Epstein-Barr virus (EBV) transformed lymphoblastoid cell lines (LCLs), we identified four EBV oncoproteins and five EBV-activated NF-κB subunits co-occupying ~1800 enhancer sites. Of these, 187 had markedly higher and broader histone H3K27ac signals characteristic of super-enhancers, and were designated “EBV super-enhancers”. EBV super-enhancer-associated genes included the MYC and BCL2 oncogenes, enabling LCL proliferation and survival. EBV super-enhancers were enriched for B cell transcription factor motifs and had a high co-occupancy of the transcription factors STAT5 and NFAT. EBV super-enhancer-associated genes were more highly expressed than other LCL genes. Disrupting EBV super-enhancers by the bromodomain inhibitor, JQ1 or conditionally inactivating an EBV oncoprotein or NF-κB decreased MYC or BCL2 expression and arrested LCL growth. These findings provide insight into mechanisms of EBV-induced lymphoproliferation and identify potential therapeutic interventions. PMID:25639793
Coping with living in the soil: the genome of the parthenogenetic springtail Folsomia candida.
Faddeeva-Vakhrusheva, Anna; Kraaijeveld, Ken; Derks, Martijn F L; Anvar, Seyed Yahya; Agamennone, Valeria; Suring, Wouter; Kampfraath, Andries A; Ellers, Jacintha; Le Ngoc, Giang; van Gestel, Cornelis A M; Mariën, Janine; Smit, Sandra; van Straalen, Nico M; Roelofs, Dick
2017-06-28
Folsomia candida is a model in soil biology, belonging to the family of Isotomidae, subclass Collembola. It reproduces parthenogenetically in the presence of Wolbachia, and exhibits remarkable physiological adaptations to stress. To better understand these features and adaptations to life in the soil, we studied its genome in the context of its parthenogenetic lifestyle. We applied Pacific Bioscience sequencing and assembly to generate a reference genome for F. candida of 221.7 Mbp, comprising only 162 scaffolds. The complete genome of its endosymbiont Wolbachia, was also assembled and turned out to be the largest strain identified so far. Substantial gene family expansions and lineage-specific gene clusters were linked to stress response. A large number of genes (809) were acquired by horizontal gene transfer. A substantial fraction of these genes are involved in lignocellulose degradation. Also, the presence of genes involved in antibiotic biosynthesis was confirmed. Intra-genomic rearrangements of collinear gene clusters were observed, of which 11 were organized as palindromes. The Hox gene cluster of F. candida showed major rearrangements compared to arthropod consensus cluster, resulting in a disorganized cluster. The expansion of stress response gene families suggests that stress defense was important to facilitate colonization of soils. The large number of HGT genes related to lignocellulose degradation could be beneficial to unlock carbohydrate sources in soil, especially those contained in decaying plant and fungal organic matter. Intra- as well as inter-scaffold duplications of gene clusters may be a consequence of its parthenogenetic lifestyle. This high quality genome will be instrumental for evolutionary biologists investigating deep phylogenetic lineages among arthropods and will provide the basis for a more mechanistic understanding in soil ecology and ecotoxicology.
Shi, Pibiao; Guy, Kateta Malangisha; Wu, Weifang; Fang, Bingsheng; Yang, Jinghua; Zhang, Mingfang; Hu, Zhongyuan
2016-04-12
The plant-specific TCP transcription factor family, which is involved in the regulation of cell growth and proliferation, performs diverse functions in multiple aspects of plant growth and development. However, no comprehensive analysis of the TCP family in watermelon (Citrullus lanatus) has been undertaken previously. A total of 27 watermelon TCP encoding genes distributed on nine chromosomes were identified. Phylogenetic analysis clustered the genes into 11 distinct subgroups. Furthermore, phylogenetic and structural analyses distinguished two homology classes within the ClTCP family, designated Class I and Class II. The Class II genes were differentiated into two subclasses, the CIN subclass and the CYC/TB1 subclass. The expression patterns of all members were determined by semi-quantitative PCR. The functions of two ClTCP genes, ClTCP14a and ClTCP15, in regulating plant height were confirmed by ectopic expression in Arabidopsis wild-type and ortholog mutants. This study represents the first genome-wide analysis of the watermelon TCP gene family, which provides valuable information for understanding the classification and functions of the TCP genes in watermelon.
Pancreatic islet enhancer clusters enriched in type 2 diabetes risk-associated variants.
Pasquali, Lorenzo; Gaulton, Kyle J; Rodríguez-Seguí, Santiago A; Mularoni, Loris; Miguel-Escalada, Irene; Akerman, İldem; Tena, Juan J; Morán, Ignasi; Gómez-Marín, Carlos; van de Bunt, Martijn; Ponsa-Cobas, Joan; Castro, Natalia; Nammo, Takao; Cebola, Inês; García-Hurtado, Javier; Maestro, Miguel Angel; Pattou, François; Piemonti, Lorenzo; Berney, Thierry; Gloyn, Anna L; Ravassard, Philippe; Skarmeta, José Luis Gómez; Müller, Ferenc; McCarthy, Mark I; Ferrer, Jorge
2014-02-01
Type 2 diabetes affects over 300 million people, causing severe complications and premature death, yet the underlying molecular mechanisms are largely unknown. Pancreatic islet dysfunction is central in type 2 diabetes pathogenesis, and understanding islet genome regulation could therefore provide valuable mechanistic insights. We have now mapped and examined the function of human islet cis-regulatory networks. We identify genomic sequences that are targeted by islet transcription factors to drive islet-specific gene activity and show that most such sequences reside in clusters of enhancers that form physical three-dimensional chromatin domains. We find that sequence variants associated with type 2 diabetes and fasting glycemia are enriched in these clustered islet enhancers and identify trait-associated variants that disrupt DNA binding and islet enhancer activity. Our studies illustrate how islet transcription factors interact functionally with the epigenome and provide systematic evidence that the dysregulation of islet enhancers is relevant to the mechanisms underlying type 2 diabetes.
Biosynthetic Genes for the Tetrodecamycin Antibiotics
Gverzdys, Tomas
2016-01-01
ABSTRACT We recently described 13-deoxytetrodecamycin, a new member of the tetrodecamycin family of antibiotics. A defining feature of these molecules is the presence of a five-membered lactone called a tetronate ring. By sequencing the genome of a producer strain, Streptomyces sp. strain WAC04657, and searching for a gene previously implicated in tetronate ring formation, we identified the biosynthetic genes responsible for producing 13-deoxytetrodecamycin (the ted genes). Using the ted cluster in WAC04657 as a reference, we found related clusters in three other organisms: Streptomyces atroolivaceus ATCC 19725, Streptomyces globisporus NRRL B-2293, and Streptomyces sp. strain LaPpAH-202. Comparing the four clusters allowed us to identify the cluster boundaries. Genetic manipulation of the cluster confirmed the involvement of the ted genes in 13-deoxytetrodecamycin biosynthesis and revealed several additional molecules produced through the ted biosynthetic pathway, including tetrodecamycin, dihydrotetrodecamycin, and another, W5.9, a novel molecule. Comparison of the bioactivities of these four molecules suggests that they may act through the covalent modification of their target(s). IMPORTANCE The tetrodecamycins are a distinct subgroup of the tetronate family of secondary metabolites. Little is known about their biosynthesis or mechanisms of action, making them an attractive subject for investigation. In this paper we present the biosynthetic gene cluster for 13-deoxytetrodecamycin in Streptomyces sp. strain WAC04657. We identify related clusters in several other organisms and show that they produce related molecules. PMID:27137499
Butyrate production in phylogenetically diverse Firmicutes isolated from the chicken caecum
Eeckhaut, Venessa; Van Immerseel, Filip; Croubels, Siska; De Baere, Siegrid; Haesebrouck, Freddy; Ducatelle, Richard; Louis, Petra; Vandamme, Peter
2011-01-01
Summary Sixteen butyrate‐producing bacteria were isolated from the caecal content of chickens and analysed phylogenetically. They did not represent a coherent phylogenetic group, but were allied to four different lineages in the Firmicutes phylum. Fourteen strains appeared to represent novel species, based on a level of ≤ 98.5% 16S rRNA gene sequence similarity towards their nearest validly named neighbours. The highest butyrate concentrations were produced by the strains belonging to clostridial clusters IV and XIVa, clusters which are predominant in the chicken caecal microbiota. In only one of the 16 strains tested, the butyrate kinase operon could be amplified, while the butyryl‐CoA : acetate CoA‐transferase gene was detected in eight strains belonging to clostridial clusters IV, XIVa and XIVb. None of the clostridial cluster XVI isolates carried this gene based on degenerate PCR analyses. However, another CoA‐transferase gene more similar to propionate CoA‐transferase was detected in the majority of the clostridial cluster XVI isolates. Since this gene is located directly downstream of the remaining butyrate pathway genes in several human cluster XVI bacteria, it may be involved in butyrate formation in these bacteria. The present study indicates that butyrate producers related to cluster XVI may play a more important role in the chicken gut than in the human gut. PMID:21375722
Andrés-Benito, Pol; Moreno, Jesús; Aso, Ester; Povedano, Mónica; Ferrer, Isidro
2017-01-01
Transcriptome arrays identifies 747 genes differentially expressed in the anterior horn of the spinal cord and 2,300 genes differentially expressed in frontal cortex area 8 in a single group of typical sALS cases without frontotemporal dementia compared with age-matched controls. Main up-regulated clusters in the anterior horn are related to inflammation and apoptosis; down-regulated clusters are linked to axoneme structures and protein synthesis. In contrast, up-regulated gene clusters in frontal cortex area 8 involve neurotransmission, synaptic proteins and vesicle trafficking, whereas main down-regulated genes cluster into oligodendrocyte function and myelin-related proteins. RT-qPCR validates the expression of 58 of 66 assessed genes from different clusters. The present results: a. reveal regional differences in de-regulated gene expression between the anterior horn of the spinal cord and frontal cortex area 8 in the same individuals suffering from sALS; b. validate and extend our knowledge about the complexity of the inflammatory response in the anterior horn of the spinal cord; and c. identify for the first time extensive gene up-regulation of neurotransmission and synaptic-related genes, together with significant down-regulation of oligodendrocyte- and myelin-related genes, as important contributors to the pathogenesis of frontal cortex alterations in the sALS/frontotemporal lobar degeneration spectrum complex at stages with no apparent cognitive impairment. PMID:28283675
A gene network bioinformatics analysis for pemphigoid autoimmune blistering diseases.
Barone, Antonio; Toti, Paolo; Giuca, Maria Rita; Derchi, Giacomo; Covani, Ugo
2015-07-01
In this theoretical study, a text mining search and clustering analysis of data related to genes potentially involved in human pemphigoid autoimmune blistering diseases (PAIBD) was performed using web tools to create a gene/protein interaction network. The Search Tool for the Retrieval of Interacting Genes/Proteins (STRING) database was employed to identify a final set of PAIBD-involved genes and to calculate the overall significant interactions among genes: for each gene, the weighted number of links, or WNL, was registered and a clustering procedure was performed using the WNL analysis. Genes were ranked in class (leader, B, C, D and so on, up to orphans). An ontological analysis was performed for the set of 'leader' genes. Using the above-mentioned data network, 115 genes represented the final set; leader genes numbered 7 (intercellular adhesion molecule 1 (ICAM-1), interferon gamma (IFNG), interleukin (IL)-2, IL-4, IL-6, IL-8 and tumour necrosis factor (TNF)), class B genes were 13, whereas the orphans were 24. The ontological analysis attested that the molecular action was focused on extracellular space and cell surface, whereas the activation and regulation of the immunity system was widely involved. Despite the limited knowledge of the present pathologic phenomenon, attested by the presence of 24 genes revealing no protein-protein direct or indirect interactions, the network showed significant pathways gathered in several subgroups: cellular components, molecular functions, biological processes and the pathologic phenomenon obtained from the Kyoto Encyclopaedia of Genes and Genomes (KEGG) database. The molecular basis for PAIBD was summarised and expanded, which will perhaps give researchers promising directions for the identification of new therapeutic targets.
Zhao, Qian; Ma, Dongna; Huang, Yuping; He, Weiyi; Li, Yiying; Vasseur, Liette; You, Minsheng
2018-04-01
Transcription factors (TFs), which play a vital role in regulating gene expression, are prevalent in all organisms and characterization of them may provide important clues for understanding regulation in vivo. The present study reports a genome-wide investigation of TFs in the diamondback moth, Plutella xylostella (L.), a worldwide pest of crucifers. A total of 940 TFs distributed among 133 families were identified. Phylogenetic analysis of insect species showed that some of these families were found to have expanded during the evolution of P. xylostella or Lepidoptera. RNA-seq analysis showed that some of the TF families, such as zinc fingers, homeobox, bZIP, bHLH, and MADF_DNA_bdg genes, were highly expressed in certain tissues including midgut, salivary glands, fat body, and hemocytes, with an obvious sex-biased expression pattern. In addition, a number of TFs showed significant differences in expression between insecticide susceptible and resistant strains, suggesting that these TFs play a role in regulating genes related to insecticide resistance. Finally, we identified an expansion of the HOX cluster in Lepidoptera, which might be related to Lepidoptera-specific evolution. Knockout of this cluster using CRISPR/Cas9 showed that the egg cannot hatch, indicating that this cluster may be related to egg development and maturation. This is the first comprehensive study on identifying and characterizing TFs in P. xylostella. Our results suggest that some TF families are expanded in the P. xylostella genome, and these TFs may have important biological roles in growth, development, sexual dimorphism, and resistance to insecticides. The present work provides a solid foundation for understanding regulation via TFs in P. xylostella and insights into the evolution of the P. xylostella genome.
Kovina, A P; Petrova, N V; Razin, S V; Yarovaia, O V
2016-01-01
In warm-blooded vertebrates, the α- and β-globin genes are organized in domains of different types and are regulated in different fashion. In cold-blooded vertebrates and, in particular, the tropical fish Danio rerio, the α- and β-globin genes form two gene clusters. A major D. rerio globin gene cluster is in chromosome 3 and includes the α- and β-globin genes of embryonic-larval and adult types. The region upstream of the cluster contains c16orf35, harbors the main regulatory element (MRE) of the α-globin gene domain in warm-blooded vertebrates. In this study, transient transfection of erythroid cells with genetic constructs containing a reporter gene under the control of potential regulatory elements of the domain was performed to characterize the promoters of the embryonic-larval and adult α- and β-globin genes of the major cluster. Also, in the 5th intron of c16orf35 in Danio reriowas detected a functional analog of the warm-blooded vertebrate MRE. This enhancer stimulated activity of the promoters of both adult and embryonic-larval α- and β-globin genes.
Histone and ribosomal RNA repetitive gene clusters of the boll weevil are linked in a tandem array.
Roehrdanz, R; Heilmann, L; Senechal, P; Sears, S; Evenson, P
2010-08-01
Histones are the major protein component of chromatin structure. The histone family is made up of a quintet of proteins, four core histones (H2A, H2B, H3 & H4) and the linker histones (H1). Spacers are found between the coding regions. Among insects this quintet of genes is usually clustered and the clusters are tandemly repeated. Ribosomal DNA contains a cluster of the rRNA sequences 18S, 5.8S and 28S. The rRNA genes are separated by the spacers ITS1, ITS2 and IGS. This cluster is also tandemly repeated. We found that the ribosomal RNA repeat unit of at least two species of Anthonomine weevils, Anthonomus grandis and Anthonomus texanus (Coleoptera: Curculionidae), is interspersed with a block containing the histone gene quintet. The histone genes are situated between the rRNA 18S and 28S genes in what is known as the intergenic spacer region (IGS). The complete reiterated Anthonomus grandis histone-ribosomal sequence is 16,248 bp.
Song, Xiaowei; Wang, Yajun; Tang, Yezhong
2013-01-01
As one of the most conserved genes in vertebrates, FoxP2 is widely involved in a number of important physiological and developmental processes. We systematically studied the evolutionary history and functional adaptations of FoxP2 in teleosts. The duplicated FoxP2 genes (FoxP2a and FoxP2b), which were identified in teleosts using synteny and paralogon analysis on genome databases of eight organisms, were probably generated in the teleost-specific whole genome duplication event. A credible classification with FoxP2, FoxP2a and FoxP2b in phylogenetic reconstructions confirmed the teleost-specific FoxP2 duplication. The unavailability of FoxP2b in Danio rerio suggests that the gene was deleted through nonfunctionalization of the redundant copy after the Otocephala-Euteleostei split. Heterogeneity in evolutionary rates among clusters consisting of FoxP2 in Sarcopterygii (Cluster 1), FoxP2a in Teleostei (Cluster 2) and FoxP2b in Teleostei (Cluster 3), particularly between Clusters 2 and 3, reveals asymmetric functional divergence after the gene duplication. Hierarchical cluster analyses of hydrophobicity profiles demonstrated significant structural divergence among the three clusters with verification of subsequent stepwise discriminant analysis, in which FoxP2 of Leucoraja erinacea and Lepisosteus oculatus were classified into Cluster 1, whereas FoxP2b of Salmo salar was grouped into Cluster 2 rather than Cluster 3. The simulated thermodynamic stability variations of the forkhead box domain (monomer and homodimer) showed remarkable divergence in FoxP2, FoxP2a and FoxP2b clusters. Relaxed purifying selection and positive Darwinian selection probably were complementary driving forces for the accelerated evolution of FoxP2 in ray-finned fishes, especially for the adaptive evolution of FoxP2a and FoxP2b in teleosts subsequent to the teleost-specific gene duplication.
Song, Xiaowei; Wang, Yajun; Tang, Yezhong
2013-01-01
As one of the most conserved genes in vertebrates, FoxP2 is widely involved in a number of important physiological and developmental processes. We systematically studied the evolutionary history and functional adaptations of FoxP2 in teleosts. The duplicated FoxP2 genes (FoxP2a and FoxP2b), which were identified in teleosts using synteny and paralogon analysis on genome databases of eight organisms, were probably generated in the teleost-specific whole genome duplication event. A credible classification with FoxP2, FoxP2a and FoxP2b in phylogenetic reconstructions confirmed the teleost-specific FoxP2 duplication. The unavailability of FoxP2b in Danio rerio suggests that the gene was deleted through nonfunctionalization of the redundant copy after the Otocephala-Euteleostei split. Heterogeneity in evolutionary rates among clusters consisting of FoxP2 in Sarcopterygii (Cluster 1), FoxP2a in Teleostei (Cluster 2) and FoxP2b in Teleostei (Cluster 3), particularly between Clusters 2 and 3, reveals asymmetric functional divergence after the gene duplication. Hierarchical cluster analyses of hydrophobicity profiles demonstrated significant structural divergence among the three clusters with verification of subsequent stepwise discriminant analysis, in which FoxP2 of Leucoraja erinacea and Lepisosteus oculatus were classified into Cluster 1, whereas FoxP2b of Salmo salar was grouped into Cluster 2 rather than Cluster 3. The simulated thermodynamic stability variations of the forkhead box domain (monomer and homodimer) showed remarkable divergence in FoxP2, FoxP2a and FoxP2b clusters. Relaxed purifying selection and positive Darwinian selection probably were complementary driving forces for the accelerated evolution of FoxP2 in ray-finned fishes, especially for the adaptive evolution of FoxP2a and FoxP2b in teleosts subsequent to the teleost-specific gene duplication. PMID:24349554
Leveraging long sequencing reads to investigate R-gene clustering and variation in sugar beet
USDA-ARS?s Scientific Manuscript database
Host-pathogen interactions are of prime importance to modern agriculture. Plants utilize various types of resistance genes to mitigate pathogen damage. Identification of the specific gene responsible for a specific resistance can be difficult due to duplication and clustering within R-gene families....
Khosravi, Claire; Kun, Roland Sándor; Visser, Jaap; Aguilar-Pontes, María Victoria; de Vries, Ronald P; Battaglia, Evy
2017-11-06
The genes of the non-phosphorylative L-rhamnose catabolic pathway have been identified for several yeast species. In Schefferomyces stipitis, all L-rhamnose pathway genes are organized in a cluster, which is conserved in Aspergillus niger, except for the lra-4 ortholog (lraD). The A. niger cluster also contains the gene encoding the L-rhamnose responsive transcription factor (RhaR) that has been shown to control the expression of genes involved in L-rhamnose release and catabolism. In this paper, we confirmed the function of the first three putative L-rhamnose utilisation genes from A. niger through gene deletion. We explored the identity of the inducer of the pathway regulator (RhaR) through expression analysis of the deletion mutants grown in transfer experiments to L-rhamnose and L-rhamnonate. Reduced expression of L-rhamnose-induced genes on L-rhamnose in lraA and lraB deletion strains, but not on L-rhamnonate (the product of LraB), demonstrate that the inducer of the pathway is of L-rhamnonate or a compound downstream of it. Reduced expression of these genes in the lraC deletion strain on L-rhamnonate show that it is in fact a downstream product of L-rhamnonate. This work showed that the inducer of RhaR is beyond L-rhamnonate dehydratase (LraC) and is likely to be the 2-keto-3-L-deoxyrhamnonate.
Ancient genomic architecture for mammalian olfactory receptor clusters
Aloni, Ronny; Olender, Tsviya; Lancet, Doron
2006-01-01
Background Mammalian olfactory receptor (OR) genes reside in numerous genomic clusters of up to several dozen genes. Whole-genome sequence alignment nets of five mammals allow their comprehensive comparison, aimed at reconstructing the ancestral olfactory subgenome. Results We developed a new and general tool for genome-wide definition of genomic gene clusters conserved in multiple species. Syntenic orthologs, defined as gene pairs showing conservation of both genomic location and coding sequence, were subjected to a graph theory algorithm for discovering CLICs (clusters in conservation). When applied to ORs in five mammals, including the marsupial opossum, more than 90% of the OR genes were found within a framework of 48 multi-species CLICs, invoking a general conservation of gene order and composition. A detailed analysis of individual CLICs revealed multiple differences among species, interpretable through species-specific genomic rearrangements and reflecting complex mammalian evolutionary dynamics. One significant instance involves CLIC #1, which lacks a human member, implying the human-specific deletion of an OR cluster, whose mouse counterpart has been tentatively associated with isovaleric acid odorant detection. Conclusion The identified multi-species CLICs demonstrate that most of the mammalian OR clusters have a common ancestry, preceding the split between marsupials and placental mammals. However, only two of these CLICs were capable of incorporating chicken OR genes, parsimoniously implying that all other CLICs emerged subsequent to the avian-mammalian divergence. PMID:17010214
antiSMASH 3.0-a comprehensive resource for the genome mining of biosynthetic gene clusters.
Weber, Tilmann; Blin, Kai; Duddela, Srikanth; Krug, Daniel; Kim, Hyun Uk; Bruccoleri, Robert; Lee, Sang Yup; Fischbach, Michael A; Müller, Rolf; Wohlleben, Wolfgang; Breitling, Rainer; Takano, Eriko; Medema, Marnix H
2015-07-01
Microbial secondary metabolism constitutes a rich source of antibiotics, chemotherapeutics, insecticides and other high-value chemicals. Genome mining of gene clusters that encode the biosynthetic pathways for these metabolites has become a key methodology for novel compound discovery. In 2011, we introduced antiSMASH, a web server and stand-alone tool for the automatic genomic identification and analysis of biosynthetic gene clusters, available at http://antismash.secondarymetabolites.org. Here, we present version 3.0 of antiSMASH, which has undergone major improvements. A full integration of the recently published ClusterFinder algorithm now allows using this probabilistic algorithm to detect putative gene clusters of unknown types. Also, a new dereplication variant of the ClusterBlast module now identifies similarities of identified clusters to any of 1172 clusters with known end products. At the enzyme level, active sites of key biosynthetic enzymes are now pinpointed through a curated pattern-matching procedure and Enzyme Commission numbers are assigned to functionally classify all enzyme-coding genes. Additionally, chemical structure prediction has been improved by incorporating polyketide reduction states. Finally, in order for users to be able to organize and analyze multiple antiSMASH outputs in a private setting, a new XML output module allows offline editing of antiSMASH annotations within the Geneious software. © The Author(s) 2015. Published by Oxford University Press on behalf of Nucleic Acids Research.
Campa, Ana; Giraldez, Ramón; Ferreira, Juan José
2011-06-01
Resistance to the eight races (3, 7, 19, 31, 81, 449, 453, and 1545) of the pathogenic fungus Colletotrichum lindemuthianum (anthracnose) was evaluated in F(3) families derived from the cross between the anthracnose differential bean cultivars Kaboon and Michelite. Molecular marker analyses were carried out in the F(2) individuals in order to map and characterize the anthracnose resistance genes or gene clusters present in Kaboon. The analysis of the combined segregations indicates that the resistance present in Kaboon against these eight anthracnose races is determined by 13 different race-specific genes grouped in three clusters. One of these clusters, corresponding to locus Co-1 in linkage group (LG) 1, carries two dominant genes conferring specific resistance to races 81 and 1545, respectively, and a gene necessary (dominant complementary gene) for the specific resistance to race 31. A second cluster, corresponding to locus Co-3/9 in LG 4, carries six dominant genes conferring specific resistance to races 3, 7, 19, 449, 453, and 1545, respectively, and the second dominant complementary gene for the specific resistance to race 31. A third cluster of unknown location carries three dominant genes conferring specific resistance to races 449, 453, and 1545, respectively. This is the first time that two anthracnose resistance genes with a complementary mode of action have been mapped in common bean and their relationship with previously known Co- resistance genes established.
Tsai, Yu-Cheng; Cooke, Nancy E.; Liebhaber, Stephen A.
2016-01-01
Abstract The relationships of higher order chromatin organization to mammalian gene expression remain incompletely defined. The human Growth Hormone (hGH) multigene cluster contains five gene paralogs. These genes are selectively activated in either the pituitary or the placenta by distinct components of a remote locus control region (LCR). Prior studies have revealed that appropriate activation of the placental genes is dependent not only on the actions of the LCR, but also on the multigene composition of the cluster itself. Here, we demonstrate that the hGH LCR ‘loops’ over a distance of 28 kb in primary placental nuclei to make specific contacts with the promoters of the two GH genes in the cluster. This long-range interaction sequesters the GH genes from the three hCS genes which co-assemble into a tightly packed ‘hCS chromatin hub’. Elimination of the long-range looping, via specific deletion of the placental LCR components, triggers a dramatic disruption of the hCS chromatin hub. These data reveal a higher-order structural pathway by which long-range looping from an LCR impacts on local chromatin architecture that is linked to tissue-specific gene regulation within a multigene cluster. PMID:26893355
Conditional clustering of temporal expression profiles
Wang, Ling; Montano, Monty; Rarick, Matt; Sebastiani, Paola
2008-01-01
Background Many microarray experiments produce temporal profiles in different biological conditions but common cluster techniques are not able to analyze the data conditional on the biological conditions. Results This article presents a novel technique to cluster data from time course microarray experiments performed across several experimental conditions. Our algorithm uses polynomial models to describe the gene expression patterns over time, a full Bayesian approach with proper conjugate priors to make the algorithm invariant to linear transformations, and an iterative procedure to identify genes that have a common temporal expression profile across two or more experimental conditions, and genes that have a unique temporal profile in a specific condition. Conclusion We use simulated data to evaluate the effectiveness of this new algorithm in finding the correct number of clusters and in identifying genes with common and unique profiles. We also use the algorithm to characterize the response of human T cells to stimulations of antigen-receptor signaling gene expression temporal profiles measured in six different biological conditions and we identify common and unique genes. These studies suggest that the methodology proposed here is useful in identifying and distinguishing uniquely stimulated genes from commonly stimulated genes in response to variable stimuli. Software for using this clustering method is available from the project home page. PMID:18334028
Wang, Chenyin; Saar, Valeria; Leung, Ka Lai; Chen, Liang; Wong, Garry
2018-01-01
Alzheimer's disease (AD) is a progressive neurodegenerative disorder characterized by the presence of extracellular amyloid plaques consisting of Amyloid-β peptide (Aβ) aggregates and neurofibrillary tangles formed by aggregation of hyperphosphorylated microtubule-associated protein tau. We generated a novel invertebrate model of AD by crossing Aβ1-42 (strain CL2355) with either pro-aggregating tau (strain BR5270) or anti-aggregating tau (strain BR5271) pan-neuronal expressing transgenic Caenorhabditis elegans. The lifespan and progeny viability of the double transgenic strains were significantly decreased compared with wild type N2 (P<0.0001). In addition, co-expression of these transgenes interfered with neurotransmitter signaling pathways, caused deficits in chemotaxis associative learning, increased protein aggregation visualized by Congo red staining, and increased neuronal loss. Global transcriptomic RNA-seq analysis revealed 248 up- and 805 down-regulated genes in N2 wild type versus Aβ1-42+pro-aggregating tau animals, compared to 293 up- and 295 down-regulated genes in N2 wild type versus Aβ1-42+anti-aggregating tau animals. Gene set enrichment analysis of Aβ1-42+pro-aggregating tau animals uncovered up-regulated annotation clusters UDP-glucuronosyltransferase (5 genes, P<4.2E-4), protein phosphorylation (5 genes, P<2.60E-02), and aging (5 genes, P<8.1E-2) while the down-regulated clusters included nematode cuticle collagen (36 genes, P<1.5E-21). RNA interference of 13 available top up-regulated genes in Aβ1-42+pro-aggregating tau animals revealed that F-box family genes and nep-4 could enhance life span deficits and chemotaxis deficits while Y39G8C.2 (TTBK2) could suppress these behaviors. Comparing the list of regulated genes from C. elegans to the top 60 genes related to human AD confirmed an overlap of 8 genes: patched homolog 1, PTCH1 (ptc-3), the Rab GTPase activating protein, TBC1D16 (tbc-16), the WD repeat and FYVE domain-containing protein 3, WDFY3 (wdfy-3), ADP-ribosylation factor guanine nucleotide exchange factor 2, ARFGEF2 (agef-1), Early B-cell Factor, EBF1 (unc-3), d-amino-acid oxidase, DAO (daao-1), glutamate receptor, metabotropic 1, GRM1 (mgl-2), prolyl 4-hydroxylase subunit alpha 2, P4HA2 (dpy-18 and phy-2). Taken together, our C. elegans double transgenic model provides insight on the fundamental neurobiologic processes underlying human AD and recapitulates selected transcriptomic changes observed in human AD brains. Copyright © 2017 Elsevier Inc. All rights reserved.
DOE Office of Scientific and Technical Information (OSTI.GOV)
Duncan, Katherine R.; Crüsemann, Max; Lechner, Anna
Genome sequencing has revealed that bacteria contain many more biosynthetic gene clusters than predicted based on the number of secondary metabolites discovered to date. While this biosynthetic reservoir has fostered interest in new tools for natural product discovery, there remains a gap between gene cluster detection and compound discovery. In this paper, we apply molecular networking and the new concept of pattern-based genome mining to 35 Salinispora strains, including 30 for which draft genome sequences were either available or obtained for this study. The results provide a method to simultaneously compare large numbers of complex microbial extracts, which facilitated themore » identification of media components, known compounds and their derivatives, and new compounds that could be prioritized for structure elucidation. Finally, these efforts revealed considerable metabolite diversity and led to several molecular family-gene cluster pairings, of which the quinomycin-type depsipeptide retimycin A was characterized and linked to gene cluster NRPS40 using pattern-based bioinformatic approaches.« less
Duncan, Katherine R.; Crüsemann, Max; Lechner, Anna; ...
2015-04-09
Genome sequencing has revealed that bacteria contain many more biosynthetic gene clusters than predicted based on the number of secondary metabolites discovered to date. While this biosynthetic reservoir has fostered interest in new tools for natural product discovery, there remains a gap between gene cluster detection and compound discovery. In this paper, we apply molecular networking and the new concept of pattern-based genome mining to 35 Salinispora strains, including 30 for which draft genome sequences were either available or obtained for this study. The results provide a method to simultaneously compare large numbers of complex microbial extracts, which facilitated themore » identification of media components, known compounds and their derivatives, and new compounds that could be prioritized for structure elucidation. Finally, these efforts revealed considerable metabolite diversity and led to several molecular family-gene cluster pairings, of which the quinomycin-type depsipeptide retimycin A was characterized and linked to gene cluster NRPS40 using pattern-based bioinformatic approaches.« less