Method for screening genes influencing drought adaptability of Hu sheep
Through transcriptome and genome-wide methylation analysis, the LOC121817420 gene was screened, which solved the gap in the selection of drought-adaptive genes of Lake Sheep, revealed the adaptive molecular mechanism of Lake Sheep in the Hotan region of Xinjiang, verified its expression differences in different regions, and provided theoretical support.
Patent Information
- Application Number
- CN202510487380.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-22
AI Technical Summary
There is a lack of effective gene screening methods in the prior art to study the adaptive changes of lake sheep in a drought environment after being introduced from the humid Taihu area to the Hotan area in Xinjiang. The genes that affect the drought adaptability of lake sheep are still unclear.
Transcriptome and genome-wide methylation analysis methods were used, combined with real-time fluorescence quantitative PCR, and the gene LOC121817420 that affects the drought adaptability of lake sheep was screened out, and the differences in expression levels in different regions were verified through high-throughput sequencing technology and bioinformatics analysis.
It was confirmed that the LOC121817420 gene was high in lake sheep in Hotan area, Xinjiang, and is the main candidate gene different from lake sheep in Taihu Basin. It revealed the molecular mechanism of drought adaptability of lake sheep and provides a theoretical basis for the quality and seed maintenance and efficient use of lake sheep.
Smart Images

Figure CN120350130A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of molecular breeding, and in particular to a method for screening genes affecting the drought adaptability of Hu sheep. Background Art
[0002] Environmental adaptability is a complex and delicate formation process. During this process, multiple genes play roles and regulate and influence each other. Eventually, multiple gene networks and metabolic pathways are formed to perform functions. The adaptation of animals to the environment is a combination of phenotypic adaptation and evolutionary adaptation. The former is caused by short-term physiological regulation, and the latter is generated by long-term genetic variation. Studying the role of functional genes in biological evolution helps to reveal the molecular mechanism of animal adaptation to the environment. The plateau environment has characteristics such as low oxygen, strong ultraviolet rays, and high cold, which pose unique challenges to the physiology and gene expression of animals. Many scholars have applied transcriptomics to the study of plateau adaptability. DNA methylation modification plays an important role in the adaptation of animals to the environment, especially when dealing with extreme environments such as temperature changes, food shortages, and high altitudes. In recent years, researchers have used advanced high-throughput sequencing technologies and genomic analysis methods to reveal the complex mechanism of DNA methylation in the environmental adaptability of mammals.
[0003] Environmental changes and human selection are accelerating the evolutionary adaptation of various organisms to the environment. In various research technologies regarding animal environmental adaptability, they have become increasingly mature, and the combined use of multiple biological technologies can more comprehensively analyze animal environmental adaptability. Hu sheep has been introduced from its humid original place to the extremely arid Hotan area, but there are relatively few studies on the drought environmental adaptability and resistance of Hu sheep. With the arrival of the "post-era" of Hu sheep, the quality preservation and efficient utilization of Hu sheep in Xinjiang have become a new challenge. Conducting research on the environmental adaptability and resistance of Hu sheep and using high-throughput sequencing technology to mine genes for high-quality traits of Hu sheep can not only provide good reference value and theoretical basis for the introduction of other sheep breeds in Xinjiang in the future, but also provide research and exploration value for the utilization of local sheep breed resources.
[0004] Currently, relevant genes regarding environmental adaptability have been reported, but it is not clear about the genetic changes that occur after Hu sheep is introduced from the Taihu Lake area to Hotan area in Xinjiang. There is an urgent need for a method for screening genes affecting the drought adaptability of Hu sheep in the existing technology. Summary of the Invention
[0005] In view of this, the embodiments of the present invention provide a method for screening genes affecting the drought adaptability of Hu sheep, and further explore the influence of the epigenetic modification regulation mechanism of TPL on the proliferation of MM cells through miR-381 / HMGB1.
[0006] To achieve the above object, the present invention mainly provides the following technical solutions:
[0007] On the one hand, an embodiment of the present invention provides a method for screening genes affecting the drought adaptability of Hu sheep, and the method includes the following steps:
[0008] Step 1. Selection and treatment of experimental sheep: Six rams with a body weight of 39-41 kg and good growth status are selected from Hotan area, Xinjiang and Taihu area respectively. 6 mL of blood (containing anticoagulant EDTA2Na, 22 mg) is collected from the jugular vein, and then the sheep are slaughtered by cervical exsanguination. Liver tissue samples are collected and quickly placed into cryotubes and stored in liquid nitrogen, and then taken back to the laboratory and placed in an ultra-low temperature refrigerator at -80 °C.
[0009] Step 2. Transcriptome analysis: RNA extraction, library construction and sequencing of Hu sheep whole blood, DNA extraction, library construction and sequencing of Hu sheep whole blood, transcriptome reference genome alignment and differential gene identification;
[0010] Step 3. Whole-genome methylation analysis: Sheep genomic DNA extraction and quality inspection, DNA library construction and quality inspection, on-machine sequencing and data filtering, reference genome alignment and DMRs identification, DNA methylation level and correlation analysis;
[0011] Step 4. Joint analysis of transcriptome and whole-genome methylation: Genes affecting the drought adaptability of Hu sheep are screened, and joint analysis of gene expression levels and methylation levels, methylation level distribution of genes with different expression patterns, intersection enrichment of RNA differential genes and promoter genes of differentially methylated regions, and intersection enrichment of RNA differential genes and gene bodies of differentially methylated regions are carried out.
[0012] Preferably, it further includes Step 5. Verification of genes affecting the drought adaptability of Hu sheep by real-time fluorescence quantitative PCR.
[0013] Furthermore, Step 2 is specifically as follows:
[0014] (I) RNA extraction and detection
[0015] Agarose gel electrophoresis: Analyze the integrity of sample RNA and whether there is DNA contamination; NanoPhotometer spectrophotometer: Detect the purity of RNA (OD260 / 280 and OD260 / 230 ratios); Agilent 2100 bioanalyzer: Accurately detect the integrity of RNA.
[0016] (II) Library construction and quality inspection
[0017] The starting RNA for library construction is total RNA, with a total amount ≥ 800 ng. The library construction kit used in library construction is the VAHTS Universal V10 RNA Sequencing Library Construction Kit. Poly(A)-tailed mRNA is enriched using Oligo(dT) magnetic beads, and then the obtained mRNA is randomly fragmented using divalent cations in NEB Fragmentation Buffer. Using the fragmented mRNA as a template and random oligonucleotides as primers, the first strand of cDNA is synthesized in the M-MuLV reverse transcriptase system. Subsequently, the RNA strand is degraded using RNase H enzyme, and the second strand of cDNA is synthesized using dNTPs as raw materials in the DNA polymerase I system.
[0018] The purified double-stranded cDNA is subjected to end repair, A-tailing, and ligation of sequencing adapters. AMPure XP magnetic beads are used to screen for cDNA fragments of approximately 200 bp. After PCR amplification, AMPure XP magnetic beads are used again to purify the PCR products, and finally a sequencing library is obtained.
[0019] After the library construction is completed, first use a Qubit 2.0 Fluorometer to perform a preliminary quantification of the library, and dilute the library to 1.5 ng / μL. Subsequently, use an Agilent 2100 bioanalyzer to detect the insert size of the library. After confirming that the insert size meets the expectations, real-time fluorescence quantitative PCR (qRT-PCR) is used to accurately quantify the effective concentration of the library (requiring the effective concentration of the library to be higher than 2 nM) to ensure that the library quality meets the sequencing requirements.
[0020] (III) Sequencing on the machine
[0021] After the library quality inspection is qualified, different libraries are pooled according to their effective concentrations and the requirements of the target sequencing data volume, and then sequenced using the Illumina NovaSeq X Plus sequencing platform to generate 150 bp paired-end reads. Sequencing uses the principle of "Sequencing by Synthesis" (SBS) technology. The specific process is as follows: Four fluorescently labeled dNTPs, DNA polymerase, and adapter primers are added to the sequencing flow cell for bridge amplification. When each sequencing cluster extends the complementary strand, a corresponding fluorescent signal is released every time a fluorescently labeled dNTP is incorporated. The sequencer captures these fluorescent signals in real time, and the computer software converts the optical signals into sequencing peak diagrams, and finally obtains the sequence information of the fragment to be tested.
[0022] (4) Data Quality Control
[0023] The image data of sequencing fragments measured by the high-throughput sequencer is converted into sequence reads (reads) through base recognition by CASAVA, and the file is in fastq format, which mainly contains the sequence information of the sequencing fragments and their corresponding sequencing quality information. A small number of reads with sequencing adapters or low sequencing quality are included in the raw data obtained by sequencing. In order to ensure the quality and reliability of data analysis, it is necessary to filter the raw data. This mainly includes removing reads with adapters, removing reads containing N (N indicates that the base information cannot be determined), and removing low-quality reads (data where the bases with Qphred <= 20 account for more than 50% of the entire read length). fastp (0.23.2) is used for raw data filtering, and the parameters are "--cut_front --cut_tail --cut_window_size 4 -a auto --cut_mean_quality 20 --length_required 36". All subsequent analyses are high-quality analyses based on clean data.
[0024] (5) Sequence Alignment to the Reference Genome
[0025] The reference genome and gene model annotation file are directly downloaded from the genome website. HISAT2 (v2.2.1) is used to build the index of the reference genome, and HISAT2 (v2.2.1) is used to align the paired-end clean reads with the reference genome, with the parameters "--phred33 --no-mixed --no-discordant".
[0026] (6) Gene Expression Level Quantification
[0027] Feature Counts (v2.0.1) is used to calculate the reads mapped to each gene. Then, the FPKM of each gene is calculated according to the gene length, and the reads mapped to the gene are calculated. FPKM refers to the expected number of transcripts per kilobase of exon model per million mapped reads. The influence of sequencing depth and gene length on read counting is considered, and it is currently the most commonly used method for estimating gene expression levels.
[0028] (7) Differential Expression Analysis
[0029] For those with biological replicates, differential expression analysis between two comparison groups was performed using the DESeq2 software (1.16.1). DESeq2 provides statistical procedures that use a model based on the negative binomial distribution to determine differential expression in digital gene expression data. The Benjamini and Hochberg method was used to adjust the resulting P-values to control the false discovery rate. Genes with padj < 0.05 found by DESeq2 were considered differentially expressed.
[0030] For those without biological replicates, the edgeR software was used. Before performing differential gene expression analysis, for each sequencing library, the read counts were adjusted through an edgeR package using a scaling normalization factor. Differential expression analysis between two conditions was performed using the edgeR R package (3.18.1). The P-values were adjusted using the Benjamini & Hochberg method. The adjusted P-values and |log2foldchange| were used as the thresholds for significantly differential expression.
[0031] (VIII) Differential gene enrichment analysis
[0032] GO enrichment analysis of differentially expressed genes was implemented using the clusterProfiler (4.2.0) R software, and GO terms with adjusted P-values less than 0.05 were regarded as GO terms significantly enriched by differentially expressed genes. KEGG is a database resource for understanding the high-level functions and utilities of biological systems from molecular-level information, especially large-scale molecular datasets generated by genome sequencing and other high-throughput databases, such as cells, organisms, and ecosystems. We used the clusterProfiler R software to analyze the statistical enrichment of differentially expressed genes in KEGG pathways.
[0033] (IX) Analysis of differential gene protein network interactions
[0034] Analysis of the network interactions of differentially expressed genes was based on the STRING database of known and predicted protein-protein interactions. For species present in the database, we constructed a network by extracting the list of target genes from the database; the target gene sequences were aligned with the selected reference protein sequences using diamond (v2.0.14.152), and then a network was established based on the known interactions of the selected reference species.
[0035] Furthermore, step 3 specifically is:
[0036] (I) Extracting genomic DNA
[0037] For tissue samples, take an appropriate amount of the sample and grind it in liquid nitrogen or homogenize it in a homogenizer; after adding 2% CTAB to the tissue, place it in a 65°C water bath for 1 h; add phenol-chloroform-isoamyl alcohol for extraction, after centrifugation, aspirate the upper aqueous phase and discard the protein layer; add isopropanol to precipitate for 10 min, centrifuge at 12,000 g at 4°C for 10 min; discard the supernatant, wash twice with 75% ethanol, and dry the precipitate; add 50 μl of TE (containing RNase) buffer to dissolve the precipitate, and incubate at 37°C for 30 min; for the dissolved DNA, take out a part for quality inspection and measure its OD value and integrity.
[0038] (II) DNA library construction
[0039] The starting amount of DNA for library construction should be greater than 1000 nanograms (>1000 ng), and at the same time, unmethylated λDNA used to evaluate the transformation efficiency is added. After the intact DNA is fragmented by sonication to obtain fragments of 100 - 500 bp, the DNA is then treated with bisulfite to convert unmethylated cytosine to uracil. The treated DNA needs to be end-repaired, including phosphorylation modification at the 5' end and addition of a dA tail at the 3' end. The end-repaired fragments need to be ligated with adapter oligonucleotides, and the DNA is purified using AM Pure XP magnetic beads. The adapter-ligated products after purification or size selection are subjected to PCR amplification, and after amplification, the PCR products are purified again using AM Pure XP magnetic beads to finally obtain the library. After the library construction is completed, first use a Qubit 2.0 fluorometer for preliminary concentration determination, and then detect the insert size of the library through an Agilent 2100 bioanalyzer. When the insert size meets the expectation, quantitative real-time fluorescence PCR (qRT-PCR) is used to accurately determine the effective concentration of the library (required to be higher than 2 nanomoles per liter) to ensure that the library quality meets the standard.
[0040] (III) Sequencing on the machine
[0041] After passing the library inspection, different libraries are mixed according to the requirements of the effective concentration and the target output data volume and then sequenced on an Illumina NovaSeq X Plus.
[0042] (IV) Quality control and alignment with the reference genome
[0043] The fastp software is used to perform quality control on the raw sequencing data to remove adapter sequences and low-quality bases (Q value < 20). The bisulfite-treated sequences are aligned to the reference genome through the BSMAP software. Calculate the bisulfite conversion rate, that is, the proportion of unmethylated cytosine (C) converted to thymine (T).
[0044] (V) Methylation detection
[0045] Extracting methylation information of cytosine sites based on the cytosine-thymine conversion principle: Unmethylated cytosine is converted to thymine, while methylated cytosine remains unchanged. Cytosine sites are classified into CpG, CHG, and CHH contexts (where H represents A, T, or C). The methylation status is determined by binomial distribution test, and cytosine sites with coverage >4× and false discovery rate (FDR) <0.05 are identified as methylated sites. The methylation level of a single cytosine is calculated by the formula ML = mC / (mC + umC), where ML is the methylation level, and mC and umC represent the reads supporting methylated C and unmethylated C, respectively. The metilene software is used to identify differentially methylated regions (DMRs), and regions with adj.p-value (q-value) <0.05 and absolute value of methylation difference (|diffMethy|) >0.2 are considered significant DMRs. Genes annotated by DMRs are called differentially methylated genes (DMGs).
[0046] Further, step 4 is specifically as follows:
[0047] (I) Joint analysis of gene expression level and methylation level
[0048] Data statistics: Statistics are performed on gene expression level (RNA_fpkm) and methylation levels (meth_level_genebody and meth_level_promoter), and a table is generated to show the expression level and methylation level of each gene. Bivariate relationship plot: A bivariate relationship plot of gene expression level and methylation levels in the gene promoter region and gene body region is drawn to show the distribution and relationship of the two groups of variables. The correlation between gene expression level and methylation level is calculated using the pearson correlation coefficient.
[0049] (II) Distribution of genebody methylation levels of genes with different expression patterns
[0050] Methylation level distribution plot: A plot of the average methylation levels of genes with high-level expression, medium-level expression, low-level expression, and non-expression patterns in the gene body and 2K upstream of TSS (transcription start site) and 2K downstream of TES (transcription end site) is drawn to show the methylation level distribution of genes with different expression patterns.
[0051] (III) Intersection enrichment of RNA differential genes and promoters of differentially methylated regions
[0052] Venn diagram: Draw a Venn diagram of the transcriptionally differentially expressed genes (DEGs) and promoters of differentially methylated regions (DMRs) to show the common genes between the two. GO functional enrichment analysis: Perform GO functional enrichment analysis on the intersection of the differentially expressed genes and the promoters of differentially methylated regions, generate a GO enrichment result table, and draw a scatter plot to show the significantly enriched functional terms. KEGG pathway enrichment analysis: Perform KEGG pathway enrichment analysis on the intersection of the differentially expressed genes (DEGs) and the genes related to the differentially methylated regions (DMRs) in the promoter region, generate a KEGG enrichment result table, and draw a scatter plot to show the significantly enriched pathway entries.
[0053] Further, step 5 is specifically as follows:
[0054] (1) Completely melt the template RNA, specific primers, 2×FastKing RT-qPCR Buffer (SYBRGreen), 50×ROX Reference Dye, and RNase-Free ddH2O. After brief centrifugation, place them on ice bath.
[0055] (2) Prepare the reaction solution under ice bath conditions according to Table 1 below:
[0056] Table 1
[0057] Composition Volume / Reaction 2×FastKing RT-qPCR Buffer (SYBR Green) 25 μl 25×RT-PCR Enzyme Mix 2 μl Forward Specific Primer (10 μM) 1.25 μl*1 Reverse Specific Primer (10 μM) 1.25 μl*1 RNA Template 10 pg - 1 μg total RNA <![CDATA[50×ROX Reference Dye 2 > Add according to different instruments <![CDATA[RNase-Free ddH2O]]> Make up the volume to 50 μl with water Total System 50 μl
[0058] (3) Perform Real Time One Step RT-qPCR reaction
[0059] Centrifuge the PCR reaction tube instantaneously with a centrifuge and then put it into a fluorescence quantitative PCR instrument for Real Time PCR reaction. It is recommended to adopt the standard PCR reaction procedure shown in Table 1. If good experimental results cannot be obtained using this procedure, then optimize the PCR conditions.
[0060] On the other hand, the embodiment of the present invention provides the gene affecting the drought adaptability of Hu sheep obtained by the screening method described above. The sequence of this gene is shown as SEQ ID NO.1.
[0061] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0062] By using the transcriptome and whole-genome methylation methods of the present invention, the LOC121817420 gene is mined, and the expression level of this gene in the whole blood of Hu sheep in two regions is detected by real-time fluorescence quantitative method. The expression level of this gene is high in Hu sheep in Hotan area of Xinjiang, which proves that this gene is one of the main candidate genes that distinguish Xinjiang Hu sheep from Hu sheep in the Taihu Lake Basin. Description of the Drawings
[0063] Figure 1RNA agarose gel electrophoresis pattern, where THSB1 - 6 are samples of Hu sheep from Taihu Lake area, and HTSB1 - 6 are samples of Hu sheep from Hotan area;
[0064] Figure 2 DNA agarose gel electrophoresis pattern, where note: THSB1 - 3 are samples of Hu sheep from Taihu Lake area, HTSB1 - 3 are samples of Hu sheep from Hotan area, and M is DL Marker 10000;
[0065] Figure 3 Distribution trend diagram of the average methylation level of gene functional elements;
[0066] Figure 4 Correlation analysis diagram of different samples, where (A) CG level; (B) CHH level; (C) CHG level;
[0067] Figure 5 Schematic diagram of q - PCR verification results. Detailed implementation manners
[0068] To further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following takes preferred embodiments to specifically describe in detail the specific implementation manners, technical solutions, features and their effects of the application according to the present invention. The specific features, structures, or characteristics in the following descriptions of multiple embodiments can be combined in any suitable form.
[0069] Example 1: Transcriptome analysis of Hu sheep in Hotan area, Xinjiang and Taihu Lake area
[0070] (1) Total blood RNA extraction, library construction and sequencing of Hu sheep
[0071] Total RNA of whole blood from Hu sheep in different regions was extracted by the Trizol method. The concentration and purity of the extracted total RNA were determined using a Nano Drop 2000 instrument, and the integrity RIN of the extracted total RNA was determined using an Agient 2100 instrument. RNA with RIN > 8.5 and 28 / 18S > 0.7 was selected for library construction. After digesting the total RNA with DNaseI, mRNA was aggregated using Oligo(dT) magnetic beads; the mRNA was fragmented; double-stranded cDNA was synthesized and purified using the mRNA as a template; end repair was performed, A bases were added, and sequencing adapters were ligated, and the fragment size was selected; the cDNA library was enriched by PCR; the concentration of the library and the inserted fragments were quality checked to ensure that the effective concentration of the library > 2 nM; PE150 sequencing was performed using the Illumina NovaSeq 6000 platform (Beijing Biomarker). The cDNA libraries of Hu sheep blood samples from different regions were sequenced using the Illumina high-throughput sequencing platform, generating a large amount of raw data (Raw Data), and after filtering to remove reads containing adapters and low-quality reads, high-quality clean data were obtained. Detection and analysis of the total RNA of whole blood from Hu sheep in different regions using a nucleic acid-protein detector showed that the OD260 / 280 value was 1.8 - 2.0; the results of 1% agarose gel electrophoresis showed three bands ( Figure 1 ), and the three bands from top to bottom were 28S rRNA, 18S rRNA, and 5S rRNA, indicating that the quality of RNA extraction from each sample was good.
[0072] (2) Extraction, library construction, and sequencing of Hu sheep whole blood DNA
[0073] Extract genomic DNA from the peripheral blood of Hu sheep in different regions according to the instructions of the blood sample DNA extraction kit. Use Nano Drop 2000 to detect the DNA purity (OD260 / 280 > 1.80), and use Qubit 2.0 to measure the DNA concentration. Use the EZDNAMethylation-Gold kit to perform bisulfite conversion on the DNA, and add Lambda phage DNA to evaluate the conversion efficiency. Fragment the genomic DNA to 300 - 500 bp with Covaris S220, repair the ends and ligate sequencing adapters, and generate a DNA library by PCR amplification. Quantify with Qubit 2.0, detect the insert fragment length with Agilent 2100, and finally detect the library with fluorescence quantitative PCR. After the samples pass the quality inspection and are processed according to the concentration and sequencing requirements, use the Hiseq2500 platform for paired-end sequencing (Beijing Biomarker Biotechnology Co., Ltd.). After obtaining the raw data, use the fastp software for quality control filtering to obtain Clean Data. After sequencing and quality inspection, finally obtain 82.43 G of Clean Data. The sequencing data volume of each sample is between 6.07 G and 7.83 G, Q30 is between 92.07% and 93.80%, Q20 is between 96.82% and 97.71%. The sequencing quality is good, the transcriptome coverage rate is high, and the sample coverage rate is above 92.98%.
[0074] (3) Transcriptome reference genome alignment and differential gene identification
[0075] Use the HISAT2 software to accurately align the Clean Reads with the sheep reference genome (ARS_UI_Ramb_v3.0) to obtain the positioning information of the Reads on the reference genome, such as the distribution of genes on chromosomes, the alignment rate, and gene structures, etc. Then use StringTie to assemble the aligned Reads, reconstruct the transcriptome for subsequent analysis. Normalize the number and transcript length of the Mapped Reads in the samples, and use StringTie through the maximum flow algorithm and normalize with FPKM as an indicator to measure the gene expression level. Based on the Count values of genes in each sample, use the DESeq2 differential analysis software to screen for differentially expressed genes. The screening criteria are: Fold Change ≥ 1.5 and Pvalue < 0.05. A total of 24,493 genes were detected in Hu sheep in the Taihu Lake region and Hu sheep in the Hotan region, and 107 genes were differentially expressed, including 6 up-regulated genes and 101 down-regulated genes. The information of up-regulated genes in Hu sheep in the Hotan region compared with Hu sheep in the Taihu Lake region is shown in Table 2.
[0076] Table 2 Expression abundance of differentially expressed genes with up-regulated expression levels
[0077]
[0078] Example 2: Genome-wide Methylation Analysis of Hu Sheep in Hotan Region of Xinjiang and Taihu Region
[0079] (1) Extraction and Quality Inspection of Sheep Genome DNA
[0080] Extract genomic DNA from the peripheral blood of Hu sheep in different regions according to the instructions of the blood sample DNA extraction kit. Use Nano Drop 2000 to detect the DNA purity (OD260 / 280 > 1.80), and use Qubit 2.0 to measure the DNA concentration. The genomic DNA bands are bright and there is no degradation ( Figure 2 ), and the quality of the genomic DNA is good, which can be used for the next step of bisulfite treatment, library construction and subsequent sequencing analysis.
[0081] (2) DNA Library Construction and Quality Inspection
[0082] Use the EZDNAMethylation-Gold kit to perform bisulfite conversion on the DNA, and add Lambda phage DNA to evaluate the conversion efficiency. Fragment the genomic DNA to 300 - 500 bp with Covaris S220, repair the ends and ligate the sequencing adapters, and generate a DNA library by PCR amplification. Quantify with Qubit 2.0, detect the insert fragment length with Agilent 2100, and finally detect the library with fluorescence quantitative PCR. The DNA extraction results of all 6 samples are qualified (Table 3).
[0083] Table 3 DNA Quality Detection Results
[0084]
[0085] (3) Sequencing on the Machine and Data Filtering
[0086] After the samples with qualified quality inspection are processed according to the concentration and sequencing requirements, use the Hiseq2500 platform for paired-end sequencing (Beijing Biomarker Technologies Co., Ltd.). After obtaining the raw data, use the fastp software for quality control filtering to obtain Clean Data. Methylation sequencing was performed on the peripheral blood of 6 Hu sheep from different regions. The conversion efficiency after bisulfite treatment exceeded 99%, indicating good treatment effect and meeting the sequencing requirements. After sequencing and quality inspection, 619.11 G of Clean bases data was obtained. The Clean Data percentage of all samples was above 79.91%, the average alignment efficiency with the reference genome was about 70.80%, and the Q30 value was between 96.63% and 97.37%, indicating good sequencing quality and being suitable for subsequent bioinformatics analysis. The genome coverage rate of all samples exceeded 98.71%
[0087] (4) Reference genome alignment and DMRs identification
[0088] The Clean Data was aligned with the sheep reference genome (ARS_UI_Ramb_v3.0) and bioinformatics analysis was performed using the Bsmap software. After sorting the aligned bam file, the sambamba software was used to remove redundancy. After obtaining the redundancy-removed alignment results, MethylDackel was used to extract methylation sites. Subsequently, the metilene software was used to identify differentially methylated regions (DMRs). The detailed methylation status of methylated C sites in different sequence contexts (CpG, CHH, CHG, where H represents C, A, or T bases) is shown in Table 4. Among them, the methylation proportion of CpG sites in the total C sites is the highest, with the percentage range of 6 samples being 69.47%-72.85%, while the proportion of methylated C sites in the remaining sequence contexts is small.
[0089] Table 4 Genome-wide methylation status of sheep peripheral blood
[0090]
[0091] (5) DNA methylation level and correlation analysis
[0092] The DNA methylation levels of various functional genomic elements were deeply analyzed. The promoter region was defined as the 2-kb range upstream of the transcription start site of each gene. Each bin represents the average methylation level of mC sites in different sequencing environment types in this region. The average methylation levels of gene functional elements in 6 samples are as Figure 3 shown. In all samples, the mC methylation level trends in the three sequence backgrounds are similar, and there is no significant difference among different functional elements. Among them, the methylation level in the 5'UTR region is the lowest, while the methylation levels in the intron, exon, and 3'UTR regions are the highest. In the promoter region, the methylation level at the distal end is the highest, followed by the middle region, and the DNA methylation level at the proximal end is the lowest. In the 5'UTR region, the methylation level at the distal region is the lowest, and the methylation level at the proximal region is the highest. In the CHH and CHG backgrounds, the methylation level in the middle region of exons is higher than that in the near 5'-end and 3'-end regions. In the CG background, the methylation levels of mC tend to be consistent. The correlation analysis results show that the methylation levels among samples within the same group are relatively close and the repeatability is good; while the methylation levels between different groups are significantly separated and the differences are large. This indicates that there are differences in the genome-wide methylation levels between the two regions ( Figure 4 ).
[0093] (6) Screening and analysis of DMRs
[0094] There are 5,745 regions with higher peripheral blood methylation levels in Hu sheep in Hotan region than in Taihu region, and these differential regions are distributed in 2,076 corresponding genes. There are 29,996 regions with lower peripheral blood methylation levels in Hu sheep in Hotan region than in Taihu region, and these differential regions are distributed in 5,247 corresponding genes. Among them, the methylation differences of genes such as ERCC4, PCCA, SUGP1, SMURF2, NLGN4X, and IMPACT are relatively large between the two groups.
[0095] Example 3: Integrated analysis of transcriptome and whole-genome methylation
[0096] Analysis of the methylation levels of genes related to methylation regions and the differential gene expression levels of their related transcriptomes found that there were a total of 7,732 genes with negative correlation between methylation levels and mRNA expression levels, and 10,881 genes with positive correlation between methylation levels and their mRNA expression levels. Among the negatively correlated genes, genes with significant differences in both methylation and transcriptome expression levels were selected, and a total of 118 genes were found. Among them, there were 13 genes with high methylation and down-regulated expression, and 11 genes with low methylation and up-regulated expression (Table 5). Among them, the gene LOC121817420 had low methylation and increased expression.
[0097] Table 5 Genes with negative correlation between methylation level and expression level
[0098]
[0099]
[0100] Example 4: Verification of gene LOC121817420 by real-time fluorescence quantitative PCR
[0101] qPCR primers were designed using Primerpremier 6, and the details of the primer sequences are shown in Table 6. Using the β-actin gene as an internal reference, the relative expression levels of each gene in the whole blood of sheep in different regions were calculated by the 2-△△CT method. The q-PCR verification results are as Figure 5 shown.
[0102] Among them, the reverse transcription system is as follows: 1 mg of total RNA, 2 μL of 10×RT buffer, 0.8 μL of 25×dNTP Mix, 2 μL of 10×RT Random Prime, 1 μL of MultiScnbe Reverse Transcriptase, supplemented with ddH2O to a total volume of 20 μL. The qRT-PCR reaction system is as follows: 10 μL of qPCR Master Mix, 1 μL of cDNA, 0.5 μL of each upstream and downstream primer, 0.1 μL of ROX, and 7.9 μL of ddH2O.
[0103] The amplification conditions are as follows: 95°C for 2 min; 95°C for 5 s, 60°C for 30 s, 72°C for 30 s, for 40 cycles; 72°C for 5 min.
[0104] Table 6 Primer sequences
[0105]
[0106] The mRNA level of the Hu sheep LOC121817420 gene in Hotan area was extremely significantly higher than that of Hu sheep in Taihu area.
[0107] For the parts not described in detail in the embodiments of the present invention, those skilled in the art can select them from the prior art.
[0108] The specific embodiments of the present invention disclosed above are only for illustration purposes, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the above-mentioned claims.
Claims
1. A method for screening genes affecting the drought adaptability of Hu sheep, characterized in that, The method comprises the following steps: Step 1. Selection and treatment of experimental sheep: 6 rams weighing 39-41 kg and in good growth condition were selected from Hotan and Taihu areas of Xinjiang, 6 mL of blood was collected from the jugular vein, and then the sheep were bled from the neck and slaughtered. Liver tissue samples were collected and quickly placed in cryopreservation tubes and placed in liquid nitrogen, and then brought back to the laboratory and placed in a -80°C ultra-low temperature refrigerator; Step 2, transcriptome analysis: extraction of RNA from Hu sheep whole blood, library construction and sequencing, extraction of DNA from Hu sheep whole blood, library construction and sequencing, transcriptome reference genome comparison and differential gene identification; Step 3: Whole-genome methylation analysis: sheep genomic DNA extraction and quality inspection, DNA library construction and quality inspection, machine sequencing and data filtering, reference genome comparison and DMRs identification, DNA methylation level and correlation analysis; Step 4: Joint analysis of transcriptome and whole genome methylation: Screen out genes that affect the drought adaptability of Hu sheep, jointly analyze gene expression and methylation levels, distribution of methylation levels in genes with different expression patterns, intersection enrichment of RNA differential genes and promoter genes in differentially methylated regions, and intersection enrichment of RNA differential genes and gene bodies in differentially methylated regions.
2. The gene screening method for influencing the drought adaptability of Hu sheep according to claim 1, wherein The method also includes step 5, using real-time fluorescence quantitative PCR to verify the genes that affect the drought adaptability of Hu sheep.
3. The gene screening method for affecting the drought adaptability of Hu sheep according to claim 1, characterized in that, Step 2 is as follows: 21) RNA extraction and detection Agarose gel electrophoresis: analyze the integrity of sample RNA and whether there is DNA contamination; 22) Library construction and quality control The starting RNA for library construction was total RNA, with a total amount of >= 800 ng. Oligo(dT) magnetic beads were used to enrich mRNA with polyA tails, and then divalent cations were used to randomly shear the obtained mRNA in NEB fragmentation buffer. The fragmented mRNA was used as a template and random oligonucleotides were used as primers to synthesize the first-chain cDNA in the M-MuLV reverse transcriptase system. The RNA chain was then degraded with RNase H, and the second-chain cDNA was synthesized with dNTPs as raw materials in the DNA polymerase I system. The purified double-stranded cDNA was end-repaired, A-tailed, and connected to a sequencing adapter. AMPure XP magnetic beads were used to screen a 200bp cDNA fragment. After PCR amplification, the PCR product was purified again using AMPure XP magnetic beads to finally obtain a sequencing library. After the library was constructed, the library was quality checked. 23) Sequencing After the library quality inspection is qualified, different libraries are mixed according to their effective concentration and target sequencing data volume requirements, and then sequenced using the Illumina NovaSeq X Plus sequencing platform to generate 150bp double-end sequencing read length; 24) Data quality control The image data of the sequencing fragments measured by the high-throughput sequencer were converted into sequence reads by CASAVA base recognition, and the raw data were filtered; 25) Sequence alignment to reference genome The index of the reference genome was constructed using HISAT2 (v2.2.1), and the paired-end clean reads were aligned to the reference genome using HISAT2 (v2.2.1). 26) Gene expression level quantification Feature Counts (v2.0.1) was used to calculate the reads mapped to each gene, and then the FPKM of each gene was calculated according to the gene length, and the reads mapped to the gene were calculated. 27) Differential expression analysis For those with biological replicates, DESeq2 software (1.16.1) was used for differential expression analysis between two comparison groups. For those without biological replicates, edgeR software was adopted. Before differential gene expression analysis, for each sequencing library, the read counts were adjusted through an edgeR package by a scaling normalization factor. The differential expression analysis between two conditions was performed using the edgeR R package (3.18.1). 28) Differential gene enrichment analysis The GO enrichment analysis of differentially expressed genes was implemented through the clusterProfiler (4.2.0) R software, and the GO terms with corrected P-values less than 0.05 were regarded as the GO terms significantly enriched by differentially expressed genes. 29) Protein network interaction analysis of differential genes The network interaction analysis of differentially expressed genes was based on the STRING database of known and predicted protein-protein interactions; for the species existing in the database, the network was constructed by extracting the target gene list from the database; diamond (v2.0.14.152) was used to align the target gene sequences with the selected reference protein sequences, and then the network was established according to the known interactions of the selected reference species.
4. The gene screening method for influencing the drought adaptability of Hu sheep according to claim 1, characterized in that, Step 3 is specifically as follows: 31) Extract genomic DNA For tissue samples, an appropriate amount of the sample was ground in liquid nitrogen or homogenized in a homogenizer; after adding 2% CTAB to the tissue, it was placed in a 65°C water bath for 1 h; phenol-chloroform-isoamyl alcohol was added for extraction, and after centrifugation, the upper aqueous phase was aspirated and the protein layer was discarded; isopropanol was added for precipitation for 10 min, and centrifuged at 12,000 g at 4°C for 10 min; the supernatant was discarded, washed twice with 75% ethanol, and the precipitate was dried; 50 μl of TE buffer was added to dissolve the precipitate, and incubated at 37°C for 30 min; for the dissolved DNA, a part was taken for quality inspection, and its OD value and integrity were measured. 32) DNA library construction The starting DNA amount for library construction > 1000 ng. At the same time, unmethylated λDNA used to evaluate transformation efficiency is added. After the intact DNA is fragmented by sonication, fragments of 100 - 500 bp are obtained. Subsequently, the DNA is treated with bisulfite to convert unmethylated cytosine to uracil. The treated DNA needs to be end-repaired, including phosphorylation modification at the 5' end and addition of a dA tail at the 3' end. The end-repaired fragments need to be ligated with adapter adaptors, and the DNA is purified using AM Pure XP magnetic beads. The adapter-ligated products after purification or size selection are subjected to PCR amplification. After amplification, the PCR products are purified again using AM Pure XP magnetic beads. Finally, a library is obtained. After the library construction is completed, first, a preliminary concentration measurement is performed using a Qubit 2.0 fluorescence quantifier. Subsequently, the size of the library insert fragments is detected by an Agilent 2100 bioanalyzer. When the insert fragments meet the expectations, quantitative real-time fluorescence PCR is used to determine the effective concentration of the library; 33) Sequencing on the machine After passing the library inspection, different libraries are mixed according to the requirements of the effective concentration and the target output data volume, and then sequenced on an Illumina NovaSeq X Plus; 34) Quality control and alignment with the reference genome The fastp software is used to perform quality control on the raw sequencing data, removing adapter sequences and low-quality bases. The bisulfite-treated sequences are aligned to the reference genome through the BSMAP software to calculate the bisulfite conversion rate; 35) Methylation detection Extract the methylation information at cytosine sites. Unmethylated cytosine is converted to thymine, while methylated cytosine remains unchanged. The metilene software is used to identify differentially methylated regions.
5. The gene screening method for influencing the drought adaptability of Hu sheep according to claim 1, characterized in that, Step 4 is specifically as follows: 41) Joint analysis of gene expression levels and methylation levels Data statistics: Statistics are performed on gene expression levels and methylation levels, generating a table to display the expression levels and methylation levels of each gene. A bivariate relationship graph of gene expression levels and methylation levels in the gene promoter region and gene body region is drawn to show the distribution and relationship of the two groups of variables. The pearson correlation coefficient is used to calculate the correlation between gene expression levels and methylation levels; 42) Methylation level distribution of genes with different expression patterns in the gene body Draw a graph showing the average methylation levels of genes with high-level expression, medium-level expression, low-level expression, and non-expression patterns at the gene body, transcription start site, and transcription termination site to display the methylation level distribution of genes with different expression patterns; 43) Enrichment of the intersection of RNA differential genes and promoters of differentially methylated regions Draw a Venn diagram of transcriptome differentially expressed genes and promoters of differentially methylated regions to show the common genes between the two. Perform GO functional enrichment analysis on the intersection of differential genes and promoters of differentially methylated regions, generating a GO enrichment result table, and draw a scatter plot to show significant functional enrichment items; perform KEGG pathway enrichment analysis on the intersection of differentially expressed genes and genes related to differentially methylated regions in the promoter region, generating a KEGG enrichment result table, and draw a scatter plot to show significantly enriched pathway entries.
6. The gene screening method for influencing the drought adaptability of Hu sheep according to claim 2, characterized in that Step 5 is specifically as follows: 51), Completely melt the template RNA, specific primers, 2×FastKing RT-qPCR Buffer (SYBR Green), 50×ROX Reference Dye and RNase-Free ddH2O, centrifuge and then place on ice bath; 52), Prepare the reaction solution under ice bath conditions; 53), Perform Real Time One Step RT-qPCR reaction For the PCR reaction tube, centrifuge it instantaneously with a centrifuge and then put it into a fluorescence quantitative PCR instrument for Real Time PCR reaction.
7. The gene affecting the drought adaptability of Hu sheep obtained by the gene screening method according to any one of claims 1-6.