Method and apparatus for identifying non-coding regions of convergent selection

By integrating population genome resequencing data and epigenetic markers, eliminating coding regions, and using protein-coding genes as anchors, cross-species conservation verification was achieved, solving the problem of inaccurate identification of non-coding regions in existing technologies, and improving the accuracy and reliability of convergent selection of non-coding regions.

CN121354660BActive Publication Date: 2026-03-31ZHEJIANG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies rely excessively on protein-coding regions when analyzing the genetic mechanisms of convergent evolution, neglecting the role of non-coding regulatory regions. Furthermore, they struggle to effectively integrate population genome resequencing data and epigenetic data, leading to inaccurate identification of selected regions in non-coding areas.

Method used

Genetic variation maps were constructed by integrating population genome resequencing data, coding regions were removed, population data were combined to reduce single genome errors, epigenetic markers were used to screen for differential regions, and protein-coding genes were used as anchors to identify conserved non-coding regions across species. A dual verification strategy of intra-species overlap and cross-species matching was adopted.

Benefits of technology

Precisely locating non-coding selected regions improves the accuracy and reliability of identifying convergent selection non-coding regions, overcomes the limitations of traditional reliance on coding regions, correlates with epigenetic regulatory information, and focuses on core non-coding regions involved in convergent evolution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121354660B_ABST
    Figure CN121354660B_ABST
Patent Text Reader

Abstract

The application provides a method and device for identifying convergent selection non-coding regions, comprising the following steps: obtaining a first non-coding selection region of a first species, and obtaining a second non-coding selection region of a second feature; obtaining a first species difference epigenetic marker region of the first species, and obtaining a second species difference epigenetic marker region of the second species; obtaining a cross-species conserved non-coding region of the first species and the second species; and taking a part of the first species and the second species located in the cross-species conserved non-coding region as a convergent selection non-coding region of the first species and the second species. The scheme breaks through the limitation of traditional dependence on coding regions by constructing a genetic variation map through integration of population genome resequencing data and eliminating coding regions, accurately locks the non-coding selection region, and simultaneously reduces accidental errors of a single genome by using population data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biotechnology, and in particular to a method and apparatus for identifying non-coding regions of convergent selection. Background Technology

[0002] Convergent evolution is an important phenomenon in the field of biological evolution, referring to the process by which different species independently evolve similar phenotypic traits under the same selective pressure. The molecular regulatory mechanisms of this phenomenon are diverse, including single nucleotide substitutions, parallel shifts in the evolutionary rates of protein-coding genes, and repetitive alterations of regulatory sequences. With the development of sequencing technology, high-quality reference genomes of various species have been assembled, providing crucial support for elucidating the genetic basis of convergent evolution. Current research often involves comparing protein-coding genes from different species across their entire genomes, and screening for key genes regulating convergent traits based on non-synonymous mutations in homologous genes.

[0003] However, existing research has significant technical limitations: on the one hand, it relies excessively on non-synonymous mutations in protein-coding regions, neglecting the regulatory role of mutations in non-coding regulatory regions, such as upstream and downstream gene regulatory regions and intron regions, on convergent traits. In fact, non-coding regions play a crucial role in phenotypic differentiation and adaptive evolution by regulating gene expression patterns, but there is currently a lack of systematic analytical methods for convergent selection in non-coding regions. On the other hand, existing methods mostly rely on the alignment of complete protein sequences from a single reference genome, making it difficult to utilize massive population genome resequencing data, unable to eliminate random errors from individual genomes, and failing to combine population genetic structure for precise sample classification, thus limiting the accuracy of selected regions identification.

[0004] With the widespread adoption of high-throughput sequencing technology, massive amounts of genome resequencing and epigenetic sequencing data have become commonplace in research. These data contain information on the genetic diversity and epigenetic regulation of species populations, making them important resources for elucidating convergent evolution mechanisms. However, there is currently a lack of technical solutions that can effectively integrate these two types of data, combining population genetic analysis with cross-species conservation verification. This makes it difficult to efficiently screen for conserved non-coding regions in different species that possess both "selection signals" and "epiggenetic differences," resulting in key technological gaps in elucidating the genetic mechanisms of convergent evolution. Summary of the Invention

[0005] This application provides a method and apparatus for identifying non-coding regions of convergent selection. By integrating population genome resequencing data to construct a genetic variation map and removing coding regions, it overcomes the limitations of traditional reliance on coding regions, accurately identifies non-coding selected regions, and reduces random errors from a single genome by utilizing population data.

[0006] In a first aspect, embodiments of this application provide a method for identifying non-coding regions through convergence selection, the method comprising:

[0007] Obtain a first target population, a first control population, and a first reference genome of a first species; obtain a second target population, a second control population, and a second reference genome of a second species, wherein the first target population exhibits convergent selection traits compared to the first control population, and the second target population exhibits convergent selection traits compared to the second control population.

[0008] Genetic variation map of a first species was obtained based on the genome resequencing data of a first target population and a first control population, and the genetic variation sites of a first reference genome. The non-coding selected regions in the genetic variation map of the first species were then identified as the first non-coding selected regions. Genetic variation map of a second species was obtained based on the genome resequencing data of a second target population and a second control population, and the genetic variation sites of a second reference genome. The non-coding selected regions in the genetic variation map of the second species were then identified as the second non-coding selected regions.

[0009] Epigenetic markers of the first target population and the first control population were obtained based on the first reference genome, and the regions of difference between the epigenetic markers of the first target population and the epigenetic markers of the first control population were taken as the first species-differential epigenetic marker region; epigenetic markers of the second target population and the second control population were obtained based on the second reference genome, and the regions of difference between the epigenetic markers of the second target population and the epigenetic markers of the second control population were taken as the second species-differential epigenetic marker region;

[0010] Using the corresponding protein-coding genes as anchors, non-coding regions with similar sequences in the first and second reference genomes were obtained as cross-species conserved non-coding regions.

[0011] The overlapping region between the non-coding selected region of the first species and the differential epigenetic marker region of the first species is obtained as the candidate non-coding region of the first species. The overlapping region between the non-coding selected region of the second species and the differential epigenetic marker region of the second species is obtained as the candidate non-coding region of the second species. The portion of the candidate non-coding region of the first species and the candidate non-coding region of the second species that is simultaneously located in the cross-species conserved non-coding region is taken as the non-coding region of convergent selection between the first species and the second species.

[0012] Secondly, embodiments of this application provide a convergence selection non-coding region identification device, comprising:

[0013] The acquisition unit is used to acquire a first target population, a first control population, and a first reference genome of a first species, and to acquire a second target population, a second control population, and a second reference genome of a second species, wherein the first target population has the trait of convergent selection compared to the first control population, and the second target population has the trait of convergent selection compared to the second control population.

[0014] The non-coding selected region acquisition unit acquires a genetic variation map of a first species based on the genome resequencing data of a first target population and a first control population, and the genetic variation sites of a first reference genome, and acquires the non-coding selected region in the genetic variation map of the first species as the first non-coding selected region; it acquires a genetic variation map of a second species based on the genome resequencing data of a second target population and a second control population, and the genetic variation sites of a second reference genome, and acquires the non-coding selected region in the genetic variation map of the second species as the second non-coding selected region;

[0015] The differential epigenetic marker region acquisition unit acquires epigenetic markers from a first target population and a first control population based on a first reference genome, and uses the regions of difference between the epigenetic markers of the first target population and the epigenetic markers of the first control population as the first species differential epigenetic marker region; and acquires epigenetic markers from a second target population and a second control population based on a second reference genome, and uses the regions of difference between the epigenetic markers of the second target population and the epigenetic markers of the second control population as the second species differential epigenetic marker region;

[0016] The cross-species conserved non-coding region acquisition unit uses the corresponding protein-coding gene as an anchor point to acquire non-coding regions with similar sequences in the first and second reference genomes as cross-species conserved non-coding regions.

[0017] The non-coding region acquisition unit of convergent selection acquires the overlapping region of the non-coding selected region of the first species and the differential epigenetic marker region of the first species as the candidate non-coding region of the first species, acquires the overlapping region of the non-coding selected region of the second species and the differential epigenetic marker region of the second species as the candidate non-coding region of the second species, and takes the portion of the candidate non-coding region of the first species and the candidate non-coding region of the second species that is simultaneously located in the cross-species conserved non-coding region as the non-coding region of convergent selection of the first species and the second species.

[0018] Thirdly, embodiments of this application provide an electronic device including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform a convergence selection non-coding region identification method.

[0019] Fourthly, embodiments of this application provide a readable storage medium storing a computer program, which, when executed by a processor, implements a convergence selection non-coding region identification method.

[0020] The main contributions and innovations of this invention are as follows:

[0021] This approach integrates population genome resequencing data to construct a genetic variation map and eliminates coding regions, overcoming the limitations of traditional coding-based methods by eliminating the need for whole-genome alignment of complete protein sequences. It precisely identifies non-coding regions selected for convergent evolution and reduces random errors from single genomes by utilizing population data. The approach combines epigenetic markers from the target and control populations to screen for differentially expressed regions, thereby linking epigenetic regulatory information and enhancing the correlation between candidate non-coding regions and convergent traits. This approach uses protein-coding genes as anchors to identify cross-species conserved non-coding regions, excluding species-specific regions through cross-species conservation verification, and focusing on the core non-coding regions truly involved in convergent evolution. This approach employs a dual verification strategy combining intra-species overlap screening and cross-species matching, narrowing the scope layer by layer and significantly improving the accuracy and reliability of identifying convergent-selective non-coding regions.

[0022] Details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects and advantages of this application more readily apparent. Attached Figure Description

[0023] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0024] Figure 1 This is a flowchart of a convergence selection non-coding region identification method according to an embodiment of this application;

[0025] Figure 2 This is a schematic diagram illustrating the acquisition of a first selected region and a second selected region according to an embodiment of this application;

[0026] Figure 3 This is a schematic diagram illustrating the acquisition of a first species-differential epigenetic marker region and a second species-differential epigenetic marker region according to an embodiment of this application;

[0027] Figure 4 This is a schematic diagram illustrating a non-coding region for obtaining convergence selection according to an example of this application;

[0028] Figure 5 This is a structural block diagram of a convergence selection non-coding region identification device according to an embodiment of this application;

[0029] Figure 6 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0030] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.

[0031] It should be noted that the steps of the corresponding methods are not necessarily performed in the order shown and described in this specification in other embodiments. In some other embodiments, the methods may include more or fewer steps than described in this specification. Furthermore, a single step described in this specification may be broken down into multiple steps in other embodiments; and multiple steps described in this specification may be combined into a single step in other embodiments.

[0032] Example 1

[0033] This application provides a method for identifying non-coding regions of convergent selection. By integrating population genome resequencing data to construct a genetic variation map and removing coding regions, it overcomes the limitations of traditional methods that rely on coding regions, accurately locating non-coding regions subject to selection. Simultaneously, it utilizes population data to reduce random errors from single genomes. Specifically, refer to... Figure 1 The method includes:

[0034] Obtain a first target population, a first control population, and a first reference genome of a first species; obtain a second target population, a second control population, and a second reference genome of a second species, wherein the first target population exhibits convergent selection traits compared to the first control population, and the second target population exhibits convergent selection traits compared to the second control population.

[0035] Genetic variation map of a first species was obtained based on the genome resequencing data of a first target population and a first control population, and the genetic variation sites of a first reference genome. The non-coding selected regions in the genetic variation map of the first species were then identified as the first non-coding selected regions. Genetic variation map of a second species was obtained based on the genome resequencing data of a second target population and a second control population, and the genetic variation sites of a second reference genome. The non-coding selected regions in the genetic variation map of the second species were then identified as the second non-coding selected regions.

[0036] Epigenetic markers of the first target population and the first control population were obtained based on the first reference genome, and the regions of difference between the epigenetic markers of the first target population and the epigenetic markers of the first control population were taken as the first species-differential epigenetic marker region; epigenetic markers of the second target population and the second control population were obtained based on the second reference genome, and the regions of difference between the epigenetic markers of the second target population and the epigenetic markers of the second control population were taken as the second species-differential epigenetic marker region;

[0037] Using the corresponding protein-coding genes as anchors, non-coding regions with similar sequences in the first and second reference genomes were obtained as cross-species conserved non-coding regions.

[0038] The overlapping region between the non-coding selected region of the first species and the differential epigenetic marker region of the first species is obtained as the candidate non-coding region of the first species. The overlapping region between the non-coding selected region of the second species and the differential epigenetic marker region of the second species is obtained as the candidate non-coding region of the second species. The portion of the candidate non-coding region of the first species and the candidate non-coding region of the second species that is simultaneously located in the cross-species conserved non-coding region is taken as the non-coding region of convergent selection between the first species and the second species.

[0039] In the current embodiment, a first target population, a first control population, and a first reference genome of a first species, and a second target population, a second control population, and a second reference genome of a second species are obtained from a public database.

[0040] For example, if the first species in this scheme is maize and the second species is rice, then the domesticated maize population is obtained from the public database NCBI or NGDC as the first target population, the wild maize population is obtained as the first control population, the maize B73 genome is obtained as the first reference genome, the domesticated rice population is obtained as the second target population, the wild rice population is obtained as the second control population, and the rice Nipponbare genome is obtained as the second reference genome.

[0041] Specifically, the convergent traits between the first and second target groups refer to the fact that although the domesticated maize population and the domesticated rice population are different species, they have evolved the same traits, such as increased yield, due to artificial or natural factors.

[0042] Specifically, taking the convergent trait of increased yield as an example, the convergent selection trait of the first target group compared to the first control group means that the yield of the first target group is higher than that of the first control group; similarly, the convergent selection trait of the second target group compared to the second control group means that the yield of the second target group is higher than that of the second control group.

[0043] For example, taking maize as the first species and rice as the second species, the genome resequencing data of the first target population, the first control population, the second target population, and the second control population are shown in Table 1:

[0044] Table 1. Genome resequencing data of the first target population, the first control population, the second target population, and the second control population.

[0045]

[0046] Specifically, genome resequencing data from different populations can be obtained directly from public databases, or whole-genome resequencing data can be obtained by providing their own data samples and using high-throughput sequencing platforms such as Illumina, PacBio, Nanopore, or BGI.

[0047] Specifically, after obtaining the genome resequencing data of the first target population and the first control population, quality control software is used to perform quality control on the genome resequencing data of the first target population and the first control population, thereby removing bases and adapter sequences with low sequencing quality.

[0048] For example, FastQC is used to perform quality checks on genome resequencing data from different populations, and Trimmomatics is used to remove bases and adapter sequences with low sequencing quality to complete quality control.

[0049] Specifically, by performing quality control on all genome resequencing data, poor-quality and redundant adapter sequences in the sequencing are removed, thereby retaining usable high-quality data.

[0050] In the current embodiment, genetic variation sites of each individual in the first target population and the first reference genome are identified based on genome resequencing data of the first target population. The genetic variation sites of all individuals in the first target population are merged to obtain a genetic variation map of the first target population. Similarly, genetic variation sites of each individual in the first control population and the first reference genome are identified based on genome resequencing data of the first control population. The genetic variation sites of all individuals in the first control population are merged to obtain a genetic variation map of the first control population. Finally, the genetic variation maps of the first target population and the first control population are integrated to obtain the genetic variation map of the first species. Genetic variation map: Based on the genome resequencing data of the second target population, the genetic variation sites of each individual in the second target population and the second reference genome are identified. The genetic variation sites of all individuals in the second target population are merged to obtain the genetic variation map of the second target population. Based on the genome resequencing data of the second control population, the genetic variation sites of each individual in the second control population and the second reference genome are identified. The genetic variation sites of all individuals in the second control population are merged to obtain the genetic variation map of the second control population. The genetic variation map of the second target population and the genetic variation map of the second control population are integrated to obtain the genetic variation map of the second species.

[0051] Specifically, taking the acquisition of the genetic variation map of the first target population as an example, the acquisition of the genetic variation maps of the other groups—the first control group, the second target group, and the second control group—is the same as that of the first target group. The steps for acquiring the genetic variation map of the first target group include:

[0052] Step 1: Use BWA to map the genome resequencing data of each individual in the first target population to the first reference genome, and use sambamba to remove duplicate mapping results;

[0053] Step 2: Use GATK4 to identify the genetic variation sites in each corresponding result obtained in Step 1. In this scheme, the genetic variation sites are single nucleotide polymorphisms (SNPs).

[0054] Step 3: Merge all the genetic variation sites from Step 2, and use BEAGLE to fill in the gaps in the merged results to obtain the genetic variation map of the first target population.

[0055] Specifically, in step 1, the genome resequencing data of each individual in the first target population consists of incomplete fragment data. Therefore, BWA is used to map these fragment data one by one to the complete sequence of the first reference genome. In addition, since there may be two individuals with the same fragment, they will be mapped repeatedly to the first reference genome. In order to reduce such duplicate data, Sambamba is used to remove the same duplicate mapping results.

[0056] Specifically, in step 2, single nucleotide polymorphisms represent single base differences in the same sequence among different individuals. GATK4 is used to identify single base differences between each individual and the first reference genome to determine the genetic variation site.

[0057] Furthermore, in the step of integrating the genetic variation maps of the first target population and the first control population to obtain the genetic variation map of the first species by integrating the genetic variation maps of the first target population and the first control population to obtain the first original genetic variation map, the first original genetic variation map is filled with genotypes and genotypes with a genotype missing rate greater than a set threshold are removed to obtain the genetic variation map of the first species. Similarly, in the step of integrating the genetic variation maps of the second target population and the first control population to obtain the genetic variation map of the second species by integrating the genetic variation maps of the second target population and the second control population to obtain the second original genetic variation map, the second original genetic variation map is filled with genotypes and genotypes with a genotype missing rate greater than a set threshold are removed to obtain the genetic variation map of the second species.

[0058] Specifically, this scheme uses BEAGLE software to fill in missing genotypes. During the genotype filling process, the missing genotypes can be inferred based on the same gene locations in other individuals in the population.

[0059] Specifically, the genotype deletion rate indicates the completeness of the genotype. If the genotype deletion rate is too high, it means that the reliability of the corresponding genotype is insufficient, so genotypes with a genotype deletion rate greater than a set threshold are removed.

[0060] In the current embodiment, the linkage degree of adjacent genetic variation sites in the genetic variation map of the first species is calculated. If the linkage degree of adjacent genetic variation sites is greater than a set threshold, any one of the adjacent genetic variation sites is removed. The linkage degree of adjacent genetic variation sites in the genetic variation map of the second species is calculated. If the linkage degree of adjacent genetic variation sites is greater than a set threshold, any one of the adjacent genetic variation sites is removed.

[0061] Specifically, PLINK is used to calculate the linkage degree between adjacent genetic variation sites, with a threshold of 0.2 within 50 bp. That is, if the linkage degree between two adjacent genetic variation sites within 50 bp is greater than the set threshold, then any one of the adjacent genetic variation sites is removed.

[0062] Specifically, in gene sequences, genetic variation sites that are too close together have highly repetitive information, which can interfere with the analysis. Therefore, PLINK is used to remove genetic variation sites that are too close together.

[0063] In the current embodiment, phylogenetic analysis, population structure analysis, and principal component analysis are performed on each individual in the genetic variation map of the first species, and individuals with consistent classification results in phylogenetic analysis, population structure analysis, and principal component analysis are retained; phylogenetic analysis, population structure analysis, and principal component analysis are performed on each individual in the genetic variation map of the second species, and individuals with consistent classification results in phylogenetic analysis, population structure analysis, and principal component analysis are retained.

[0064] Specifically, for individuals in the genetic variation map of the first species, it is necessary to ensure whether the individual belongs to the first target population or the first control population. Therefore, only individuals whose classification results in phylogenetic analysis, population structure analysis, and principal component analysis are all from the first target population / second target population are retained to avoid errors in subsequent calculations. The same applies to the genetic variation map of the second species.

[0065] In the current embodiment, a first selected region is obtained based on the genetic variation map of a first species, and a non-coding selected region of the first species is obtained by removing the coding region from the selected region of the first species based on a first reference genome; a second selected region is obtained based on the genetic variation map of a second species, and a non-coding selected region of the second species is obtained by removing the coding region from the selected region of the second species based on a second reference genome.

[0066] Specifically, the first selected region and the second selected region are as follows: Figure 2As shown, since the genetic variation map of the first species records the whole genome sequence of the first species containing all genetic variation sites, the population to which each individual belongs, and the genotype of each individual at each genetic variation site; and the genetic variation map of the second species records the whole genome sequence of the second species containing all genetic variation sites, the population to which each individual belongs, and the genotype of each individual at each genetic variation site, the genetic variation map of the first species is input into XP-CLR and GenWin software to obtain the selection intensity of each set interval, and a set number of intervals with selection intensity from high to low are selected as the first selected region. Similarly, the genetic variation map of the second species is input into XP-CLR and GenWin software to obtain the selection intensity of each set interval, and a set number of intervals with selection intensity from high to low are selected as the second selected region.

[0067] In other words, for the first selected region, the whole genome sequence of the first species is divided into multiple first intervals of a preset interval size. The selection intensity of the corresponding first interval is calculated based on the allele frequency of each genetic variation site in the first target population and the allele frequency of each genetic variation site in the first control population. Similarly, for the second selected region, the whole genome sequence of the second species is divided into multiple second intervals of a preset interval size. The selection intensity of the corresponding second interval is calculated based on the allele frequency of each genetic variation site in the second interval in the second target population and the allele frequency of each genetic variation site in the second control population.

[0068] For example, XP-CLR was used to calculate the degree of allele frequency differentiation at loci in maize and rice populations composed of genetically well-defined materials. Specifically, regions in both maize and rice populations with XP-CLR values ​​in the top 10% of the whole genome, and where genetic diversity in wild species was significantly higher than in domesticated species, were considered as high-confidence selection regions for maize and rice, respectively.

[0069] Further, the first gene region in the first reference genome is obtained, and the part overlapping with the gene region in the first selected region is removed to obtain the non-coding selected region of the first species; the second gene region in the second reference genome is obtained, and the part overlapping with the gene region in the second selected region is removed to obtain the non-coding selected region of the second species.

[0070] Specifically, the first gene region is obtained based on the annotation information of the first reference genome. Since the gene region includes coding and non-coding regions, the non-coding selected region of the first species is obtained by removing the part of the first selected region that overlaps with the gene region.

[0071] For example, the substract module in bedtools is used to remove the portion of the first selected region that overlaps with the gene region, and the same applies to the acquisition of the non-coding selected region of the second species.

[0072] In the current embodiment, epigenetic sequencing data of a first target population and a first control population are obtained respectively. The epigenetic sequencing data of the first target population is compared with a first reference genome to obtain epigenetic markers of the first target population, and the epigenetic sequencing data of the first control population is compared with the first reference genome to obtain epigenetic markers of the first control population. Epigenetic sequencing data of a second target population and a second control population are obtained respectively. The epigenetic sequencing data of the second target population is compared with a second reference genome to obtain epigenetic markers of the second target population, and the epigenetic sequencing data of the second control population is compared with a second reference genome to obtain epigenetic markers of the second control population.

[0073] Specifically, after obtaining the epigenetic sequencing data of the first target population and the first control population, quality control software is used to perform quality control on the epigenetic sequencing data to remove low-quality bases and adapter sequences. Similarly, the epigenetic sequencing data of the second target population and the second control population are also subjected to quality control. For example, FastQC and Fastp are used for quality control.

[0074] For example, whole-genome bisulfite sequencing data of domesticated maize was obtained from public databases as the epigenetic sequencing data of the first target population, and whole-genome bisulfite sequencing data of wild maize was obtained as the epigenetic sequencing data of the first control population; whole-genome bisulfite sequencing data of domesticated rice was obtained from public databases as the epigenetic sequencing data of the second target population, and whole-genome bisulfite sequencing data of wild rice was obtained as the epigenetic sequencing data of the second control population, as shown in Table 2:

[0075] Table 2. Epigenetic sequencing data of the first target population, the first control population, the second target population, and the second control population.

[0076]

[0077] For example, epigenetic sequencing data of wild maize, domesticated maize, wild rice, and domesticated rice were aligned to their respective reference genomes using bismark. The deduplicate_bismark module of bismark was used to identify repetitive sequences, and the methylation level of individual sites was identified using bismark_methylation_extractor. The unionbedg module in bedtools was used to merge the CG, CHG, and CHH methylation sites of individual individuals, thereby obtaining the epigenetic markers of the first target population, the first control population, the second target population, and the second control population.

[0078] In this scheme, the differential methylation intervals between the epigenetic markers of the first target population and the epigenetic markers of the first control population are calculated as the first species-differential epigenetic marker region; the differential methylation intervals between the epigenetic markers of the second target population and the epigenetic markers of the second control population are calculated as the second species-differential epigenetic marker region.

[0079] For example, using metilene to calculate differentially methylated regions, the calculated differential epigenetic marker regions of the first species and the differential epigenetic marker regions of the second species are as follows: Figure 3 As shown.

[0080] In the current embodiment, the corresponding coding protein gene pairs in the first reference genome and the second reference genome are obtained, and the similarity of the sequence corresponding to each protein gene pair is calculated using a dynamic programming algorithm. Sequences with similarity greater than a set threshold are selected as cross-species conserved non-coding regions.

[0081] Specifically, for each protein-coding gene, its corresponding sequence is the sequence of the upstream and downstream segments of the protein-coding gene and the sequence of the intron region. That is, the corresponding sequence of the protein-coding gene is the sequence of the corresponding promoter region, terminator region and intron region.

[0082] For example, for the B73 and Nipponbare reference genomes, KAT was used to calculate the k-mers of the genomes. When the k-mer length was set to 20 bp, non-redundant genome k-mer sequence information was generated. Based on the annotation information of B73, minimap2 was used to align the CDS sequence of B73 to the rice reference genome. Finally, dCNS was used to obtain conserved collinear regions of maize and rice as cross-species conserved non-coding regions.

[0083] In the current embodiment, intra-species screening is first performed by obtaining the overlapping region between the non-coding selected region and the differential epigenetic marker region of the first species. This allows for the identification of non-coding regions associated with different traits in the first species as candidate non-coding regions for the first species. Similarly, non-coding regions associated with different traits in the second species are identified as candidate non-coding regions for the second species. Then, cross-species matching is performed to identify the portion where the candidate non-coding regions of the first and second species simultaneously lie in a conserved collinear region. This portion represents the convergent trait. In other words, the portion of the candidate non-coding regions of the first and second species simultaneously located in a cross-species conserved non-coding region is identified as the convergent selection non-coding region for the first and second species. A schematic diagram of obtaining the convergent selection non-coding region is shown below. Figure 4 As shown.

[0084] For example, the intersect module in bedtools is used to filter non-coding regions for convergent selection.

[0085] In the current embodiment, after obtaining the non-coding region of convergent selection, targets are designed on maize and rice using technologies such as CRISPR and CRISPRi, and phenotypic surveys of transgenic materials are conducted over many years at multiple locations to verify the function of the non-coding region of convergent selection.

[0086] Example 2

[0087] Based on the same concept, referencing Figure 5 This application also proposes a convergence selection non-coding region identification device, comprising:

[0088] The acquisition unit is used to acquire a first target population, a first control population, and a first reference genome of a first species, and to acquire a second target population, a second control population, and a second reference genome of a second species, wherein the first target population and the second target population have convergent traits, the first control population does not have convergent traits of the first target population, and the second control population does not have convergent traits of the second target population.

[0089] The non-coding selected region acquisition unit acquires a genetic variation map of a first species based on the genome resequencing data of a first target population and a first control population, and the genetic variation sites of a first reference genome, and acquires the non-coding selected region in the genetic variation map of the first species as the first non-coding selected region; it acquires a genetic variation map of a second species based on the genome resequencing data of a second target population and a second control population, and the genetic variation sites of a second reference genome, and acquires the non-coding selected region in the genetic variation map of the second species as the second non-coding selected region;

[0090] The differential epigenetic marker region acquisition unit acquires epigenetic markers from a first target population and a first control population based on a first reference genome, and uses the regions of difference between the epigenetic markers of the first target population and the epigenetic markers of the first control population as the first species differential epigenetic marker region; and acquires epigenetic markers from a second target population and a second control population based on a second reference genome, and uses the regions of difference between the epigenetic markers of the second target population and the epigenetic markers of the second control population as the second species differential epigenetic marker region;

[0091] The cross-species conserved non-coding region acquisition unit uses the corresponding protein-coding gene as an anchor point to acquire the non-coding region with the same sequence in the first reference genome and the second reference genome as the cross-species conserved non-coding region.

[0092] The non-coding region acquisition unit of convergent selection acquires the overlapping region of the non-coding selected region of the first species and the differential epigenetic marker region of the first species as the candidate non-coding region of the first species, acquires the overlapping region of the non-coding selected region of the second species and the differential epigenetic marker region of the second species as the candidate non-coding region of the second species, and takes the portion of the candidate non-coding region of the first species and the candidate non-coding region of the second species that is simultaneously located in the cross-species conserved non-coding region as the non-coding region of convergent selection of the first species and the second species.

[0093] Example 3

[0094] This embodiment also provides an electronic device, see reference. Figure 6 It includes a memory 404 and a processor 402, the memory 404 storing a computer program and the processor 402 being configured to run the computer program to perform the steps in any of the above method embodiments.

[0095] Specifically, the processor 402 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.

[0096] Memory 404 may include a mass storage device for data or instructions. For example, and not limitingly, memory 404 may include a hard disk drive (HDD), a floppy disk drive, a solid-state drive (SSD), flash memory, an optical disk drive, a magneto-optical disk drive, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 404 may include removable or non-removable (or fixed) media. Where appropriate, memory 404 may be internal or external to a data processing device. In a particular embodiment, memory 404 is non-volatile memory. In a particular embodiment, memory 404 includes read-only memory (ROM) and random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable read-only memory (PROM), an erasable read-only memory (EPROM), an electrically erasable read-only memory (EEPROM), an electrically alterable read-only memory (EAROM), or flash memory, or a combination of two or more of these. Where appropriate, the RAM can be Static Random-Access Memory (SRAM) or Dynamic Random-Access Memory (DRAM). DRAM can be Fast Page Mode Dynamic Random-Access Memory (FPMDRAM), Extended Data Out Dynamic Random-Access Memory (EDODRAM), Synchronous Dynamic Random-Access Memory (SDRAM), etc.

[0097] The memory 404 can be used to store or cache various data files that need to be processed and / or communicated, as well as possible computer program instructions executed by the processor 402.

[0098] The processor 402 reads and executes computer program instructions stored in the memory 404 to implement any of the convergence selection non-coding region identification methods in the above embodiments.

[0099] Optionally, the electronic device may further include a transmission device 406 and an input / output device 408, wherein the transmission device 406 is connected to the processor 402, and the input / output device 408 is connected to the processor 402.

[0100] The transmission device 406 can be used to receive or send data via a network. Specific examples of the network described above may include wired or wireless networks provided by the communication provider of the electronic device. In one example, the transmission device includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 406 may be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0101] Input / output device 408 is used to input or output information. In this embodiment, the input information may be a first reference genome, a second reference genome, etc., and the output information may be a convergent selection non-coding region, etc.

[0102] Optionally, in this embodiment, the processor 402 can be configured to perform the following steps via a computer program:

[0103] Obtain the first target population, the first control population, and the first reference genome of the first species; obtain the second target population, the second control population, and the second reference genome of the second species. The first target population and the second target population have convergent traits, the first control population does not have the convergent traits of the first target population, and the second control population does not have the convergent traits of the second target population.

[0104] Genetic variation map of a first species was obtained based on the genome resequencing data of a first target population and a first control population, and the genetic variation sites of a first reference genome. The non-coding selected regions in the genetic variation map of the first species were then identified as the first non-coding selected regions. Genetic variation map of a second species was obtained based on the genome resequencing data of a second target population and a second control population, and the genetic variation sites of a second reference genome. The non-coding selected regions in the genetic variation map of the second species were then identified as the second non-coding selected regions.

[0105] Epigenetic markers of the first target population and the first control population were obtained based on the first reference genome, and the regions of difference between the epigenetic markers of the first target population and the epigenetic markers of the first control population were taken as the first species-differential epigenetic marker region; epigenetic markers of the second target population and the second control population were obtained based on the second reference genome, and the regions of difference between the epigenetic markers of the second target population and the epigenetic markers of the second control population were taken as the second species-differential epigenetic marker region;

[0106] Using the corresponding protein-coding genes as anchors, non-coding regions with identical sequences in the first and second reference genomes are obtained as cross-species conserved non-coding regions.

[0107] The overlapping region between the non-coding selected region of the first species and the differential epigenetic marker region of the first species is obtained as the candidate non-coding region of the first species. The overlapping region between the non-coding selected region of the second species and the differential epigenetic marker region of the second species is obtained as the candidate non-coding region of the second species. The portion of the candidate non-coding region of the first species and the candidate non-coding region of the second species that is simultaneously located in the cross-species conserved non-coding region is taken as the non-coding region of convergent selection between the first species and the second species.

[0108] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated here.

[0109] Generally, various embodiments can be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. Some aspects of the invention can be implemented in hardware, while others can be implemented by firmware or software executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. Although various aspects of the invention may be shown and described as block diagrams, flowcharts, or using some other graphical representation, it should be understood that, by way of non-limiting example, these blocks, apparatuses, systems, techniques, or methods described herein can be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.

[0110] Embodiments of the present invention can be implemented by computer software, which may be executable by a data processor of a mobile device, such as a processor entity, or by hardware, or by a combination of software and hardware. Computer software or programs (also referred to as program products) including software routines, applets, and / or macros can be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. The computer program product may include one or more computer-executable components configured to perform the embodiments when the program is run. The one or more computer-executable components may be at least one piece of software code or a portion thereof. Additionally, it should be noted in this respect that, as Figure 6 Any box in the logical flow can represent a program step, or interconnected logic circuits, boxes and functions, or a combination of program steps and logic circuits, boxes and functions. Software can be stored on physical media such as memory chips or blocks of storage implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as DVDs and their data variants, CDs, etc. The physical medium is a non-transient medium.

[0111] Those skilled in the art should understand that the technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments have been described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0112] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for identifying non-coding regions through convergence selection, characterized in that, The method comprises the following steps: obtaining a first target population, a first control population and a first reference genome of a first species, and obtaining a second target population, a second control population and a second reference genome of a second species, wherein the first target population has a convergent selection trait compared with the first control population, and the second target population has a convergent selection trait compared with the second control population; obtaining a first species genetic variation map based on the genetic variation sites of the first reference genome and the genome resequencing data of the first target population and the first control population, and obtaining a non-coding selection region in the first species genetic variation map as a first non-coding selection region; and obtaining a second species genetic variation map based on the genetic variation sites of the second reference genome and the genome resequencing data of the second target population and the second control population, and obtaining a non-coding selection region in the second species genetic variation map as a second non-coding selection region; obtaining epigenetic markers of the first target population and the first control population based on the first reference genome, and taking a difference region of the epigenetic markers of the first target population and the first control population as a first species difference epigenetic marker region; obtaining epigenetic markers of the second target population and the second control population based on the second reference genome, and taking a difference region of the epigenetic markers of the second target population and the second control population as a second species difference epigenetic marker region, wherein the epigenetic sequencing data of the first target population and the first control population are obtained respectively, the epigenetic sequencing data of the first target population is aligned with the first reference genome to obtain the epigenetic markers of the first target population, and the epigenetic sequencing data of the first control population is aligned with the first reference genome to obtain the epigenetic markers of the first control population; the epigenetic sequencing data of the second target population and the second control population are obtained respectively, the epigenetic sequencing data of the second target population is aligned with the second reference genome to obtain the epigenetic markers of the second target population, and the epigenetic sequencing data of the second control population is aligned with the second reference genome to obtain the epigenetic markers of the second control population; taking a non-coding region with a sequence similarity greater than a set threshold in the first reference genome and the second reference genome as a cross-species conserved non-coding region, with the corresponding coding protein gene as an anchor point; taking an overlapping region of the first species non-coding selection region and the first species difference epigenetic marker region as a first species candidate non-coding region, and taking an overlapping region of the second species non-coding selection region and the second species difference epigenetic marker region as a second species candidate non-coding region, and taking a part of the first species candidate non-coding region and the second species candidate non-coding region simultaneously located in the cross-species conserved non-coding region as a convergent selection non-coding region of the first species and the second species.

2. The method of claim 1, wherein the non-coding region is identified by convergent selection. The genetic variation sites of each individual in the first target group and the first reference genome are identified based on the genome resequencing data of the first target group, the genetic variation sites of all individuals in the first target group are combined to obtain a genetic variation map of the first target group, the genetic variation sites of each individual in the first control group and the first reference genome are identified based on the genome resequencing data of the first control group, the genetic variation sites of all individuals in the first control group are combined to obtain a genetic variation map of the first control group, and the genetic variation map of the first target group and the genetic variation map of the first control group are integrated to obtain a first-species genetic variation map; the genetic variation sites of each individual in the second target group and the second reference genome are identified based on the genome resequencing data of the second target group, the genetic variation sites of all individuals in the second target group are combined to obtain a genetic variation map of the second target group, the genetic variation sites of each individual in the second control group and the second reference genome are identified based on the genome resequencing data of the second control group, the genetic variation sites of all individuals in the second control group are combined to obtain a genetic variation map of the second control group, and the genetic variation map of the second target group and the genetic variation map of the second control group are integrated to obtain a second-species genetic variation map.

3. The method of claim 2, wherein the non-coding region is identified by convergent selection. The genetic variation map of the first target group and the genetic variation map of the first control group are integrated to obtain a first original genetic variation map, genotype filling is performed on the first original genetic variation map, and genotypes with a genotype missing rate greater than a set threshold value are removed to obtain a first-species genetic variation map; the genetic variation map of the second target group and the genetic variation map of the second control group are integrated to obtain a second original genetic variation map, genotype filling is performed on the second original genetic variation map, and genotypes with a genotype missing rate greater than a set threshold value are removed to obtain a second-species genetic variation map.

4. The method of claim 1, wherein the non-coding region is selected by convergent selection. A first selected region is obtained based on the first-species genetic variation map, and a coding region in the first-species selected region is removed based on the first reference genome to obtain a first-species non-coding selected region. A second selected region is obtained based on the second-species genetic variation map, and a coding region in the second-species selected region is removed based on the second reference genome to obtain a second-species non-coding selected region.

5. The method of claim 4, wherein the non-coding region is identified by the step of, A first gene region in the first reference genome is obtained, and a part overlapping with the gene region in the first selected region is removed to obtain a first-species non-coding selected region. A second gene region in the second reference genome is obtained, and a part overlapping with the gene region in the second selected region is removed to obtain a second-species non-coding selected region.

6. The method of claim 1, wherein the non-coding region is identified by convergent selection. Protein gene pairs corresponding to the first reference genome and the second reference genome are obtained, a dynamic programming algorithm is used to calculate the similarity of the sequences corresponding to each protein gene pair, and sequences with a similarity greater than a set threshold value are selected as cross-species conserved non-coding regions.

7. An apparatus for identifying a non-coding region of convergent selection, comprising: Comprise: The acquisition unit is used for acquiring a first target group, a first control group and a first reference genome of a first species, and acquiring a second target group, a second control group and a second reference genome of a second species, wherein the first target group has a convergent selection trait compared with the first control group, and the second target group has a convergent selection trait compared with the second control group; The non-coding region under selection area acquisition unit is used for acquiring a genetic variation map of the first species based on the genomic resequencing data of the first target group and the first control group and genetic variation sites of the first reference genome, and acquiring a non-coding region under selection in the genetic variation map of the first species as a first non-coding region under selection; and acquiring a genetic variation map of the second species based on the genomic resequencing data of the second target group and the second control group and genetic variation sites of the second reference genome, and acquiring a non-coding region under selection in the genetic variation map of the second species as a second non-coding region under selection; The differential epigenetic marker area acquisition unit is used for acquiring epigenetic markers of the first target group and the first control group based on the first reference genome, and acquiring a differential area of the epigenetic markers of the first target group and the first control group as a differential epigenetic marker area of the first species; and acquiring epigenetic markers of the second target group and the second control group based on the second reference genome, and acquiring a differential area of the epigenetic markers of the second target group and the second control group as a differential epigenetic marker area of the second species, wherein the epigenetic sequencing data of the first target group and the first control group are respectively acquired, the epigenetic sequencing data of the first target group is compared with the first reference genome to acquire the epigenetic markers of the first target group, and the epigenetic sequencing data of the first control group is compared with the first reference genome to acquire the epigenetic markers of the first control group; and the epigenetic sequencing data of the second target group and the second control group are respectively acquired, the epigenetic sequencing data of the second target group is compared with the second reference genome to acquire the epigenetic markers of the second target group, and the epigenetic sequencing data of the second control group is compared with the second reference genome to acquire the epigenetic markers of the second control group; The cross-species conserved non-coding region acquisition unit is used for taking the non-coding regions with a sequence similarity greater than a set threshold in the first reference genome and the second reference genome as cross-species conserved non-coding regions with the corresponding coding protein genes as anchors; The convergent selection non-coding region acquisition unit is used for acquiring an overlapping area of the non-coding region under selection of the first species and the differential epigenetic marker area of the first species as a first species candidate non-coding region, acquiring an overlapping area of the non-coding region under selection of the second species and the differential epigenetic marker area of the second species as a second species candidate non-coding region, and taking a part of the first species candidate non-coding region and the second species candidate non-coding region simultaneously located in the cross-species conserved non-coding region as a convergent selection non-coding region of the first species and the second species. 8.An electronic device comprising a memory and a processor, the electronic device comprising: The memory stores a computer program, and the processor is configured to run the computer program to perform the method for identifying a convergent selection non-coding region according to any one of claims 1-6.

9. A readable storage medium, characterized by, The readable storage medium stores a computer program, and the computer program is executed by the processor to realize the method for identifying the non-coding region of the convergent selection according to any one of claims 1-6.

Citation Information

Patent Citations

  • Method and device for identifying conservative non-coding elements (CNEs) in genome

    CN120781143A