A method for locating genomic reproductive regions

By leveraging published plant genomes and annotation files, the method efficiently identifies genomic reproductive regions through sequence and gene expression analysis, addressing the complexity and time constraints of traditional methods.

CN118824375BActive Publication Date: 2025-07-15CHINA AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410909361.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-08
Publication Date
2025-07-15
Estimated Expiration
2044-07-08

AI Technical Summary

Technical Problem

Traditional genomic reproductive region localization methods are complex and time-consuming, requiring the construction of large-scale isolation of groups and conducting large-scale sequencing work, which consumes a lot of time and resources.

Method used

By comparing the genome sequence similarity, repeat sequence annotation and gene density of target species and related species, combining the gene expression data of flower tissue, reproductive-related regions were screened out, and the published plant genome and annotation files were used to achieve rapid localization.

Benefits of technology

The reproductive regions of target species are achieved efficiently and quickly positioned within a day, and results can be obtained within a few weeks if the genome is complex, simplifying traditional methods and promoting plant reproductive research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118824375B_ABST
    Figure CN118824375B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for locating genomic reproductive regions. This method specifically includes two steps: locating genomic recombination suppression regions and locating reproduction-related regions. Among them, the location of genomic recombination suppression regions includes three steps: filtering intervals with low sequence similarity, high repetitive sequence content, and low gene density, to obtain the recombination suppression intervals of the target species; on this basis, extracting the set of genes specifically expressed in the reproductive organs of the target species and counting their distribution in each recombination suppression interval, the reproduction-related regions of the angiosperm genome can be obtained, helping scientific researchers to accelerate the research on plant reproductive regions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bioinformatics technology, and specifically relates to a method for locating genomic reproductive regions. Background Art

[0002] The fruits and seeds produced by angiosperms are the main food sources for many animals (including humans). Especially in agriculture and horticulture, crops are the basis of the human food supply. Fruits and vegetables provide rich vitamins and minerals to promote health, while flowers enhance the quality of human life and spiritual pleasure through their ornamental and aesthetic values.

[0003] Angiosperms mainly achieve sexual reproduction through flowering and seeds. Among them, stamens produce pollen, and ovules in the pistil receive pollen for fertilization to form seeds. Compared with asexual reproduction, genetic recombination in the process of sexual reproduction generates new gene combinations, improving genetic diversity. This enables more individuals to potentially have adaptive advantages when facing environmental changes and resisting pests and diseases, thereby increasing the survival rate of the species.

[0004] Traditional methods for locating genomic reproductive regions are complex and cumbersome. They require researchers to invest a large amount of time and space in the early stage to construct large-scale segregating populations, and then estimate and compare the positions of genetic maps and physical maps; or conduct a large amount of sequencing work after investigating phenotypes, and then perform trait mapping, which is a time-consuming and laborious task. Among them, constructing segregating populations and investigating phenotypes are usually carried out in units of growing seasons or even years. For some species with a long juvenile period, it may take an entire career of scientific researchers to construct a suitable population. Summary of the Invention

[0005] Aiming at the defects existing in the prior art, the purpose of the present invention is to provide a method for locating genomic reproductive regions. This method utilizes a large number of published plant genomes and annotation files, and based on the sequence characteristics and gene expression of reproductive regions, achieves the purpose of efficiently and quickly locating reproductive regions, effectively simplifies the cumbersome traditional location method, and promotes the research of plant reproduction by researchers.

[0006] To achieve the above purpose, the technical solution adopted by the present invention is:

[0007] A method for locating genomic reproductive regions, characterized by comprising the following steps:

[0008] Step 1, locating genomic recombination suppression regions;

[0009] Step 2, locating reproduction-related regions.

[0010] On the basis of the above solution, the specific steps of the said Step 1 are:

[0011] Step 1-1: Compare the genomic sequences of the target species and its related species, and calculate the sequence similarity of the whole-genome sequences of the target species and its related species, as well as the sequence similarity between different interval sequence fragments. Detect whether the difference between the sequence similarity of each of the above interval sequence fragments and the whole-genome sequence similarity reaches a significant level, and retain the intervals where the sequence similarity is significantly lower than the average value of the whole genome (P<0.05).

[0012] Step 1-2: Conduct repeat sequence annotation on the genomic sequence of the target species, extract location information, and on the basis of filtering the sequence similarity results in Step 1-1, calculate the repeat sequence content in each interval sequence, and detect whether the difference between the repeat sequence content in each of the above interval sequence fragments and the whole-genome repeat sequence content reaches a significant level, and retain the intervals where the repeat sequence content is significantly higher than the whole-genome repeat sequence content (P<0.05).

[0013] Step 1-3: Format the gene annotation file of the target species and retain the necessary location information. On the basis of filtering the sequence similarity in Step 1-1 and filtering the repeat sequence density in Step 1-2, calculate the gene density in each interval sequence fragment, and detect whether the difference between the gene density in each of the above interval sequence fragments and the whole-genome gene density reaches a significant level, and retain the intervals where the gene density is significantly lower than the gene density in the whole-genome sequence (P<0.05), that is, the genomic recombination suppression regions are obtained.

[0014] On the basis of the above scheme, the specific steps of Step 2 are as follows:

[0015] Step 2-1: Obtain the gene expression levels of the flower tissue and other tissues of the target species, screen the genes that are only expressed in the flower tissue but not in other tissues, and obtain the set of genes specifically expressed in the reproductive organs of the target species; the other tissues are the root tissue, stem tissue, leaf tissue, fruit tissue, tendril tissue, seedling tissue and rosette tissue of the target species.

[0016] In this step, the gene expression levels of the flower tissue and other tissues of the target species are from public databases or calculated using conventional procedures. The conventional procedure for calculating gene expression levels mainly includes two steps. First, align the original sequencing data to the reference genome to obtain the aligned bam file, which can be completed using the bioinformatics analysis software hisat2 and samtools. Then, calculate the gene expression levels using the bam file and the gene annotation file of the reference genome, which can be completed using the bioinformatics analysis software stringtie.

[0017] Step 2-2: Count the distribution of each gene in the set of genes specifically expressed in reproductive organs obtained in Step 2-1 in each recombination suppression region of the genome of the target species obtained in Step 1, and select the regions containing a large number of the above-mentioned genes specifically expressed in reproductive organs and having a high proportion of the above-mentioned genes specifically expressed in reproductive organs, that is, obtain the reproductive-related regions on the genome.

[0018] The beneficial effect of the method for locating the reproductive region of the genome according to the present invention is as follows:

[0019] This method utilizes a large number of published plant genomes and annotation files, and based on the sequence characteristics and gene expression of the reproductive region, achieves the purpose of efficiently and quickly locating the reproductive region. This method can obtain the candidate reproductive regions of the target species in as fast as one day. If the genome of the target species is very complex, the results can also be obtained within a few weeks by splitting the chromosomes. This method effectively simplifies the cumbersome traditional location method and promotes the research of researchers on plant reproduction. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The present invention has the following drawings:

[0021] Figure 1 is the location process of the recombination suppression region;

[0022] Figure 2 is a schematic diagram of the downloaded genome sequence file;

[0023] Figure 3 is a schematic diagram of the downloaded gene annotation file;

[0024] Figure 4 is a schematic diagram of the sequence alignment result;

[0025] Figure 5 is a schematic diagram of the result after filtering sequence similarity;

[0026] Figure 6 is a schematic diagram of the annotation result of the simplified repetitive sequence;

[0027] Figure 7 is a schematic diagram of the result after filtering the repetitive sequence content;

[0028] Figure 8 is a schematic diagram of the simplified gene annotation file;

[0029] Figure 9 is a schematic diagram of the result after filtering the gene density;

[0030] Figure 10 is a schematic diagram of the final recombination suppression interval;

[0031] Figure 11 is a schematic diagram of the set of genes specifically expressed in reproductive organs;

[0032] Figure 12 Schematic diagram of the final genomic reproductive region;

[0033] Figure 13 Schematic diagram of the localization of the reproductive region on the 'Golden Delicious' genome;

[0034] Figure 14 Schematic diagram of the localization of the reproductive region on the Arabidopsis thaliana genome;

[0035] Figure 15 Schematic diagram of the localization of the reproductive region on the cultivated grape genome; Specific implementation mode

[0036] The present invention will be further described in detail below with reference to the accompanying drawings.

[0037] As Figure 1 shown, the method for localizing the genomic reproductive region of the present invention includes the following steps:

[0038] Input the genomic sequences of the target species and related species, use ncmer for sequence alignment, and then calculate the sequence similarity of the entire genomes of the target species and related species and the sequence similarity between sequence fragments in different intervals, and retain the intervals where the sequence similarity is significantly lower than the average value of the whole genome;

[0039] Input the genomic sequence of the target species, use RepeatModeler and RepeatMasker to annotate the repetitive sequences of the genomic sequence of the target species, and on the basis of filtering the sequence similarity results, count the content of repetitive sequences in each interval sequence, and retain the intervals where the content of repetitive sequences is significantly higher than the content of repetitive sequences in the whole genome;

[0040] Input the genomic sequence of the target species and the gene annotation file, and on the basis of filtering the sequence similarity and repetitive sequence density, count the gene density in each interval sequence fragment, and retain the intervals where the gene density is significantly lower than the gene density of the whole genome, that is, the genomic recombination suppression region is obtained.

[0041] As Figures 2-3 shown, the formats of the genomic sequence file and gene annotation file required for the method for localizing the genomic reproductive region of the present invention are as shown in the figure.

[0042] As Figures 4-9 shown, the formats of the intermediate files for filtering low sequence similarity, high repetitive sequence content and low gene density in the method for localizing the genomic reproductive region of the present invention are as shown in the figure.

[0043] As Figures 10-12As shown in the figure, the recombination suppression interval, the set of genes specifically expressed in reproductive organs, and the genomic reproductive region result file obtained by combining the two in the genomic reproductive region positioning method of the present invention are as shown in the figure.

[0044] The following combines the attached Figures 13-15 , and describes the specific implementation manners of the present invention in detail. The examples are used to illustrate the specific content of the present invention, but not to limit the scope of the present invention.

[0045] Example 1

[0046] (1) Experimental materials

[0047] The target species is the 'Golden Delicious' apple, and the related species is the wild apple

[0048] (2) Data download

[0049] The genomic sequence and gene annotation file can be downloaded from the public database GDR (https: / / www.rosaceae.org / ). The transcriptome data of each tissue of the apple can be downloaded from the public database NCBI (https: / / www.ncbi.nlm.nih.gov / ), and its BioProject numbers are PRJNA263198 and PRJNA482033.

[0050] (3) Localization of the genomic recombination suppression region of 'Golden Delicious'

[0051] The genome size of 'Golden Delicious' without chromosome 0 is 656.83 Mb. By aligning the genomic sequences of 'Golden Delicious' and the wild apple, the sequence similarity of the whole genome is 71.07%. Intervals with significantly lower sequence similarity than this value (P<0.05) are filtered, with a size of 370.69 Mb, accounting for 56.40% of the whole genome.

[0052] On this basis, repetitive sequence annotation is performed on the 'Golden Delicious' genome. The repetitive sequence content of the whole genome is 52.85%. Intervals with significantly higher repetitive sequence content than this value (P<0.05) are filtered, with a size of 211.80 Mb, accounting for 32.25% of the whole genome.

[0053] The 'Golden Delicious' genome annotation contains 43,488 genes, and the gene density of the whole genome is 66.21 genes / Mb. Intervals with significantly lower gene density than this value (P<0.05) are filtered, with a size of 176.20 Mb, accounting for 26.83% of the whole genome, which is the recombination suppression region of the 'Golden Delicious' genome.

[0054] (4) Localization of the genomic reproductive-related region of 'Golden Delicious'

[0055] Using the conventional process to calculate the transcriptome data of various tissues of apples, after merging, screening, and filtering, a total of 4,785 expressed genes were finally obtained within the recombination suppression region of the 'Golden Delicious' genome, among which 303 were specifically expressed genes in reproductive organs. MdoRS91 is the region of 29,400,001 - 32,200,000 on chromosome 17 of apples, containing a total of 26 specifically expressed genes in reproductive organs, including the gametophytic self-incompatibility female determinant (S-RNase) and the male determinant (F-box). It has the largest number among all recombination suppression regions, accounting for 27.96% of the expressed genes in MdoRS91, and is the self-incompatibility locus of apples. In addition, MdoRS25 is the region of 14,800,001 - 15,800,000 on chromosome 5 of apples, containing 5 specifically expressed genes in reproductive organs, including pollen-specific leucine-rich repeat extensin-like protein (PELP), and is the region related to pollen development and fertilization recognition.

[0056] Example 2

[0057] (1) Experimental materials

[0058] The target species is Arabidopsis thaliana, and the related species is Arabidopsis lyrata.

[0059] (2) Data download

[0060] The genome sequence and gene annotation files can be downloaded from the public database TAIR (https: / / www.arabidopsis.org / ). The transcriptome data of various tissues of Arabidopsis thaliana can be downloaded from the public database NCBI (https: / / www.ncbi.nlm.nih.gov / ), and its BioProject numbers are PRJNA1060775, PRJNA1083752, and PRJNA1071999.

[0061] (3) Localization of the recombination suppression region of the Arabidopsis thaliana genome

[0062] The genome size of Arabidopsis thaliana is 119.15 Mb. By comparing the genome sequences of Arabidopsis thaliana and Arabidopsis lyrata, the sequence similarity of the whole genome is 46.33%. Intervals with significantly lower sequence similarity than this value (P<0.05) are filtered, with a size of 57.19 Mb, accounting for 48.00% of the whole genome. On this basis, repetitive sequence annotation is carried out on the Arabidopsis thaliana genome. The repetitive sequence content of the whole genome is 16.53%. Intervals with significantly higher repetitive sequence content than this value (P<0.05) are filtered, with a size of 38.54 Mb, accounting for 32.35% of the whole genome. The Arabidopsis thaliana genome annotation contains 28,309 genes, and the gene density of the whole genome is 237.59 genes / Mb. Intervals with significantly lower gene density than this value (P<0.05) are filtered, with a size of 35.29 Mb, accounting for 29.62% of the whole genome, which is the recombination suppression region of the Arabidopsis thaliana genome.

[0063] (4) Localization of the reproduction-related region of the Arabidopsis thaliana genome

[0064] Using the conventional process to calculate the transcriptome data of each tissue of Arabidopsis thaliana, after merging, screening and filtering, finally 2,770 expressed genes are contained in the recombination suppression region of the Arabidopsis thaliana genome, among which 322 are genes specifically expressed in reproductive organs. AthRS98 is the region from 11,340,001 to 11,420,000 on chromosome 4 of Arabidopsis thaliana, containing a total of 4 genes specifically expressed in reproductive organs, including the sporophytic self-incompatibility female determinant (SRK) and male determinant (SCR), and is the self-incompatibility locus of Arabidopsis thaliana. AthRS55 is the region from 7,680,001 to 7,810,000 on chromosome 3 of Arabidopsis thaliana, containing 20 genes specifically expressed in reproductive organs including receptor-like serine / threonine protein kinases (RLKs), and has the largest number of genes among all recombination suppression regions. AthRS5 is the region from 9,890,001 to 9,950,000 on chromosome 1 of Arabidopsis thaliana, containing 3 genes specifically expressed in reproductive organs including the plant self-incompatibility S1 protein (S1) and pollen coat protein (PCP). Both AthRS55 and AthRS5 are regions related to self-incompatibility and pollination and fertilization of Arabidopsis thaliana.

[0065] Example 3

[0066] (1) Experimental materials

[0067] The target species is cultivated grape, and the related species is another assembled haplotype of cultivated grape.

[0068] (2) Data download

[0069] Genomic sequences and gene annotation files can be downloaded from the public database Zenodo (https: / / zenodo.org / ). Transcriptome data of various tissues of cultivated grapes can be downloaded from the public database NCBI (https: / / www.ncbi.nlm.nih.gov / ), with BioProject numbers PRJNA549877 and PRJNA593045.

[0070] (3) Localization of the genomic recombination suppression region in cultivated grapes

[0071] The genome size of cultivated grapes is 449.79 Mb. By aligning different genomic sequences of cultivated grapes, the sequence similarity of the whole genome is 67.70%. Intervals with significantly lower sequence similarity than this value (P < 0.05) are filtered, with a size of 209.66 Mb, accounting for 46.61% of the whole genome. On this basis, repetitive sequence annotation is carried out on the cultivated grape genome. The repetitive sequence content of the whole genome is 53.67%. Intervals with significantly higher repetitive sequence content than this value (P < 0.05) are filtered, with a size of 160.95 Mb, accounting for 35.78% of the whole genome. The cultivated grape genome annotation contains 27,684 genes, and the gene density of the whole genome is 61.55 genes / Mb. Intervals with significantly lower gene density than this value (P < 0.05) are filtered, with a size of 141.18 Mb, accounting for 31.39% of the whole genome, which is the genomic recombination suppression region of cultivated grapes.

[0072] (4) Localization of the genomic reproduction-related region in cultivated grapes

[0073] Vitis plants include dioecious species and monoecious species, that is, they have bisexual and unisexual flowers, and unisexual flowers include female and male flowers. Using the conventional process to calculate the transcriptome data of various tissues of cultivated grapes, after merging, screening, and filtering, a total of 2,524 expressed genes are finally obtained in the genomic recombination suppression region of cultivated grapes, among which 385 are reproductive organ-specific expressed genes. VviRS52 is the region of 4,890,001 - 4,990,000 on chromosome 2 of grapes, where the PLATZ transcription factor is a reproductive organ-specific expressed gene, and this region is the sex determination region of grapes. VviRS321 is the region of 15,520,001 - 15,700,000 on chromosome 7 of grapes, containing 6 reproductive organ-specific expressed genes such as rapid alkalinization factor (RALF) and serine / threonine protein kinase, and is a region related to flower development and fertilization recognition.

[0074] The specific embodiments of the present invention disclosed above are only for illustration, but the protection scope of the present invention is not limited thereto. Any changes and substitutions that can be easily made by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

[0075] The content not described in detail in this specification belongs to the prior art well-known to those skilled in the art.

Claims

1. A method for locating genomic reproductive regions, characterized in that, It includes the following steps: Step 1, localizing genomic recombination suppression regions; Step 2, localizing reproduction-related regions; The specific steps of Step 1 are as follows: Step 1-1, comparing the genomic sequences of the target species and its related species, calculating the sequence similarity of the whole-genome sequences of the target species and its related species, and the sequence similarity between different interval sequence fragments; detecting whether the difference between the sequence similarity of each interval sequence fragment and the whole-genome sequence similarity reaches a significant level, and retaining the intervals with significantly lower sequence similarity than the average of the whole genome; Step 1-2, performing repeat sequence annotation on the genomic sequence of the target species, extracting position information, and on the basis of filtering the sequence similarity results in Step 1-1, counting the repeat sequence content in each interval sequence, detecting whether the difference between the repeat sequence content in each interval sequence fragment and the whole-genome repeat sequence content reaches a significant level, and retaining the intervals with significantly higher repeat sequence content than the whole-genome repeat sequence content; Step 1-3, formatting the gene annotation file of the target species and retaining the necessary position information; on the basis of filtering sequence similarity in Step 1-1 and filtering repeat sequence density in Step 1-2, counting the gene density in each interval sequence fragment, detecting whether the difference between the gene density in each interval sequence fragment and the whole-genome gene density reaches a significant level, and retaining the intervals with significantly lower gene density than the whole-genome gene density, which are the genomic recombination suppression regions; The specific steps of Step 2 are as follows: Step 2-1, obtaining the gene expression levels of the flower tissue and other tissues of the target species, screening the genes that are only expressed in the flower tissue but not in other tissues, and obtaining the set of genes specifically expressed in the reproductive organs of the target species; the other tissues are the root tissue, stem tissue, leaf tissue, fruit tissue, tendril tissue, seedling tissue, and rosette tissue of the target species; Step 2-2, counting the distribution of each gene in the set of genes specifically expressed in the reproductive organs obtained in Step 2-1 in each genomic recombination suppression region of the target species obtained in Step 1, and selecting the regions with a large number of the above genes specifically expressed in the reproductive organs and a high proportion of the above genes specifically expressed in the reproductive organs, which are the reproduction-related regions on the genome.