Bioinformatics-based genetic data mining system for breeding of astragalus sinicus
By simultaneously collecting and processing the genomic data of the host and rhizobium symbiotic microorganisms of milkvetch, identifying and eliminating cross-contamination sites, and constructing a trait association analysis model, the problem of microbial interference in milkvetch breeding was solved, and the reliability and efficiency of precision breeding decisions were improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- FUJIAN AGRI FERTILE SOIL BIOTECHNOLOGY CO LTD
- Filing Date
- 2025-10-15
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies have failed to effectively identify and remove genetic information interference from the genomes of rhizobium symbiotic microorganisms in milkvetch breeding, resulting in insufficient accuracy of host genome trait association analysis and affecting the reliability of breeding decisions.
By collecting and synchronizing sequencing data of the host genome of *Astragalus membranaceus* and the genome of its symbiotic microorganisms *Rhizobium*, sequence quality control and redundancy removal were performed, cross-contamination sites were identified and eliminated, purified host genome data free from microbial interference was generated, a trait association analysis model was constructed, and a set of genetic markers for precision breeding was output.
It significantly improved data accuracy, avoided the bias of host genome analysis caused by microbial interference, quantified the contribution of microbial genomes to target traits, established a more comprehensive and accurate trait association analysis model, and improved the reliability and efficiency of breeding decisions.
Smart Images

Figure CN120954505B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data mining technology, and more specifically, to a genetic data mining system for milkvetch breeding based on bioinformatics. Background Technology
[0002] Milk clover is widely used to improve soil fertility and is cultivated as a forage crop; traditional genetic breeding methods mainly rely on the host plant's own genome information to conduct trait association analysis in order to achieve variety selection and improvement.
[0003] Existing technologies have failed to effectively identify and remove genetic information interference from the genomes of rhizobium symbiotic microorganisms during the genetic data mining process for milkvetch breeding. This results in insufficient accuracy of trait association analysis of the host genome, thereby affecting the reliability of precision breeding decisions for milkvetch. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a genetic data mining system for milkvetch breeding based on bioinformatics to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides the following technical solution:
[0006] A bioinformatics-based genetic data mining system for milkvetch breeding includes:
[0007] Data processing module: Collects and synchronizes the raw sequencing data of the host genome of Astragalus membranaceus and the genome of the rhizobium symbiotic microorganism, and generates a standardized genetic information dataset through sequence quality control and redundancy removal.
[0008] Contamination identification module: Based on a standardized genetic information dataset, it identifies sequence cross-contamination sites between the host genome and the symbiotic microbial genome, and outputs the location information of the cross-contamination sites;
[0009] Interference Removal Module: Based on the location information of cross-contamination sites, it identifies and removes genetic information interference from microorganisms, generating purified host genome data free of microbial interference;
[0010] Contribution Analysis Module: Based on the location information of cross-contamination sites, analyze the contribution of the microbial genome to the target traits of milkvetch and output the microbial effect weights;
[0011] Association Modeling Module: Integrates purified host genome data with microbial interference removed and microbial effect weights to construct an association analysis model for milkvetch traits and outputs trait association site information;
[0012] Marker screening module: Based on trait association locus information, it performs screening analysis of candidate breeding markers and outputs a set of genetic markers for precision breeding.
[0013] In a preferred embodiment, raw sequencing data of the *Astragalus membranaceus* host genome and the rhizobium symbiotic microorganism genome are collected and synchronized. Through sequence quality control and redundancy removal, a standardized genetic information dataset is generated, specifically:
[0014] Raw sequencing data of the host genome of milkvetch and raw sequencing data of the genomes of rhizobium microorganisms that coexist with milkvetch were collected.
[0015] Time synchronization processing was performed on the raw sequencing data of the host genome of *Astragalus membranaceus* and the raw sequencing data of the microbial genome of *Rhizobium*.
[0016] The raw sequencing data of the host genome of *Astragalus membranaceus* and the raw sequencing data of the microbial genome of *Rhizobium* after time synchronization were processed to remove sequence redundancy and form a standardized genetic information dataset.
[0017] In a preferred embodiment, based on a standardized genetic information dataset, sequence cross-contamination sites between the host genome and the symbiotic microbial genome are identified, and the location information of the cross-contamination sites is output, specifically:
[0018] Each gene sequence fragment in the standardized genetic information dataset was aligned with the reference sequence of the host genome of *Astragalus membranaceus* and the reference sequence of the microbial genome of *Rhizobium*, and the alignment coordinates of each gene sequence fragment on the reference sequence of the host genome of *Astragalus membranaceus* and the reference sequence of the microbial genome of *Rhizobium* were recorded.
[0019] Gene sequence fragments that simultaneously meet a set similarity threshold in both the reference genome sequence of *Astragalus membranaceus* host and the reference genome sequence of *Rhizobium* microorganisms;
[0020] Coordinate clustering analysis is performed on the corresponding alignment coordinates of gene sequence fragments to output the location information of cross-contamination sites, including chromosome number, start coordinates, and end coordinates.
[0021] In a preferred embodiment, based on the location information of cross-contamination sites, genetic information interference from microorganisms is identified and removed to generate purified host genome data free of microbial interference, specifically as follows:
[0022] The alignment coordinates of each gene sequence fragment in the standardized genetic information dataset on the reference sequence of the milkvetch host genome are compared with the location information of the cross-contamination site;
[0023] Identify gene sequence fragments whose alignment coordinates fall within the range of cross-contamination site location information and label them as genetic information interference fragments of microbial origin;
[0024] The genetic information interference fragments marked as originating from microorganisms are removed from the standardized genetic information dataset to form purified host genome data free from microbial interference.
[0025] In a preferred embodiment, based on the location information of cross-contamination sites, the contribution of the microbial genome to the target trait of *Astragalus membranaceus* is analyzed, and the microbial effect weight is output, specifically as follows:
[0026] Based on the location information of cross-contamination sites, gene sequence fragments whose alignment coordinates are located within the reference sequence of the rhizobium microbial genome are extracted from the standardized genetic information dataset;
[0027] The frequency of gene sequence fragments in target trait samples of different milkvetch varieties was statistically analyzed to determine the association strength between each gene sequence fragment and the target trait.
[0028] The contribution of gene sequence fragments to the target trait is calculated based on the association strength, and the weight of microbial effect is output.
[0029] In a preferred embodiment, purified host genome data free from microbial interference is integrated with microbial effect weights to construct a milkvetch trait association analysis model, outputting trait association site information, specifically:
[0030] Based on purified host genome data free from microbial interference, an initial association model between the host genotype of milkvetch and the target trait was constructed.
[0031] The corrected association model is obtained by adjusting the association parameters in the initial association model based on the weights of microbial effects.
[0032] Based on the corrected association model, gene sequence fragments associated with the target trait in the host genome of milkvetch are located, and trait association site information including host chromosome number, start coordinates and end coordinates is output.
[0033] In a preferred embodiment, based on trait-associated locus information, candidate breeding markers are screened and analyzed to output a set of genetic markers for precision breeding, specifically:
[0034] Based on trait-associated locus information, genetic marker loci corresponding to each gene sequence fragment in the host genome of Astragalus membranaceus were extracted;
[0035] The frequency of each genetic marker locus in different target trait sample populations was statistically analyzed, and the correlation coefficient between each genetic marker locus and the target trait was calculated.
[0036] Genetic marker loci with correlation coefficients reaching a set threshold are selected to form a genetic marker set for precision breeding.
[0037] The technical effects and advantages of the bioinformatics-based genetic data mining system for milkvetch breeding in this invention are as follows:
[0038] By simultaneously collecting sequencing data of the host genome of milkvetch and the genome of rhizobium microorganisms, and performing sequence quality control and redundancy removal, the accuracy of the data was significantly improved. Microbial genetic information interference was identified and eliminated, avoiding biases in host genome analysis caused by microbial interference. Simultaneously, the contribution of the microbial genome to the target trait was quantified, establishing a more comprehensive and accurate trait association analysis model. Ultimately, a precise set of breeding genetic markers was obtained, significantly improving the reliability and efficiency of milkvetch breeding decisions and promoting the breeding process of new milkvetch varieties. Attached Figure Description
[0039] Figure 1 This is a schematic diagram of the structure of the bioinformatics-based genetic data mining system for milkvetch breeding according to the present invention. Detailed Implementation
[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention. Example 1
[0041] Figure 1 This invention presents a bioinformatics-based genetic data mining system for milkvetch breeding, comprising:
[0042] Data processing module: Collects and synchronizes the raw sequencing data of the host genome of Astragalus membranaceus and the genome of the rhizobium symbiotic microorganism, and generates a standardized genetic information dataset through sequence quality control and redundancy removal.
[0043] Contamination identification module: Based on a standardized genetic information dataset, it identifies sequence cross-contamination sites between the host genome and the symbiotic microbial genome, and outputs the location information of the cross-contamination sites;
[0044] Interference Removal Module: Based on the location information of cross-contamination sites, it identifies and removes genetic information interference from microorganisms, generating purified host genome data free of microbial interference;
[0045] Contribution Analysis Module: Based on the location information of cross-contamination sites, analyze the contribution of the microbial genome to the target traits of milkvetch and output the microbial effect weights;
[0046] Association Modeling Module: Integrates purified host genome data with microbial interference removed and microbial effect weights to construct an association analysis model for milkvetch traits and outputs trait association site information;
[0047] Marker screening module: Based on trait association locus information, it performs screening analysis of candidate breeding markers and outputs a set of genetic markers for precision breeding.
[0048] Raw sequencing data of the host genome of *Astragalus membranaceus* and the genomes of rhizobia symbiotic microorganisms were collected and synchronized. Through sequence quality control and redundancy removal, a standardized genetic information dataset was generated, including:
[0049] Raw sequencing data of the host genome of milkvetch and raw sequencing data of the genomes of rhizobium microorganisms that coexist with milkvetch were collected.
[0050] The method for collecting raw sequencing data of the milkvetch host genome is as follows: A population of target milkvetch plants is planted in an experimental field or laboratory. When the plants reach a specific maturity stage, such as flowering or seed formation, several plants are randomly selected, for example, twenty plants. Leaf or bud tissue samples are taken, and high-throughput sequencing instruments, such as short-sequence sequencing platforms, are used to obtain the complete raw sequencing data of the host genome for each milkvetch plant. The method for collecting raw sequencing data of the rhizobium microbial genome is as follows: Root nodule samples are dug from the roots of milkvetch plants. The number of samples corresponds to the number of plants collected during the milkvetch host genome collection, for example, five root nodule samples are collected from each plant. All root nodule samples are surface-sterilized, and the genomic DNA of the rhizobium microorganisms is extracted by culturing or direct disruption. The DNA samples are then sequenced using the same high-throughput sequencing instruments as those used for milkvetch host genome sequencing to obtain the raw sequencing data of the rhizobium microbial genome.
[0051] Time synchronization processing was performed on the raw sequencing data of the host genome of *Astragalus membranaceus* and the raw sequencing data of the microbial genome of *Rhizobium*.
[0052] Since the acquisition of different sequencing data usually involves a certain time interval, and the sequencing time of different samples within the same batch may also differ due to variations in instrument operating efficiency, it is necessary to perform unified time-scale correction processing on all raw sequencing data of the *Astragalus membranaceus* host genome and all raw sequencing data of the *Rhizobium* microbial genome. This includes: mapping all raw sequencing data to the same experimental period according to the labeling information of the sampled plants, for example, using the number of days of growth of the sampled plants as a benchmark, recording the exact time length from sampling to obtaining raw sequencing data for each sample, and then using the longest sequencing data acquisition time as a unified benchmark to map earlier sequencing data to a unified time benchmark, so as to achieve temporal consistency between the raw sequencing data of the *Astragalus membranaceus* host genome and the raw sequencing data of the *Rhizobium* microbial genome, thereby eliminating potential errors caused by time differences.
[0053] Sequence redundancy removal was performed on the raw sequencing data of the host genome of *Astragalus membranaceus* and the raw sequencing data of the microbial genome of *Rhizobium* after time synchronization processing to form a standardized genetic information dataset.
[0054] During high-throughput sequencing, experimental procedures may result in the repeated occurrence of certain gene sequence fragments from the same sample. Therefore, it is necessary to identify and merge these duplicated fragments. The raw sequencing data of the *Astragalus membranaceus* host genome and *Rhizobium* microbial genome are segmented into gene sequence fragments of a certain length, for example, each fragment is set to a length of several base units. Global sequence alignment is performed between each segmented gene sequence fragment, using a sequence similarity algorithm. This involves comparing the two gene sequences one by one to determine the number of identical sites, and then calculating the ratio of the number of identical sites to the fragment length, which is the similarity score. For fragments that meet alignment threshold conditions, such as those with similarity scores exceeding a set threshold, they are identified as redundant sequences and merged, i.e., removing duplicate fragments and retaining only the unique representative fragment. Finally, the processed gene sequence fragments are reassembled to form a standardized genetic information dataset. A standardized genetic information dataset refers to a data set where, after sequence redundancy removal, there are no longer any redundant relationships between the sequence fragments.
[0055] Based on a standardized genetic information dataset, sequence cross-contamination sites between the host genome and the symbiotic microbial genome are identified, and the location information of the cross-contamination sites is output, including:
[0056] Each gene sequence fragment in the standardized genetic information dataset was aligned with the reference sequence of the host genome of *Astragalus membranaceus* and the reference sequence of the microbial genome of *Rhizobium*, and the alignment coordinates of each gene sequence fragment on the reference sequence of the host genome of *Astragalus membranaceus* and the reference sequence of the microbial genome of *Rhizobium* were recorded.
[0057] The reference sequence for the host genome of *Astragalus membranaceus* is a standard genome sequence determined through multiple sequencing and genome assembly processes in previous studies and experiments. For example, a specific *Astragalus membranaceus* variety is selected as a reference variety and obtained after multiple sequencing and assembly processes. Similarly, the reference sequence for the rhizobium microbial genome is a standard genome sequence obtained through multiple high-throughput sequencing and assembly processes. The rhizobium microbial genome reference sequence is typically obtained by extracting the genome of *Astragalus membranaceus* plant rhizobium strains and assembling them. Sequence alignment specifically involves taking each gene sequence fragment from the standardized genetic information dataset and aligning it positionally and base-by-base with the *Astragalus membranaceus* host genome reference sequence, moving the window sequentially from one end of the genome reference sequence. The gene sequence fragments were compared base-by-base with the gene sequences within the reference sequence window, and the number of identical base positions was recorded. The proportion of identical base positions to the total length of the gene sequence fragment was used as the sequence similarity. After the alignment with the reference sequence of the *Astragalus membranaceus* host genome was completed, the same gene sequence fragment was then compared base-by-base with the reference sequence of the *Rhizobium* microbial genome, and the alignment results of the gene sequence fragments on the *Rhizobium* microbial genome reference sequence were recorded. Through the above methods, the alignment position information of each gene sequence fragment on the reference sequence of the *Astragalus membranaceus* host genome and the reference sequence of the *Rhizobium* microbial genome were obtained, and the corresponding alignment coordinate information was formed.
[0058] Gene sequence fragments that simultaneously meet a set similarity threshold in both the reference genome sequence of *Astragalus membranaceus* host and the reference genome sequence of *Rhizobium* microorganisms;
[0059] A pre-set sequence similarity threshold, such as a specific ratio, is determined by dividing the number of identical base positions by the total length of the gene sequence fragment. When a gene sequence fragment is aligned with the *Astragalus membranaceus* host genome reference sequence, if the similarity obtained is equal to or exceeds the set sequence similarity threshold, it is recorded as meeting the threshold condition. Then, the same gene sequence fragment is aligned with the *Rhizobium* microbial genome reference sequence for similarity assessment. If the similarity with the *Rhizobium* microbial genome reference sequence also meets or exceeds the set sequence similarity threshold, it indicates that the gene sequence fragment simultaneously meets the similarity requirements of both the *Astragalus membranaceus* host genome reference sequence and the *Rhizobium* microbial genome reference sequence, and is therefore recorded as a cross-contamination candidate sequence fragment. Conversely, if a gene sequence fragment only meets the sequence similarity threshold with one of the reference genome sequences, it is not recorded as a cross-contamination candidate sequence fragment. Through this method, a batch of gene sequence fragments that simultaneously meet the set similarity threshold conditions with both the *Astragalus membranaceus* host genome reference sequence and the *Rhizobium* microbial genome reference sequence are screened from the standardized genetic information dataset.
[0060] Coordinate clustering analysis is performed on the corresponding alignment coordinates of gene sequence fragments to output the location information of cross-contamination sites, including chromosome number, start coordinates, and end coordinates;
[0061] For each gene sequence fragment that simultaneously meets the similarity threshold condition, the alignment coordinates of the start and end positions on the *Astragalus membranaceus* host genome reference sequence and the *Rhizobium* microbial genome reference sequence are extracted sequentially. The alignment coordinates of all gene sequence fragments on the reference sequences are then marked on the corresponding chromosome map of the reference sequences. Based on the density of the alignment positions on the chromosome map, multiple gene sequence fragments whose distribution positions are within a specific distance threshold range are grouped into the same cluster. For example, when the distance between two adjacent alignment positions is less than a pre-set distance threshold, the two adjacent alignment positions are divided into the same cluster. Other adjacent positions are then merged to form several coordinate clusters. Through the above method, several coordinate clusters are formed on the *Astragalus membranaceus* host genome reference sequence and the *Rhizobium* microbial genome reference sequence, with each cluster representing a cross-contamination site.
[0062] Based on coordinate clusters, the location information of cross-contamination sites is determined. Specifically, the start and end positions of all gene sequence fragments within each coordinate cluster are recorded using chromosome numbers, start coordinates, and end coordinates, respectively, ultimately forming a database of cross-contamination site location information containing chromosome numbers, start coordinates, and end coordinates. Each record in the database represents the precise location of the cross-contamination site on the reference sequences of the *Astragalus membranaceus* host genome and the *Rhizobium* microbial genome. The location information of cross-contamination sites will be used to identify, label, and remove genetic information interference from microorganisms, thereby ensuring the reliability of data analysis.
[0063] Based on the location information of cross-contamination sites, genetic information interference from microorganisms is identified and eliminated to generate purified host genome data free from microbial interference, including:
[0064] The alignment coordinates of each gene sequence fragment in the standardized genetic information dataset on the reference sequence of the milkvetch host genome are compared with the location information of the cross-contamination site;
[0065] For each gene sequence fragment in the standardized genetic information dataset, a gene sequence fragment is randomly selected from the dataset, for example, a specific start and end position on a chromosome of the *Actinidia cuspidata* host genome. This specific start and end position is then compared one by one with each cross-contamination site location record in the cross-contamination site location information database to determine whether the selected gene sequence fragment's coordinates fall within the chromosome coordinate range of the cross-contamination site location information. This comparison is performed on each gene sequence fragment in the standardized genetic information dataset, and the results are recorded each time, until all gene sequence fragment coordinates in the standardized genetic information dataset have been compared with the cross-contamination site location information.
[0066] Identify gene sequence fragments whose alignment coordinates fall within the range of cross-contamination site location information and label them as genetic information interference fragments of microbial origin;
[0067] When the alignment coordinates of a gene sequence fragment are found to be within the range of cross-contamination site location information after one-to-one alignment, it indicates that the gene sequence fragment may simultaneously possess genetic characteristics of both the milkvetch host genome and the rhizobium microbial genome. Therefore, the gene sequence fragment is defined as a microbial-derived genetic information interference fragment. For example, if the alignment result of a gene sequence fragment on the milkvetch host genome reference sequence is within a specific coordinate range of a certain chromosome, corresponding to the interval defined by the start and end coordinates of a certain cross-contamination site in the cross-contamination site location information database, then the gene sequence fragment is marked as a microbial-derived genetic information interference fragment, and a special marker is added to the standardized genetic information dataset to distinguish it from gene sequence fragments that are not identified as microbial-derived genetic information interference fragments. All gene sequence fragments whose coordinates are within the range of cross-contamination site location information are identified and marked.
[0068] Remove genetic information interference fragments labeled as originating from microorganisms from standardized genetic information datasets to form purified host genome data free from microbial interference;
[0069] Based on the identified and labeled microbial-derived genetic interference fragments, all gene sequence fragments recorded in the standardized genetic information dataset are retrieved one by one. Gene sequence fragments labeled as microbial-derived genetic interference fragments are deleted from the standardized genetic information dataset. After deleting all labeled microbial-derived genetic interference fragments, the remaining gene sequence fragments are those unaffected by microbial-derived genetic information interference. The unaffected gene sequence fragments are then recombined according to their original coordinate positions to form purified host genome data free from microbial interference. Each gene sequence fragment in the purified host genome data free from microbial interference has been excluded from microbial-derived genetic information interference.
[0070] Based on the location information of cross-contamination sites, the contribution of microbial genomes to the target traits of milkvetch is analyzed, and the microbial effect weights are output, including:
[0071] Based on the location information of cross-contamination sites, gene sequence fragments whose alignment coordinates are located within the reference sequence of the rhizobium microbial genome are extracted from the standardized genetic information dataset;
[0072] For each cross-contamination location information record, based on the chromosome number, start coordinates, and end coordinates, all gene sequence fragments whose alignment coordinates on the Rhizobium microbial genome reference sequence fall within the cross-contamination location information range are extracted from the standardized genetic information dataset. For example, if the cross-contamination location information record contains a specific interval from a specific start coordinate to a specific end coordinate on a certain chromosome of the Rhizobium microbial genome, all gene sequence fragments are retrieved from the standardized genetic information dataset. When the alignment coordinates of a gene sequence fragment on the Rhizobium microbial genome reference sequence fall within the specific interval, the gene sequence fragment is extracted. The above extraction process is repeated until all gene sequence fragments whose alignment coordinates on the Rhizobium microbial genome reference sequence are extracted from the standardized genetic information dataset.
[0073] The frequency of gene sequence fragments in target trait samples of different milkvetch varieties was statistically analyzed to determine the association strength between each gene sequence fragment and the target trait.
[0074] Multiple target trait sample populations of milkvetch varieties were established, such as selecting high-yielding and low-yielding milkvetch varieties as representative sample populations, each containing a certain number of individual plants. The distribution of extracted gene sequence fragments in different target trait sample populations was analyzed one by one. Specifically, for each milkvetch variety target trait sample population, the host genome data of each plant within the population was retrieved, and the frequency of occurrence of the extracted gene sequence fragment in the plant's host genome was recorded. For example, a specific gene sequence fragment appeared 18 times in 20 milkvetch plants in the high-yielding sample population and 5 times in 20 plants in the low-yielding sample population; the frequency of occurrence of the specific gene sequence fragment in different sample populations was recorded respectively. Using the above method, all extracted gene sequence fragments were distributed among different target trait sample populations. The frequency of each gene sequence fragment in each target trait sample population was obtained through statistical analysis. The association strength between each gene sequence fragment and the target trait was analyzed based on the frequency of occurrence. Specifically, for each gene sequence fragment, the frequency of occurrence of the gene sequence fragment in a sample population of milkvetch varieties with good performance in the target trait was compared with the frequency of occurrence of milkvetch varieties with poor performance in the target trait. The association strength between the gene sequence fragment and the target trait was obtained by calculating the difference in frequency between the two sample populations and the ratio of the difference in frequency to the total frequency. For example, if the frequency of occurrence of the gene sequence fragment in the good-performing population is much higher than that in the poor-performing population, the association strength between the gene sequence fragment and the target trait is considered high; if there is no difference in the frequency of occurrence of the gene sequence fragment in the two populations, the association strength with the target trait is considered low.
[0075] The contribution of gene sequence fragments to the target trait is calculated based on the association strength, and the weight of microbial effect is output.
[0076] The association strength between each gene sequence fragment and the target trait is normalized by dividing the association strength of each gene sequence fragment by the sum of the association strengths of all gene sequence fragments to obtain the relative contribution of the gene sequence fragment to the target trait. For example, if the association strength of a gene sequence fragment accounts for a specific proportion of the total association strength of all gene sequence fragments, then that specific proportion is the contribution of the gene sequence fragment to the target trait. The value of the contribution of each gene sequence fragment is then defined as the microbial effect weight. The microbial effect weights of all gene sequence fragments are recorded.
[0077] By integrating purified host genome data free from microbial interference and microbial effect weights, a trait association analysis model for *Actinidia cuspidata* was constructed, outputting trait association site information, including:
[0078] Based on purified host genome data free from microbial interference, an initial association model between the host genotype of milkvetch and the target trait was constructed.
[0079] Several target traits are selected as the analysis objects, such as plant yield, growth rate, or stress resistance. Each target trait requires the establishment of an independent association model. Each gene sequence fragment in the purified host genome data is examined. For example, using several milkvetch varieties or individual plants as the research object, association relationships are established using phenotypic and genotypic data. Specifically, each milkvetch plant possesses a target trait, such as plant yield or growth rate. The target trait and its corresponding host genome data are compared, and an association model is established by analyzing the relationship between gene sequences and target traits. For example, if the gene sequence... The presence of gene sequence fragments with higher frequencies in high-yielding milkvetch plants and lower frequencies in low-yielding milkvetch plants preliminarily establishes a potential association between gene sequence fragments and yield traits. Similarly, the same method is used to analyze growth rate traits, identifying gene sequence fragments in the host genome that may have potential associations with the target trait through gene-by-gene sequence fragment analysis. Finally, all potentially associated gene sequence fragments are recorded to construct an initial association model between the milkvetch host genotype and the target trait. The initial association model records the potential association strength between each gene sequence fragment and the target trait.
[0080] The corrected association model is obtained by adjusting the association parameters in the initial association model based on the weights of microbial effects.
[0081] The microbial effect weight represents the relative contribution of each gene sequence fragment extracted from the rhizobium microbial genome sequence to the target trait of *Astragalus membranaceus*. The association parameters in the initial association model are the potential association strength between each gene sequence fragment and the target trait, determined during the construction of the initial association model; for example, the ratio of the difference in frequency of each gene sequence fragment among different *Astragalus membranaceus* varietal populations. To obtain more accurate association analysis results, it is necessary to adjust the association parameters in the initial association model by incorporating the microbial effect weight. This involves integrating the contribution of microbial gene sequence fragments to the target trait. The adjustment method includes: for each gene sequence fragment in the initial association model, first determining the microbial effect weight corresponding to each gene sequence fragment; then adding the value of the microbial effect weight to the initial association model. Specifically, the association strength is determined as follows: For example, if the initial association strength of a gene sequence fragment is a specific value and the microbial effect weight is also a specific value, then the initial association strength and the microbial effect weight are numerically combined and the weights are adjusted by addition or multiplication. The new value after adjustment of the microbial effect weight is defined as the corrected association parameter. The association parameters in all initial association models are adjusted until all association parameters are corrected. All the adjusted association parameters constitute a new corrected association model. The corrected association model records the adjusted association strength between each gene sequence fragment and the target trait. Compared with the initial association model, it takes into account the contribution of the microbial genome, making it more accurate and objective. Finally, the corrected association model is formed, which contains the accurate association strength of each gene sequence fragment after weight adjustment.
[0082] Based on the corrected association model, gene sequence fragments associated with the target trait in the host genome of milkvetch are located, and trait association site information including host chromosome number, start coordinate and end coordinate is output;
[0083] Define a specific association strength threshold, which can be a specific proportion of the overall association strength distribution. Then, compare the association strength of each gene sequence fragment with the specific association strength threshold. If the association strength of a gene sequence fragment exceeds the threshold, it is determined that the gene sequence fragment is associated with the target trait and is recorded as a target trait association site. Record the chromosome number, start coordinate, and end coordinate of the gene sequence fragment determined as a target trait association site on the host genome reference sequence. For example, a specific gene sequence fragment is located within a specific start coordinate to a specific end coordinate range of a specific chromosome number in the host genome. Record all trait association site information after threshold determination in the trait association site database. Each record includes the chromosome number, start coordinate, and end coordinate. The trait association site database can display the location information of all gene sequence fragments in the milkvetch host genome that affect a specific target trait. For example, the trait association site database can accurately retrieve which specific chromosome locations are closely associated with the yield or growth rate traits of milkvetch.
[0084] Based on trait-related locus information, candidate breeding markers are screened and analyzed to output a set of genetic markers for precision breeding, including:
[0085] Based on trait-associated locus information, genetic marker loci corresponding to each gene sequence fragment in the host genome of Astragalus membranaceus were extracted;
[0086] Genetic marker sites are specific DNA sequences in the host genome that can be stably passed on to offspring for trait tracking and genetic improvement. The method for extracting genetic marker sites involves: recording information for each trait-associated site, for example, extracting the corresponding gene sequence fragment from the host genome reference sequence based on a specific interval determined by chromosome number, start coordinates, and end coordinates; screening and identifying genetic marker sites in the extracted gene sequence fragments, including single nucleotide polymorphism (SNP) sites, insertion / deletion sites, or simple sequence repeat sites; for example, within a specific gene sequence fragment within the range of a target trait-associated site, there may be multiple SNP sites or insertion / deletion sites, all of which may serve as potential genetic marker sites; analyzing all possible genetic marker sites within each gene sequence fragment and recording the positions of all sites; the position recording method is the chromosome number of the *Astragalus membranaceus* host genome, the start coordinates of the genetic marker site, and the end coordinates of the genetic marker site; for example, a SNP site is located at a specific coordinate position on a specific chromosome in the *Astragalus membranaceus* host genome; and recording the positions of all genetic marker sites through the above site-by-site analysis.
[0087] The frequency of each genetic marker locus in different target trait sample populations was statistically analyzed, and the correlation coefficient between each genetic marker locus and the target trait was calculated.
[0088] Multiple target trait sample populations were selected, such as a high-yielding milkvetch variety population and a low-yielding milkvetch variety population; each target trait population contained a specific number of individual milkvetch plants; for example, each target trait population contained twenty or more milkvetch plants. The frequency of each genetic marker locus in different target trait sample populations was statistically analyzed. Specifically, for a specific genetic marker locus, such as a single nucleotide polymorphism locus on a specific chromosome number, the host genome data of all milkvetch plants in each target trait sample population was retrieved to determine whether the genetic marker locus exists in the genome data of each plant; if the genetic marker locus is detected in eighteen out of twenty milkvetch plants in a specific population, the frequency is eighteen divided by twenty; through the above... The method involves recording the frequency of each genetic marker locus in different target trait populations. After determining the frequency of each genetic marker locus in different target trait populations, the correlation coefficient between each genetic marker locus and the target trait is calculated. The correlation coefficient is calculated as follows: first, the target trait within each sample population is determined, for example, the yield per plant in a high-yield population and the yield per plant in a low-yield population. Statistical analysis is then performed between the frequency of the genetic marker locus and the target trait, specifically by multiplying the covariance of the frequency of the genetic marker locus and the target trait by their respective standard deviations. Using this method, the correlation coefficient between each genetic marker locus and the target trait is calculated. The correlation coefficient represents the strength of the correlation between the genetic marker locus and the target trait.
[0089] Genetic marker loci with correlation coefficients reaching a set threshold are selected to form a genetic marker set for precision breeding;
[0090] A correlation coefficient threshold is pre-set. The threshold can be determined by selecting a high-proportion location in the overall correlation coefficient distribution to ensure a sufficiently strong correlation between the selected genetic marker loci and the target trait. For example, if the correlation coefficient between a genetic marker locus and a yield trait exceeds the threshold, the genetic marker locus is considered a valid genetic marker locus for precision breeding. By analyzing the correlation coefficients of all genetic marker loci, the coordinate information of all genetic marker loci exceeding the set threshold is recorded, forming a precision breeding genetic marker set. The genetic marker set contains genetic marker locus information, including the chromosome number, start coordinates, and end coordinates of each marker locus. Each record in the genetic marker set provides explicit genetic information support for precision breeding. Through the genetic marker set, breeding work can achieve genetic selection and improvement of target traits, such as using marker loci in the genetic marker set for selection when breeding high-yielding milkvetch varieties.
[0091] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0092] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0093] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0094] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0095] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0096] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.
[0097] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0098] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0099] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A bioinformatics-based genetic data mining system for milkvetch breeding, characterized in that, include: Data processing module: Collects and synchronizes the raw sequencing data of the host genome of Astragalus membranaceus and the genome of the rhizobium symbiotic microorganism, and generates a standardized genetic information dataset through sequence quality control and redundancy removal. The contamination identification module, based on a standardized genetic information dataset, identifies cross-contamination sites between the host genome and the symbiotic microbial genome, and outputs the location information of these cross-contamination sites. Specifically, it aligns each gene sequence fragment in the standardized genetic information dataset with both the *Astragalus membranaceus* host genome reference sequence and the *Rhizobium* microbial genome reference sequence, recording the alignment coordinates of each gene sequence fragment on both sequences. It then filters gene sequence fragments that simultaneously meet a set similarity threshold in both sequences. Finally, it performs coordinate clustering analysis on the corresponding alignment coordinates of the gene sequence fragments, outputting the location information of the cross-contamination sites, including chromosome number, start coordinates, and end coordinates. Interference Removal Module: Based on the location information of cross-contamination sites, this module identifies and removes genetic information interference from microorganisms, generating purified host genome data free of microbial interference. Specifically, it compares the alignment coordinates of each gene sequence fragment in the standardized genetic information dataset with the location information of the cross-contamination site on the reference sequence of the *Actinidia cuspidata* host genome; identifies gene sequence fragments whose alignment coordinates fall within the range of the cross-contamination site location information and marks them as genetic information interference fragments from microorganisms; and removes the genetic information interference fragments marked as genetic information interference from microorganisms from the standardized genetic information dataset, forming purified host genome data free of microbial interference. Contribution Analysis Module: Based on the location information of cross-contamination sites, analyze the contribution of the microbial genome to the target traits of milkvetch and output the microbial effect weights; Specifically, based on the location information of cross-contamination sites, gene sequence fragments whose alignment coordinates are located within the reference sequence of the rhizobium microbial genome are extracted from the standardized genetic information dataset; the frequency of occurrence of gene sequence fragments in target trait samples of different milkvetch varieties is counted to determine the association strength between each gene sequence fragment and the target trait; The contribution of gene sequence fragments to the target trait is calculated based on the association strength, and the weight of microbial effect is output. Association Modeling Module: Integrates purified host genome data with microbial interference removed and microbial effect weights to construct an association analysis model for milkvetch traits and outputs trait association site information; Marker screening module: Based on trait association locus information, it performs screening analysis of candidate breeding markers and outputs a set of genetic markers for precision breeding.
2. The bioinformatics-based genetic data mining system for milkvetch breeding according to claim 1, characterized in that, Raw sequencing data of the host genome of *Astragalus membranaceus* and the genomes of rhizobia symbiotic microorganisms were collected and synchronized. Through sequence quality control and redundancy removal, a standardized genetic information dataset was generated, specifically: Raw sequencing data of the host genome of milkvetch and raw sequencing data of the genomes of rhizobium microorganisms that coexist with milkvetch were collected. Time synchronization processing was performed on the raw sequencing data of the host genome of *Astragalus membranaceus* and the raw sequencing data of the microbial genome of *Rhizobium*. The raw sequencing data of the host genome of *Astragalus membranaceus* and the raw sequencing data of the microbial genome of *Rhizobium* after time synchronization were processed to remove sequence redundancy and form a standardized genetic information dataset.
3. The bioinformatics-based genetic data mining system for milkvetch breeding according to claim 1, characterized in that, By integrating purified host genome data free from microbial interference and microbial effect weights, a trait association analysis model for milkvetch was constructed, outputting trait association site information, specifically: Based on purified host genome data free from microbial interference, an initial association model between the host genotype of milkvetch and the target trait was constructed. The corrected association model is obtained by adjusting the association parameters in the initial association model based on the weights of microbial effects. Based on the corrected association model, gene sequence fragments associated with the target trait in the host genome of milkvetch are located, and trait association site information including host chromosome number, start coordinates and end coordinates is output.
4. The bioinformatics-based genetic data mining system for milkvetch breeding according to claim 3, characterized in that, Based on trait-related locus information, candidate breeding markers are screened and analyzed to output a set of genetic markers for precision breeding, specifically: Based on trait-associated locus information, genetic marker loci corresponding to each gene sequence fragment in the host genome of Astragalus membranaceus were extracted; The frequency of each genetic marker locus in different target trait sample populations was statistically analyzed, and the correlation coefficient between each genetic marker locus and the target trait was calculated. Genetic marker loci with correlation coefficients reaching a set threshold are selected to form a genetic marker set for precision breeding.