Method and system for finding optimal matching individual for breeding improvement, and storage medium
By constructing a quantitative matching model based on the contribution of functional sites and the characteristics of linkage regions, the problems of trait determination and high cost and long cycle in traditional breeding methods are solved. Direct matching of genotype differences is achieved, which improves breeding efficiency and accuracy and supports rapid breeding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-03-24
AI Technical Summary
Traditional breeding methods face challenges in identifying traits when searching for individuals with different target traits, resulting in high costs and long cycles, making it difficult to quickly adapt to market demands.
By integrating genomics, bioinformatics, and statistical methods, a quantitative matching model based on functional site contribution and linkage region characteristics is constructed. This model directly matches based on genotype differences and uses PVE value weighting and block difficulty coefficient correction to achieve a shift from empirical selection to quantitative decision-making.
Precise selection of breeding individuals significantly reduces field trial costs, shortens the breeding cycle, improves matching accuracy and breeding efficiency, and supports the rapid development of agricultural breeding.
Smart Images

Figure CN121725874A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of gene breeding technology, specifically to a method, system, and storage medium for finding the best-matching individuals for breeding improvement. Background Technology
[0002] Breeding, as a core means of targeted improvement of key economic traits such as crop yield, occupies a pivotal position in agricultural development. Its conventional operation involves hybridizing different types of crop individuals with the same trait to improve the target trait in offspring. For example, when faced with a low-yielding plant like plant A, traditional breeding strategies typically involve hybridizing it with a high-yielding plant B, hoping that through gene recombination, the offspring of the originally low-yielding plant A will possess high-yielding characteristics, thereby achieving the goal of improving the yield trait of plant A.
[0003] However, this process faces numerous challenging limitations. First and foremost is that the individual plants in the improved strain and the plants used for matching must exhibit significant phenotypic differences in the target trait. The occurrence and development of traits are essentially the result of the combined effects of environment and genetic material. However, trait changes induced by environmental factors are often unstable. Taking rice as an example, in different years, due to fluctuations in environmental factors such as day length, rainfall, and temperature, even when planting the same variety of rice, its yield, plant height, grain shape, and other traits will differ. This environmentally induced trait variation is like a fog, making accurately determining trait differences when searching for plants that match the improved individual an extremely challenging task.
[0004] To accurately identify individuals exhibiting differences in target traits, traditional methods rely on large-scale phenotypic statistical work. This means breeders need to meticulously measure and record various traits of a large number of plants, covering aspects such as plant height, leaf morphology and color, and fruit size and quality. Furthermore, given the interference of environmental factors on trait expression, obtaining stable and reliable phenotypic data to accurately determine the degree of trait difference often requires continuous field observation over several years. For example, to determine the differences in lodging resistance of a wheat variety, it may be necessary to plant it under different seasons and climatic conditions, recording lodging under different environmental stresses such as strong winds and rainfall.
[0005] This traditional breeding method is undoubtedly a massive drain on human and material resources. A large amount of manpower is devoted to field management, data measurement and recording, while huge sums of money are spent on land leasing, seed purchases, fertilizers, and pesticides. More importantly, the lengthy breeding cycle becomes a bottleneck restricting agricultural development. From the initial selection and breeding of a new variety to its final market launch, it can take years or even decades, making it difficult for agricultural production to quickly adapt to changes in market demand and to provide farmers with higher-quality, higher-yield crop varieties in a timely manner.
[0006] In summary, the challenges of trait determination in traditional breeding methods for finding individuals with different target traits, as well as the resulting high costs and long cycles, urgently require an innovative solution to overcome these difficulties and propel agricultural breeding technology to new heights. Summary of the Invention
[0007] To address the shortcomings of existing technologies, this invention provides a method, system, and storage medium for finding optimal matching individuals for breeding improvement. By integrating genomics, bioinformatics, and statistical methods, this invention achieves a technological breakthrough in precisely screening breeding-matching individuals at the genotype level, and constructs a quantitative matching model based on functional locus contribution and linkage region characteristics. Compared with traditional breeding methods, the technical advantages of this invention are reflected in three dimensions: First, it overcomes the limitations of phenotypic identification, directly matching based on genotype differences and avoiding interference from environmental factors; second, through PVE value weighting and block difficulty coefficient correction, it achieves a shift from empirical selection to quantitative decision-making, improving matching accuracy; and third, it shortens the breeding cycle and significantly reduces the cost of field trials.
[0008] To achieve the above objectives, the technical solution of the present invention is: a method for finding the best-matched individuals for breeding improvement, the method comprising the following steps: S1. Obtain basic SNP tags; S2. Collect genotype, location, and contribution pve information of known functional loci to establish a functional loci database; S3. Based on the basic SNP tags in step S1, calculate the chain length between basic SNP tags to obtain the location of the chain region block, and establish a block library; S4. Based on the location and genotype information of the functional sites in step S2, obtain the functional sites, dominant genotypes and inferior genotypes contained in the material to be improved, and generate genotype test results. S5. Perform functional site genotypic difference retrieval on the material to be matched, compare the genotypic differences between the material to be matched and the material to be improved, and calculate the matching score between the material to be matched and the material to be improved based on the contribution pve value, block length coefficient and the similarity of base sequences within the block. S6. Based on the total score calculated in step S5, select the material with the highest total score as the best matching individual for breeding the material to be improved.
[0009] Further, step S1 includes the following steps: S101. Obtain WGS sequencing data for all samples; S102. Use FASTP software to perform quality control on WGS sequencing data to obtain valid data; S103. Use BWA software to compare and obtain the bam file; S104. Use GATK software to perform mutation detection to obtain SNP markers; S105. Based on sequencing depth, deletion rate and minimum allele frequency, SNP markers are filtered to obtain basic SNP markers.
[0010] Furthermore, in step S2, the functional site locations of different reference genome versions are uniformly aligned to the same reference genome using BWA software to determine the marked locations.
[0011] Furthermore, in step S3, the PopLDdecay software is used to calculate the chain length between basic SNP tags, and the plink software is used to obtain the location of the chain region block.
[0012] Furthermore, in step S5, the matching score formula for the material to be matched and the material to be improved is: Match score = score of functional sites outside the block + Σ (score of functional sites within each block); Outside-block functional locus score = Σ(dominant difference locus pve) - Σ(disadvantageous difference locus pve); Block functional locus score = [Σ(dominant differential loci pve) - Σ(disadvantageous differential loci pve)] × block difficulty coefficient; Where in the formula Dominant differential loci: The genotypes of the material to be matched and the material to be improved are different, and the material to be matched can introduce functional loci of the dominant genotype; Disadvantageous differential sites: The genotypes of the material to be matched and the material to be improved are different, and the material to be matched will introduce functional sites of the disadvantageous genotype.
[0013] Furthermore, the block difficulty coefficient = (block length coefficient × weight 1) + (similarity of base sequences within the block × weight 2). Block length coefficient = 1 - [(length of each block - minimum block length) / (length of the longest block - minimum block length)]; Similarity of base sequences within a block = Number of identical sites between the material to be matched and the material to be improved within the block ÷ Block length; Weight 1 + Weight 2 = 1.
[0014] Furthermore, the values of weight 1 and weight 2 are both 0.5.
[0015] A system for implementing the above-described method for finding the best-matched individuals for breeding improvement, the system comprising: The data acquisition module is used to acquire basic SNP markers, known functional site information, and block data; The genotyping module is used to detect the genotype of the material to be improved and generate genotyping results. The matching and scoring module is used to calculate the total score of the material to be matched and the material to be improved. The results output module is used to output the best matching individual information based on the total score.
[0016] Furthermore, it also includes a storage module for storing basic SNP markers, functional site information, block data, genotype test results, and total score data.
[0017] A storage medium storing program instructions that, when executed, perform the method described above for finding the best-matched individuals for breeding improvement.
[0018] The beneficial effects achieved by this invention are as follows: 1. This invention directly analyzes the genotypes of functional loci, overcoming the limitations of traditional breeding methods that rely on phenotypic determination and effectively avoiding interference from environmental factors on phenotypic expression. It establishes a functional locus database using known genotypes, locations, and contribution PVE information of functional loci, and combines this with analysis of linkage regions (blocks) to accurately assess the genetic complementarity between materials to be matched and those to be improved. In calculating the matching score, it comprehensively considers the contribution of functional loci, the length coefficient of linkage regions, and the similarity of base sequences. In particular, by introducing a block difficulty coefficient, it scientifically quantifies the ease with which different linkage regions can achieve gene recombination during breeding, making the matching score more objective and accurate.
[0019] 2. This invention can quickly screen out the best matching individuals from a large number of materials to be matched, significantly reducing the consumption of manpower and material resources in the breeding process, greatly shortening the breeding cycle, providing efficient and precise technical support for agricultural breeding, helping to accelerate the cultivation and promotion of superior varieties, and promoting the sustainable development of agricultural production.
[0020] 3. This invention is applicable to the genetic improvement of various crops, and shows strong adaptability, especially in hybridization breeding, backcrossing and gene aggregation. Through the joint analysis of functional sites and block regions, it can effectively avoid unfavorable linkage, improve the transfer efficiency of target traits, enhance the predictability and controllability of breeding design, and improve the accuracy of genetic gain prediction through the synergistic analysis of functional sites and linkage regions.
[0021] 4. This invention innovatively employs a dual calculation logic of positive and negative summation. While statistically analyzing the positive contribution of dominant differential loci (PVE), it simultaneously deducts the negative impact of inferior differential loci (PVE), accurately quantifying the net improvement value of matched materials, effectively reducing breeding losses caused by the transmission of inferior genotypes, and improving matching accuracy. Based on the objective biological law that recombination is easier outside of a block than inside, a block difficulty coefficient is constructed to quantify the difficulty of recombination within the block. Instead of blindly pursuing the breaking of linkage, it prioritizes screening matched materials with low recombination resistance and easy transmission of dominant loci. The block difficulty coefficient incorporates two key dimensions: block length and base similarity. The longer the block, the higher the recombination difficulty; the higher the base sequence similarity, the lower the recombination probability. Through dual-dimensional weighted calculation, it comprehensively reflects the true impact of blocks on recombination. Through a clear mathematical formula (matching score = external functional locus score + Σ internal functional locus score) and standardized parameters, parental selection is transformed from experience-driven to data-driven, promoting the standardization and intelligent upgrading of breeding technology.
[0022] 5. Through modular system design, this invention achieves full automation of the entire process from data acquisition, genotype analysis, matching scoring to result output, thereby improving the standardization and intelligence of breeding decisions. Combined with expandable storage modules and a programmed instruction execution environment, it is easy to promote and apply in different crops and populations, and has good compatibility and practicality. Attached Figure Description
[0023] Figure 1 This is a flowchart illustrating a method for finding the best-matched individuals for breeding improvement. Detailed Implementation
[0024] To better understand the purpose, structure, and function of this invention, the following detailed description, in conjunction with the accompanying drawings, provides a method, system, and storage medium for finding the best-matched individuals for breeding improvement.
[0025] like Figure 1 As shown, a method for finding the best-matched individuals for breeding improvement includes the following steps: S1. Obtain basic SNP markers. Basic SNP markers serve as reference points for the genome and are used for subsequent analysis and calculations. They are the basic data for constructing linkage regions and evaluating functions. S2. Collect detailed information on known functional loci, including their genotype, specific location in the genome, and contribution (pve) data. Based on this information, establish a functional loci database to provide a comprehensive reference for subsequent identification and evaluation. S3. Based on the basic SNP markers obtained in step S1, calculate the linkage length between basic SNP markers to determine the location of linked region blocks in the genome, and establish a corresponding block library to support subsequent matching and similarity analysis. S4. Based on the functional site library established in step S2, combined with the genotype data of the material to be improved, identify the functional sites carried by the material to be improved, distinguish between dominant and inferior genotypes, generate a genotype health check report, and provide a basis for breeding selection. S5. Perform a systematic search for functional site genotypic differences in the materials to be matched, compare the genotypes of the materials to be matched and the materials to be improved, and comprehensively consider the contribution value (pve) of the functional site, the length coefficient of the linkage region block, and the similarity of the base sequence inside the block. The matching score between the two is calculated by weighting. S6. Based on the total matching score obtained in step S5, the higher the matching score, the stronger the complementary potential between the material to be matched and the material to be improved in terms of target trait improvement. Select the material to be matched with the highest score from the candidate materials and determine it as the best breeding match for the material to be improved, so as to provide a scientific basis for subsequent breeding strategies.
[0026] Step S1 includes the following steps: S101. Obtain WGS sequencing data for all samples. The relevant data comes from the high-throughput sequencing results of different samples and serves as the basic input material for subsequent analysis. S102. Use FASTP software to perform quality control processing on the acquired WGS sequencing data, including removing low-quality sequences, adapter contamination, and trimming low-quality bases at the ends, to finally obtain high-quality and effective sequencing data, providing a reliable guarantee for subsequent analysis. S103. Use BWA software to align the quality-controlled valid sequencing data with the reference genome. Through sequence matching and localization processing, generate a bam file containing alignment information. The bam file records the specific location of each sequencing fragment on the genome. S104. Use GATK software to perform variant detection and analysis, and obtain SNP markers by identifying single nucleotide polymorphism (SNP) gene variants. S105. Based on sequencing depth, deletion rate, and minimum allele frequency, the detected SNP markers are screened and filtered to remove unreliable or low-quality sites, resulting in high-quality and reliable basic SNP markers for subsequent genetic analysis or related research.
[0027] In step S2, BWA software is used to perform a unified alignment process for the functional site locations in different reference genome versions, accurately mapping all sites to the same standard reference genome. This effectively determines the specific location information of each marker, ensuring the consistency and comparability of subsequent analyses and providing reliable basic data support for subsequent gene function studies and variation analysis.
[0028] In step S3, PopLDdecay software is used to perform a detailed analysis of the linkage degree between basic SNP markers, calculate and plot linkage decay curves to determine the linkage length between basic SNP markers, and use plink software to further process and analyze the data, identify and locate linkage region blocks in the genome, and accurately extract and record the location information of linkage region blocks.
[0029] In step S5, the matching score formula for the material to be matched and the material to be improved is: Match score = score of functional sites outside the block + Σ (score of functional sites within each block); Outside-block functional locus score = Σ(dominant difference locus pve) - Σ(disadvantageous difference locus pve); Block functional locus score = [Σ(dominant differential loci pve) - Σ(disadvantageous differential loci pve)] × block difficulty coefficient; Block difficulty coefficient = (block length coefficient × weight 1) + (similarity of base sequences within the block × weight 2). Block length coefficient = 1 - [(length of each block - minimum block length) / (length of longest block - minimum block length)] Similarity of base sequences within a block = Number of identical sites between the material to be matched and the material to be improved within the block ÷ Block length; Weight 1 + Weight 2 = 1; Where in the formula Dominant differential loci: The genotypes of the material to be matched and the material to be improved are different, and the material to be matched can introduce functional loci of the dominant genotype; Disadvantageous differential sites: The genotypes of the material to be matched and the material to be improved are different, and the material to be matched will introduce functional sites of the disadvantageous genotype; The values of weight 1 and weight 2 are both 0.5.
[0030] A system for finding the best-matched individuals for breeding improvement includes a data acquisition module, a genotyping module, a matching scoring module, and a results output module. These modules work collaboratively to automate the entire process from raw SNP data to matching individual selection. The data acquisition module acquires basic SNP markers, known functional loci information, and block data, standardizing SNP information generated from different platforms. The genotyping module performs genotyping of functional loci based on a reference genome, identifying advantageous and disadvantageous loci between the material to be improved and the material to be matched, generating genotyping results. The matching scoring module, based on the genotyping results and combining block structure information and functional locus PVE values, calculates a comprehensive matching score for each material to be matched and the material to be improved according to a set formula, prioritizing the individual with the highest score as the optimal match. The results output module presents key parameters such as matching scores, distribution of differential loci, and intra-block similarity in a visual report, supporting more precise and efficient breeding decisions.
[0031] This system's modules work collaboratively to automate the entire process from raw SNP data to matching individual screening. Standardized workflows ensure the reproducibility and accuracy of analysis results, support batch processing of multiple materials, and improve breeding decision-making efficiency. After users input genotype data for materials to be matched and those to be improved, the system automatically performs locus alignment, linkage block identification, and functional locus effect assessment. Based on the matching score formula, it calculates the optimal combination and outputs recommended individuals and their corresponding scores, providing precise guidance for molecular design breeding. The matching scoring module introduces a block difficulty coefficient, fully considering the impact of linkage region length and sequence conservation on recombination difficulty, making the scoring more closely reflect actual breeding needs. The results output module generates a visual report including matching score ranking, a heatmap of dominant locus distribution, and similarity curves within blocks, facilitating breeders' intuitive assessment of the feasibility of potential improvement schemes. The system supports batch processing of thousands of candidate materials, making it suitable for parental screening and aggregation breeding design in large-scale breeding projects, providing strong technical support for modern molecular breeding. The system employs a distributed computing architecture, significantly improving data processing speed and stability, and can update the latest functional locus database in real time, ensuring the cutting-edge nature and accuracy of the analysis results.
[0032] A storage medium stores program instructions that, when executed by a processor, enable the processor to implement the aforementioned method for finding the best-matched individuals for breeding improvement, including the entire process of data acquisition, genotypic examination, matching scoring, and result output. The storage medium is a non-transitory computer-readable storage medium, supporting local deployment and cloud access, ensuring the compatibility and reproducibility of the method under different computing environments. The program instructions, through a modular design, enable the orderly invocation of each analysis step, ensuring efficient and stable data flow. During execution, the system can automatically verify the integrity of input data and dynamically adjust parameter configurations to adapt to the genomic characteristics of different species, improving the method's versatility. It also supports user-defined functional site weights and scoring thresholds to meet diverse breeding objectives. The entire process requires no manual intervention, significantly reducing operational complexity and providing a standardized and intelligent solution for genome-assisted breeding. The system has good scalability, supporting the integration of phenotypic data and environmental factor information to further optimize the predictive capabilities of the matching model. Through a continuous iterative learning mechanism, the system can adjust the scoring algorithm based on historical breeding performance feedback, improving recommendation accuracy. All analysis processes follow standardized procedures, ensuring that results are traceable and verifiable.
[0033] Example 1 The specific process of this embodiment is as follows: Genotypic data of the materials to be improved and candidate materials were obtained. All WGS sequencing data were subjected to strict quality control using FASTP software to ensure high-quality and effective sequencing data. BWA software was used to align the high-quality data with the reference genome to generate corresponding BAM format alignment files. GATK software was used to perform variant detection on the BAM files to identify single nucleotide polymorphism (SNP) markers. By setting a series of filtering criteria, including key parameters such as sequencing depth, deletion rate, and minimum allele frequency, the initially detected SNP markers were screened to obtain a set of high-quality and reliable basic SNP markers for subsequent genetic analysis research.
[0034] During the acquisition of functional loci, detailed information on known functional loci needs to be collected, including their corresponding genotype, pve, and location key data. For locus location information from different reference genome versions, BWA software is used for unified processing. By comparing it with the same standard reference genome, the accuracy and consistency of the location information of all markers are ensured. Table 1 shows the information of known functional loci for reference and practical use.
[0035] Table 1 Information on known functional sites
[0036] Note: Chrom: The chromosome number where the functional site is located; Pos: Functions to indicate the physical location of a point on the genome; Adv: The dominant genotype of a functional locus (the genotype with better phenotypic expression). Phe: The phenotype associated with this functional site; Pve: The contribution of this functional site to its associated phenotype.
[0037] During block data acquisition, PopLDdecay and plink software were used for linkage analysis. Based on the basic SNP marker data obtained in the preceding steps, the length and location of the blocks were determined by calculating the decay rate of linkage disequilibrium. Specifically, PopLDdecay software was used to assess the trend of linkage disequilibrium with physical distance, while plink software was used to identify and delineate specific block regions. This process allows for accurate acquisition of the start and end positions of each block on the genome. The generated block files contain detailed block information, as shown in Table 2, which illustrates the chromosomal location, start and end points, and related statistical indicators for each block.
[0038] Table 2 Block Information
[0039] Note: CHROM: Chromosome number where the block is located; START: The starting position of the block; END: The end position of the block.
[0040] Genotyping, based on the specific location of functional loci in the genome and their corresponding genotype information, obtains various functional loci contained in the biological material to be improved, identifies dominant and suboptimal genotypes in the material, and thus provides a scientific basis and data support for subsequent genetic improvement and breeding work. The results of genotyping are usually presented in document form, recording the detected information. Table 3 provides an example of a genotyping document for reference.
[0041] Table 3. Examples of Genotyping Test Results
[0042] Note: Phe: Target trait to be improved Sample: ID of the sample to be improved; Chr: Chromosome ID; Pos: The physical location of the functional site; Beneficial_allele: The dominant genotype for the target trait at this functional locus; Genotype: The genotype of the material to be improved at this location; Flag: Is the material to be improved the dominant genotype at this position? Yes; no.
[0043] During parental mating, genotypic differences at functional sites are searched in the materials to be matched, and sites that differ from the material to be improved in genotypic form at functional sites are considered. Specifically: for functional sites outside the block, sites with genotypic differences that can introduce a dominant genotype are positively summed based on PVE values, while sites with differences that can introduce a suboptimal genotype are negatively summed based on PVE values. Simultaneously, block length and the similarity of base sequences within the block are considered. Therefore, for functional sites located within the block, the sum of PVE values is multiplied by the product of the block length coefficient (the value after standardizing all block lengths) and its weight, and by the product of the similarity of base sequences within the block (the number of identical sites within the block / block length) and its weight (the weights of both are summed to 1, with a default value of 0.5 for each). After these calculations, each material to be matched receives a score. The higher the score, the greater the potential for improvement in the offspring after hybridization with the material to be improved. Example files are shown in Table 4.
[0044] Table 4. Examples of Matching Score Results
[0045] Note: Sample_D: ID of the material to be improved Sample_H: The material to be matched. Score: Matching score between the material to be matched and the material to be improved. It is understood that the present invention has been described through some embodiments, and those skilled in the art will recognize that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the invention. Furthermore, under the teachings of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the protection scope of the present invention.
Claims
1. A method for finding the best-matched individuals for breeding improvement, characterized in that, The method includes the following steps: S1. Obtain basic SNP tags; S2. Collect genotype, location, and contribution pve information of known functional loci to establish a functional loci database; S3. Based on the basic SNP tags in step S1, calculate the chain length between basic SNP tags to obtain the location of the chain region block, and establish a block library; S4. Based on the location and genotype information of the functional sites in step S2, obtain the functional sites, dominant genotypes and inferior genotypes contained in the material to be improved, and generate genotype test results. S5. Perform functional site genotypic difference retrieval on the material to be matched, compare the genotypic differences between the material to be matched and the material to be improved, and calculate the matching score between the material to be matched and the material to be improved based on the contribution pve value, block length coefficient and the similarity of base sequences within the block. S6. Based on the total score calculated in step S5, select the material with the highest total score as the best matching individual for breeding the material to be improved.
2. The method for finding the best-matched individual for breeding improvement according to claim 1, characterized in that, Step S1 includes the following steps: S101. Obtain WGS sequencing data for all samples; S102. Use FASTP software to perform quality control on WGS sequencing data to obtain valid data; S103. Use BWA software to compare and obtain the bam file; S104. Use GATK software to perform mutation detection to obtain SNP markers; S105. Based on sequencing depth, deletion rate and minimum allele frequency, SNP markers are filtered to obtain basic SNP markers.
3. The method for finding the best-matched individual for breeding improvement according to claim 1, characterized in that, In step S2, the functional site locations of different reference genome versions are uniformly aligned to the same reference genome using BWA software to determine the location of the marker.
4. The method for finding the best-matched individual for breeding improvement according to claim 1, characterized in that, In step S3, the PopLDdecay software is used to calculate the linkage length between basic SNP tags, and the plink software is used to obtain the location of the linkage region block.
5. The method for finding the best-matched individual for breeding improvement according to claim 1, characterized in that, In step S5, the matching score formula for the material to be matched and the material to be improved is: Match score = score of functional sites outside the block + Σ (score of functional sites within each block); Outside-block functional locus score = Σ(dominant difference locus pve) - Σ(disadvantageous difference locus pve); Block functional locus score = [Σ(dominant differential loci pve) - Σ(disadvantageous differential loci pve)] × block difficulty coefficient; Where: Dominant differential loci: The genotypes of the material to be matched and the material to be improved are different, and the material to be matched can introduce functional loci of the dominant genotype; Disadvantageous differential sites: The genotypes of the material to be matched and the material to be improved are different, and the material to be matched will introduce functional sites of the disadvantageous genotype.
6. The method for finding the best-matched individual for breeding improvement according to claim 5, characterized in that, Block difficulty coefficient = (block length coefficient × weight 1) + (similarity of base sequences within the block × weight 2). Block length coefficient = 1 - [(length of each block - minimum block length) / (length of the longest block - minimum block length)]; Similarity of base sequences within a block = Number of identical sites between the material to be matched and the material to be improved within the block ÷ Block length; Weight 1 + Weight 2 = 1.
7. The method for finding the best-matched individual for breeding improvement according to claim 6, characterized in that, The values of weight 1 and weight 2 are both 0.
5.
8. A system for implementing a method for finding the best-matched individuals for breeding improvement as described in any one of claims 1 to 7, characterized in that, The system includes: The data acquisition module is used to acquire basic SNP markers, known functional site information, and block data; The genotyping module is used to detect the genotype of the material to be improved and generate genotyping results. The matching and scoring module is used to calculate the total score of the material to be matched and the material to be improved. The results output module is used to output the best matching individual information based on the total score.
9. The system for finding the best-matching individuals for breeding improvement according to claim 8, characterized in that, It also includes a storage module for storing basic SNP markers, functional site information, block data, genotype test results, and total score data.
10. A storage medium storing program instructions, characterized in that, When the program instructions are executed, they perform a method for finding the best-matched individuals for breeding improvement as described in any one of claims 1 to 7.