A Method for Constructing Mendelian Mismatch Matrices Based on Mendel's Law of Separation

By constructing a Mendelian mismatch matrix based on Mendel's law of segregation, the problems of model dependence, missing data, and subjectivity of grading thresholds in genetic evaluation of crop germplasm resources were solved, thus achieving objective grading of germplasm resources and fair evaluation of genetic differences.

CN122493948APending Publication Date: 2026-07-31INST OF FRUIT & TEA HUBEI ACAD OF AGRI SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INST OF FRUIT & TEA HUBEI ACAD OF AGRI SCI
Filing Date
2026-05-08
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies in crop germplasm resource banks suffer from model dependence, difficulties in handling missing data, and subjectivity in grading thresholds, leading to biases in genetic assessment and subjective results.

Method used

Based on Mendel's law of segregation, a Mendelian mismatch matrix is ​​constructed. By converting the allele set format, constructing the Mendelian mismatch matrix and the effective locus statistical matrix, and using the binary vector inner product to calculate the mismatch relationship, genetic classification without the need for a preset threshold is achieved.

Benefits of technology

This approach enables objective and reliable diversity assessment and grading of germplasm resources, eliminates the impact of missing sequencing data, and ensures the fairness of genetic difference assessment and the reproducibility of results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493948A_ABST
    Figure CN122493948A_ABST
Patent Text Reader

Abstract

This invention relates to the field of Mendelian research and discloses a method for constructing a Mendelian mismatch matrix based on Mendel's law of segregation. The method includes the following steps: obtaining codominant marker genotype data of individuals in a target germplasm resource population; preprocessing the codominant marker genotype data; determining locus mismatches based on the intersection of allele sets, and constructing a mismatch matrix and an effective locus statistical matrix representing pairwise correspondences among the data samples; performing pairwise element-level normalization correction on the original Mendelian mismatch matrix based on the effective locus statistical matrix, converting the absolute mismatch number into a relative mismatch rate in the 0-1 interval, ultimately obtaining an unbiased Mendelian mismatch matrix; analyzing the Mendelian mismatch matrix data to divide the target germplasm resource population into multiple levels, thereby achieving diversity evaluation and hierarchical screening of the target germplasm population. This invention compensates for missing genotype data and efficiently obtains germplasm resource evaluation data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioinformatics, and more specifically to a method for constructing a Mendelian mismatch matrix based on Mendel's law of segregation. Background Technology

[0002] In the construction, management, and molecular breeding research of crop germplasm banks, accurately assessing genetic diversity within populations, identifying redundant and repetitive germplasm, and elucidating parental derivation relationships are core tasks for maintaining the efficient operation of germplasm banks. With the widespread adoption of high-throughput sequencing technology and codominant molecular markers such as simple sequence repeats (SSRs), single nucleotide polymorphisms (SNPs), and insertion-deletion markers (InDels), researchers can obtain large-scale population genotypic data at a lower cost. However, traditional population genetic analysis methods face the following technical bottlenecks:

[0003] First, there is the problem of model dependency. Existing methods, such as kinship matrix (GRM) and Structure Bayes clustering, all rely on population genetics assumptions such as Hardy-Weinberg equilibrium and random mating. However, actual crop germplasm resource banks are composed of complex artificial hybrid pedigrees, geographical isolation, and asexual reproduction offspring, which do not conform to the above theoretical premises, leading to systematic evaluation bias.

[0004] Second, there is the challenge of handling missing data. Due to the limited quality of DNA samples in actual sequencing, a large amount of missing data is inevitably generated. Traditional algorithms either directly remove sites with missing values, resulting in different actual site numbers and dimensional imbalances for different germplasm pairings; or they use population mean imputation, which introduces false genetic information. The industry urgently needs a germplasm evaluation method that does not rely on population model assumptions and can reasonably compensate for missing genotype data.

[0005] Third, there is the issue of subjectivity in the grading threshold. Existing genetic distance clustering methods require manual pre-setting of the number of groups or distance thresholds, and the results change with parameter adjustments, lacking objective reproducibility.

[0006] Mendel's law of segregation is a fundamental axiom of genetics. In the sexual reproduction of diploid organisms, each allele in the offspring must have one and only one allele from the father and one from the mother. Therefore, it can be deduced that real direct parents and offspring must share at least one allele at any codominant marker site that conforms to Mendel's inheritance. It is impossible for alleles to be completely non-overlapping. This invention is strictly based on this rigid constraint to construct a Mendel mismatch matrix system that combines mathematical certainty, data missing compensation capability, and hierarchical endogeneity. Summary of the Invention

[0007] The purpose of this invention is to provide a method for constructing a Mendelian mismatch matrix based on Mendel's law of separation, in order to solve the above-mentioned problems.

[0008] The above-mentioned technical objective of the present invention is achieved through the following technical solution:

[0009] In a first aspect, the present invention provides a method for constructing a Mendelian mismatch matrix based on Mendel's law of separation, comprising the following steps:

[0010] S1. Obtain the codominant marker genotype data of sample individuals in the target germplasm resource population, wherein the codominant markers include SSR, SNP, and InDel;

[0011] S2. Preprocess the codominant marker genotype data to convert the genotype data into an allele set format, wherein homozygotes are converted into single-element sets, heterozygotes are converted into two-element sets, and missing sites are converted into empty sets.

[0012] S3, Construction The elements of a Mendelian mismatch matrix M of order M are... Defined as the number of sites for which the intersection of the allele sets of two qualities i and j at all sites is empty and there are no missing alleles;

[0013] S4. Construct the effective site statistical matrix L, whose elements This represents the number of valid loci for non-deletion alignments between two types of samples.

[0014] S5. Division by elements = / Obtain the standardized mismatch matrix M' and set the effective site threshold. >0.8 As a criterion for determining statistical validity;

[0015] S6. Analyze the mismatch matrix data and determine the minimum mismatch number based on matrix M. Perform germplasm Layered genetic grading, when At that time, germplasm i belongs to The grading is endogenously determined by the population's genetic structure, without the need for a preset threshold, thus enabling diversity evaluation and grading screening of the target germplasm population.

[0016] S7, to Grade germplasm execution Two-step judgment: First, filter =0 Mendelian compatible pairings, then verify their mismatch spectrum consistency—if for all k≠i,j, the following holds true: = If the genotype is identical, it is determined to be a duplicate genotype; otherwise, a parent-child relationship is suspected.

[0017] A further feature of the present invention is that the construction of the Mendelian mismatch matrix in step S3 is accelerated by binary vector inner product, the allele set is mapped to binary vector, and the rigid mismatch at the site level is determined by vector inner product operation, thereby reducing the algorithm complexity from O(N²LA) to O(N²L), where A is the average string length of the allele.

[0018] A further configuration of the present invention is as described in step S6. Genetic stratification satisfies mutual exclusion, exhaustiveness, and endogeneity, and defines cumulative stratification. Hierarchical descriptions used for genetic specificity.

[0019] A further setting of the present invention is: the step S7 described The condition for consistency of mismatch patterns in the two-step determination is: the two germplasms not only match each other... =0, and is completely consistent with the mismatch pattern of all other germplasms in the population, which is the same genotype in a set sense.

[0020] A further provision of the present invention is that the method for constructing the Mendelian mismatch matrix in S3 includes:

[0021]

[0022] in This represents the total number of codominant markers at site l. This represents the mismatch matrix of germplasm individuals i and j at site l;

[0023]

[0024] in and This refers to the set of alleles of individuals i and j at locus l.

[0025] A further provision of the present invention is that the method for constructing the effective site statistical matrix in S3 includes:

[0026]

[0027] in, This is an indicator function that takes the value 1 when the condition is met and 0 when the condition is not met.

[0028] A further provision of the present invention is that the method for constructing the normalized mismatch matrix in S4 includes: based on the constructed effective site statistics matrix. For mismatch matrix Corrections were made to eliminate the bias in results caused by differences in the number of effective marker sites among different germplasms:

[0029]

[0030] Only when the number of effective loci between paired germplasms meets the requirement Only when the condition is met will the correction result of the pair be determined to have statistical power and retained; otherwise, an invalid message will be returned.

[0031] A further provision of the present invention is that the Mendelian mismatch matrix and the effective site statistics matrix also need to undergo compliance verification, and the determination method includes:

[0032] S1, Symmetry: The matrix is ​​a strictly symmetric matrix, and all germplasm pairings satisfy... ;

[0033] S2, Boundedness: All off-diagonal elements satisfy... , No value exceeding the total number of markers;

[0034] S3. Non-negative integer property: All off-diagonal elements are non-negative integers, with no floating-point numbers, negative numbers, or invalid or missing values;

[0035] S4, Missing Regularity: The diagonal elements of the matrix are set to placeholders NA, which represent no cross-comparison meaning, and have no valid values. The diagonal settings are consistent.

[0036] The compliance verification passes only if the mismatch matrix and the effective site statistics matrix satisfy all of the above determination methods; otherwise, an invalid message is returned.

[0037] The present invention also provides a germplasm resource evaluation and grading system based on a full-population mismatch matrix, the system comprising modules for implementing the method of any one of claims 1 to 7:

[0038] The data preprocessing module is used to receive population-level codominant marker genotype data and perform allele set format conversion and generate binary encoding matrices.

[0039] A matrix acceleration construction module is used to output a mismatch matrix and an effective site statistics matrix according to the formula using the binary encoding matrix.

[0040] The evaluation and grading output module incorporates a standardized correction algorithm, matrix compression strategy, compliance verification mechanism, and multi-level mismatch rate analysis threshold, which is used to perform highly reliable redundancy screening and diversity grading of the entire population of germplasm resources.

[0041] The present invention also provides a computer-readable storage medium having stored thereon a computer instruction program executable by a processor, characterized in that the program, when executed, performs the method as described in any one of claims 1 to 7.

[0042] The beneficial effects of this invention are:

[0043] 1. Based on Mendel's law of segregation, a mismatch matrix is ​​constructed to... The dimensional matrix form represents a large number of mismatch relationships between pairs of germplasm, forming a mismatch lineage of each germplasm individual with respect to all other individuals in the population. This allows researchers to quickly classify the germplasm resources to be explored through the mismatch lineage.

[0044] 2. By constructing an effective site statistical matrix This process allows for the construction of a standardized mismatch matrix. By correcting for errors caused by measurement failures through element division between matrices, the impact of missing sequencing data on the results is avoided, ensuring the fairness of the evaluation of genetic differences among germplasms. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0046] Figure 1 This is a flowchart of a method for constructing a Mendelian mismatch matrix based on Mendel's law of separation, according to the present invention.

[0047] Figure 2 This is a table of original parameters for a demonstration germplasm population provided in Embodiment 2 of the present invention;

[0048] Figure 3 This is a polymorphism parameter table for the reserved sites provided in Embodiment 4 of the present invention;

[0049] Figure 4 The apple germplasm is implemented using a genetic grading table as provided in Embodiment 4 of the present invention. Detailed Implementation

[0050] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” used herein may also mean including the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in the specific order described or illustrated, unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.

[0051] The technical solution of the present invention will be clearly and completely described below with reference to specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0052] Example 1: Preprocessing of Mendel's mismatched matrix

[0053] The target germplasm resource population was collected using conventional techniques employed by those skilled in the art, such as whole-genome sequencing, targeted capture sequencing, or traditional PCR. The original genotype dataset consists of N diploid individuals and includes, but is not limited to, microsatellite DNA markers, single nucleotide polymorphisms, or CAPS markers. Since the original data is usually stored in the form of text strings such as "A / T" or "101 / 105", it is not only inefficient for computers to directly process strings, but also difficult to perform algebraic operations. Therefore, this invention converts the data into a data type that is easy for computers to read and operate on by converting the allele set format.

[0054] For diploid organisms, whose genotype at any l-th marker locus is composed of two alleles on two homologous chromosomes, the text is first converted into a set according to the following rules. :

[0055] If an individual is homozygous at a certain locus, such as "101 / 101" or "A / A", it is converted into a set containing only one element, such as {101} or {A}; if it is heterozygous, such as "101 / 105" or "A / T", it is converted into a two-element set, such as {101,105} or {A,T}; if the locus cannot be genotyped or the data is incorrect due to sequencing quality issues, it is uniformly converted into an empty set with no elements. And it is marked as NA in the encoding.

[0056] Subsequently, the ensemble data is converted into a binary encoding matrix. Assuming that in the entire population of N individuals, the l-th locus is detected with a total of [number missing] [items missing]. Different alleles constitute the allele set at that locus. .

[0057] The system will include individuals The set at this site Mapped to a length of binary row vectors The mapping method is: traversing the entire set. Each allele in (k ranges from 1 to n), if Then the vector The k-th element is assigned the value 1; if If the value is 0, it is assigned a value of 0; if the data is missing and the set is empty, it is directly mapped to a vector of all zeros. .

[0058] Subsequently, a mismatch matrix and a valid locus statistical matrix are obtained through matrix operations. Mendel's laws state that in diploid offspring, the two alleles at any valid gene locus must be inherited from their respective parents. Therefore, if two germplasms do not share any alleles at a certain codominant locus, it can be directly determined that these two germplasms are Mendelian mismatches. In the binary vector space constructed in this invention, the operation of set intersection is transformed into the calculation of the inner product of two vectors. Let... Let be the inner product of the vectors corresponding to germplasm i and germplasm j at the l-th position, then we have Based on this formula, the mismatch matrix of germplasm individuals i and j at site l can be defined. With effective site statistics submatrix Relationship:

[0059] 1. Invalid comparison: If or A vector consisting entirely of zeros indicates that the pair encountered missing data at this site, rendering the alignment invalid. It neither generates a mismatch nor accumulates as a valid site; in this case, a value is assigned. =0, ;

[0060] 2. Absolute mismatch: If two vectors are not all zero, but their inner product is exactly the same. This means that the two have no overlap in the same dimension, that is, they do not share any alleles. In this case, a rigid mismatch occurs, and the assignment is... =1, 1;

[0061] 3. Compatibility Status: If both vectors are not all zero vectors, and their inner product is... This indicates the existence of shared alleles, consistent with the possible transmission patterns of Mendelian inheritance, and assigns a value. =0, 1.

[0062] Traversing all germplasm For combinations and all After identifying each site, the corresponding mismatch matrix and valid site statistics matrix are output through addition: , To maintain the integrity of matrix operations and avoid meaningless data interference from diagonal elements representing comparisons between germplasm individuals and themselves, all diagonal elements of the matrix are assigned invalid placeholders. .

[0063] Example 2

[0064] To facilitate understanding of the entire algebraic transformation and derivation process described above, this embodiment constructs a low-dimensional but complete demonstration population, which contains only 4 germplasm accessions, denoted as [missing information]. Three codominant marker genotype loci were detected and set as [missing information]. Each locus contains 2 to 4 alleles; see details below. Figure 2 Based on the aforementioned rules, the text is converted into a set, and then mapped into a binary vector according to the order. Based on the above calculation rules, it can be deduced that... and The pairing data is as follows:

[0065] exist site, This indicates that there are shared genes and no mismatch occurs. =0, 1;

[0066] exist site, This indicates that there are shared genes and no mismatch occurs. =0, 1;

[0067] exist site, This indicates that there are shared genes and no mismatch occurs. =0, 1;

[0068] Summing yields the representations of primes i and j in the mismatch matrix. In the effective site statistics matrix ;

[0069] For datasets without shared genes, such as Because of the mismatch, therefore =1, 1;

[0070] By analogy, all pairwise pairing data of germplasm can be calculated, generating the mismatch matrix and the statistical matrix of effective loci as follows:

[0071]

[0072] If only the matrix is ​​observed ,because It may be mistaken for germplasm Both with They are suspected to be highly related, and also... They are suspected to be highly related, but if the matrix is ​​compared... You will find that, due to Genotype deletion exists. and The number of basic comparisons is less than and The lack of a clear basis for comparison leads to unfairness in the comparison. To ensure that the corrected results are statistically significant, a validation threshold needs to be set, allowing the comparison to proceed only if the number of valid loci between each pair of paired germplasms meets the required standard. The system only accepts and records the correction results of a pair when the effective alignment rate accounts for 80% or more of the total number of gene markers; for those pairs below the threshold, the system marks them as low confidence and prompts researchers to supplement sequencing data.

[0073] After construction is complete Afterwards, compliance checks are required to prevent errors caused by computer problems such as memory overflow or floating-point calculations during the computation of large amounts of data. Only matrices that pass all checks can be considered reliable.

[0074] 1. Symmetry test: Matrix and It must be a matrix that is strictly symmetric along its diagonal, and all germplasm pairings must satisfy the following condition. This ensures that relationship evaluation is two-way;

[0075] 2. Value range compliance check: All off-diagonal elements satisfy... , The value should not exceed the total number of markers to avoid errors caused by memory overflow.

[0076] 3. Non-negative integer check: All off-diagonal elements are non-negative integers, with no floating-point numbers, negative numbers, or invalid missing values.

[0077] 4. Missing rule check: All diagonal elements of the matrix are set to placeholders NA, indicating no valid values. The diagonal settings are consistent to prevent falling into an infinite loop during the calculation process.

[0078] Example 3

[0079] To verify the suitability and application value of the computational method of this invention in real agricultural biological big data, this section presents an implementation example using local pear varieties from Henan Province. This example utilizes 12 SSR loci to genotype 199 local pear varieties from Henan Province. After removing duplicates, genetic diversity analysis is conducted based on 128 unique genotypes. The aim is to verify the suitability of the method of this invention in the evaluation and management of germplasm resource diversity. The detailed execution steps and output report of the computer system are as follows:

[0080] Step 1: Import the genotype data of 12 SSR markers from 199 pear resources into the evaluation system of this invention according to the required format;

[0081] Step Two: The system maps allele information containing text fragments such as "120 / 124" and "135 / 135" into a high-dimensional binary vector matrix. Because it departs from the traditional character-by-character search and comparison mechanism, it instead utilizes the matrix inner product instruction set of the computer processor, reducing the complexity of matrix construction from... Reduced to This can speed up computer processing.

[0082] Step 3: Finally, the compliance verification matrix is ​​parsed, and an audit report is output: This report is displayed for the matrix... In the mismatch verification, a total of 71 samples were found to be duplicates, forming 29 duplicate groups. For example, the wild sand pear series formed the largest duplicate group of 21 samples. After manual confirmation, these individuals can be merged for management, which can release a large amount of greenhouse and nursery resources.

[0083] After removing redundancy and obtaining a total of 128 unique genotypes, a total of 35 pairs of potential high-confidence parent-offspring derivation relationships with zero mismatches were obtained, providing molecular evidence for the artificial domestication history of sand pear;

[0084] After removing redundancy and obtaining a total of 128 unique genotypes, 50 G2 or D2+ grade germplasms that do not depend on the lineage of current mainstream cultivated germplasms were screened and purified for those that met the filtering conditions. These germplasms were verified by Mendel's laws under natural conditions to have special gene sources or genetic bases and can serve as important materials for the next generation of pear resource breeding.

[0085] Example 4

[0086] This embodiment uses apple germplasm genotype data provided by the German Fruit Tree Germplasm Bank Public Data Warehouse to verify the applicability, computational efficiency, and biological reliability of the Mendelian mismatch matrix construction method and germplasm genetic grading system described in this invention on large-scale real datasets. The original data covers the genotyping results of 1404 apple germplasms based on 17 nuclear genome SSR markers. All markers have been identified as core markers by the European Plant Genetic Resources Cooperation Program Apple / Pear Working Group. This standardized marker system provides a unified genotype language for cross-center comparisons.

[0087] According to the aforementioned rules of this invention, the original genotype data is preprocessed as follows:

[0088] 1. Triploid and higher polyploid genotypes are excluded. Polyploid germplasm has complex allele dosage effects and does not meet the rigid determination premise of Mendel's law of segregation for diploids in this invention.

[0089] 2. The frequency of invalid alleles at each locus was inferred using methods known to those skilled in the art, and some loci were eliminated. The frequencies of invalid alleles at these loci all exceeded the quality control threshold, which may lead to systematic false mismatches.

[0090] 3. Preserve marker combinations with no significant linkage disequilibrium between loci to ensure that each locus provides independent genetic information;

[0091] Ultimately, 1085 diploid accessions were preserved, with complete genotypes for 15 SSR loci. The missing data rate was 0.055%, and the marker data quality met the input requirements of this invention. The polymorphism parameters of the 15 preserved loci are as follows: Figure 3 ;

[0092] Subsequently, the genotypes of 15 loci from 1085 germplasms were converted into allele sets and further mapped to binary vectors. For any pairing (i, j), locus compatibility was calculated using vector inner product operations. After traversing all 588070 germplasm combinations and all 15 loci, a mismatch matrix M and a valid locus statistics matrix L were output, with the diagonal elements uniformly assigned the value NA. The construction of the 1085×1085×15 dimension mismatch matrix and valid locus statistics matrix took only 3.55 seconds, with a time complexity of O(N²L) and a space complexity of O(N²). This achieved batch parallel computation of binary vector inner products, avoiding the O(N²LA) complexity caused by character-by-character string comparison, and satisfying the requirements of... Compliance verification confirms the data is valid;

[0093] Furthermore, according to the rules in S6 of this invention, genetic grading is performed on apple germplasm. Figure 4 The grading results present a pyramid structure. The fact that 73.7% of the varieties are of a certain grade reflects that in the history of apple breeding, modern cultivated varieties originated from a limited number of wild ancestors, and after thousands of years of artificial selection, the gene pool has become highly homogenized.

[0094] Example 5

[0095] To support the execution of the aforementioned algorithm, this invention also provides a germplasm resource evaluation and grading system based on a mismatch matrix. This system architecture is deployed in a high-performance scientific computing cluster or a distributed cloud computing environment, and includes the following functional modules in its physical and logical architecture:

[0096] The data preprocessing module, as the data front-end, is used to receive population-level codominant marker genotype data and perform allele set format conversion and generate binary encoding matrices.

[0097] A matrix acceleration construction module is used to output a mismatch matrix and an effective site statistics matrix according to the formula using the binary encoding matrix.

[0098] The evaluation and grading output module incorporates a standardized correction algorithm, matrix compression strategy, compliance verification mechanism, and multi-level mismatch rate analysis thresholds. This allows for highly reliable redundancy screening and diversity grading of the entire germplasm population.

[0099] Furthermore, this invention is not only embodied in a computing method and system, but also, in industrial applications, includes a storage medium for computer-readable instructions, such as a non-volatile solid-state drive or a large disk array. It also provides an electronic device comprising a large-capacity high-speed random access memory (RAM) and a processor. When the device is awakened to execute the system instructions stored in the aforementioned medium, it can reproduce all implementation details of this invention, from gene data analysis to mismatch matrix hierarchical differentiation, end-to-end, providing a solution for digital breeding in agriculture and forestry.

[0100] This application is described with reference to flowchart illustrations and / or block diagrams of methods, systems, and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0101] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0102] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0103] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features, and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A Mendelian mismatch matrix construction method based on Mendel's law of segregation, characterized by, Includes the following steps: S1. Obtain the codominant marker genotype data of sample individuals in the target germplasm resource population, wherein the codominant markers include SSR, SNP, and InDel; S2. Preprocess the codominant marker genotype data to convert the genotype data into an allele set format, wherein homozygotes are converted into single-element sets, heterozygotes are converted into two-element sets, and missing sites are converted into empty sets. S3, construct a Mendelian mismatch matrix M, whose elements is defined as the count of loci where the allelic sets of two species i, j have empty intersection and are both non-deleted. S4. Construct the effective site statistical matrix L, whose elements This represents the number of valid loci for non-deletion alignments between two types of samples. S5. Division by elements = / Obtain the standardized mismatch matrix M' and set the effective site threshold. >0.8 As a criterion for determining statistical validity; S6. Analyze the mismatch matrix data and determine the minimum mismatch number based on matrix M. Perform germplasm Layered genetic grading, when At that time, germplasm i belongs to The grading is endogenously determined by the population's genetic structure, without the need for a preset threshold, thus enabling diversity evaluation and grading screening of the target germplasm population. S7, to Grade germplasm execution Two-step judgment: First, filter =0 Mendelian compatible pairings, then verify their mismatch spectrum consistency—if for all k≠i,j, the following holds true: = If the genotype is identical, it is determined to be a duplicate genotype; otherwise, a parent-child relationship is suspected.

2. The method for constructing a Mendelian mismatch matrix based on Mendel's law of segregation according to claim 1, characterized in that: The construction of the Mendelian mismatch matrix in step S3 is accelerated by using binary vector inner product. The allele set is mapped to a binary vector, and the rigid mismatch at the site level is determined by the vector inner product operation, reducing the algorithm complexity from O(N²LA) to O(N²L), where A is the average string length of the allele.

3. The method for constructing a Mendelian mismatch matrix based on Mendel's law of segregation according to claim 1, characterized in that: The steps described in step S6 Genetic stratification satisfies mutual exclusion, exhaustiveness, and endogeneity, and defines cumulative stratification. Hierarchical descriptions used for genetic specificity.

4. The method for constructing a Mendelian mismatch matrix based on Mendel's law of segregation according to claim 1, characterized in that... The steps described in step S7 The condition for consistency of mismatch patterns in the two-step determination is: the two germplasms not only match each other... =0, and is completely consistent with the mismatch pattern of all other germplasms in the population, which is the same genotype in a set sense.

5. The method for constructing a Mendelian mismatch matrix based on Mendel's law of segregation according to claim 1, characterized in that, The method for constructing the Mendelian mismatch matrix described in S3 includes: in This represents the total number of codominant markers at site l. This represents the mismatch matrix of germplasm individuals i and j at site l; in and This refers to the set of alleles of individuals i and j at locus l.

6. The method for constructing a Mendelian mismatch matrix based on Mendel's law of segregation according to claim 4, characterized in that, The method for constructing the effective site statistics matrix described in S3 includes: in, This is an indicator function that takes the value 1 when the condition is met and 0 when the condition is not met.

7. The method for constructing a Mendelian mismatch matrix based on Mendel's law of segregation according to claim 5, characterized in that, The methods for constructing the standardized mismatch matrix in S4 include: based on the constructed effective site statistics matrix. For mismatch matrix Corrections were made to eliminate the bias in results caused by differences in the number of effective marker sites among different germplasms: Only when the number of effective loci between paired germplasms meets the requirement Only when the condition is met will the correction result of the pair be determined to have statistical power and retained; otherwise, an invalid message will be returned.

8. The method for constructing a Mendelian mismatch matrix based on Mendel's law of segregation according to claim 1, characterized in that, The Mendelian mismatch matrix and the statistical matrix of valid sites also need to undergo compliance verification, and the determination methods include: S1, Symmetry: The matrix is ​​a strictly symmetric matrix, and all germplasm pairings satisfy... ; S2, Boundedness: All off-diagonal elements satisfy... , No value exceeding the total number of markers; S3. Non-negative integer property: All off-diagonal elements are non-negative integers, with no floating-point numbers, negative numbers, or invalid or missing values; S4, Missing Regularity: The diagonal elements of the matrix are set to placeholders NA, which represent no cross-comparison meaning, and have no valid values. The diagonal elements are set consistently. The compliance verification passes only if the mismatch matrix and the effective site statistics matrix satisfy all of the above determination methods; otherwise, an invalid message is returned.

9. A germplasm resource evaluation and grading system based on a full-population mismatch matrix, characterized in that, The system includes modules for implementing the method of any one of claims 1 to 7: The data preprocessing module is used to receive population-level codominant marker genotype data and perform allele set format conversion and generate binary encoding matrices. A matrix acceleration construction module is used to output a mismatch matrix and an effective site statistics matrix according to the formula using the binary encoding matrix. The evaluation and grading output module incorporates a standardized correction algorithm, matrix compression strategy, compliance verification mechanism, and multi-level mismatch rate analysis threshold, which is used to perform highly reliable redundancy screening and diversity grading of the entire population of germplasm resources.

10. A computer-readable storage medium having a program of computer instructions executable by a processor stored thereon, characterized in that, When the program is executed, it performs the method as described in any one of claims 1 to 7.