Ultrahigh-precision AI variety identification system and adaptive optimization method thereof

The ultra-high precision AI variety identification system, which combines multiple machine learning algorithms and phylogenetic trees, solves the problems of long identification cycles and low accuracy, achieving efficient and accurate variety identification and reducing breeding costs.

CN120975272APending Publication Date: 2025-11-18TIANJIN BIAOHUAXING GENE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511136267.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing methods for variety identification suffer from problems such as long identification cycles, low accuracy, and cumbersome operation. In particular, the PCA+random forest-based scheme lacks hyperparameter optimization, resulting in inaccurate identification results.

Method used

An ultra-high precision AI variety identification system is adopted. Through data acquisition, preprocessing, dimensionality reduction, training, labeling and screening, and multiple trial voting, combined with PCA, Random Forest, XGBoost and LightGBM algorithms, and adaptive optimization using Bayesian optimization and phylogenetic trees, the final model is constructed and the variety is identified.

Benefits of technology

It achieves high precision, high efficiency and wide applicability in variety identification, significantly reduces breeding time and computational resource consumption, and improves the accuracy and efficiency of variety identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120975272A_ABST
    Figure CN120975272A_ABST
Patent Text Reader

Abstract

The invention discloses an ultrahigh-precision AI variety identification system and an adaptive optimization method thereof, and the method comprises the steps: obtaining a high-quality tag data set, carrying out the preprocessing, obtaining a tag data matrix, carrying out the dimension reduction processing through PCA and singular value decomposition, determining a tag reduction set, and carrying out the recognition of the tag reduction set. Three machine learning algorithms of Random Forest, XGBoost and LightGBM are combined with cross validation and Bayesian optimization to screen a final marker, a final model is trained through a random forest, and ultrahigh-precision identification of varieties is realized through multiple test voting and phylogenetic tree auxiliary correction. The method can solve the problems of long identification period, low accuracy, tedious operation and the like in a traditional variety identification method, effectively improves the variety identification efficiency and accuracy, reduces various costs, and can be widely applied to the fields of agriculture, gardening, ecology and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of variety identification technology, specifically to an ultra-high precision AI variety identification system and its adaptive optimization method. Background Technology

[0002] With the continuous development of biotechnology and in-depth exploration in the field of plant breeding, variety identification has become an important task in agricultural production, plant resource protection, and genetic research. Variety identification is of great significance for ensuring the quality of agricultural products, protecting the interests of variety rights holders, and promoting breeding innovation.

[0003] Traditional methods for variety identification primarily rely on morphological characteristics and molecular marker techniques. Morphological identification distinguishes different varieties by observing external morphological features such as plant height, leaf shape, and flower color. This method is relatively simple to operate, but it has a long identification cycle and its accuracy is greatly affected by the environment, as different growing environments may cause the same variety to exhibit different morphological characteristics. Molecular marker techniques, on the other hand, identify varieties based on differences at the DNA level, such as restriction fragment length polymorphism (RFLP) and random amplified polymorphic DNA (RAPD). Although these techniques offer improved accuracy compared to morphological methods, they are cumbersome to operate, require highly skilled technicians, and have lower detection efficiency.

[0004] In recent years, with the rapid development of machine learning technology, it has achieved remarkable results in fields such as image recognition and natural language processing, providing new ideas and methods for variety identification. Machine learning is a data-driven algorithm and technology that automatically learns and identifies the inherent patterns and characteristics of data by training models. In the field of variety identification, machine learning technology can automatically extract the characteristic information of plant varieties by learning and analyzing data from a large number of plant samples, and achieve accurate classification and identification of unknown samples.

[0005] Currently, several machine learning-based variety identification schemes exist. For example, Random Forest combines the predictions of multiple decision trees. In variety identification, Random Forest can handle datasets with a large number of features (such as SNPs, InDels, SVs, and CNVs from resequencing and microarray data) and automatically select feature labels that significantly influence variety identification. By training a large number of decision trees, Random Forest can reduce the risk of overfitting, improve the model's generalization ability, and make variety identification results more reliable.

[0006] Extreme Gradient Boosting Machine (XGBoost) improves a model's predictive ability by iteratively training multiple decision trees. In variety identification, XGBoost efficiently handles large datasets (such as SNPs, InDels, SVs, and CNVs markers from resequencing and microarrays) and prevents overfitting through automatic handling of missing values ​​and built-in regularization techniques. Furthermore, XGBoost supports custom loss functions, allowing the model to be flexibly adjusted according to the specific needs of variety identification. Therefore, XGBoost achieves good predictive performance in variety identification, providing an effective tool for this purpose.

[0007] Lightweight Gradient Boosting Machine (LightGBM) is a histogram-based gradient boosting algorithm that optimizes model performance by reducing algorithm complexity and increasing training speed. In variety identification, LightGBM efficiently handles high-dimensional feature data and constructs an efficient model structure using a leaf node-based decision tree algorithm. This gives LightGBM a significant advantage in handling complex variety identification problems, enabling it to quickly and accurately identify the characteristics of different varieties.

[0008] Principal component analysis (PCA) is an unsupervised dimensionality reduction technique that projects high-dimensional genotypic data (such as millions of SNP loci) into a low-dimensional space (usually 2D or 3D) through linear transformation, preserving the largest variance information in the data.

[0009] A phylogenetic tree is a tree diagram representing the evolutionary relationships between biological groups or genes, showing the differentiation process from a common ancestor through its branching structure. Nodes represent evolutionary events (such as speciation) or a common ancestor; branches represent evolutionary paths, the length of which can reflect genetic distance or time span; leaf nodes represent extant or extinct groups / individuals; and monophyletic groups consist of branches composed of a common ancestor and all its descendants.

[0010] Currently, variety identification uses a PCA + random forest approach. However, PCA is time-consuming for screening and labeling, and it also requires chromosome splitting. Furthermore, random forest uses specified parameters and does not incorporate finding optimal hyperparameters to improve the accuracy of variety identification. Another approach uses PCA with three models trained to produce a single result. Although Bayesian optimization is used, this single result can lead to significant errors in variety prediction, and the model may misclassify, resulting in inaccurate variety identification. Therefore, it is necessary to design an ultra-high-precision AI variety identification system and its adaptive optimization method. Summary of the Invention

[0011] The purpose of this invention is to provide an ultra-high precision AI variety identification system and its adaptive optimization method to solve the problems mentioned in the background art.

[0012] To achieve the above objectives, the present invention provides the following technical solution: an ultra-high precision AI variety identification system, comprising:

[0013] The data acquisition module is used to acquire a high-quality labeled dataset, which is in VCF format;

[0014] The preprocessing module is used to preprocess the labeled dataset to obtain a labeled data matrix;

[0015] The dimensionality reduction module is used to reduce the dimensionality of the labeled data matrix, determine the optimal eigenvalues, and calculate the reduced labeled set.

[0016] The training module is used to train the labeled reduced set using three machine learning algorithms, including cross-validation and Bayesian optimization, to obtain the label feature importance values ​​of the model with the optimal parameters;

[0017] The label filtering module is used to filter the final labels based on the importance values ​​of the label features corresponding to three machine learning algorithms;

[0018] The final model building module is used to train the optimal parameters to build the final model by using the random forest algorithm, including cross-validation and Bayesian optimization algorithm, on the final labels, and then evaluate the model and visualize it.

[0019] The variety identification module is used to identify the varieties to be identified using the final model.

[0020] The multiple trial voting module is used to repeatedly test the steps from the training module to the variety identification module with different random seeds n times, and vote on the candidate varieties n times to confirm the final model prediction results.

[0021] The auxiliary correction module is used to construct phylogenetic trees for the sites before and after PCA filtering, respectively. The evolutionary relationship of the phylogenetic trees is used to assist in correcting the results obtained by the voting module in multiple trials, so as to obtain the final result of variety identification.

[0022] Preferably, the data acquisition module is connected to the preprocessing module, the preprocessing module is connected to the dimensionality reduction module, the dimensionality reduction module is connected to the training module, the training module is connected to the labeling and filtering module, the labeling and filtering module is connected to the final model building module, the final model building module is connected to the variety identification module, the variety identification module is connected to the multiple trial voting module, and the multiple trial voting module is connected to the auxiliary correction module.

[0023] Preferably, the preprocessing module processes the genotypes of the selected high-quality markers into a 0-1-2 format, with the wild-type homozygous genotype AA recorded as 0, the heterozygous genotype AB recorded as 1, and the mutant homozygous genotype BB recorded as 2; and converts the genotype file into an m×n matrix, where m is the number of samples and n is the number of markers.

[0024] Preferably, the dimensionality reduction module uses PCA and singular value decomposition to calculate the covariance of each marker site among different samples, forming a covariance matrix T. n x m The contribution of the variance of a particular principal component is equal to the corresponding eigenvalue δ of the original index correlation matrix. i The variance contribution rate of the i-th principal component is given by the following formula:

[0025]

[0026] T i The larger the value, the stronger the ability of the corresponding principal component to reflect comprehensive information; calculate the eigenvalues ​​and eigenvectors of the covariance matrix; by sorting the eigenvalues, when the size of the (n1+1)th eigenvalue is significantly lower than that of the n1th eigenvalue, that is, when the size of the (n1+1)th eigenvalue is twice or more than that of the n1th eigenvalue, retain the eigenvectors of the first n1 eigenvalues, where n1 is a natural number less than n; that is, in PC n1 and PC n1+1 Between these points, the variance contribution rate of the principal components decreased significantly. The first n1 principal components were retained to reduce the number of SNP sites. Then, according to the formula: If no significant decrease is found, the eigenvectors of the first 3 eigenvalues ​​are retained by default. The score of each label is calculated using the eigenvectors of the first n1 eigenvalues. The first k labels are selected for training the machine learning algorithm model. The labels are filtered using PCA and SVD methods to obtain a filtered m×k matrix of labels.

[0027] Preferably, in the training module, the three machine learning algorithms are Random Forest, XGBoost, and LightGBM; all three algorithms are trained using k-fold cross-validation, with 5-fold cross-validation used by default; Bayesian optimization is used to find their respective optimal hyperparameters; and F1-score, precision, recall, and accuracy are used to evaluate the model's fit.

[0028] Preferably, the label filtering module takes the intersection or union of the feature importance of each label calculated by the three models as the final label set, with the intersection being used by default.

[0029] Preferably, an adaptive optimization method for an ultra-high precision AI variety identification system includes the following steps:

[0030] Step a: Obtain a high-quality labeled dataset;

[0031] Step b: Preprocess the labeled dataset to obtain a labeled data matrix;

[0032] Step c: Perform dimensionality reduction on the labeled data matrix, determine the optimal eigenvalues, and calculate the reduced labeled set;

[0033] Step d: Train the labeled reduced set using three machine learning algorithms, including cross-validation and Bayesian optimization, to obtain the label feature importance values ​​of the model with the optimal parameters;

[0034] Step e: Filter the final labels using the label feature importance values ​​corresponding to the three machine learning algorithms;

[0035] Step f: Train the optimal parameters to build the final model by using the random forest algorithm, including cross-validation and Bayesian optimization, on the final labels, and then evaluate and visualize the model.

[0036] Step g: Use the final model to identify the variety to be identified;

[0037] Step h: Repeat the experiment n times with different random seeds in step dg, vote n times on the candidate varieties, and confirm the final model prediction results;

[0038] Step i: Construct phylogenetic trees using the sites before and after PCA filtering, respectively. Use the evolutionary relationships of the phylogenetic trees to help correct the h results and obtain the final variety identification results.

[0039] Preferred,

[0040] In step a, if the marker file is not of high quality, use vcftools software to filter the marker sites. The filtering conditions are as follows:

[0041] Remove missing sites from the SNP dataset;

[0042] Remove sites with a minimum allele frequency of less than 5%;

[0043] Remove sites with a read depth less than 3; use the default parameters for the rest.

[0044] Beneficial effects:

[0045] (1) This invention achieves the goal of “high precision, high efficiency and wide applicability” in variety identification through the whole process optimization of “efficient screening - multi-model collaboration - multiple verification - evolution correction”, effectively reducing the time cost, manpower cost and computing resource consumption of breeding, and providing a revolutionary technical solution for variety identification in agriculture, horticulture, ecology and other fields.

[0046] (2) This invention uses PCA combined with singular value decomposition for dimensionality reduction, avoiding the cumbersome step of splitting chromosomes required in traditional PCA marker screening, thus significantly shortening the marker screening time. At the same time, the number of principal components to be retained is automatically determined by eigenvalues, eliminating the need for manual intervention and further improving screening efficiency.

[0047] (3) This invention calculates the importance of marker features using three algorithms: Random Forest, XGBoost, and LightGBM. It then selects the final markers by taking the intersection or union of the results. This avoids the bias of a single model in judging features and ensures that the selected markers contribute the most to variety identification, thereby improving the accuracy of identification from the source of data.

[0048] The above description is merely an overview of the technical solutions of the embodiments of this application. In order to better understand the technical means of the embodiments of this application and to implement them in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the embodiments of this application more apparent and understandable, specific implementation methods of this application are described below. Attached Figure Description

[0049] Figure 1 This is a block diagram illustrating the principle of the present invention;

[0050] Figure 2 This is a flowchart of the workflow of the present invention;

[0051] Figure 3 This is a schematic diagram illustrating the variance contribution rate of all SNP sites on the first n principal components in the example.

[0052] Figure 4 A schematic diagram illustrating the final importance of SNP sites selected for this example;

[0053] Figure 5 The following is a diagram showing the four evaluation metrics of the final model in the example implementation;

[0054] Figure 6 This is a schematic diagram illustrating the distribution and linear equation of observed and predicted values ​​in an example.

[0055] Figure 7 This is a schematic diagram of the phylogenetic tree of the site training samples and prediction samples after data cleaning, as shown in the example.

[0056] Figure 8This is a schematic diagram of the phylogenetic tree of the predicted samples after PCA filtering in an example. Detailed Implementation

[0057] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims and drawings of this application are intended to cover non-exclusive inclusion.

[0059] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0060] Please see Figures 1-8 This invention discloses an ultra-high precision AI variety identification system, comprising:

[0061] Data acquisition module 1 is used to acquire a high-quality labeled dataset, wherein the labeled dataset is in VCF format;

[0062] Preprocessing module 2 is used to preprocess the labeled dataset to obtain a labeled data matrix;

[0063] Dimensionality reduction module 3 is used to reduce the dimensionality of the labeled data matrix, determine the optimal eigenvalues, and calculate the reduced labeled set.

[0064] Training module 4 is used to train the labeled reduced set using three machine learning algorithms, including cross-validation and Bayesian optimization, to obtain the label feature importance values ​​of the model under the optimal parameters;

[0065] The label filtering module 5 is used to filter the final labels based on the importance values ​​of the label features corresponding to the three machine learning algorithms.

[0066] The final model building module 6 is used to train the optimal parameters to build the final model by using the random forest algorithm, including cross-validation and Bayesian optimization algorithm, on the final labels, and then evaluate the model and visualize it.

[0067] Variety identification module 7 is used to identify the varieties to be identified using the final model.

[0068] The multiple trial voting module 8 is used to repeatedly test the steps from the training module to the variety identification module with different random seeds n times, and vote on the candidate varieties n times to confirm the final model prediction results.

[0069] The auxiliary correction module 9 is used to construct phylogenetic trees for the sites before and after PCA filtering, respectively. The evolutionary relationship of the phylogenetic trees is used to assist in correcting the results obtained by the voting module in multiple trials, so as to obtain the final result of variety identification.

[0070] The data acquisition module 1 is connected to the preprocessing module 2, the preprocessing module 2 is connected to the dimensionality reduction module 3, the dimensionality reduction module 3 is connected to the training module 4, the training module 4 is connected to the labeling and filtering module 5, the labeling and filtering module 5 is connected to the final model construction module 6, the final model construction module 6 is connected to the variety identification module 7, the variety identification module 7 is connected to the multiple trial voting module 8, and the multiple trial voting module 8 is connected to the auxiliary correction module 9.

[0071] In this invention, the preprocessing module processes the genotypes of the selected high-quality markers into a 0-1-2 format, with wild-type homozygous genotype AA marked as 0, heterozygous genotype AB marked as 1, and mutant homozygous genotype BB marked as 2; and converts the genotype file into an m×n matrix, where m is the number of samples and n is the number of markers.

[0072] In this invention, the dimensionality reduction module uses PCA and singular value decomposition to calculate the covariance of each marker site across different samples, forming a covariance matrix T. n x m The contribution of the variance of a particular principal component is equal to the corresponding eigenvalue δ of the original index correlation matrix. i The variance contribution rate of the i-th principal component is given by the following formula:

[0073]

[0074] T i The larger the value, the stronger the ability of the corresponding principal component to reflect comprehensive information; calculate the eigenvalues ​​and eigenvectors of the covariance matrix; by sorting the eigenvalues, when the size of the (n1+1)th eigenvalue is significantly lower than that of the n1th eigenvalue, that is, when the size of the (n1+1)th eigenvalue is twice or more than that of the n1th eigenvalue, retain the eigenvectors of the first n1 eigenvalues, where n1 is a natural number less than n; that is, in PC n1 and PC n1+1Between these points, the variance contribution rate of the principal components decreased significantly. The first n1 principal components were retained to reduce the number of SNP sites. Then, according to the formula: If no significant decrease is found, the eigenvectors of the first 3 eigenvalues ​​are retained by default. The score of each label is calculated using the eigenvectors of the first n1 eigenvalues. The first k labels are selected for training the machine learning algorithm model. The labels are filtered using PCA and SVD methods to obtain a filtered m×k matrix of labels.

[0075] In this invention, the training module uses three machine learning algorithms: Random Forest, XGBoost, and LightGBM. All three algorithms are trained using k-fold cross-validation, with 5-fold cross-validation used by default. Bayesian optimization is used to find the optimal hyperparameters for each algorithm. F1-score, precision, recall, and accuracy are used to evaluate the model's fit.

[0076] In this invention, the label selection module takes the intersection or union of the feature importance of each label calculated by the three models as the final label set, with the intersection being used by default.

[0077] This invention provides an adaptive optimization method for an ultra-high precision AI variety identification system, comprising the following steps:

[0078] Step a: Obtain a high-quality labeled dataset;

[0079] Step b: Preprocess the labeled dataset to obtain a labeled data matrix;

[0080] Step c: Perform dimensionality reduction on the labeled data matrix, determine the optimal eigenvalues, and calculate the reduced labeled set;

[0081] Step d: Train the labeled reduced set using three machine learning algorithms, including cross-validation and Bayesian optimization, to obtain the label feature importance values ​​of the model with the optimal parameters;

[0082] Step e: Filter the final labels using the label feature importance values ​​corresponding to the three machine learning algorithms;

[0083] Step f: Train the optimal parameters to build the final model by using the random forest algorithm, including cross-validation and Bayesian optimization, on the final labels, and then evaluate and visualize the model.

[0084] Step g: Use the final model to identify the variety to be identified;

[0085] Step h: Repeat the experiment n times with different random seeds in step dg, vote n times on the candidate varieties, and confirm the final model prediction results;

[0086] Step i: Construct phylogenetic trees using the sites before and after PCA filtering, respectively. Use the evolutionary relationships of the phylogenetic trees to help correct the h results and obtain the final variety identification results.

[0087] In step a, if the marker file is not of high quality, the marker sites are filtered using vcftools software, with the following filtering conditions:

[0088] Remove missing sites from the SNP dataset;

[0089] Remove sites with a minimum allele frequency of less than 5%;

[0090] Remove sites with a read depth less than 3; use the default parameters for the rest.

[0091] Example:

[0092] The method of this invention was used to test the resequencing WGS genotype data of a certain tree variety, information on 60 samples from 10 varieties, and 15 candidate identification variety samples.

[0093] When using PCA + 3 models with Bayesian optimization and 5-fold cross-validation to train once to predict the variety of 15 samples, the results are as follows. Figure 3 It is a variance contribution rate map of all SNP loci on the first n principal components. Based on the variance contribution rate map of the principal components, k principal components are determined to calculate the SNP score. Figure 4 This is a ranking chart of the importance of SNP sites after screening to determine their final importance. Figure 5 This is a graph showing the four evaluation indicators of the final model, all of which are 1. Table 1 shows the predicted and actual variety number results. Although all indicators are 1, the prediction error rate is 6 / 15.

[0094] Table 1 Comparison of predicted varieties and actual varieties after only one training session.

[0095]

[0096]

[0097] When PCA, three Bayesian models, and 5-fold cross-validation were used for repeated training 20 times, and the 15 candidate samples were voted on 20 times to finally confirm the variety, the results are as follows. Table 2 shows the predicted and actual variety identification results. The prediction error rate was 3 / 15, which is higher than the accuracy of training once, thus improving the certainty of variety identification.

[0098] Table 2 Comparison of Predicted Varieties and Actual Varieties After 20 Training Rounds of Voting Methods

[0099]

[0100]

[0101] When PCA, three models with Bayesian optimization and 5-fold cross-validation were used for repeated training 20 times, and 15 candidate samples were voted on to finally confirm which variety it was, plus tree-assisted correction, the results are as follows. Figure 7 It is a phylogenetic tree of the site training samples and prediction samples after data cleaning. Figure 8 This is the phylogenetic tree of the training samples after PCA filtering. Table 3 shows the predicted and actual variety numbering results, which are based on the results in Table 2, and further corrected using the tree results. The final prediction error rate is 0 / 15.

[0102] Table 3 shows the results of the tree-based auxiliary correction in Table 2 – a comparison of predicted and actual varieties.

[0103]

[0104]

[0105] In summary, this invention solves the problems of long identification cycles, low accuracy, and cumbersome operation in traditional variety identification methods, effectively improving the efficiency and accuracy of variety identification and reducing various costs. Through the optimization of the entire process of "efficient screening - multi-model collaboration - multiple verifications - evolutionary correction", it achieves the goal of "high precision, high efficiency, and wide applicability" in variety identification, effectively reducing breeding time costs, labor costs, and computing resource consumption, and providing a revolutionary technical solution for variety identification in agriculture, horticulture, ecology and other fields.

[0106] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A high-precision AI variety identification system, characterized in that: include: The data acquisition module (1) is used to acquire a high-quality labeled dataset, wherein the labeled dataset is in VCF format; The preprocessing module (2) is used to preprocess the labeled dataset to obtain a labeled data matrix; The dimension reduction module (3) is used to perform dimension reduction processing on the labeled data matrix, determine the optimal eigenvalues, and calculate the reduced labeled set. Training module (4) is used to train the labeled reduced set using three machine learning algorithms, including cross-validation and Bayesian optimization, to obtain the label feature importance values ​​of the model under the optimal parameters; The label filtering module (5) is used to filter the final labels by the label feature importance values ​​corresponding to the three machine learning algorithms; The final model building module (6) is used to train the best parameters to build the final model by using the random forest algorithm, including cross-validation and Bayesian optimization algorithm on the final labels, and then evaluate the model and visualize it. Variety identification module (7) is used to identify the varieties to be identified through the final model; The multiple trial voting module (8) is used to repeatedly test the steps from the training module to the variety identification module with different random seeds n times, vote on the candidate varieties n times, and confirm the final model prediction results. The auxiliary correction module (9) is used to construct phylogenetic trees for the sites before and after PCA filtering, and to correct the results obtained by the voting module in multiple trials through the evolutionary relationship of the phylogenetic trees, so as to obtain the final result of variety identification.

2. The ultra-high precision AI variety identification system according to claim 1, characterized in that: The data acquisition module (1) is connected to the preprocessing module (2), the preprocessing module (2) is connected to the dimensionality reduction module (3), the dimensionality reduction module (3) is connected to the training module (4), the training module (4) is connected to the labeling and filtering module (5), the labeling and filtering module (5) is connected to the final model building module (6), the final model building module (6) is connected to the variety identification module (7), the variety identification module (7) is connected to the multiple trial voting module (8), and the multiple trial voting module (8) is connected to the auxiliary correction module (9).

3. The ultra-high precision AI variety identification system according to claim 1, characterized in that: The preprocessing module processes the genotypes of the selected high-quality markers into a 0-1-2 format, with wild-type homozygous genotype AA marked as 0, heterozygous genotype AB marked as 1, and mutant homozygous genotype BB marked as 2; and converts the genotype file into an m×n matrix, where m is the number of samples and n is the number of markers.

4. The ultra-high precision AI variety identification system according to claim 1, characterized in that: The dimensionality reduction module uses PCA and singular value decomposition to calculate the covariance of each marker site across different samples, forming a covariance matrix T. nxm The contribution of the variance of a particular principal component is equal to the corresponding eigenvalue δ of the original index correlation matrix. i The variance contribution rate of the i-th principal component is given by the following formula: T i The larger the value, the stronger the ability of the corresponding principal component to reflect comprehensive information; calculate the eigenvalues ​​and eigenvectors of the covariance matrix; by sorting the eigenvalues, when the size of the (n1+1)th eigenvalue is significantly lower than that of the n1th eigenvalue, that is, when the size of the (n1+1)th eigenvalue is twice or more than that of the n1th eigenvalue, retain the eigenvectors of the first n1 eigenvalues, where n1 is a natural number less than n; that is, in PC n1 and PC n1+1 Between these points, the variance contribution rate of the principal components decreased significantly. The first n1 principal components were retained to reduce the number of SNP sites. Then, according to the formula: If no significant decrease is found, the eigenvectors of the first 3 eigenvalues ​​are retained by default. The score of each label is calculated using the eigenvectors of the first n1 eigenvalues. The first k labels are selected for training the machine learning algorithm model. The labels are filtered using PCA and SVD methods to obtain a filtered m×k matrix of labels.

5. The ultra-high precision AI variety identification system according to claim 1, characterized in that: The training module employs three machine learning algorithms: Random Forest, XGBoost, and LightGBM. All three algorithms are trained using k-fold cross-validation, with 5-fold cross-validation used by default. Bayesian optimization is used to find the optimal hyperparameters for each algorithm. F1-score, precision, recall, and accuracy are used to evaluate the model's fit.

6. The ultra-high precision AI variety identification system according to claim 1, characterized in that: The label filtering module takes the intersection or union of the feature importance of each label calculated by the three models as the final label set, with the intersection being used by default.

7. An adaptive optimization method for an ultra-high precision AI variety identification system, characterized in that: Includes the following steps: Step a: Obtain a high-quality labeled dataset; Step b: Preprocess the labeled dataset to obtain a labeled data matrix; Step c: Perform dimensionality reduction on the labeled data matrix, determine the optimal eigenvalues, and calculate the reduced labeled set; Step d: Train the labeled reduced set using three machine learning algorithms, including cross-validation and Bayesian optimization, to obtain the label feature importance values ​​of the model with the optimal parameters; Step e: Filter the final labels using the label feature importance values ​​corresponding to the three machine learning algorithms; Step f: Train the optimal parameters to build the final model by using the random forest algorithm, including cross-validation and Bayesian optimization, on the final labels, and then evaluate and visualize the model. Step g: Use the final model to identify the variety to be identified; Step h: Repeat the experiment n times with different random seeds in step dg, vote n times on the candidate varieties, and confirm the final model prediction results; Step i: Construct phylogenetic trees using the sites before and after PCA filtering, respectively. Use the evolutionary relationships of the phylogenetic trees to help correct the h results and obtain the final variety identification results.

8. The adaptive optimization method for an ultra-high precision AI variety identification system according to claim 7, characterized in that: In step a, if the marker file is not of high quality, use vcftools software to filter the marker sites. The filtering conditions are as follows: Remove missing sites from the SNP dataset; Remove sites with a minimum allele frequency of less than 5%; Remove sites with a read depth less than 3; use the default parameters for the rest.