Wheat germplasm recommendation algorithm based on GA optimization clustering

By combining the GA-optimized K-Means clustering algorithm with UMAP and L-BFGS-B technologies, the matching challenge of high-dimensional heterogeneous agricultural data was solved, achieving precise matching of wheat germplasm with regional climate and improving the performance and accuracy of the recommendation system.

CN120873660APending Publication Date: 2025-10-31HENAN AGRICULTURAL UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511014370.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing clustering algorithms face problems of initial sensitivity, noise, and redundant features when processing high-dimensional heterogeneous agricultural data, making it difficult to achieve accurate matching between wheat germplasm and regional climate.

Method used

A wheat germplasm recommendation system was constructed using a genetic algorithm (GA)-based K-Means clustering algorithm, combined with UMAP dimensionality reduction and L-BFGS-B optimization techniques, and a hybrid initialization strategy and multilayer perceptron model. Cosine similarity was used to screen matching varieties.

Benefits of technology

It significantly improves the performance and robustness of clustering algorithms, enhances the ability to process high-dimensional heterogeneous agricultural data, achieves precise matching of wheat germplasm with regional climate, and provides higher accuracy, precision, and recall.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873660A_ABST
    Figure CN120873660A_ABST
Patent Text Reader

Abstract

The invention provides a wheat germplasm recommendation algorithm based on GA optimization clustering, the algorithm utilizes the genetic algorithm (GA) to optimize and improve a K-Means clustering algorithm, combines a UMAP dimension reduction technology and an L-BFGS-B algorithm to construct a wheat germplasm recommendation model, and studies that a wheat germplasm recommendation model is constructed by integrating wheat germplasm data of Henan Province and urban meteorological data. A hybrid initialization strategy and a refined optimization method are designed, so that the performance of the clustering algorithm is remarkably improved; experimental results show that the optimized algorithm is superior to a traditional method in accuracy, precision rate and recall rate; in addition, an ablation experiment verifies the key effects of UMAP dimension reduction, hybrid initialization and an L-BFGS-B algorithm on the model performance; finally, the two-stage recommendation model (MLP classification + cosine similarity sorting) constructed by research can provide accurate wheat germplasm recommendation for breeding experts, assists in regional adaptive variety breeding, provides a new thought for efficient utilization of agricultural germplasm resources, and is of great significance to guarantee grain safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of precise screening and efficient utilization of crop germplasm resources, specifically to a wheat germplasm recommendation method based on an optimized and improved clustering algorithm. Background Technology

[0002] Against the backdrop of escalating global climate change and increasingly urgent food security needs, the precise selection and efficient utilization of crop germplasm resources have become key issues in modern agricultural development. Traditional breeding methods rely on manual experience for selection, which suffers from limitations such as low efficiency and strong subjectivity when faced with ever-increasing multidimensional germplasm data and complex environmental factors. In recent years, the in-depth application of machine learning technology in the agricultural field has provided a new technical path for building data-driven intelligent breeding decision-making systems. However, existing clustering methods still face two challenges when processing high-dimensional heterogeneous agricultural data: firstly, irrelevant features and noisy data can mislead clustering algorithms; secondly, the curse of dimensionality can cause high-dimensional data points to become sparse, making it difficult to discover any structure in the data.

[0003] Both domestic and international researchers have made significant efforts in crop recommendation algorithms, aiming to establish models of the relationship between varieties and traits to improve the accurate identification of important characteristics such as stress resistance and high yield. International research focuses on multi-source data fusion and optimization of classical algorithms. For example, Pakistani scholar Kiruthika combined climate, soil, and pest data with weighted support vector machines, while a Bangladeshi team used K-means clustering to process multi-dimensional agricultural data. Although these methods improve model adaptability, they are not efficient in feature selection for high-dimensional heterogeneous data and do not fully consider the nonlinear relationship between germplasm resources and regional environment. Domestic research focuses on innovation in recommendation algorithms, such as fine-grained feature interaction networks and dynamic interest modeling techniques, but they face challenges such as strong dependence on data quality and limited model generalization ability in practical applications. It is particularly noteworthy that existing research generally has the following limitations: (1) Traditional clustering algorithms are sensitive to initial centers and are prone to getting trapped in local optima; (2) Noise and redundant features in high-dimensional germplasm data reduce model performance; (3) There is a lack of refined optimization methods for the adaptability of wheat germplasm to regional climate.

[0004] In view of the above, this study proposes a wheat germplasm recommendation algorithm based on GA optimized clustering. This method aims to overcome the bottlenecks of traditional algorithms in terms of initial sensitivity and the curse of dimensionality, and provide a new technical path for the accurate matching of wheat germplasm with regional climate. Summary of the Invention

[0005] To address the above issues and overcome the shortcomings of existing technologies, this invention provides a wheat germplasm recommendation algorithm based on GA optimized clustering. By innovatively integrating genetic algorithms and K-Means clustering, combined with UMAP dimensionality reduction and L-BFGS-B optimization techniques, an intelligent wheat germplasm recommendation system is constructed.

[0006] The wheat germplasm recommendation algorithm based on GA optimized clustering is characterized by the following steps:

[0007] S1: Obtain wheat germplasm data of the target province from the existing crop germplasm information network, and obtain meteorological data of cities in the target province for the past two years from the existing meteorological website. For missing data in the entire column of the two datasets, use deletion method; for missing partial data, use KNN to fill in; for useful non-numerical data, use label encoding to convert it into numerical data; for outliers, use box-line method to find them and replace them with the average value.

[0008] S2: Classify wheat germplasm data from two dimensions: resistance characteristics and growth characteristics, and select important features. Construct a high-dimensional feature space through feature interaction, superposition of meteorological features, standardization, and urban coding to enhance feature expression capabilities.

[0009] S3: For the final high-dimensional features, the UMAP dimensionality reduction strategy is adopted. UMAP is used to capture non-linear structures and map high-dimensional features to low-dimensional space, providing a low-noise and highly interpretable feature representation for subsequent genetic algorithm optimization of clustering.

[0010] S4: The K-means++ clustering algorithm is optimized and improved by optimizing the cluster centers through a hybrid initialization strategy. The hybrid initialization strategy and the finite memory quasi-Newton algorithm (L-BFGS-B) are designed to iteratively optimize the clustering algorithm, further improving its performance and providing high-quality labels for the multilayer perceptron.

[0011] S5: A wheat germplasm recommendation model is established based on multilayer perceptron and cosine similarity. According to the climate conditions of the selected city and the needs of breeding experts, the dimensionality-reduced features are associated with clustering labels. First, the multilayer perceptron model is used to predict the potential adaptation category of the target variety, and then cosine similarity is used to screen the most matching varieties within the same category.

[0012] The beneficial effects of the above technical solution are as follows:

[0013] (1) This study proposes a K-Means clustering algorithm based on genetic algorithm (GA) optimization, and combines UMAP dimensionality reduction technology, hybrid initialization strategy and L-BFGS-B algorithm to construct a wheat germplasm recommendation model. By integrating wheat germplasm data of the target province and urban meteorological data, a hybrid initialization strategy and refined optimization method are designed, which significantly improves the performance of the clustering algorithm. Experimental results show that the algorithm is significantly better than the traditional method in terms of accuracy, precision and recall, indicating that the algorithm has stronger robustness and generalization ability when processing high-dimensional heterogeneous agricultural data.

[0014] (2) This study verified the key role of UMAP dimensionality reduction, hybrid initialization and L-BFGS-B algorithm in model performance through ablation experiments. UMAP dimensionality reduction can extract the essential features of data through nonlinear manifold learning, the hybrid initialization strategy effectively avoids the local optimum problem, and the L-BFGS-B algorithm significantly improves the optimization efficiency.

[0015] (3) The two-stage recommendation model (MLP classification + cosine similarity ranking) constructed in this study shows through experimental results that the recommended germplasm has significant advantages in stress resistance and growth quality indicators, providing a scientific basis for the breeding of regionally adaptable varieties and providing precise wheat germplasm recommendations for breeding experts. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the overall technical process of the present invention;

[0017] Figure 2 This is a schematic diagram of wheat characteristics in a specific embodiment of the present invention;

[0018] Figure 3 This is a schematic diagram of meteorological characteristics in a specific embodiment of the present invention;

[0019] Figure 4 This is a schematic diagram of the initialization process of the hybrid strategy algorithm of the present invention;

[0020] Figure 5 This is a schematic diagram of the L-BFGS-B algorithm flow of the present invention;

[0021] Figure 6 This is a schematic diagram illustrating the accuracy analysis of each algorithm in an experimental verification of a specific embodiment of the present invention;

[0022] Figure 7 This is a schematic diagram illustrating the accuracy analysis of each algorithm in an experimental verification of a specific embodiment of the present invention;

[0023] Figure 8 This is a schematic diagram illustrating the recall analysis of each algorithm in a specific embodiment of the present invention.

[0024] Figure 9 This is a schematic diagram of the technical process of the wheat germplasm recommendation model of the present invention;

[0025] Figure 10 This is a schematic diagram comparing known meteorological recommendations for cities in a specific embodiment of the present invention;

[0026] Figure 11 This is a schematic diagram comparing the recommended models for cities with unknown meteorological conditions in a specific embodiment of the present invention. Detailed Implementation

[0027] The foregoing and other technical contents, features and effects of the present invention will be clearly presented in the following detailed description of the embodiments with reference to the accompanying drawings. All contents mentioned in the following embodiments are based on the accompanying drawings.

[0028] Example 1, as Figure 1 As shown, this study focuses on the correlation analysis between wheat germplasm resources and urban climate conditions in the target province. Based on the improved genetic algorithm, the K-Means clustering method is used to explore the characteristics of wheat varieties. A multilayer perceptron classifier is established to predict the performance of wheat varieties in specific urban environments. The aim is to provide more accurate wheat variety recommendations for breeding experts in different cities in the target province.

[0029] The main research contents are as follows:

[0030] (1) Data acquisition and processing: wheat germplasm data of the target province was obtained from the existing crop germplasm information network, and meteorological data of each city in the target province within 2 years were obtained from the existing meteorological website;

[0031] For missing data in an entire column in both datasets, deletion is used; for missing data in a partial column, KNN is used to fill in the missing data.

[0032] For useful non-numerical data, label encoding is used. Specifically, label encoding achieves this transformation by assigning a unique integer value to each category. For example, for the field of "resistance level", we can encode "high", "medium", and "low" as "1", "2", and "3" respectively.

[0033] Outliers are identified using the box method and replaced with the average value. Specifically, the threshold for outliers is determined by calculating the quartiles (Q1, Q2, Q3) and interquartile range (IQR) of the dataset. Generally, any data point less than Q1 - 1.5 * IQR or greater than Q3 + 1.5 * IQR is considered an outlier, and we choose to replace outliers with the average value of the dataset.

[0034] (2) Multi-source data fusion and feature engineering: Important features are selected from the resistance and growth characteristics of wheat. A high-dimensional feature space is constructed by combining feature interaction, superimposing meteorological features, standardization (Note: Standardization is an existing technology that aims to eliminate the dimensional differences between different features) and city coding (Note: City coding is also converted into numerical values) to enhance the feature expression capability.

[0035] (3) High-dimensional data dimensionality reduction technology: UMAP dimensionality reduction strategy is adopted for the final high-dimensional features. UMAP captures non-linear structure and maps high-dimensional features to low-dimensional space, providing low-noise and highly interpretable feature representation for subsequent genetic algorithm optimization clustering.

[0036] (4) Improved K-means based on GA optimization: The cluster centers are optimized by a hybrid initialization strategy. The design adopts a hybrid initialization strategy and L-BFGS-B algorithm for iterative optimization, which further improves the performance of the clustering algorithm and provides high-quality labels for the multilayer perceptron.

[0037] (5) Establish a wheat germplasm recommendation model based on multilayer perceptron and cosine similarity: According to the climate conditions of the selected city and the needs of breeding experts, the dimensionality-reduced features are associated with cluster labels. First, the multilayer perceptron model is used to predict the potential adaptation category of the target variety, and then cosine similarity is used to screen the most matching variety within the same category.

[0038] Example 2: Data Acquisition and Preprocessing. This study integrated 843 wheat germplasm records from the target province from the existing crop germplasm resource database with 11,712 urban meteorological observation data from the target province published by the National Meteorological Service Center. This dataset covers comprehensive information on key traits of wheat growth and development and regional climate characteristics, providing solid data support for crop-climate correlation analysis.

[0039] The wheat germplasm dataset includes two main dimensions: growth characteristics and resistance characteristics.

[0040] Growth characteristics: crude protein content (%), lysine content (%), sedimentation value (ml), and hardness index, reflecting the nutritional quality and processing characteristics of wheat grains;

[0041] Resistance characteristics: drought resistance, waterlogging resistance, bud salt tolerance, and field cold resistance are assessed using a 1-5 level quantitative scoring system, where level 1 indicates the most sensitive resistance and level 5 indicates the strongest resistance.

[0042] Table 1 shows the statistical information of key traits of wheat germplasm resources in the target province. Among them, growth characteristic indicators exhibit significant genetic diversity (crude protein content ranges from 11.2% to 17.8%, sedimentation value ranges from 22.5% to 45.3 ml), and the distribution of resistance characteristic scores indicates significant differences in the adaptability of different varieties to adverse conditions. This dataset provides a multi-dimensional research perspective for analyzing the relationship between wheat germplasm characteristics and regional climate adaptability. Future research will utilize feature fusion and pattern recognition technologies to construct a climate-intelligent wheat germplasm recommendation model.

[0043] Table 1. Statistical information on key traits of wheat germplasm resources in the target province.

[0044] Variety Name Crude protein / % Lysine / % Precipitation value / ml hardness drought resistance Flood resistance Salt tolerance during bud stage Field cold resistance Jilin rat 10.01 0.34 8.5 17 1 4 5 4 Monk's head 17.8 0.5 31.5 11.9 3 1 1 1 Apricots sweep the day 19.48 0.49 27 26 5 1 3 1 Bo Shi Ba Mai 12.32 0.4 14 47 3 5 4 5 Red gourd head 10.01 0.34 8.5 17 1 4 5 4 Red centipede head 17.8 0.49 19.5 14.8 3 1 5 1 White gourd head 18.23 0.45 29 11 3 5 4 2 White-breasted centipede 14.75 0.46 26.3 18.2 3 1 3 1 Buddha's Hand Wheat 11.84 0.38 19 13.3 3 2 4 2 White Mango 17.8 0.5 31.5 12.5 3 1 1 1 Small Red Mango 16.62 0.47 34 10.5 4 5 5 2 Centipede wheat 17.31 0.45 29 10.6 4 1 3 1 Frost-covered wheat 10.01 0.34 8.5 17 1 4 5 4 Bamboo Stalk Green 17.79 0.44 24 17.7 3 4 1 1 Little Buddha's Hand 16.4 0.45 27 11.2 4 3 3 3 Zhang Fei Hu 10.01 0.34 8.5 17 1 4 2 4 Red bald head 10.62 0.36 5 10.4 4 4 5 1 Small Red Mango 16.54 0.44 26 10.8 3 5 5 5 White gourd 18.23 0.45 29 11.4 3 5 3 3 Red yeast rice 17.04 0.47 37 10.8 3 2 4 3 Blind Stone Eight 17.1 0.48 33.2 11 3 5 5 2 Red Mango 10.38 0.34 8 11 3 4 5 2

[0045] The urban meteorological dataset records daily climate parameters for major cities in the target province over a two-year period:

[0046] Temperature characteristics: daily average temperature (°C), daily maximum temperature (°C), and daily minimum temperature (°C) characterize the distribution of heat resources;

[0047] Moisture characteristics: daily average relative humidity (%) and daily total precipitation (mm), reflecting the degree of environmental humidity;

[0048] Wind speed characteristics: Daily average wind speed (m / s) affects the regulation of farmland microclimate.

[0049] Table 2 shows relevant meteorological information for some cities in the target province. Among them, the daily average temperature, daily maximum temperature, daily minimum temperature, daily average relative humidity, daily average wind speed, and daily total precipitation are the most important characteristics for evaluating the environmental factors affecting the performance of wheat varieties.

[0050] Table 2 Meteorological Information of Some Cities in the Target Province

[0051] date City Average daily temperature / ℃ Daily maximum temperature / ℃ Daily minimum temperature / ℃ Daily average relative humidity / % Daily average wind speed (m / s) Total daily precipitation / mm 2023 / 2 / 20 Sanmenxia City 4 10 -3.2 41.38 6.62 0 2023 / 2 / 21 Sanmenxia City 2.84 4.2 0.5 88.62 8.5 0.5 2023 / 2 / 22 Sanmenxia City 2.01 4.8 -0.3 95.5 3.88 19.8 2023 / 2 / 23 Sanmenxia City 4.46 10.4 0.7 82.62 4.5 0 2023 / 2 / 24 Sanmenxia City 5.99 14.7 -1 71.5 6.31 0 2023 / 2 / 25 Sanmenxia City 3.38 4 2.3 85.75 8.12 0 2023 / 2 / 26 Sanmenxia City 6.34 9.2 2.6 66.5 8.75 0 2023 / 2 / 27 Sanmenxia City 8.76 15.4 3.9 53.88 8.75 0 2023 / 2 / 28 Sanmenxia City 10.5 17.2 5.4 47.38 8.25 0 2023 / 3 / 1 Sanmenxia City 9.91 14.5 5.9 23.5 9.12 0 2023 / 3 / 2 Sanmenxia City 7.45 16.1 -1.3 35.88 6.25 0 2023 / 3 / 3 Sanmenxia City 10.4 18.7 3.2 34.12 9 0 2023 / 3 / 4 Sanmenxia City 10.39 22.1 0.6 29.38 7.25 0 2023 / 3 / 5 Sanmenxia City 13 24.4 3.1 33.5 5.62 0 2023 / 3 / 6 Sanmenxia City 14.62 24.4 7.5 38.88 6.38 0 2023 / 3 / 7 Sanmenxia City 15.38 26.2 5.7 31 6 0

[0052] Example 3, as Figure 2 As shown, feature selection was performed on wheat germplasm data. In constructing a climate-smart wheat variety recommendation model, feature engineering, as a crucial step, requires a systematic consideration of the entire life cycle characteristics of wheat germplasm. This study innovatively constructed a seven-dimensional feature space encompassing resistance and growth characteristics, achieving a comprehensive characterization of wheat's overall performance.

[0053] The environmental adaptability evaluation system comprises drought resistance, waterlogging tolerance, and field cold resistance, using a quantitative scoring mechanism (levels 1-5) to accurately characterize the variety's ability to withstand drought, waterlogging, and low-temperature stress. Against the backdrop of intensifying climate change, these characteristics have become core indicators for ensuring regional yield stability, directly impacting the ecological suitability range of varieties.

[0054] Growth quality characteristics: In terms of nutritional quality, crude protein content determines the baking characteristics of flour, lysine, as the first limiting amino acid, affects nutritional value, and precipitation value reflects the ability to form a gluten network. These three factors together constitute the "golden triangle" evaluation system for wheat processing quality. In terms of processing characteristics, the hardness index predicts the suitability of the end product through the physical characteristics of the grain and is an important basis for grading bread wheat and noodle wheat.

[0055] This feature system achieves full-chain coverage from environmental adaptation to end-use, satisfying both breeders' requirements for broad-spectrum adaptability and the market's demand for high-quality specialty wheat. Notably, significant biological interactions exist between features. For example, drought resistance may indirectly alter protein content by affecting physiological activities during the grain-filling stage, and waterlogging tolerance may be related to root development, thus influencing sedimentation values. Future work will introduce interactive feature modeling, employing polynomial feature expansion (note: polynomial feature expansion is an existing technology; its core idea is to capture interactions between features by generating product combinations of existing features, thereby improving the model's expressive power) or automatic neural network crossover methods to deeply explore the nonlinear correlation mechanisms between features, further enhancing the prediction accuracy and generalization ability of the recommendation model.

[0056] Example 4, as Figure 3 As shown, urban meteorological characteristics were selected. Since climate factors, as important environmental variables, profoundly influence the physiological processes and varietal adaptability of wheat, this study constructed a climate index system including temperature, moisture, wind speed, and three-dimensional characteristics to provide a quantitative basis for the varietal-climate matching mechanism.

[0057] (1) Temperature dimension: Daily average temperature and its extreme values ​​(T_avg, T_max, T_min) constitute a triple representation of thermal conditions. Extreme temperature events (such as T_max > 35℃ or T_min < -5℃) are direct causes of yield fluctuations, while effective accumulated temperature (GDD) determines the length of the growing season. Temperature sensitivity analysis can reveal the photothermal response characteristics of varieties and provide a physiological basis for ecozonal division.

[0058] (2) Moisture Dimension: Daily average relative humidity (RH_avg) affects water use efficiency and disease resistance by influencing transpiration rate and disease occurrence threshold. Studies have shown that when RH_avg > 85%, the incidence of Fusarium head blight increases exponentially, while RH_avg < 40% may induce water stress. This indicator provides environmental constraint parameters for disease-resistant breeding and irrigation system optimization. Daily total precipitation (P_total) determines the water supply pattern, and its seasonal distribution and coupling degree with the critical period of wheat water requirement (jointing-grain-filling stage) directly affect yield formation. The coefficient of variation of precipitation (CV_p) can be used to assess drought / flood risk, providing an environmental pressure indicator for the selection of drought / flood resistance traits.

[0059] (3) Wind speed dimension: Daily average wind speed (WS_avg) affects photosynthesis and mechanical stability by altering the canopy microclimate. Actual measurement data shows that when WS_avg > 4 m / s, the risk of lodging increases by 32%, while moderate ventilation (WS_avg = 1.5-2.5 m / s) can reduce the incidence of powdery mildew. This parameter provides a basis for wind effect assessment for lodging-resistant breeding and field management.

[0060] This climate index system not only provides a refined description of regional climate characteristics, but more importantly, it establishes a response mechanism between environmental variables and varietal traits. By constructing a climate-trait correlation matrix, the recommendation model can dynamically assess the adaptability of varieties under different climate scenarios, providing decision support for the breeding of regional specialty varieties and the optimization of planting layout.

[0061] Example 5: Selection of Dimensionality Reduction Techniques. In the feature engineering stage of the wheat germplasm recommendation system, this study innovatively introduces the latest achievement in manifold learning—the UMAP dimensionality reduction algorithm. Based on Riemannian geometry and fuzzy topology theory, this algorithm can adaptively capture the inherent nonlinear manifold structure of the data, making it particularly suitable for biofeature analysis with complex interaction effects. Compared to traditional methods, UMAP exhibits superior computational efficiency and dimensionality compression capabilities while preserving the data topology, providing a more robust feature representation for subsequent germplasm resource recommendation.

[0062] To comprehensively evaluate the applicability of different dimensionality reduction strategies, this study designed a systematic comparative experimental scheme: 1) linear dimensionality reduction based on PCA; 2) nonlinear dimensionality reduction using UMAP; and 3) a hybrid strategy combining PCA preprocessing with UMAP refined dimensionality reduction (PCA+UMAP). This multi-faceted comparative analysis not only verifies the superiority of the UMAP algorithm on biometric data but also provides a scientific basis for selecting dimensionality reduction methods in different application scenarios.

[0063] This study uses multi-dimensional clustering effectiveness indicators to quantitatively evaluate the dimensionality reduction results, specifically including the following four core indicators:

[0064] (1) Best number of clusters (Best K): The intrinsic clustering dimension of the data structure is determined by combining the elbow rule and the silhouette coefficient analysis, which reflects the natural grouping characteristics of the data in the dimensionality reduction space.

[0065] (2) Silhouette coefficient (SC): It comprehensively measures the compactness of samples within a cluster and the separation between clusters. Its value range is [-1,1]. SC>0.5 indicates the existence of a clear cluster structure, and SC>0.7 indicates a strong clustering pattern.

[0066] (3) Davies-Bouldin index (DBI): The clustering quality is evaluated based on the ratio of intra-cluster divergence to inter-cluster distance. DBI∈[0,∞). The smaller the value, the higher the inter-cluster discrimination. Usually, DBI<0.5 is considered a good clustering.

[0067] (4) Calinski-Harabasz index (CHI): The clustering effectiveness is measured by the ratio of between-group dispersion to within-group dispersion. CHI∈[0,∞), and the larger the value, the more significant the clustering structure. CHI>1000 usually indicates high-quality clustering.

[0068] By quantitatively evaluating the clustering quality, feature retention rate, and model prediction accuracy of the dimensionality-reduced data, the performance comparison results are shown in Table 3.

[0069] Table 3 Comparison of Dimensionality Reduction Effects

[0070] Dimensionality reduction methods Best K SC CHI DBI PCA 2 0.454 162.358 0.839 UMAP 2 0.717 617.452 0.247 PCA+UMAP 4 0.703 3144.102 0.368

[0071] The experimental results show that UMP has a much better dimensionality reduction effect than PCA. Although it has a higher CHI value (3144.10) than the PCA+UMAP hybrid method, UMAP exhibits better biological interpretability while maintaining the optimal number of clusters with K=2, avoiding over-segmentation that would increase the complexity of subsequent interpretation.

[0072] Note: UMAP dimensionality reduction is a manifold learning method designed to accurately represent the local structure of data and better integrate the global structure. In machine learning and data analysis, clustering is an important unsupervised learning method widely used in tasks such as image classification, bioinformatics, recommender systems, and anomaly detection. However, with the expansion of data scale and the increase in data dimensionality, traditional clustering algorithms face challenges such as the curse of dimensionality, low computational efficiency, and difficulty in feature selection. Existing linear dimensionality reduction methods (such as PCA) struggle to handle nonlinear data structures, while nonlinear methods (such as t-SNE) suffer from high computational complexity or insufficient preservation of global structure. UMAP, as an efficient nonlinear dimensionality reduction method, can find lower-dimensional manifolds that are easier to cluster, thereby improving the performance of traditional clustering algorithms. Its core idea is to map high-dimensional data to a lower-dimensional space while preserving the local and global topological structure of the data. UMAP consists of two key steps: fuzzy topological representation in the high-dimensional space and optimized embedding in the low-dimensional space. In the high-dimensional space, UMAP ensures the continuity of the local structure by calculating the k nearest neighbors of each point (controlled by the parameter n_neighbors) and constructing a fuzzy topological structure. In low-dimensional space, UMAP optimizes the cross-entropy loss function to make the fuzzy topological structure of the low-dimensional embedding as close as possible to the structure of the high-dimensional space. UMAP's advantages lie in its high computational efficiency, strong global structure preservation ability, and good parameter interpretability. Compared to t-SNE, UMAP can process large-scale data faster and balance local and global structures by adjusting n_neighbors. Its mathematical foundation is based on algebraic topology and geometric theory, making it suitable for tasks such as visualization, cluster preprocessing, and feature extraction.

[0073] Example 6, as Figure 4 , Figure 5 As shown, this paper presents an optimized K-Means clustering algorithm based on Genetic Algorithm (GA). Traditional genetic algorithms often suffer from parameter sensitivity and the trap of getting stuck in local optima. Parameter sensitivity arises because the performance of genetic algorithms is highly dependent on the selection of their internal parameters, such as population size, crossover probability, and mutation probability. Inappropriate parameter settings may lead to low algorithm efficiency or failure to find the global optimum. This study innovatively proposes an integrated optimization framework that drives clustering analysis through a genetic algorithm (GA), deeply integrates a hybrid initialization strategy of K-Means++ and Gaussian Mixture Model (GMM), and combines the L-BFGS-B Newton algorithm for fine-tuning the objective function.

[0074] Note: A genetic algorithm is a search heuristic algorithm that mimics the biological evolutionary process and is used to solve optimization and search problems. It constructs a simulated environment where candidate solutions ("individuals") compete for survival through fitness evaluation, with solutions having higher fitness having a greater chance of reproduction. Through this mechanism, the algorithm seeks to find the optimal or feasible solution within a given problem space. The core process includes: initializing the population and randomly generating N individuals, each represented as a chromosome (usually a binary string, real number vector, etc.); fitness evaluation, calculating the fitness value of each individual to reflect the quality of the solution; selecting high-quality individuals for the next generation based on fitness using roulette wheel or tournament methods; generating new offspring through crossover of the selected individuals; and randomly changing gene values ​​with probabilities to increase diversity through mutation. By repeatedly performing these steps, the genetic algorithm can gradually improve the quality of the solution after multiple generations of iteration, approaching the optimal solution. The fitness function typically increases with each generation, indicating the algorithm's effectiveness in solving specific problems.

[0075] (1) Hybrid initialization strategy:

[0076] Traditional genetic algorithms (GA) employ a completely random initialization strategy in clustering optimization tasks. While this mechanism ensures population diversity, it can easily lead to the initial population getting trapped in local optima, thus reducing convergence efficiency. To address this drawback, such as... Figure 4 As shown, this study proposes an innovative hybrid initialization framework that deeply integrates the probability density estimation advantages of the K-Means++ clustering algorithm and the Gaussian mixture model (GMM), and generates a high-quality initial population through a dual-mode collaborative initialization mechanism.

[0077] The core idea of ​​this strategy is to construct a dual-path initialization channel: when the population initialization process is started, each newborn individual determines its initialization path through a uniformly distributed random number generator based on a preset mixing ratio threshold (0.5 in this study).

[0078] K-Means++ initialization path:

[0079] If random number Individuals with values ​​less than the threshold of 0.5 are initialized using K-Means++. K-Means++ effectively avoids distribution bias caused by randomness by intelligently selecting initial cluster centers.

[0080] In K-Means++, suppose there is a dataset Where n is the number of samples and d is the number of features. Initially, a sample point is randomly selected. As the first cluster center, subsequent center points Perform distance-weighted selection:

[0081]

[0082] in Indicates sample Distance to the nearest selected center.

[0083] Through iterative optimization using a distance-weighted selection mechanism, the algorithm eventually converges to the optimal solution space represented by K cluster centers, forming a set of individuals with clear geometric meaning. This process dynamically adjusts sample weights, allowing the cluster centers to gradually approximate the essential structure of the data distribution.

[0084] GMM initialization path:

[0085] If random number Individuals enter the GMM initialization channel. This study assumes that the observed data are generated by a linear combination of the probability density functions of K heterogeneous Gaussian distributions. That is, the observed data... The probability density function is:

[0086]

[0087] in It is the first The mixing coefficients of the Gaussian components satisfy the following conditions: ,and The mean is The covariance matrix is The Gaussian probability density function.

[0088] The model parameter estimation employs the Expectation-Maximization (EM) algorithm, which consists of two alternating steps:

[0089] I. Expected Step (E-Step):

[0090] Based on the current parameter estimation Calculate each sample point Posterior probability of belonging to the k-th Gaussian component Construct the soft allocation matrix:

[0091]

[0092] II. Maximization Step (M-Step):

[0093] Based on the probability distribution obtained from E-Step Update the mean vector of the Gaussian components. Covariance matrix and mixing coefficient This maximizes the model's log-likelihood function. The update formulas are as follows:

[0094]

[0095]

[0096]

[0097] Through iterative optimization of the EM algorithm, GMM eventually converges to a set of K cluster centers (the mean vector of each Gaussian component). The optimal solution space is formed by .

[0098] Both GMM and K-Means++ share the same optimization objective: to uncover the intrinsic structural features of the data distribution. Their difference lies in their implementation paths: GMM models data using a probabilistic framework, while K-Means++ optimizes data using a geometric distance metric. The two initialization results are weighted and fused to form the final population, combining the fast convergence of K-Means++ with the deep adaptability of GMM to different data distributions.

[0099] In summary, when the random number is less than the threshold, individuals are initialized using K-Means++, which effectively avoids distribution bias caused by randomness by intelligently selecting initial cluster centers. Otherwise, they enter the GMM initialization channel, utilizing the Gaussian mixture model's ability to finely model the probability distribution of the sample space to generate initial solutions that conform to the inherent structure of the data. The two initialization results are weighted and fused to form the final population, retaining the fast convergence characteristics of K-Means++ while incorporating the deep adaptability of GMM to data distribution. The algorithm flow of the hybrid strategy is described in [link to algorithm description]. Figure 4 .

[0100] Note: K-means is a classic unsupervised learning clustering algorithm. Its core idea is to iteratively optimize the data into K clusters, ensuring that each data point belongs to the cluster corresponding to its nearest cluster centroid. This is achieved by alternating between assignment and update until convergence, ultimately minimizing the objective function. In assignment, each data point is assigned to the cluster corresponding to its nearest cluster centroid; in update, the centroid of each cluster is recalculated, and the mean of all points within that cluster is taken. Assume we have a dataset... Applying to the K-Means algorithm, during initialization, three initial centroids are randomly selected. The distance from each point to the centroid is calculated, and the point is assigned to the nearest cluster. The centroid positions are then recalculated and updated. This is one round of assignment and update, and subsequent assignments and updates are continuously performed until convergence. The K-Means algorithm is simple, easy to understand, and highly efficient. However, K-means also has some drawbacks: it requires pre-specifying the cluster K, and improper selection can lead to distorted results; it is sensitive to the initial cluster centers, and random initialization may result in local optima; it performs poorly on non-spherical clusters or clusters with large differences in size / density. K-means++ is an improved version of K-Means.

[0101] (2) L-BFGS-B algorithm:

[0102] During the evolutionary process of genetic algorithms, even high-quality individuals may still exhibit subtle suboptimal solutions. Traditional mutation operations struggle to achieve fine-tuning in this regard. Figure 5 As shown, the finite memory quasi-Newton algorithm (L-BFGS-B) is introduced as a local optimizer. Its advantages are that it can use second derivative information to accelerate convergence and support variable boundary constraints.

[0103] Assume the optimal population obtained from mixed initialization is After calculating fitness and crossover variation, the top 10% of individuals are selected. As the optimization target, for each elite individual, the cluster center matrix needs to be calculated first. Flattening as an optimization variable Let the cluster center matrix be... The size is The flattened optimization variables It is a length of The vector.

[0104] Next, boundary constraints are established. Assume that each dimension of each cluster center has upper and lower bounds. For the ... The first cluster center Dimension, its lower bound is The upper boundary is Then the optimization variable The boundary constraints are:

[0105]

[0106] Define the objective function It is related to the quality of the clustering results, and the calculation formula is:

[0107]

[0108] in, It belongs to the first A set of samples for each cluster. It is the first The center of each cluster.

[0109] Calculate the gradient of the objective function To iteratively update and optimize variables Let the current iteration point be... Then the next iteration point The update formula is based on the idea of ​​the quasi-Newton method:

[0110]

[0111] in, The step size is determined by the line search method, ensuring that the objective function decreases sufficiently in that direction. It is the search direction, derived from the inverse of the approximate Hessian matrix. and negative gradient Calculated, i.e. .

[0112] After each iteration, a convergence check is performed. A common convergence condition is that the change in the objective function is less than a certain threshold. ,Right now:

[0113]

[0114] Or the norm of the gradient is less than a certain threshold, such as:

[0115]

[0116] When the convergence condition is met, the iteration stops, and the optimal individual is obtained. If the condition is not met, the parameter vector is updated again, and the gradient of the objective function is calculated again. The optimization process is as follows: Figure 4 As shown.

[0117] Note: L-BFGS-B is an integrated algorithm that combines the quasi-Newton method (BFGS), the finite memory strategy (L-BFGS), and boundary constraint handling (B) techniques to efficiently solve complex optimization problems. In large-scale optimization problems, the objective function typically has millions of variables and must satisfy physical constraints. Traditional gradient descent methods struggle to handle these constraints, while Newton's method is computationally expensive. The quasi-Newton method solves the problem iteratively by constructing an approximate quadratic model of the objective function. While the BFGS algorithm updates the approximate Hessian inverse matrix using a correction formula, traditional BFGS has high memory overhead. L-BFGS stores the gradient difference and variable difference from the most recent preset iterations and uses a double-loop recursive formula to calculate the approximate search direction from the initial approximate matrix and the current gradient, significantly reducing memory requirements. By determining the effective set—the set of variable indices at the boundary whose search direction violates the constraints—the quadratic programming subproblem is solved under these constraints: the search direction that satisfies the constraints is obtained, ensuring that the iteration point is within the feasible region. The core of the L-BFGS-B algorithm lies in updating the simulated second derivative information using BFGS with limited memory, while utilizing projected gradients to ensure that iteration points do not violate boundary constraints. Taking the training of a neural network with a sigmoid output as an example, the weights need to be limited to [-10, 10] to prevent numerical overflow. L-BFGS-B converges quickly while ensuring constraints by truncating out-of-bounds weights and adjusting the search direction, whereas ordinary gradient descent may cause oscillations due to frequent out-of-bounds errors.

[0118] Example 7, as Figure 6 , Figure 7 , Figure 8 As shown, based on Examples 1-6, experiments and result analysis were conducted on the optimized and improved clustering algorithm of this study.

[0119] (1) Simulation experiment:

[0120] To systematically verify the performance of the proposed optimization algorithm in the wheat germplasm recommendation task, this study designed a set of rigorous comparative experiments to verify the effectiveness of the algorithm. Experimental environment: Windows 10 Home Chinese Edition, 11th Gen Intel(R) Core(TM) i7-11800H 2.30 GHz, 16.0 GB memory; program runtime environment: Python 3.10.11.

[0121] In the experiments, the GA-optimized clustering algorithm proposed in this study was compared with three advanced methods: UMAP dimensionality reduction and fusion method, simulated annealing optimization algorithm, and particle swarm optimization algorithm. All comparative experiments were set with the same termination condition, namely a maximum number of iterations of 500, to ensure the convergence of the algorithm.

[0122] This study uses a three-dimensional metric to evaluate model performance: accuracy, precision, and recall. The experimental design, through a controlled variable approach, systematically compares the clustering quality and recommendation accuracy of various algorithms in the wheat germplasm feature space within a unified computational framework, providing quantitative evidence for subsequent algorithm optimization and application promotion.

[0123] After the same number of iterations, the experimental results are as follows: As shown.

[0124] Table 4. Experimental Comparison Results

[0125] Algorithm Name Accuracy / (%) Precision / (%) Recall / (%) UMAP dimensionality reduction 80.80% 83.11% 80.81% Simulated annealing algorithm 95.50% 95.56% 95.50% Particle Swarm Optimization Algorithm 98.58% 98.61% 98.58% This research algorithm 99.41% 99.41% 99.40%

[0126] Experimental results show that the proposed GA-optimized clustering algorithm achieves significant performance improvements across all evaluation metrics. The algorithm achieves an accuracy of 99.41%, a 0.83 percentage point improvement over the second-best particle swarm optimization algorithm (98.58%). The algorithm maintains a high precision of 99.41%, expanding its advantage over the particle swarm optimization algorithm (98.61%) to 0.80 percentage points. Furthermore, the algorithm leads in recall with a 99.40% performance, exceeding the particle swarm optimization algorithm (98.58%) by 0.82 percentage points.

[0127] Experimental data fully demonstrate that the algorithm presented in this paper has comprehensive performance advantages in representation learning and recommendation prediction in wheat germplasm feature space, providing more accurate intelligent support for the efficient utilization of crop germplasm resources.

[0128] (2) Analysis and Discussion:

[0129] Accuracy represents the proportion of samples that the model correctly predicts out of the total sample. For example, if 99 out of 100 wheat varieties are correctly classified, the accuracy is 99%. This metric is the most intuitive and can reflect the overall predictive ability of the model.

[0130] like Figure 6 As shown, the algorithm in this study (red) has the largest proportion, forming a clear advantage area, while the FgFisNet algorithm (blue) has the smallest area and lags significantly behind. Other algorithms show a gradient distribution. The performance ranking is as follows: algorithm in this study > particle swarm optimization > simulated annealing > fusion UMAP dimensionality reduction > FgFisNet model.

[0131] Precision measures how many samples predicted as positive are actually positive. For example, if a model predicts 100 high-quality wheat varieties, and 95 of them are indeed high-quality, the precision is 95%. This metric reflects the reliability of the recommendations.

[0132] like Figure 7 As shown in the heatmap, the algorithm in this study has the darkest color block (99.41%, dark red), while the fused UMAP dimensionality reduction has the lightest color block (83.11%, light yellow). Performance ranking: Algorithm in this study > Particle Swarm Optimization > Simulated Annealing > Fusion UMAP Dimensionality Reduction.

[0133] Recall measures the percentage of true positive samples that are correctly predicted. The higher these three metrics, the better the model's performance. For example, if there are 100 high-quality wheat varieties and the model identifies 90, the recall rate is 90%. This metric reflects the comprehensiveness of the search.

[0134] like Figure 8 As shown, during the 0th to 500th iterations, the algorithm in this study (red line) consistently maintained the highest recall rate and achieved stable convergence. In contrast, the fused UMAP dimensionality reduction (gray line) exhibited greater fluctuations. Performance ranking: This algorithm > Particle Swarm Optimization > Simulated Annealing > Fusion UMAP Dimensionality Reduction.

[0135] (3) Ablation test:

[0136] To delve into the performance contributions of each key module in the proposed algorithm, this study designed a systematic ablation study. This study precisely controls the activation states of the algorithm components, constructing model variants with different configurations to quantitatively analyze the gain effect of each module on the overall recommendation performance.

[0137] The specific experimental design includes:

[0138] I. Basic Model: All algorithm components are fully retained, including UMAP feature dimensionality reduction, hybrid initialization strategy, and L-BFGS-B optimizer;

[0139] II. Variant Model 1: Remove the UMAP dimensionality reduction module and directly perform cluster analysis on the original feature space;

[0140] III. Variant Model 2: Replace the hybrid initialization strategy with traditional random initialization, while keeping other modules unchanged;

[0141] IV. Variant Model 3: Disable the L-BFGS-B optimization algorithm and use the standard gradient descent method for parameter updates.

[0142] Table 5 shows the results of the ablation experiment. By comparing the performance differences of each variant model with the base model in terms of accuracy, precision, and recall, the performance contribution of each algorithm component can be accurately analyzed. This experimental method follows the principle of controlling variables, providing a fine-grained analytical perspective for algorithm optimization and helping to understand the synergistic mechanism between modules.

[0143] Table 5 Ablation Experiment Results

[0144] Algorithm Name Accuracy / (%) Precision / (%) Recall / (%) UMAP Removal Dimensionality Reduction 92.89% 93.06% 92.89% Remove mixed initialization 99.05% 99.07% 99.05% Remove L-BFGS-B algorithm 99.05% 99.06% 99.05% This research algorithm 99.41% 99.41% 99.41%

[0145] Based on the experimental results, it can be seen that different components of the algorithm in this study have different impacts on the final performance.

[0146] Removing UMAP dimensionality reduction: Model accuracy decreased to 92.89%, while precision and recall decreased simultaneously to 93.06% and 92.89%, respectively. This verifies the crucial role of UMAP dimensionality reduction in feature space optimization, which effectively extracts essential data features through nonlinear manifold learning, providing more discriminative feature representations for subsequent clustering analysis.

[0147] Removing the hybrid initialization strategy resulted in a 0.36 percentage point decrease in model accuracy, precision, and recall (99.41% → 99.05%). This indicates that the hybrid initialization strategy, by combining the global exploration capabilities of K-Means++ with the probability density modeling advantages of GMM, provides a high-quality initial solution space for the optimization process, effectively avoiding the problem of traditional random initialization easily getting trapped in local optima.

[0148] Removing the L-BFGS-B optimization also resulted in a 0.36 percentage point decrease in model performance. This demonstrates that the L-BFGS-B algorithm played a crucial role in the parameter optimization phase. It achieved rapid convergence of the objective function through a quasi-Newton method, while the boundary constraint handling mechanism ensured the numerical stability of the optimization process, thereby improving the overall performance of the model.

[0149] The algorithm in this study achieved an excellent performance of 99.41% across all metrics, fully demonstrating the synergistic effect among the modules. UMAP dimensionality reduction constructs an effective feature space, hybrid initialization provides a robust starting point, and L-BFGS-B ensures efficient parameter optimization.

[0150] Example 8, as Figure 9 As shown, a wheat germplasm recommendation model was constructed and the recommendation results were analyzed.

[0151] (1) Recommendation model construction:

[0152] In this wheat germplasm recommendation system, a two-stage fusion strategy is employed to achieve accurate recommendations. First, the system checks if a city exists in the database. If not, it inputs the city's meteorological characteristics and matches them against existing city meteorological data in the database using cosine similarity. Second, after dimensionality reduction using UMAP, a multilayer perceptron (MLP) classification model is constructed. Supervised learning establishes a nonlinear mapping relationship between the feature space and wheat variety categories. This model uses a deep neural network structure to capture complex feature interactions and, after training, can effectively predict the probability distribution of the variety cluster to which an input sample belongs.

[0153] To further improve recommendation accuracy, a nearest neighbor retrieval mechanism based on cosine similarity is introduced. Specifically, for a sample to be recommended, its high-dimensional feature representation is first obtained through a trained MLP model, and then the cosine similarity between this feature and the pre-stored feature vectors of the cluster centers of each variety is calculated. Cosine similarity, as a standardized metric for measuring the consistency of vector direction, can effectively capture potential association patterns between features.

[0154] The final recommendation results are generated using a Top-K sorting strategy: candidate varieties are sorted in descending order based on similarity scores, and the top 5 varieties with the highest similarity are selected to form the recommendation list. The wheat recommendation algorithm process is as follows: Figure 9 As shown.

[0155] This hybrid recommendation paradigm of "coarse-grained classification + fine-grained similarity ranking" not only utilizes the global classification capability of the MLP model, but also enhances the personalization of the recommendation results through local similarity calculation, achieving a good balance between recommendation accuracy and computational efficiency.

[0156] (2) Recommendation Results and Analysis:

[0157] To validate the performance of the recommendation system, this study constructed a test set by randomly sampling 800 wheat germplasm records from the merged database. The climate type of the target city was input into the system as both a known and unknown condition. The key trait indicators and similarity scores of the Top-5 candidate germplasm generated by the recommendation system are shown in Tables 6 and 7.

[0158] Table 6. Recommended Results Based on Known Urban Meteorological Information

[0159] drought resistance Flood resistance Field cold resistance Crude protein / % Lysine / % Precipitation value / ml hardness Similarity Input germplasm 3.0 3.0 3.0 12.5 0.45 30 15 \ Top-1 2.0 3.0 3.0 16.32 0.45 37.0 21.5 99.72% Top-2 3.0 4.0 3.0 17.10 0.48 33.2 16.6 99.70% Top-3 2.0 2.0 2.0 16.11 0.46 35.5 16.3 99.70% Top-4 3.0 3.0 3.0 16.40 0.45 31.2 16.1 99.67% Top-5 3.0 3.0 3.0 16.87 0.51 32.0 16.7 99.65%

[0160] Table 7 Recommendation results for cities with unknown meteorological conditions

[0161] drought resistance Flood resistance Field cold resistance Crude protein / % Lysine / % Precipitation value / ml hardness / Similarity Input germplasm 3.0 2.0 2.0 15 0.5 25 20 \ Top-1 3.0 2.0 3.0 16.60 0.46 28.0 22.5 99.97% Top-2 3.0 2.0 3.0 16.80 0.48 27.2 21.7 99.96% Top-3 4.0 2.0 3.0 16.69 0.46 28.0 22.7 99.96% Top-4 4.0 3.0 3.0 16.69 0.47 28.0 22.7 99.95% Top-5 3.0 3.0 3.0 17.62 0.50 28.2 22.4 99.94%

[0162] To visually demonstrate the comparison between the recommended results and the various attributes of the input wheat germplasm, the table above is shown below. Figure 10 , Figure 11 As shown.

[0163] like Figure 10As shown, the system successfully captured the multi-dimensional trait association patterns of wheat germplasm. The crude protein content (16.32%) and sedimentation value (37.0 ml) of the top-1 recommended germplasm were significantly higher than those of the input germplasm, which corresponds to its highest similarity score of 99.72%, indicating that the system can recommend new germplasm with potential improvement value while maintaining trait similarity. Notably, the top-2 to top-5 candidate germplasms exhibited rich genetic diversity in indicators such as lysine content (0.45%-0.51%) and hardness index (16.1-16.7), providing a multi-objective optimization selection space for breeding decisions.

[0164] like Figure 11 As shown, the crude protein content (16.60%-17.62%) and hardness (21.7-22.7) of all recommended germplasms were significantly improved, with Top-5 showing the best performance in crude protein (17.62%) and lysine (0.50%). The drought resistance of Top-3 and Top-4 improved to 4.0, and the waterlogging resistance of Top-4 and Top-5 improved to 3.0. It is particularly noteworthy that Top-4 maintained the stability of other traits while improving two stress resistance indicators (drought resistance +1.0 and waterlogging resistance +1.0).

[0165] Experimental data demonstrate that the recommendation system can not only accurately match adaptive trait combinations under target climatic conditions, but also effectively mine germplasm resources with complementary traits, which is of great significance for improving the regional adaptability and quality of wheat cultivation.

[0166] Based on the embodiments 1-8 above, this study innovatively integrates genetic algorithms and K-Means clustering, combined with UMAP dimensionality reduction and L-BFGS-B optimization techniques, to construct an intelligent wheat germplasm recommendation system. At the theoretical level, a hybrid initialization strategy and boundary constraint optimization method are proposed, effectively solving the initial sensitivity and local optima problems of traditional clustering algorithms. At the application level, the generated recommendation model can accurately match wheat varieties with regional climate characteristics, providing scientific decision support for breeding experts. This research not only promotes interdisciplinary innovation between agricultural informatics and crop science, but also provides a practical technical solution for achieving efficient utilization of germplasm resources and ensuring national food security, possessing significant practical value for promoting the development of smart agriculture.

[0167] The above description is only for illustrating the present invention and should be understood as not being limited to the above embodiments. Various modifications that conform to the spirit of the present invention are within the protection scope of the present invention.

Claims

1. A wheat germplasm recommendation algorithm based on GA optimized clustering, characterized in that, Includes the following steps: S1: Obtain wheat germplasm data of the target province from the existing crop germplasm information network, and obtain meteorological data of cities in the target province for the past two years from the existing meteorological website. For missing data in the entire column of the two datasets, use deletion method; for missing partial data, use KNN to fill in; for useful non-numerical data, use label encoding to convert it into numerical data; for outliers, use box-line method to find them and replace them with the average value. S2: Classify wheat germplasm data from two dimensions: resistance characteristics and growth characteristics, and select important features. Construct a high-dimensional feature space through feature interaction, superposition of meteorological features, standardization, and urban coding to enhance feature expression capabilities. S3: For the final high-dimensional features, the UMAP dimensionality reduction strategy is adopted. UMAP is used to capture non-linear structures and map high-dimensional features to low-dimensional space, providing a low-noise and highly interpretable feature representation for subsequent genetic algorithm optimization of clustering. S4: The K-means++ clustering algorithm is optimized and improved by optimizing the cluster centers through a hybrid initialization strategy. The hybrid initialization strategy and the finite memory quasi-Newton algorithm (L-BFGS-B) are designed to iteratively optimize the clustering algorithm, further improving its performance and providing high-quality labels for the multilayer perceptron. S5: A wheat germplasm recommendation model is established based on multilayer perceptron and cosine similarity. According to the climate conditions of the selected city and the needs of breeding experts, the dimensionality-reduced features are associated with clustering labels. First, the multilayer perceptron model is used to predict the potential adaptation category of the target variety, and then cosine similarity is used to screen the most matching varieties within the same category.

2. The wheat germplasm recommendation algorithm based on GA optimized clustering according to claim 1, characterized in that, The feature interactions in step S2 include: By employing multinomial feature expansion or deep neural network automatic cross-multiplication methods, interactive feature modeling is performed to deeply explore the nonlinear correlation mechanism between features, thereby further improving the prediction accuracy and generalization ability of the recommendation model.

3. The wheat germplasm recommendation algorithm based on GA optimized clustering according to claim 1, characterized in that, The hybrid initialization strategy in step S4 includes: A hybrid initialization framework is proposed, which deeply integrates the probability density estimation advantages of K-Means++ clustering algorithm and Gaussian mixture model (GMM) to generate a high-quality initial population through a dual-mode collaborative initialization mechanism. When the population initialization process is started, if the random number of each new individual is less than the preset mixing ratio threshold, the individual adopts the K-Means++ initialization path. This method effectively avoids the distribution deviation caused by randomness by intelligently selecting the initial cluster center. Conversely, the GMM initialization path is used, which utilizes the Gaussian mixture model's ability to finely model the probability distribution of the sample space to generate an initial solution that conforms to the inherent structure of the data. The two types of initialization results are combined through a weighted fusion mechanism to form a hybrid initialization optimal population, which retains the fast convergence characteristics of K-Means++ and incorporates the deep adaptability of GMM to data distribution.

4. The wheat germplasm recommendation algorithm based on GA optimized clustering according to claim 3, characterized in that, The finite-memory quasi-Newton algorithm (L-BFGS-B) in step S4 includes: S4-1: Assume the optimal population obtained from mixed initialization is... After calculating fitness and crossover variation, the top 10% of individuals are selected. As an optimization target; S4-2: For each elite individual, the cluster center matrix needs to be calculated first. Flattening as an optimization variable Let the cluster center matrix be... The size is The flattened optimization variables It is a length of ; S4-3: Establish boundary constraints. Assume that each dimension of each cluster center has upper and lower bounds. For the i-th... The first cluster center Dimension, its lower bound is The upper boundary is Then the optimization variable The boundary constraints are: (1) Define the objective function It is related to the quality of the clustering results, and the calculation formula is: (2) in, It belongs to the first A set of samples for each cluster. It is the first The center of each cluster; S4-4: Calculate the gradient of the objective function To iteratively update and optimize variables Let the current iteration point be... Then the next iteration point The update formula is based on the idea of ​​the quasi-Newton method: (3) in, The step size is determined using a line search method, ensuring that the objective function decreases sufficiently in that direction. It is the search direction, derived from the inverse of the approximate Hessian matrix. and negative gradient Calculated, i.e. ; S4-5: After each iteration, perform a convergence check. A common convergence condition is that the change in the objective function is less than a certain threshold. ,Right now: (4) Or the norm of the gradient is less than a certain threshold, such as: (5) When the convergence condition is met, the iteration stops and the optimal individual is obtained. If the condition is not met, the parameter vector is updated again and the gradient of the objective function is calculated again.

Citation Information

Cited By

  • SDGs-based irrigation area water grain ecosystem sustainable development evaluation system method

    CN121660256A