Agricultural park planting and breeding planning method and system based on multi-model fusion

By employing a multi-model fusion approach and utilizing VAE, GMM, and LightGBM algorithms, the problems of reliance on experience and insufficient feature mining in agricultural park planting and breeding planning were solved. This approach enabled accurate identification of resource endowments and quantification of industrial potential, thereby improving the scientific nature of planning and the efficiency of resource utilization.

CN121860148APending Publication Date: 2026-04-14SICHUAN AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SICHUAN AGRI UNIV
Filing Date
2026-02-04
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing agricultural park planning methods rely on expert experience and lack objective quantitative standards. Traditional algorithms struggle to extract deep niche characteristics and lack theoretical-current situation comparisons, resulting in highly subjective planning results, insufficient feature mining, weak model generalization ability, and difficulty in accurately identifying resource endowment patterns and quantifying industrial potential.

Method used

A multi-model fusion approach is adopted, including variational autoencoder (VAE) to extract deep niche features, Gaussian mixture model (GMM) for resource endowment pattern recognition, lightweight gradient booster (LightGBM) for industry suitability prediction, and gap analysis to diagnose supply and demand mismatch and generate decision support charts.

Benefits of technology

It enables precise identification of resource endowment patterns and prediction of industry suitability for agricultural parks, breaking through the linear limitations of traditional methods, providing objective zoning governance and quantitative potential mining, and improving the scientific nature of planning and the efficiency of resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121860148A_ABST
    Figure CN121860148A_ABST
Patent Text Reader

Abstract

The invention discloses an agricultural park planting and breeding planning method and system based on multi-model fusion, and relates to the technical field of intelligent agriculture. The method comprises the following steps: firstly, collecting multi-source heterogeneous data such as climate, terrain and soil of a park; deep nonlinear ecological niche features are extracted by using a variational auto-encoder (VAE), and a resource endowment mode is objectively identified through a Gaussian mixture model (GMM); and then constructing a LightGBM multi-label prediction model to calculate the industrial theoretical suitability, executing supply and demand mismatch diagnosis and difference analysis (Gap analysis) in combination with the actual planting situation, and quantitatively mining the industrial planting expansion potential. According to the method, the problems of high subjectivity, insufficient feature mining and difficulty in potential quantification of traditional planning are effectively solved, a scientific decision from objective resource classification to precise industrial layout is realized, and the agricultural resource allocation efficiency is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of smart agriculture and agricultural big data processing technology, specifically to a method for agricultural production layout planning and decision support using artificial intelligence technology.

[0002] More specifically, this invention relates to an agricultural park planting and breeding planning method and system that integrates variational autoencoder (VAE), Gaussian mixture model (GMM) and lightweight gradient booster (LightGBM) algorithms, combined with niche theory, to identify natural resource endowments, evaluate industrial suitability and quantitatively mine development potential of agricultural parks. Background Technology

[0003] Currently, the digital economy has become a key engine driving global economic growth, but in the agricultural sector, the digital transformation process is relatively lagging. As pioneers in the development of modern agriculture, national modern agricultural industrial parks bear the important responsibility of leading the structural reform of agricultural supply and improving the level of agricultural modernization. Scientific and rational planting and breeding planning and production layout are prerequisites for leveraging the agglomeration effect of industrial parks, achieving efficient resource utilization, and ensuring sustainable industrial development.

[0004] However, existing methods for agricultural park planting and breeding planning and suitability evaluation still have the following significant problems in practical applications: 1. Reliance on expert experience and lack of objective quantitative standards. Traditional planning relies on qualitative analysis by experts, which makes it difficult to accurately analyze the complex nonlinear coupling relationships between natural elements such as climate and soil. This results in highly subjective planning outcomes that are prone to adaptation biases in different regions.

[0005] 2. Insufficient depth of algorithm mining and difficulty in feature extraction. Faced with high-dimensional, heterogeneous and noisy agricultural environmental data, traditional statistical or simple clustering algorithms (such as K-Means) are unable to extract deep niche features, resulting in inaccurate identification of resource endowment patterns.

[0006] 3. Lack of theoretical-current situation comparison makes it impossible to quantify potential. Existing evaluations only focus on theoretical suitability (basic ecological niche), neglecting quantitative comparison with actual planting (real ecological niche), and lacking an effective difference calculation mechanism (Gap). This makes it impossible to quantify the potential for industrial expansion, and decision-making is difficult to extend to the level of layout optimization.

[0007] 4. The models have weak generalization ability and lack a general framework. Most existing models are designed for single crops or specific regions and lack a general framework that can cover multiple varieties and adapt to multiple habitats. They are difficult to simultaneously take into account unsupervised resource identification and supervised suitability prediction.

[0008] Therefore, there is an urgent need for an intelligent planning method that can deeply mine multi-source heterogeneous environmental data, objectively identify resource endowment patterns, accurately predict industrial suitability, and quantitatively reveal the potential for industrial development, so as to provide strong technical support for the scientific layout and precise policy implementation of agricultural parks. Summary of the Invention

[0009] To overcome the technical problems in existing agricultural park planning, such as reliance on experience-based judgment with strong subjectivity, insufficient feature mining of high-dimensional nonlinear environmental data by traditional algorithms, and difficulty in quantifying industrial development potential due to the lack of quantitative comparison between theory and current status, this invention provides an agricultural park planting and breeding planning method and system based on multi-model fusion.

[0010] To achieve the above objectives, the present invention adopts the following technical solution: an agricultural park planting and breeding planning method based on multi-model fusion, comprising the following steps: Step S1: Acquisition and Preprocessing of Multi-Source Heterogeneous Data. Collect multi-dimensional natural environment data and actual planting and breeding data for the target agricultural park. The natural environment data includes climate data (such as average annual temperature, precipitation, and humidity), topographic data (such as altitude and slope), and soil data (such as type and texture). Standardize numerical features and perform one-hot encoding on categorical features to construct a multimodal feature matrix.

[0011] Step S2: Niche Feature Reconstruction Based on VAE. A variational autoencoder (VAE) model is constructed. The preprocessed multimodal feature matrix is ​​input into the encoder, and it is compressed into latent variables (z) in a low-dimensional latent space through nonlinear mapping. These latent variables serve as the denoised deep niche feature vector, used to characterize the complex natural resource endowment of the park.

[0012] Step S3: Resource endowment pattern identification based on Gaussian Mixture Model (GMM). The niche feature vector is input into the Gaussian Mixture Model (GMM). Based on the Bayesian Information Criterion (BIC), the optimal number of clusters is automatically selected, and probabilistic soft clustering is performed on the target park to identify multiple resource endowment patterns with significant heterogeneity.

[0013] Step S4: Industry Suitability Prediction Based on LightGBM. A multi-label classification prediction model based on the LightGBM algorithm is constructed. The original environmental features, niche feature vectors, and resource endowment pattern categories are used as joint inputs, and the actual crop categories grown in the park are used as labels for training. The model outputs the predicted probability value of each agricultural industry under the current environment, which represents the theoretical basic niche suitability of the crop.

[0014] Step S5: Supply-Demand Mismatch Diagnosis and Potential Exploration (Gap Analysis). Obtain the actual planting status of the park (i.e., the current ecological niche). Calculate the difference between the theoretical suitability probability ( ) and the actual planting status ( ) for each agricultural industry to obtain the potential index. If the Gap value is significantly positive, the industry is determined to be a vacant ecological niche with development potential, meaning the environment in the area is suitable but not yet fully developed; if the Gap value is negative or close to zero, it is determined to be of low suitability or saturated.

[0015] Step S6: Decision Support Visualization. Based on the above calculation results, generate decision support charts: use box plots to show the differences in environmental characteristics of each resource mode; use heat maps to show the industry suitability matrix under each resource mode; use bubble charts to show the distribution of potential indices for each industry, and formulate differentiated zoning governance and industry expansion strategies accordingly.

[0016] The beneficial effects of this invention are as follows: 1. Breaking through linear limitations, feature extraction is more accurate. The introduction of a variational autoencoder (VAE) replaces the traditional direct feature extraction method, effectively solving the challenges of high dimensionality, high noise, and nonlinearity in agricultural environmental data. By extracting deep latent variables, the crop's living environment (niche) can be more fundamentally characterized, significantly improving the accuracy of subsequent clustering and prediction.

[0017] 2. Objective pattern identification enables more scientific zoning management. By utilizing the Gaussian Mixture Model (GMM) combined with the BIC criterion, a fully data-driven resource endowment pattern identification was achieved, overcoming the blind spots of manual experience-based zoning. It can accurately identify physically significant resource types such as high-altitude cold-climate and hot-humid-rainy-abundant-rainmate types, providing an objective basis for differentiated agricultural zoning management.

[0018] 3. Pioneering Gap Analysis for More Quantifiable Potential Discovery. This invention innovatively proposes a Gap analysis method based on the comparison between the basic ecological niche (predicted value) and the actual ecological niche (actual value). It not only answers what is suitable to plant, but also what is lacking, and where planting is not only suitable but also currently undeveloped. By visually displaying the potential index through bubble charts, it effectively identifies resource-industry mismatch areas, providing a quantifiable tool for industrial restructuring and diversified expansion within the park.

[0019] 4. Complete decision-making loop and strong versatility. A complete technical loop is constructed, from feature perception to pattern recognition to decision optimization. This method is not limited to specific crops or regions, has strong generalization ability, and can be widely applied to agricultural planning scenarios at different scales, significantly improving the scientific nature of agricultural production layout and resource utilization efficiency. Attached Figure Description

[0020] Figure 1This is a schematic diagram of the overall process architecture of the agricultural park planting and breeding planning method and system based on multi-model fusion provided in the embodiments of the present invention; the diagram shows the technical logic from the perception layer (VAE feature extraction), the cognition layer (GMM pattern recognition) to the decision layer (LightGBM prediction and gap potential mining).

[0021] Figure 2 This is a schematic diagram of the Bayesian Information Criterion (BIC) optimization curve used to determine the optimal number of clusters in a Gaussian Mixture Model (GMM) according to an embodiment of the present invention; the figure shows the trend of the BIC value and the optimal inflection point as the number of clusters K changes.

[0022] Figure 3 This is a scatter plot of the cluster distribution of resource endowment patterns after t-SNE dimensionality reduction based on VAE latent variable features in an embodiment of the present invention; the figure shows the clustering and separation effect of different agricultural parks in the low-dimensional niche feature space.

[0023] Figure 4 This is a box plot illustrating the distribution differences of different resource endowment patterns (cluster categories) identified in this embodiment of the invention in terms of temperature, altitude, humidity, and precipitation characteristics; this plot is used to explain the specific physical environment meaning of each resource pattern.

[0024] Figure 5 This is a schematic diagram of the ROC performance evaluation curve of the LightGBM multi-label suitability prediction model on the test set in an embodiment of the present invention; the figure shows the accuracy (AUC value) of the model in predicting the suitability of various crops.

[0025] Figure 6 This is a heat map of the crop industry theoretical suitability (basic ecological niche) matrix generated under different resource endowment modes in the embodiments of the present invention; the color intensity in the figure represents the probability of theoretical suitability predicted by the model.

[0026] Figure 7 This is a schematic diagram of the distribution of agricultural industry development potential (vacant ecological niche) calculated based on gap analysis in an embodiment of the present invention; the size of the bubble in the diagram represents the expansion potential index of the industry under the corresponding resource mode. Detailed Implementation

[0027] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. These embodiments are implemented based on the technical solution of the present invention, providing detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.

[0028] Method Implementation Process Step S1: Acquisition and Preprocessing of Multi-Source Heterogeneous Data A multidimensional environmental database was constructed by selecting typical national-level modern agricultural industrial parks in my country. The collected data included climate data such as average annual temperature, annual precipitation, and average air humidity; geomorphological data such as altitude and slope; and soil data such as soil type and texture. The current dominant industries of the parks were also recorded as training labels. For the collected raw data, numerical features (such as temperature and altitude) were standardized using Z-Score to eliminate dimensional differences, and categorical features (such as soil type) were converted into numerical vectors using One-Hot encoding. Finally, the processed data from each dimension were concatenated to construct a high-dimensional multimodal input feature matrix. X .

[0029] Step S2: Niche Feature Reconstruction Based on VAE To address the issues of high noise and nonlinearity in agricultural environmental data, a variational autoencoder (VAE) is constructed to extract deep features. The feature matrix is ​​then processed... X The input encoder maps the mean and log-variance of the latent space, and introduces random noise for reparameterization sampling to obtain latent variables. These latent variables are the denoised low-dimensional niche feature vectors. The decoder then maps them back to the original space to calculate the reconstruction error. The model is trained by jointly minimizing the reconstruction error and the KL divergence.

[0030] Step S3: Resource endowment pattern recognition based on GMM Cluster analysis was performed using the extracted feature vectors and a Gaussian Mixture Model (GMM). First, the optimal number of clusters was determined based on the Bayesian Information Criterion (BIC), such as... Figure 2 As shown, the BIC value reaches the optimal inflection point when K=5, therefore the park is divided into 5 resource patterns. The clustering effect is as follows: Figure 3 As shown, the scatter plot after t-SNE dimensionality reduction demonstrates good spatial separation of the five data categories. Further combining... Figure 4 The box plots are used to analyze the physical meaning of each pattern. For example, category 1 shows significant high altitude and low temperature characteristics (high mountain and cold type), while category 4 shows high temperature and high humidity characteristics and low altitude characteristics (humid and hot plain type), thus achieving an objective classification of the park's resource endowment.

[0031] Step S4: Industry Suitability Prediction Based on LightGBM A multi-label classification model based on LightGBM was constructed, and a OneVsRest strategy was used to train binary classifiers separately for each industry. The original environmental features, latent variables, and cluster labels were concatenated as joint input. The ROC curve of the model on the test set is shown below. Figure 5 As shown, the predicted AUC values ​​for each industry remain at a high level, indicating that the model can accurately output the theoretical suitability probability (basic niche) of each industry.

[0032] Step S5: Supply-demand mismatch diagnosis and potential exploration (Gap analysis) This is the core step of the invention. Obtain the actual planting status vector of the park and calculate the potential index. Generate a decision chart based on the calculation results. Figure 6 The heatmap shows the average theoretical suitability of each resource pattern for different industries; Figure 7 The bubble chart displays the potential index of each industry. The larger the bubble, the higher the theoretical suitability of the area but the lower the actual planting area (large vacant ecological niche). Decision-makers can use this to identify key information, such as the great potential for tea cultivation in Category 2 areas, and thus formulate differentiated industry optimization strategies.

[0033] Example 2 Based on the above method, this embodiment also provides an agricultural park planting and breeding planning system, including: a data acquisition module for collecting and storing climate, topography, and planting and breeding data; a feature reconstruction module for extracting niche features using a VAE model; a pattern recognition module for identifying resource types using a GMM algorithm; a suitability prediction module for calculating theoretical suitability using a LightGBM model; a potential mining module for performing gap analysis to calculate potential indices; and a visualization and interaction module for generating and displaying, for example, Figures 2 to 7 The decision support chart shown.

[0034] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. All modifications made within the principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for agricultural park planting and breeding planning based on multi-model fusion, characterized in that, Includes the following steps: S1: Acquisition and preprocessing of multi-source heterogeneous data, collecting multi-dimensional natural environment data and actual planting and breeding data of the target agricultural park, and constructing a multi-modal feature matrix; S2: Niche feature reconstruction based on variational autoencoder: A variational autoencoder model is constructed, and the multimodal feature matrix is ​​input into the encoder, mapping it to latent variables in a low-dimensional latent space, which serve as the denoised niche feature vector. S3: Resource endowment pattern recognition based on Gaussian mixture model: The niche feature vector is input into the Gaussian mixture model, and the optimal number of clusters is determined based on the Bayesian information criterion, dividing the target agricultural park into multiple resource endowment patterns with significant heterogeneity. S4: Industry suitability prediction based on LightGBM: A multi-label suitability prediction model is constructed. Using the original multidimensional natural environment data, niche feature vectors, and resource endowment patterns as joint inputs, the theoretical suitability probability of each agricultural industry in the current environment is predicted; S5: Supply and demand mismatch diagnosis and potential mining, obtaining the actual planting and breeding status of the park, calculating the difference between the theoretical suitability probability of each agricultural industry and the actual planting and breeding status, obtaining the potential index, and identifying vacant ecological niches based on the potential index; S6: Decision support visualization, generating visual charts based on the resource endowment patterns, theoretical suitability probabilities, and potential indices to guide the optimization of agricultural production layout.

2. The agricultural park planting and breeding planning method based on multi-model fusion according to claim 1, characterized in that, In step S1, the multidimensional natural environment data includes climate data, topographic data, and soil data; the data preprocessing includes: Z-Score standardization for numerical features and one-hot encoding for categorical features.

3. The agricultural park planting and breeding planning method based on multi-model fusion according to claim 1, characterized in that, In step S2, the construction process of the variational autoencoder includes: mapping the input data to the mean vector and log-variance vector of the latent distribution using an encoder network; introducing standard normal distribution noise for sampling using a reparameterization technique to generate the latent variables; mapping the latent variables back to the original data space using a decoder network to obtain reconstructed data; and training and optimizing the model by minimizing the sum of the reconstruction error and the KL divergence. According to claim 1, the method for planning agricultural park planting and breeding based on multi-model fusion is characterized in that, in step S3, determining the optimal number of clusters based on the Bayesian information criterion specifically includes: training Gaussian mixture models within a preset range of cluster numbers; calculating the Bayesian information criterion values ​​corresponding to different cluster numbers; and selecting the cluster number corresponding to when the Bayesian information criterion value reaches an inflection point and is relatively small as the optimal number of clusters.

4. The agricultural park planting and breeding planning method based on multi-model fusion according to claim 1, characterized in that, In step S4, the multi-label suitability prediction model uses the LightGBM algorithm and combines it with the OneVsRest strategy to train a binary classifier for each agricultural industry category and output the theoretical suitability probability of that industry.

5. The agricultural park planting and breeding planning method based on multi-model fusion according to claim 1, characterized in that, In step S5, the formula for calculating the potential index is: where is the potential index, is the theoretical suitability probability, and is the actual planting and breeding state, with a value of 0 or 1; if it is positive and greater than the preset threshold, then the industry is determined to be an advantageous industry with the potential for expansion.

6. The agricultural park planting and breeding planning method based on multi-model fusion according to claim 1, characterized in that, In step S6, the visualization charts include at least: a box plot to show the differences in environmental characteristics under different resource endowment models; a heat map to show the distribution of theoretical suitability of each industry under each resource endowment model; and a bubble chart to show the size of the potential index of each industry under each resource endowment model.

7. An agricultural park planting and breeding planning system based on multi-model fusion, characterized in that, include: The data acquisition and preprocessing module is used to acquire multidimensional natural environment data and actual planting and breeding data of the target agricultural park and perform preprocessing. The module includes a feature reconstruction module with a built-in variational autoencoder for extracting deep niche feature vectors from the data; a pattern recognition module with a built-in Gaussian mixture model for identifying resource endowment patterns in the park; a suitability prediction module with a built-in LightGBM model for calculating the theoretical suitability probability of each agricultural industry; a potential mining module for performing differential analysis, calculating potential indices, and identifying vacant niches; and a visualization and interaction module for generating and displaying decision support charts.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 7.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 7.