Intelligent agricultural yield decision tree prediction method and system

By performing cluster analysis with spatial neighborhood constraints on field plots and training local decision tree models, local prediction feature maps are generated. Finally, a global decision tree ensemble model is used for prediction, which solves the problem of yield prediction bias caused by spatial heterogeneity in traditional methods and achieves accurate agricultural yield prediction.

CN121860160APending Publication Date: 2026-04-14FUJIAN AGRI VOCATIONAL & TECH COLLEGE
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-16
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Traditional smart agriculture yield decision tree prediction methods cannot effectively address the heterogeneity of yield-driving mechanisms caused by spatially continuous gradual or abrupt changes due to factors such as soil fertility, microclimate, and irrigation management, resulting in systemic biases in local areas.

Method used

By acquiring the spectral and textural features of fields during key phenological periods, spatial neighborhood-constrained clustering analysis is performed to form multiple spatially continuous field clusters. Local decision tree models are trained independently to generate local prediction feature maps. Finally, yield prediction is performed through a global decision tree ensemble model.

Benefits of technology

It enables accurate yield prediction by leveraging the contextual information of fields within complex spatial patterns under spatial heterogeneity interference, avoiding prediction bias caused by neglecting spatial relationships in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121860160A_ABST
    Figure CN121860160A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent agricultural yield decision tree prediction method and system. The method comprises the following steps: forming an original feature set of a field; performing clustering analysis on all the field parcels based on the original feature set of the field parcels and the spatial positions of the field parcels to form a plurality of field parcel clusters; training a local decision tree prediction model of each field block cluster according to the historical yield data of all the field blocks corresponding to each field block cluster; inputting the field original feature set of any to-be-predicted field in the target area into all the local decision tree prediction models in parallel, and further generating a local prediction feature map; and carrying out vector splicing on the field original feature set of the to-be-predicted field and the local prediction feature map to obtain a fusion feature vector, and further training to obtain a global decision tree integration model for whole-region yield prediction so as to complete yield prediction of the to-be-predicted field. By adopting the scheme provided by the invention, yield prediction can be carried out through the context associated information of the field in the complex spatial pattern under the interference of spatial heterogeneity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of agricultural yield prediction technology, and more specifically, to a smart agricultural yield decision tree prediction method and system. Background Technology

[0002] Agricultural yield forecasting refers to the process of dynamically monitoring and simulating farmland environmental factors and crop physiological states by integrating multi-source information such as satellite remote sensing data acquisition, meteorological monitoring networks, soil sensor arrays, and historical agricultural databases, combined with crop growth models, machine learning algorithms, and spatiotemporal statistical analysis techniques, thereby generating a quantitative yield forecast report at the regional or field scale.

[0003] In traditional smart agriculture yield prediction using decision tree models, the approach relies heavily on a uniform feature-yield relationship model for all fields. This approach struggles to effectively address the heterogeneity of yield-driving mechanisms caused by the continuous, gradual, or abrupt changes in factors such as soil fertility, microclimate, and irrigation management. For instance, traditional methods use a global decision tree model to predict yields for all fields, and the model parameters reflect the average pattern across the entire region. This model cannot accurately characterize or adapt to the unique relationships within local spatial sub-regions that deviate significantly from the average pattern. This leads to systematic biases in highly heterogeneous local areas. Therefore, how to predict yields using the contextual information of fields within complex spatial patterns under spatial heterogeneity has become a challenge for the industry. Summary of the Invention

[0004] This application provides a smart agriculture yield decision tree prediction method and system, which can predict yield by using the contextual information of fields in complex spatial patterns under spatial heterogeneity interference.

[0005] Firstly, this application provides a smart agriculture yield decision tree prediction method, comprising the following steps: The spectral and textural features of each field within the target area during key phenological periods are obtained to form the original feature set of the fields; Based on the original feature set of the fields and the spatial location of the fields, a cluster analysis with spatial neighborhood constraints is performed on all fields to form multiple spatially continuous field clusters. Based on the historical yield data of all fields corresponding to each field cluster and the corresponding original feature set of the fields, the local decision tree prediction model of each field cluster is trained independently. The original feature set of any field to be predicted within the target area is input in parallel into all trained local decision tree prediction models, and then the local prediction feature map of the corresponding field to be predicted is generated by the prediction value of each local decision tree prediction model. The original feature set of the field to be predicted is concatenated with the local predicted feature map to obtain a fused feature vector. Then, a global decision tree ensemble model for yield prediction of the whole region is trained by the fused feature vector and historical yield data to complete the yield prediction of the field to be predicted.

[0006] In some embodiments, performing clustering analysis with spatial neighborhood constraints on all fields based on the original feature set of the fields and the spatial location of the fields to form multiple spatially continuous field clusters specifically includes: Based on the spatial location of the fields, adjacency relationships are defined, and a spatial weight matrix representing spatial neighborhood constraints is generated. The original feature set of the fields is fused with the spatial weight matrix to construct a joint similarity matrix for spatially constrained clustering; Clustering and quality assessment are performed on the joint similarity matrix to determine the optimal partitioning scheme, thereby obtaining multiple spatially continuous field clusters.

[0007] In some embodiments, training a local decision tree prediction model for each field cluster independently, based on the historical yield data of all fields corresponding to each field cluster and the corresponding original feature set of the fields, specifically includes: Historical yield data corresponding to all fields within each field cluster and feature data corresponding to the original feature set of the fields are extracted to construct a local training sample set for each field cluster. Each local training sample set is preprocessed to obtain a preprocessed local training sample subset. A decision tree model is trained for each preprocessed subset of local training samples to obtain the local decision tree prediction model for each field cluster.

[0008] In some embodiments, inputting the original feature set of any field to be predicted within the target area into all trained local decision tree prediction models in parallel specifically includes: Based on the identifier of the field to be predicted, the feature data corresponding to the field to be predicted is extracted from the original feature set of the field; The feature data is synchronously input into all trained local decision tree prediction models for forward inference, thereby obtaining the predicted yield of the field to be predicted under different local decision tree prediction models.

[0009] In some embodiments, generating a local prediction feature map of the corresponding field to be predicted using the predicted values ​​of each local decision tree prediction model specifically includes: All predicted values ​​are sorted according to the fixed order of the field clusters corresponding to each local decision tree prediction model; The sorting results are associated with the spatial coordinates of the fields to be predicted; Based on the mapping results, a local prediction feature map is constructed for the corresponding field to be predicted.

[0010] In some embodiments, concatenating the original feature set of the field to be predicted with the local predicted feature map to obtain a fused feature vector specifically includes: The original feature vector is generated by using the original feature set of the field to be predicted; The prediction feature vector of the field to be predicted is determined based on the local prediction feature map; The original feature vector and the predicted feature vector are concatenated to obtain a fused feature vector.

[0011] In some embodiments, a global decision tree ensemble model for yield prediction across the entire region is trained using the fused feature vectors and historical yield data to complete the yield prediction for the field to be predicted. Specifically, this includes: The fused feature vectors corresponding to all fields with historical yield data are collected, and a global training sample set is constructed using the historical yield data as labels. A decision tree ensemble model is trained based on the global training sample set as a global decision tree ensemble model for full-region output prediction. The fused feature vector of the field to be predicted is input into the trained global decision tree ensemble model to complete the yield prediction of the field to be predicted.

[0012] In some embodiments, the critical phenological period refers to a specific growth stage in the crop life cycle in which the crop's morphology, physiology, and yield formation have a decisive influence on the final yield, and whose canopy spectral reflectance characteristics have significant distinguishability.

[0013] In some embodiments, the spectral characteristics refer to the energy reflected by ground objects to different electromagnetic bands as detected by remote sensing sensors.

[0014] Secondly, this application provides a smart agricultural yield decision tree prediction system, comprising: The acquisition module is used to acquire the spectral and textural features of each field within the target area during key phenological periods, forming the original feature set of the field. The processing module is used to perform clustering analysis with spatial neighborhood constraints on all fields based on the original feature set of the fields and the spatial location of the fields, forming multiple spatially continuous field clusters; The processing module is also used to independently train the local decision tree prediction model for each field cluster based on the historical yield data of all fields corresponding to each field cluster and the corresponding original feature set of the fields. The processing module is also used to input the original feature set of any field to be predicted in the target area into all trained local decision tree prediction models in parallel, and then generate the local prediction feature map of the corresponding field to be predicted through the prediction values ​​of each local decision tree prediction model. The execution module is used to concatenate the original feature set of the field to be predicted with the local prediction feature map to obtain a fused feature vector. Then, the fused feature vector and historical yield data are used to train a global decision tree ensemble model for yield prediction of the whole region, so as to complete the yield prediction of the field to be predicted.

[0015] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects: The smart agriculture yield decision tree prediction method and system provided in this application firstly divides the entire field area into multiple spatially continuous field clusters through cluster analysis based on spatial neighborhood constraints. This decomposes the complex, global spatial heterogeneity problem into multiple locally homogeneous sub-regions, and establishes a local decision tree prediction model for each sub-region. This process transforms the spatially continuous, gradually changing, or abruptly changing yield-driving mechanism heterogeneity into independently modelable locally homogeneous units through clustering, providing a data foundation for accurately characterizing the local feature-yield relationship. Subsequently, local decision tree prediction models are trained based on historical data of each field cluster. These models can capture the unique yield-driving patterns within their respective clusters and adapt to local soil fertility, microclimate, and other conditions. Finally, the original feature set of the field to be predicted is input in parallel into all local models to generate a local prediction feature map. This feature map quantifies the relationship between the field and different local region models. The method first determines the degree of conformity, thereby transforming abstract contextual information into concrete, quantifiable, multi-dimensional predicted value vectors, realizing the mathematical representation transformation from discrete field features to continuous spatial correlation information. Then, the original feature set of the fields is concatenated with the local predicted feature map to obtain a fused feature vector, which simultaneously contains the field's own attributes and its contextual information within the complex spatial pattern. Subsequently, a global decision tree ensemble model is trained based on the fused feature vector. This model, through an ensemble learning mechanism, uses the prediction results of the local model as enhancement features, thereby dynamically integrating spatial contextual cues during the decision-making process, avoiding prediction bias caused by neglecting spatial correlation in traditional methods. Finally, yield prediction is performed using the global decision tree ensemble model. In summary, this scheme can predict yield under spatial heterogeneity interference by utilizing the contextual information of fields within a complex spatial pattern. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating a smart agriculture yield decision tree prediction method according to some embodiments of this application.

[0017] Figure 2 This is a schematic diagram of the process for determining field clusters according to some embodiments of this application.

[0018] Figure 3 This is a schematic diagram of the process for determining the original feature vector according to some embodiments of this application.

[0019] Figure 4 This is a schematic diagram of the structure of a smart agriculture yield decision tree prediction system according to some embodiments of this application.

[0020] Figure 5 This is an internal structural diagram of a computer device for implementing a smart agriculture yield decision tree prediction method according to some embodiments of this application. Detailed Implementation

[0021] To better understand the technical solutions in this embodiment, the technical solutions in this embodiment will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0022] refer to Figure 1 The figure is a flowchart illustrating a smart agriculture yield decision tree prediction method according to some embodiments of this application. The smart agriculture yield decision tree prediction method mainly includes the following steps: In step 101, the spectral and textural features of each field within the target area during the key phenological period are obtained to form the original feature set of the field.

[0023] In practice, the process begins by acquiring multi-temporal satellite remote sensing images covering the target area during the crop growing season. Image data with cloud cover meeting requirements and corresponding to key crop phenological stages (such as jointing and heading stages) are then selected and preprocessed with radiometric calibration and atmospheric correction to obtain true reflectance data for ground features. Next, the geometrically corrected remote sensing images are spatially overlaid with the field vector boundary data of the target area. For each independent field polygon, the average reflectance of its internal pixels is calculated, and based on a pre-defined spectral band combination formula, the normalized mean reflectance (NMR) is calculated. A series of vegetation indices, such as the vegetation index and the enhanced vegetation index, are used as spectral features to characterize crop growth. At the same time, red-edge or near-infrared bands that are sensitive to the spatial structure of vegetation are selected, or grayscale images are generated directly using NDVI calculation results. The gray-level co-occurrence matrix method is used to calculate statistical quantities such as texture contrast, homogeneity, and entropy of each field as the analysis unit, which are used as texture features to characterize the spatial distribution rules of the crop canopy. Finally, all spectral and texture feature values ​​corresponding to each field are sorted and aggregated according to the field ID to construct a structured original feature set of the field.

[0024] It should be noted that the key phenological period mentioned in this application refers to a specific growth stage in the crop life cycle in which the morphology, physiology, and yield formation have a decisive influence on the final yield, and whose canopy spectral reflectance characteristics have significant distinguishability; the spectral characteristics refer to the quantitative indicators closely related to biophysical parameters such as vegetation leaf area index and chlorophyll content, generated by combining and calculating the reflected energy of ground objects to different electromagnetic bands detected by remote sensing sensors through specific mathematical formulas (such as vegetation indices); the texture characteristics refer to the texture features obtained by analyzing the spatial variation frequency and regularity of pixel gray values ​​in remote sensing images (such as through gray-level co-occurrence matrix algorithms). The statistical quantities used to quantify the spatial structural attributes of crop canopy, such as roughness and uniformity, are obtained. The growing season remote sensing images refer to a sequence of digital images, periodically captured by a satellite platform during the entire growth cycle of the target crop from sowing to harvest, containing multiple spectral bands from visible light to shortwave infrared. The original feature set of the field is a structured data table with a single field as the basic recording unit, and its spectral and texture features during key phenological periods as attribute fields. Its physical significance lies in digitally representing the crop growth status and spatial heterogeneity of each field, providing initial input for subsequent cluster analysis and yield prediction modeling.

[0025] In step 102, based on the original feature set of the fields and the spatial location of the fields, a cluster analysis with spatial neighborhood constraints is performed on all fields to form multiple spatially continuous field clusters.

[0026] In some embodiments, reference Figure 2 As shown in the figure, this is a schematic flowchart illustrating the process of determining field clusters according to some embodiments of this application. The formation of multiple spatially contiguous field clusters by performing cluster analysis with spatial neighborhood constraints on all fields based on the original feature set and spatial location of the fields can be achieved through the following steps: First, in step 1021, adjacency relationships are defined based on the spatial location of the fields, and a spatial weight matrix representing spatial neighborhood constraints is generated. Then, in step 1022, the original feature set of the field plots is fused with the spatial weight matrix to construct a joint similarity matrix for spatially constrained clustering; Finally, in step 1023, clustering and quality assessment are performed on the joint similarity matrix to determine the optimal partitioning scheme, thereby obtaining multiple spatially continuous field clusters.

[0027] In specific implementation, defining adjacency relationships based on the spatial location of the fields and generating a spatial weight matrix representing spatial neighborhood constraints can be achieved in the following way: First, calculate the geometric center coordinates of each field based on its vector boundary data; then, calculate the Euclidean distance between all pairs of fields based on the geometric center coordinates; next, determine the spatial adjacency relationship between fields using a shared boundary method or a distance threshold method. The shared boundary method means that if the vector polygons of two fields have a common edge with a length greater than zero, they are considered adjacent. The distance threshold method means that a preset distance threshold is used; if the geometric polygons of two fields have a common edge with a length greater than zero, they are considered adjacent. If the Euclidean distance between the centroids is less than the threshold, they are considered adjacent. Finally, a symmetric matrix is ​​constructed based on the spatial adjacency relationship. If field i and field j are considered adjacent, a non-zero weight value is assigned to the position of the i-th row and j-th column and the j-th row and i-th column of the matrix; otherwise, the weight value is zero. This generates the spatial weight matrix. Preferably, the non-zero weight value can be set according to the length of the common edge of the two adjacent fields or the reciprocal of the centroid distance to reflect the strength of the spatial connection. In other embodiments, the spatial adjacency relationship can also be defined based on the spatial k-nearest neighbor relationship or Delaunay triangulation. This application does not limit this.

[0028] It should be noted that the spatial weight matrix mentioned in this application refers to a symmetric matrix used to quantify the spatial proximity between all pairs of fields within the target area and to serve as the mathematical carrier of the spatial neighborhood constraint.

[0029] In specific implementation, fusing the original feature set of the fields with the spatial weight matrix to construct a joint similarity matrix for spatially constrained clustering can be achieved in the following ways: First, the original feature set of the fields is standardized, for example, using the Z-score standardization method to eliminate the influence of different feature dimensions; then, based on the standardized features, the feature similarity between fields is calculated to construct a feature similarity matrix. The calculation of feature similarity can use cosine similarity or a similarity measurement method based on Gaussian kernel function; subsequently, the feature similarity matrix is ​​fused with the spatial weight matrix. The fusion operation includes, but is not limited to, multiplying corresponding elements of the two matrices or performing a linear weighted combination of the two matrices; the joint similarity matrix is ​​obtained through the fusion operation; preferably, when performing a linear weighted combination, an adjustable fusion coefficient can be configured for feature similarity and spatial weight respectively to control their relative importance in the joint similarity; in other embodiments, a graph-based fusion method can also be used, for example, first converting the spatial weight matrix into a graph Laplacian matrix and then combining it with the feature similarity matrix. This application does not limit this.

[0030] It should be noted that the joint similarity matrix mentioned in this application refers to a comprehensive matrix obtained by fusing the feature similarity matrix representing the similarity of spectral and texture features of fields with the spatial weight matrix representing spatial proximity. Its physical significance lies in providing a unified similarity measure that simultaneously encodes feature similarity and spatial proximity between fields for spatially constrained clustering.

[0031] In specific implementation, clustering and quality assessment of the joint similarity matrix to determine the optimal partitioning scheme, thereby obtaining multiple spatially continuous field clusters, can be achieved in the following way: First, a spectral clustering algorithm is used to solve the joint similarity matrix. This algorithm includes constructing a Laplacian matrix, performing eigenvalue decomposition, and partitioning the fields in the space spanned by the first K eigenvectors. Then, the partitioning quality under different cluster numbers K is evaluated using silhouette coefficients. The calculation of the silhouette coefficients depends on the average similarity between each field and other fields in the same cluster, as well as... The average similarity of the plots in the nearest neighbor cluster is used to determine the optimal number of clusters. The optimal number of clusters is then used to execute the spectral clustering algorithm, outputting the final plot division result as the multiple spatially continuous plot clusters. Preferably, when evaluating the number of clusters, the Davidson-Bolding index or Calinski-Harabasz index can be combined with the silhouette coefficient for comprehensive judgment. In other embodiments, clustering algorithms capable of directly handling spatial constraints, such as the SKATER algorithm or ClustGeo algorithm, can be used to solve the joint similarity matrix or directly the standardized features and spatial weights; this application does not limit this. Furthermore, after obtaining the multiple spatially continuous plot clusters, their spatial continuity quality can be further verified to ensure that the clustering results meet the design goal of spatial neighborhood constraints. This can be achieved by, for example: calculating the spatial fragmentation index of each plot cluster; for each plot cluster, identifying all the plot polygons it contains, and calculating the number of independent subgraphs that are not connected in the geometric set formed by these polygons; and then combining the number of independent subgraphs with the total number of plots in the cluster. The ratio of the numbers is used as the spatial fragmentation index of the cluster. If the spatial fragmentation index of a cluster is greater than a preset threshold (e.g., 1, which means that the cluster is not spatially connected), it is determined that the cluster does not meet the spatial continuity requirement. The spatial weight strength or adjacency definition threshold in the clustering algorithm can be adjusted, and the clustering analysis can be performed again. Preferably, the preset threshold can be set to 1, which strictly requires that all fields in each field cluster must be spatially connected. In other embodiments, spatial autocorrelation test based on Moran's index can also be used to evaluate the spatial consistency within each cluster. This application does not limit this.

[0032] It should be noted that the spatially continuous field clusters mentioned in this application refer to the grouping of fields formed by the cluster analysis with spatial neighborhood constraints, in which the internal fields are adjacent or close to each other in terms of spatial geographical location. This application adaptively divides the target area into several spatially continuous sub-regions that are relatively uniform in terms of crop growth characteristics, providing a spatial framework for the subsequent establishment of a yield prediction model that adapts to local patterns.

[0033] In step 103, based on the historical yield data of all fields corresponding to each field cluster and the corresponding original feature set of the fields, the local decision tree prediction model of each field cluster is trained independently.

[0034] In some embodiments, training a local decision tree prediction model for each field cluster independently, based on the historical yield data of all fields corresponding to each field cluster and the corresponding original feature set of the fields, can be achieved through the following steps: Historical yield data corresponding to all fields within each field cluster and feature data corresponding to the original feature set of the fields are extracted to construct a local training sample set for each field cluster. Each local training sample set is preprocessed to obtain a preprocessed local training sample subset. A decision tree model is trained for each preprocessed subset of local training samples to obtain the local decision tree prediction model for each field cluster.

[0035] In specific implementation, extracting the historical yield data corresponding to all fields within each field cluster and the corresponding feature data from the original feature set of the fields to construct a local training sample set for each field cluster can be achieved in the following way: For each field cluster in multiple spatially continuous field clusters, based on the list of field identifiers contained in the field cluster, retrieve and extract the historical yield data of these fields in the corresponding year from the stored historical yield database. At the same time, based on the same list of field identifiers, index and extract the feature data of these fields in all feature dimensions from the original feature set of the fields. Subsequently, use the historical yield data extracted from each field as the yield label of the sample, and use all its corresponding feature data as the feature vector of the sample, and align them according to the field samples. Finally, gather the feature vectors of all field samples within the field cluster with their corresponding yield labels to construct a local training sample set for the field cluster.

[0036] It should be noted that the local training sample set mentioned in this application refers to the data set constructed by the historical yield data (as labels) of all fields in the cluster and the feature data (as input) corresponding to the original feature set of the fields for each spatially continuous field cluster.

[0037] In specific implementation, each local training sample set is preprocessed to obtain a preprocessed local training sample subset. This can be achieved in the following ways: First, outlier detection and processing are performed on the feature data in the local training sample set. For example, a statistical quantile-based method is used to identify potential outlier feature values, and they are corrected or removed according to preset rules. Next, scale normalization is performed on the processed feature data. For example, a minimum-maximum normalization method is used to linearly map the values ​​of each feature dimension to between zero and one, or a Z-score normalization method is used to adjust each feature dimension so that its mean is zero and its standard deviation is one. At the same time, the output labels of the samples are checked, and any obvious outliers or missing values ​​are processed accordingly. Finally, the feature data processed above and the corrected output labels are recombined to form the preprocessed local training sample subset.

[0038] In specific implementation, a decision tree model is trained based on each preprocessed subset of local training samples to obtain the local decision tree prediction model for each field cluster. This can be achieved in the following way: For each preprocessed subset of local training samples corresponding to a field cluster, all its feature data are used as model input variables, and its corresponding yield label is used as the model prediction target; a decision tree-based machine learning algorithm is selected as the base model for training, such as a classification and regression tree algorithm, or its ensemble boosting algorithm such as gradient boosting decision tree; during the training process, cross-validation is used to search and evaluate in a pre-defined model hyperparameter combination space to determine the model suitable for the current local training samples. The optimal hyperparameter combination of the subset is used, whereby the hyperparameters may include the maximum depth of the tree, the minimum number of samples required for internal node splitting, etc. The optimal hyperparameter combination is used to complete the final training of the decision tree model on the corresponding local training sample subset. After training, the obtained model is saved as the local decision tree prediction model corresponding to the field cluster. Preferably, the cross-validation method can be K-fold cross-validation, and the overfitting prevention index (such as the mean square error of the validation set) is used as the target of hyperparameter optimization. In other embodiments, the decision tree model can also be replaced by other tree-based ensemble models, such as random forests, or Bayesian optimization methods can be used for automatic hyperparameter search. This application does not limit this.

[0039] It should be noted that the local decision tree prediction model described in this application refers to a decision tree-type machine learning model specifically used to predict the yield of the field cluster and its spatially adjacent fields, which is independently trained based on the local training sample set corresponding to the spatially continuous field cluster.

[0040] In step 104, the original feature set of any field to be predicted within the target area is input in parallel to all the trained local decision tree prediction models, and then the local prediction feature map of the corresponding field to be predicted is generated by the prediction value of each local decision tree prediction model.

[0041] In some embodiments, the original feature set of any field to be predicted within the target area can be input in parallel into all trained local decision tree prediction models by the following steps: Based on the identifier of the field to be predicted, the feature data corresponding to the field to be predicted is extracted from the original feature set of the field; The feature data is synchronously input into all trained local decision tree prediction models for forward inference, thereby obtaining the predicted yield of the field to be predicted under different local decision tree prediction models.

[0042] In specific implementation, the extraction of feature data corresponding to the field to be predicted from the original feature set of the field to be predicted can be achieved in the following way, for example: First, obtain the unique field identifier of the field to be predicted; then, use the field identifier as the query key to search in the structured original feature set data table of the field to locate the unique data record row that matches the identifier; next, obtain the feature values ​​of all defined spectral and texture features from the located data record row; finally, arrange all feature values ​​according to their predetermined feature order in the original feature set of the field to form a feature data vector containing all feature dimensions, which is used as the feature data corresponding to the field to be predicted; preferably, a data integrity verification step can be added before reading the feature values, for example, checking whether there are missing feature values ​​in the record row. If so, interpolation or the feature mean of the field cluster to which the field belongs can be used to fill in the missing values; in other embodiments, if the original feature set of the field is stored in the form of a key-value database or a spatial database, the feature data can also be obtained by directly calling the query statement through the database query interface. This application does not limit this.

[0043] It should be noted that the feature data corresponding to the field to be predicted mentioned in this application refers to the set of all feature values ​​extracted from the original feature set of the structured field, based on the identifier of the specific field, which represent the spectral and textural growth status of the crop in the key phenological period. It digitally represents the crop growth status of the field to be predicted in the current growing season in vector form.

[0044] In specific implementation, the feature data is synchronously input into all trained local decision tree prediction models for forward inference, thereby obtaining the predicted yield of the field under different local decision tree prediction models. This can be achieved in the following way: First, the feature data vector of the field to be predicted is fed into each of the local decision tree prediction models; then, the prediction process of each model is started. Each model, based on its own trained decision rules (including node splitting features, splitting thresholds, and leaf node output values), performs top-down discrimination and mapping on the input feature data vector, and finally outputs a value representing the predicted yield; by waiting and collecting all models... After completing their prediction processes, each outputs its predicted yield value, resulting in a set of predicted yield values ​​equal to the total number of local decision tree prediction models. This set represents the predicted yield values ​​of the field to be predicted under different local decision tree prediction models. Preferably, before feeding the feature data vector into the model, the feature data vector can be uniformly standardized based on the same standardized parameters used during the training of each local decision tree prediction model to ensure that the distribution of input data is consistent with that of model training data. In other embodiments, the synchronous input and inference process can be implemented through multi-threading or a distributed computing framework to improve the processing efficiency for a large number of fields to be predicted, and this application does not limit this.

[0045] It should be noted that the predicted yield values ​​of the field to be predicted under different local decision tree prediction models mentioned in this application refer to the set of yield prediction values ​​independently output by each model based on its own learned local patterns after the feature data of the field to be predicted are input into each of the local decision tree prediction models.

[0046] In some embodiments, generating local prediction feature maps of the corresponding fields to be predicted from the predicted values ​​of each local decision tree prediction model can be achieved by the following steps: All predicted values ​​are sorted according to the fixed order of the field clusters corresponding to each local decision tree prediction model; The sorting results are associated with the spatial coordinates of the fields to be predicted; Based on the mapping results, a local prediction feature map is constructed for the corresponding field to be predicted.

[0047] In specific implementation, sorting all predicted values ​​according to the fixed order of the field clusters corresponding to each local decision tree prediction model can be achieved in the following way: First, obtain the preset number sequence of the multiple spatially continuous field clusters; then, according to the order of this number sequence, select the predicted values ​​output by the local decision tree prediction model corresponding to the number from the obtained set of yield prediction values; then, place the selected predicted values ​​into a new ordered set according to their corresponding number order; finally, the ordered set is the sequence of predicted values ​​sorted according to the fixed order of the field clusters; preferably, the preset number sequence of the field clusters can be generated and stored immediately after the clustering analysis step is completed; in other embodiments, if the local decision tree prediction model already contains its corresponding field cluster number when stored, it can also directly sort the model according to this number and read its output value, which is not limited in this application.

[0048] In specific implementation, the association mapping between the sorting result and the spatial coordinates of the field to be predicted can be achieved in the following way, for example: First, obtain the geometric center coordinates of the field to be predicted; then, create a data mapping entry, using the geometric center coordinates as the index key of the entry; next, store the sorted predicted value sequence as the data value associated with the index key in the data mapping entry, thereby completing the association mapping from spatial location to predicted value sequence; the geometric center coordinates can be calculated by querying the vector boundary data of the field; preferably, the data mapping entry can be implemented in the form of key-value pairs, database records, or specific data structures; in other embodiments, the geometric center coordinates and the predicted value sequence can also be merged into a record with a unified structure, which is not limited in this application.

[0049] It should be noted that the association mapping described in this application refers to the data operation process of establishing a one-to-one correspondence between the ordered sequence of yield prediction values ​​and the spatial location coordinates (such as the geometric center coordinates) of the corresponding field to be predicted.

[0050] In specific implementation, constructing a local prediction feature map of the corresponding field to be predicted based on the mapping result can be achieved in the following way, for example: First, extract the sequence of predicted values ​​corresponding to the geometric center coordinates of the field to be predicted from the result of the association mapping; regard this sequence as a multi-dimensional feature vector, where each dimension of the vector corresponds to the number of the field cluster, and its value is the predicted value corresponding to the cluster; then, encapsulate this multi-dimensional feature vector into an independent data object; finally, define this data object as the local prediction feature map of the field to be predicted; the local prediction feature map physically represents the prediction performance spectrum of the field to be predicted under various local yield models represented by different field clusters; preferably, the geometric center coordinates can be retained as metadata of the data object during encapsulation, but this application does not limit this.

[0051] It should be noted that the local prediction feature map mentioned in this application refers to a multi-dimensional feature vector composed of multiple local model prediction value sequences arranged in a fixed order corresponding to the field to be predicted. Its physical meaning is that it is a meta-feature formed by the fusion of evaluation opinions of multiple "local experts", which characterizes the comprehensive yield positioning of the field to be predicted under the global spatial heterogeneity pattern.

[0052] In step 105, the original feature set of the field to be predicted is concatenated with the local prediction feature map to obtain a fused feature vector. Then, a global decision tree ensemble model for yield prediction of the whole region is trained by the fused feature vector and historical yield data to complete the yield prediction of the field to be predicted.

[0053] In some embodiments, the original feature set of the field to be predicted and the local predicted feature map are concatenated to obtain a fused feature vector, which can be achieved by the following steps: The original feature vector is generated by using the original feature set of the field to be predicted; The prediction feature vector of the field to be predicted is determined based on the local prediction feature map; The original feature vector and the predicted feature vector are concatenated to obtain a fused feature vector.

[0054] For specific implementation, refer to Figure 3As shown in the figure, this is a flowchart illustrating the process of determining the original feature vector in some embodiments of this application. The original feature vector can be generated from the original feature set of the field to be predicted in the following ways: for example, based on the identifier of the field to be predicted, query and locate the corresponding data record row from the structured original feature set data table; read the specific values ​​of the field to be predicted in all spectral and textural features from the record row; organize these feature values ​​according to the inherent feature arrangement order defined in the original feature set of the field to form a one-dimensional numerical array as the original feature vector; preferably, when reading the feature values, numerical validity verification can be performed, and missing values ​​can be filled using the average value of the corresponding feature of the field cluster to which the field belongs; in other embodiments, if the original feature data has been pre-stored in vector form, it can also be directly called through the identifier index, which is not limited in this application.

[0055] It should be noted that the original feature vector mentioned in this application refers to a one-dimensional vector directly formed by organizing all the spectral and texture feature values ​​corresponding to the field to be predicted in the original feature set of the field in an inherent order.

[0056] In specific implementation, determining the predicted feature vector of the field to be predicted based on the local predicted feature map can be achieved in the following ways, for example: First, obtain the completed local predicted feature map of the field to be predicted; the local predicted feature map has been defined as a well-encapsulated data object in the previous steps, the core of which is an ordered multi-dimensional feature vector, with each dimension of the vector corresponding to the predicted yield value of each field cluster in order; then, parse or extract the ordered multi-dimensional feature vector from the data object; use this extracted multi-dimensional feature vector directly as the predicted feature vector; preferably, if the local predicted feature map is stored in a format with metadata, the metadata needs to be stripped during extraction, and only the numerical vector part needs to be retained; in other embodiments, if the local predicted feature map has been stored as an independent vector file when it was generated, the contents of that file can also be directly read as the predicted feature vector, and this application does not limit this.

[0057] It should be noted that the predicted feature vector mentioned in this application refers to a multi-dimensional vector composed purely of the predicted values ​​of each local model, which is parsed from the local predicted feature map of the field to be predicted.

[0058] In specific implementation, the original feature vector and the predicted feature vector are concatenated to obtain the fused feature vector. This can be achieved in the following way: for example, the original feature vector and the predicted feature vector are concatenated end-to-end in the numerical dimension. Specifically, all elements of the predicted feature vector are appended sequentially to all elements of the original feature vector to form a new one-dimensional numerical array with a dimension equal to the sum of the two. This newly formed one-dimensional numerical array is the fused feature vector. Preferably, before concatenation, the dimensional information of the two vectors can be verified to ensure the correctness of the concatenation operation in terms of data structure. In other embodiments, other vector fusion methods can also be used, such as weighting and scaling a vector before concatenation, or using a more complex interactive feature generation method. This application does not limit this to any particular method.

[0059] It should be noted that the fused feature vector mentioned in this application refers to a new feature vector formed by concatenating the original feature vector and the predicted feature vector in terms of dimension. This application constructs an enhanced feature representation that simultaneously contains field growth status observation information and multi-view local model evaluation information.

[0060] In some embodiments, a global decision tree ensemble model for yield prediction across the entire region is trained using the fused feature vectors and historical yield data. This yield prediction for the field to be predicted can be achieved through the following steps: The fused feature vectors corresponding to all fields with historical yield data are collected, and a global training sample set is constructed using the historical yield data as labels. A decision tree ensemble model is trained based on the global training sample set as a global decision tree ensemble model for full-region output prediction. The fused feature vector of the field to be predicted is input into the trained global decision tree ensemble model to complete the yield prediction of the field to be predicted.

[0061] In specific implementation, the global training sample set can be constructed by collecting the fused feature vectors corresponding to all fields with historical yield data and using the historical yield data as labels. For example, firstly, obtain the identifiers of all fields with historical yield records within the target area; for each field identifier, generate its corresponding fused feature vector according to the aforementioned steps; simultaneously, extract the yield data of the field from the historical yield database as the real label based on the field identifier and its corresponding year; subsequently, pair and align the fused feature vector corresponding to each field with its extracted real yield label; finally, summarize all the paired "fused feature vector-yield label" data combinations to construct the global training sample set. Preferably, during the construction process, a sample partitioning strategy consistent with that used when generating local prediction feature maps can be adopted to ensure that the data used for training and evaluation are independent of each other and to prevent bias in model evaluation. In other embodiments, if the data is continuous in the time dimension, the training set and test set can also be divided according to the chronological order to verify the model's temporal generalization ability. This application does not limit this.

[0062] It should be noted that the global training sample set mentioned in this application refers to a supervised learning dataset composed of the fused feature vectors corresponding to all fields with historical yield records within the target area and their actual historical yield labels.

[0063] In specific implementation, training a decision tree ensemble model based on the global training sample set as a global decision tree ensemble model for full-region output prediction can be achieved in the following way: First, the global training sample set is divided into a training subset for model parameter learning and a validation subset for evaluating generalization performance; a decision tree ensemble algorithm, such as gradient boosting decision tree, is selected as the base algorithm of the global decision tree ensemble model; then, a search space containing the key hyperparameters of the decision tree ensemble algorithm is defined, including the number of decision trees in the ensemble, the maximum depth of a single decision tree, the minimum number of samples required for a leaf node, etc.; cross-validation is used to optimize the hyperparameters in the search space. Specifically, the training subset is further divided into K folds, and K-1 folds are used sequentially. The model is trained on the data, and its performance is evaluated on the remaining 1-fold data. This process is repeated K times to obtain a stable performance estimate for each hyperparameter combination. By comparing the average performance metrics (such as negative mean squared error) of different hyperparameter combinations during cross-validation, the hyperparameter combination with the best performance is selected. Then, the decision tree ensemble algorithm is retrained on the entire training subset using the optimal hyperparameter combination to obtain the final model. After training, the obtained model is saved as the global decision tree ensemble model. Preferably, the cross-validation process can use a grid search or random search strategy to traverse the hyperparameter search space. In other embodiments, more efficient hyperparameter tuning methods such as Bayesian optimization can also be used, or regularization terms and early stopping methods can be introduced during training to control model complexity. This application does not limit this.

[0064] In specific implementation, the fused feature vector of the field to be predicted is input into the trained global decision tree ensemble model to complete the yield prediction of the field to be predicted. This can be achieved in the following way: First, prepare the fused feature vector of the field to be predicted, and ensure that its dimensions and order are completely consistent with the feature definitions used during model training; then, input the fused feature vector into the global decision tree ensemble model; after receiving the input, each base decision tree integrated within the global decision tree ensemble model independently judges and traverses the input features according to the splitting rules from the root node to the leaf node, and finally each tree outputs a preliminary yield prediction value for the field; then, the model aggregates all the preliminary prediction values ​​according to its preset ensemble rules (for regression tasks, usually the arithmetic mean of the output values ​​of all base decision trees), and calculates and generates a final yield prediction value as the yield prediction result of the field to be predicted, thereby completing the yield prediction of the field to be predicted; in other embodiments, the model can also provide contribution analysis of different feature dimensions (including original spectral texture features and local prediction features) to the current prediction result, which is not limited in this application.

[0065] It should be noted that the global decision tree ensemble model described in this application refers to a decision tree ensemble model trained based on the global training sample set, which is used to make a final yield decision by integrating the field's own characteristics and local model prediction information.

[0066] It should be noted that, to understand how the global decision tree ensemble model utilizes the fused feature vector for decision-making, feature importance analysis can be performed after model training. For example, a feature importance evaluation method based on permutation importance or Gini impurity can be used to calculate the contribution of each feature dimension (including the original spectral texture features and the predicted values ​​of each dimension from the local prediction feature map) in the fused feature vector to the final yield prediction result. Through analysis, it can be found that the feature dimensions from the local prediction feature map often have significant importance rankings, indicating that the global decision tree ensemble model effectively learns and utilizes expert opinion information provided by different spatial local models. In a physical sense, the local prediction feature map provides the global model with contextual information representing the similarity relationship between the field and multiple local regions, thereby improving the prediction accuracy and robustness of the model in complex spatial heterogeneous environments. In other embodiments, the model can also be observed under what conditions it relies more on the original features and under what conditions it relies more on the local prediction features by analyzing the specific decision paths in the global model; this application does not limit this.

[0067] Furthermore, in another aspect of this application, in some embodiments, this application provides a smart agricultural yield decision tree prediction system, with reference to... Figure 4 The figure is a schematic diagram of the structure of a smart agricultural yield decision tree prediction system according to some embodiments of this application. The smart agricultural yield decision tree prediction system 200 includes: an acquisition module 201, a processing module 202, and an execution module 203, which are described below: The acquisition module 201 in this application is mainly used to acquire the spectral and textural features of each field in the target area during the key phenological period, and to form the original feature set of the field. Processing module 202, in this application, is mainly used to perform clustering analysis with spatial neighborhood constraints on all fields based on the original feature set of the fields and the spatial location of the fields, forming multiple spatially continuous field clusters; In addition, the processing module 202 in this application is also used to independently train the local decision tree prediction model of each field cluster based on the historical yield data of all fields corresponding to each field cluster and the corresponding original feature set of the fields. In addition, the processing module 202 in this application is also used to input the original feature set of any field to be predicted in the target area into all trained local decision tree prediction models in parallel, and then generate the local prediction feature map of the corresponding field to be predicted through the prediction values ​​of each local decision tree prediction model. The execution module 203 in this application is mainly used to concatenate the original feature set of the field to be predicted with the local prediction feature map to obtain a fused feature vector, and then train a global decision tree ensemble model for the whole region yield prediction through the fused feature vector and historical yield data to complete the yield prediction of the field to be predicted.

[0068] In addition, this application also provides a computer device, the computer device including a memory and a processor, the memory storing code, and the processor being configured to acquire the code and execute the above-described smart agriculture yield decision tree prediction method.

[0069] In some embodiments, reference Figure 5 This figure is an internal structural diagram of a computer device implementing a smart agriculture yield decision tree prediction method according to some embodiments of this application. The smart agriculture yield decision tree prediction method in the above embodiments can be implemented through... Figure 5 The computer device shown is used to implement this, and the computer device 300 includes at least one processor 301, a communication bus 302, a memory 303, and at least one communication interface 304.

[0070] The processor 301 may be a general-purpose central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more devices used to control the execution of the smart agriculture yield prediction method in this application.

[0071] The communication bus 302 is used to transmit information between the aforementioned components.

[0072] Memory 303 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile optical discs, Blu-ray discs, etc.), magnetic disks or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. Memory 303 may exist independently and be connected to processor 301 via communication bus 302. Memory 303 may also be integrated with processor 301.

[0073] The memory 303 stores program code for executing the scheme of this application, and its execution is controlled by the processor 301. The processor 301 executes the program code stored in the memory 303. The program code may include one or more software modules. In the above embodiments, the smart agriculture yield decision tree prediction method can be implemented by the processor 301 and one or more software modules in the program code in the memory 303.

[0074] Communication interface 304 uses any transceiver-like device to communicate with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.

[0075] In a specific implementation, as one example, a computer device may include multiple processors, each of which may be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. Here, a processor may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0076] The aforementioned computer device can be a general-purpose computer device or a special-purpose computer device. In specific implementations, the computer device may be a desktop computer, a portable computer, a network server, a handheld digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, a communication device, or an embedded device. This application does not limit the type of computer device.

[0077] In addition, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described smart agriculture yield decision tree prediction method.

[0078] In summary, the smart agriculture yield decision tree prediction method and system disclosed in this application obtains the spectral and textural features of each field in the target area during key phenological periods to form an original feature set of the fields; based on the original feature set of the fields and the spatial location of the fields, a clustering analysis with spatial neighborhood constraints is performed on all fields to form multiple spatially continuous field clusters; according to the historical yield data of all fields corresponding to each field cluster and the corresponding original feature set of the fields, a local decision tree prediction model for each field cluster is independently trained; and the original feature set of any field to be predicted in the target area is processed in parallel. The data is input to all trained local decision tree prediction models, and then the predicted values ​​of each local decision tree prediction model are used to generate local prediction feature maps for the corresponding fields to be predicted. The original feature set of the fields to be predicted is concatenated with the local prediction feature maps to obtain a fused feature vector. Then, the fused feature vector and historical yield data are used to train a global decision tree ensemble model for yield prediction of the entire region, so as to complete the yield prediction of the fields to be predicted. Yield prediction can be performed under spatial heterogeneity interference by using the contextual association information of fields in complex spatial patterns.

[0079] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0080] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A smart agriculture yield decision tree prediction method, characterized in that, Includes the following steps: The spectral and textural features of each field within the target area during key phenological periods are obtained to form the original feature set of the fields; Based on the original feature set of the fields and the spatial location of the fields, a cluster analysis with spatial neighborhood constraints is performed on all fields to form multiple spatially continuous field clusters; Based on the historical yield data of all fields corresponding to each field cluster and the corresponding original feature set of the fields, the local decision tree prediction model of each field cluster is trained independently. The original feature set of any field to be predicted within the target area is input in parallel into all trained local decision tree prediction models, and then the local prediction feature map of the corresponding field to be predicted is generated by the prediction value of each local decision tree prediction model. The original feature set of the field to be predicted is concatenated with the local predicted feature map to obtain a fused feature vector. Then, a global decision tree ensemble model for yield prediction of the whole region is trained by the fused feature vector and historical yield data to complete the yield prediction of the field to be predicted.

2. The method as described in claim 1, characterized in that, Based on the original feature set of the fields and their spatial location, cluster analysis with spatial neighborhood constraints is performed on all fields to form multiple spatially continuous field clusters, specifically including: Based on the spatial location of the fields, adjacency relationships are defined, and a spatial weight matrix representing spatial neighborhood constraints is generated. The original feature set of the fields is fused with the spatial weight matrix to construct a joint similarity matrix for spatially constrained clustering; Clustering and quality assessment are performed on the joint similarity matrix to determine the optimal partitioning scheme, thereby obtaining multiple spatially continuous field clusters.

3. The method as described in claim 1, characterized in that, Based on the historical yield data of all fields corresponding to each field cluster and the corresponding original feature set of the fields, the local decision tree prediction model for each field cluster is trained independently, specifically including: Historical yield data corresponding to all fields within each field cluster and feature data corresponding to the original feature set of the fields are extracted to construct a local training sample set for each field cluster. Each local training sample set is preprocessed to obtain a preprocessed local training sample subset. A decision tree model is trained for each preprocessed subset of local training samples to obtain the local decision tree prediction model for each field cluster.

4. The method as described in claim 1, characterized in that, The process of inputting the original feature set of any field to be predicted within the target area into all trained local decision tree prediction models in parallel includes: Based on the identifier of the field to be predicted, the feature data corresponding to the field to be predicted is extracted from the original feature set of the field; The feature data is synchronously input into all trained local decision tree prediction models for forward inference, thereby obtaining the predicted yield of the field to be predicted under different local decision tree prediction models.

5. The method as described in claim 1, characterized in that, The generation of local prediction feature maps for the corresponding fields to be predicted based on the prediction values ​​of each local decision tree prediction model specifically includes: All predicted values ​​are sorted according to the fixed order of the field clusters corresponding to each local decision tree prediction model; The sorting results are associated with the spatial coordinates of the fields to be predicted; Based on the mapping results, a local prediction feature map is constructed for the corresponding field to be predicted.

6. The method as described in claim 1, characterized in that, The original feature set of the field to be predicted is concatenated with the local predicted feature map to obtain the fused feature vector, specifically including: The original feature vector is generated by using the original feature set of the field to be predicted; The prediction feature vector of the field to be predicted is determined based on the local prediction feature map; The original feature vector and the predicted feature vector are concatenated to obtain a fused feature vector.

7. The method as described in claim 1, characterized in that, A global decision tree ensemble model for yield prediction across the entire region is obtained by training the fused feature vectors and historical yield data. This model is used to predict the yield of the fields to be predicted. Specifically, the yield prediction includes: The fused feature vectors corresponding to all fields with historical yield data are collected, and a global training sample set is constructed using the historical yield data as labels. A decision tree ensemble model is trained based on the global training sample set as a global decision tree ensemble model for full-region output prediction. The fused feature vector of the field to be predicted is input into the trained global decision tree ensemble model to complete the yield prediction of the field to be predicted.

8. The method as described in claim 1, characterized in that, The critical phenological period refers to a specific growth stage in the crop's life cycle in which its morphology, physiology, and yield formation have a decisive influence on the final yield, and whose canopy spectral reflectance characteristics have significant distinguishability.

9. The method as described in claim 1, characterized in that, The spectral characteristics refer to the energy reflected by ground objects to different electromagnetic bands, as detected by remote sensing sensors.

10. A smart agricultural yield decision tree prediction system, characterized in that, include: The acquisition module is used to acquire the spectral and textural features of each field within the target area during key phenological periods, forming the original feature set of the field. The processing module is used to perform clustering analysis with spatial neighborhood constraints on all fields based on the original feature set of the fields and the spatial location of the fields, forming multiple spatially continuous field clusters; The processing module is also used to independently train the local decision tree prediction model for each field cluster based on the historical yield data of all fields corresponding to each field cluster and the corresponding original feature set of the fields. The processing module is also used to input the original feature set of any field to be predicted in the target area into all trained local decision tree prediction models in parallel, and then generate the local prediction feature map of the corresponding field to be predicted through the prediction values ​​of each local decision tree prediction model. The execution module is used to concatenate the original feature set of the field to be predicted with the local prediction feature map to obtain a fused feature vector. Then, the fused feature vector and historical yield data are used to train a global decision tree ensemble model for yield prediction of the whole region, so as to complete the yield prediction of the field to be predicted.

Citation Information

Patent Citations

  • Crop yield prediction method and system

    CN110414738A

  • Blueberry yield prediction method based on machine learning

    CN112906298A

  • Cotton yield prediction method, device, equipment and medium

    CN120387542A

  • Crop rotation yield prediction method and device based on neural network

    CN121072907A

  • Farmland yield prediction method and system based on heterogeneous graph neural network

    CN121146219A