A yield prediction method, device, equipment, medium and program product
By constructing a time series knowledge graph and using interpretable models for analysis, the problems of inaccurate and poor interpretability of crop yield prediction in the prior art are solved, and higher prediction accuracy and scientific guidance are achieved.
Patent Information
- Application Number
- CN202411918500.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-12-24
AI Technical Summary
The prior art has problems of insufficient accuracy and poor interpretability in crop yield prediction, and it is impossible to effectively analyze the dynamic relationship between multiple influencing factors and yield.
By constructing a time series knowledge graph between output, influencing factors and time, and using it as the input of the output time series prediction model for prediction, combined with an interpretable model for analysis, we can determine the degree of impact of output influencing factors on estimated output.
It improves the accuracy of crop yield prediction, reveals the dynamic coupling relationship between multiple influencing factors, provides technicians with reliable data reference and scientific guidance, and helps achieve steady wheat yield increase.
Smart Images

Figure CN119358775B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of crop management, and in particular to a yield prediction method, device, equipment, medium and program product. Background Art
[0002] The yield of crops is affected by the dynamic coupling of multiple factors such as variety, environmental conditions and cultivation measures. Taking wheat, one of the staple crops, as an example, the interactions between multiple factors affecting wheat yield are complex, and there is a trade-off effect. Therefore, a deep understanding of the dynamic relationship of multiple factors such as wheat genetic potential, growth mechanism, and environmental response mechanism, and revealing the dynamic relationship of multiple factors affecting wheat yield, can provide scientific guidance for wheat planting from multiple dimensions such as scientific variety selection and precise application of water and fertilizer to help steadily increase wheat production.
[0003] Existing yield prediction methods can only explore the correlation between the above factors. For example, some statistical analysis methods can only handle linear relationships, and have limited ability to handle nonlinear relationships between multiple influencing factors. Therefore, it is impossible to simultaneously correlate and analyze the dynamic correlation between multiple influencing factors and yield, resulting in inaccurate prediction of its yield. In addition, although the machine learning model used in related technologies can handle the complex nonlinear relationship between multiple influencing factors, the model has poor interpretability and cannot clearly explain the mechanism of action of each factor in the coupling relationship. Summary of the invention
[0004] In view of this, the present invention provides a yield prediction method, device, equipment, medium and program product to solve the problems of inaccurate crop yield prediction and poor interpretability.
[0005] In a first aspect, the present invention provides a method for predicting yield, the method comprising:
[0006] Obtain the historical crop yields and historical parameter values of yield-influencing factors within the target area within the historical period;
[0007] Construct a time series knowledge graph between output, influencing factors and time based on historical time, historical output and historical parameter values;
[0008] The time series knowledge graph is used as the input of the pre-built yield time series prediction model to predict the yield and obtain the estimated yield of crops in the target area.
[0009] Based on the interpretable model, an interpretable analysis is performed on the prediction results of the yield time series forecasting model to determine the influence of yield influencing factors on the estimated yield.
[0010] The yield prediction method provided by the present invention constructs a time series knowledge graph between yield, influencing factors, and time according to the historical time of the target area, the historical yield of crops within the historical time, and the historical parameters of yield influencing factors, and uses the time series knowledge graph as the input of the yield time series prediction model to perform yield prediction, obtain the estimated yield of crops within the target area, and determine the degree of influence of yield influencing factors on the estimated yield based on an interpretable model. By constructing a time series knowledge graph, and performing model yield prediction and interpretable analysis, the present invention can effectively associate and fuse the influencing factor data within a continuous time, reveal the dynamic coupling relationship between multiple influencing factors, improve the accuracy of yield prediction, provide reliable data reference and scientific guidance for technical personnel, and thus increase wheat yield.
[0011] In an optional embodiment, historical crop yields and historical parameter values of yield influencing factors within a target area within a historical period are obtained, including: determining yield influencing factors that affect crop yields within the target area based on factor analysis, the yield influencing factors including: climate factors, soil factors, cultivation management factors and variety characteristics; taking years as the unit of measurement, obtaining historical parameter values of climate factors, soil factors, cultivation management factors and variety characteristics within multiple consecutive historical years, the historical parameter values being at least one of historical parameter values or historical parameter intensities.
[0012] The present invention can determine the factors affecting crop yield through factor analysis, thereby predicting crop yield based on various influencing factors, ensuring the scientificity and reliability of the prediction results. In addition, obtaining the parameter values of the influencing factors in consecutive years can reveal the dynamic coupling relationship between the influencing factors in the subsequent analysis process, thereby improving the accuracy of the prediction results.
[0013] In an optional implementation, a temporal knowledge graph of output, influencing factors and time is constructed based on historical time, historical output and historical parameter values, including: taking historical output as the first entity, output influencing factors as the second entity, and historical years as the third entity; and constructing a temporal knowledge graph with the first entity as the central entity based on the first entity, the second entity and the third entity.
[0014] The present invention can reveal the dynamic coupling relationship between various influencing factors by constructing a time series knowledge graph between output-influence relationship-influencing factor-time, and provide scientific data support for the time series output prediction model.
[0015] In an optional implementation, a method for constructing a production time series prediction model includes:
[0016] Based on the pre-constructed time series knowledge graphs in different regional scopes, each historical year is taken as the starting year, and the ending year is determined according to the preset year interval. Data is screened from the time series knowledge graph according to the time range corresponding to different starting years to the ending years to obtain multiple training data sets; a production time series prediction model is constructed, which includes a graph neural network model, a sequence model and a multi-layer neural network model connected in sequence, and the historical production corresponding to the year after the ending year in the training data set is taken as the actual production, and the historical parameter values and historical production corresponding to each year in the training data set are taken as input parameters to train the production time series prediction model until the error between the output value of the production time series prediction model and the actual production is less than a preset threshold.
[0017] The present invention can obtain rich sample data by constructing a training data set containing relevant data of multiple years in different regions, and performs model training based on the dynamic coupling relationship between various influencing factors in the sample data, thereby obtaining a time series yield prediction model that sequentially performs spatial semantic information feature extraction, time series data fusion and yield regression prediction, thereby ensuring the accuracy of the prediction results.
[0018] In an optional embodiment, the time series knowledge graph is used as the input of a pre-built yield time series prediction model to perform yield prediction to obtain an estimated yield of crops within the target area, including: taking the year before the predicted year as the end year, and determining the start year according to a preset year interval; filtering data from the time series knowledge graph within the target area according to the time range corresponding to the start year to the end year to obtain an input data set; and inputting the input data set into the yield time series prediction model to obtain the estimated yield of the predicted year.
[0019] The present invention predicts crop yields through a time series yield prediction model, and can quickly, efficiently and accurately obtain the estimated crop yields, providing technical personnel with reliable data references.
[0020] In an optional embodiment, an interpretable analysis is performed on the prediction results of the yield time series prediction model based on an interpretable model to determine the degree of influence of the yield influencing factors on the estimated yield, including: selecting historical parameter values corresponding to a preset year from an input data set as a set of sample data, and generating multiple sets of disturbance data based on the sample data; inputting the multiple sets of disturbance data into the yield time series prediction model respectively to obtain multiple prediction values, and the number of multiple groups and multiple corresponding values is equal; fitting the multiple sets of disturbance data and the multiple prediction values based on the interpretable model to obtain the impact effect and influence weight of the yield influencing factors on the yield.
[0021] The present invention obtains the impact effect and impact weight of each influencing factor on the yield through an interpretable model, and can provide a detailed explanation and scientific basis for the dynamic coupling relationship of multiple influencing factors, thereby providing comprehensive scientific guidance for high crop yields based on the impact effects and impact weights, and ensuring that crop yields can exceed estimated yields to the greatest extent.
[0022] In a second aspect, the present invention provides a yield prediction device, the device comprising:
[0023] A data acquisition module is used to obtain the historical yield of crops and historical parameter values of yield influencing factors within the historical period of the target area;
[0024] A graph construction module is used to construct a time series knowledge graph between output, influencing factors and time based on historical time, historical output and historical parameter values;
[0025] The yield prediction module is used to use the time series knowledge graph as the input of the pre-built yield time series prediction model to perform yield prediction and obtain the estimated yield of crops within the target area;
[0026] The model interpretation module is used to perform interpretable analysis on the prediction results of the yield time series prediction model based on the interpretable model to determine the degree of influence of yield influencing factors on the estimated yield.
[0027] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor are communicatively connected to each other, computer instructions are stored in the memory, and the processor executes the output prediction method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.
[0028] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the yield prediction method of the first aspect or any corresponding embodiment thereof.
[0029] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions for causing a computer to execute the yield prediction method of the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0031] Figure 1 is a flow chart of a method for predicting output according to an embodiment of the present invention;
[0032] Figure 2 is a schematic diagram of factors affecting yield according to a method for predicting yield according to an embodiment of the present invention;
[0033] Figure 3 is a flow chart of another method for predicting yield according to an embodiment of the present invention;
[0034] Figure 4 is a schematic diagram of a time series knowledge graph of another method for predicting production according to an embodiment of the present invention;
[0035] Figure 5 is a flow chart of another method for predicting production according to an embodiment of the present invention;
[0036] Figure 6 is a schematic diagram of a forecasting model structure flow chart of another method for forecasting output according to an embodiment of the present invention;
[0037] Figure 7 is a schematic diagram for explaining prediction results of another method for predicting yield according to an embodiment of the present invention;
[0038] Figure 8 is a structural block diagram of a device for predicting output according to an embodiment of the present invention;
[0039] Fig. 9 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0040] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0041] The embodiment of the present invention is applicable to the scenario of predicting the crop yield of this year, taking wheat as an example. Although the machine learning model used in the related art can handle the complex nonlinear relationship between multiple influencing factors, the model has poor interpretability and cannot clearly explain the mechanism of action of each factor in the coupling relationship. Therefore, it is impossible to simultaneously correlate and analyze the dynamic correlation between multiple influencing factors and yield, and it is also impossible to provide scientific and clear guidance for wheat planting from multi-dimensional factors such as scientific variety selection and precise application of water and fertilizer. The embodiment of the present invention provides a yield prediction method, which achieves the effect of improving the prediction accuracy by constructing a time series knowledge graph and predicting the yield based on a time series prediction model.
[0042] According to an embodiment of the present invention, an embodiment of a method for predicting production output is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0043] In this embodiment, a method for predicting production is provided, which can be used in mobile terminals, such as mobile computers, etc. Figure 1 is a flow chart of a method for predicting output according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:
[0044] Step S101, obtaining the historical yield of crops and historical parameter values of yield influencing factors within the target area within a historical period.
[0045] Specifically, in the embodiment of the present invention, taking wheat as one of the main crops as an example, the yield influencing factors affecting wheat yield are determined in advance through literature research and actual investigation, including: climate factors, soil factors, cultivation management factors and variety characteristics, among which, Figure 2 As shown, light, temperature, and precipitation are climate factors, soil fertility and soil texture are soil factors, and irrigation and fertilization are cultivation management factors, which are only examples and are not limited to this. According to the pre-divided regional scope, for example, the regional scope corresponding to the county in a province is divided. If you want to predict the wheat yield of a certain target regional scope in the current year, you need to obtain the historical yield of wheat in the wheat planting area in the target regional scope within the historical time and the historical parameter values of various yield influencing factors. Taking the prediction of wheat yield in 2025 as an example, obtain the long sequence data corresponding to various yield influencing factors from 2004 to 2024 and use them as historical parameter values, but it is not limited to this.
[0046] Step S102, construct a time series knowledge graph between output, influencing factors and time based on historical time, historical output and historical parameter values.
[0047] Specifically, in an embodiment of the present invention, crop yield is affected by the coupling of multiple factors such as varieties, climatic conditions and cultivation measures. The influence relationship includes the influence effect (positive influence or negative influence) and the influence weight, and the interaction relationship between the influencing factors is complex, for example: the coupling between climatic factors (light and temperature, precipitation and temperature), the coupling between soil factors and climatic factors (soil fertility and precipitation, soil texture and temperature), the coupling between variety characteristics and climatic factors (drought resistance and precipitation), the coupling between cultivation management factors and variety characteristics (precipitation and drought resistance), the coupling between cultivation management factors and soil factors (fertilization amount and soil texture), etc., but not limited to this. Based on the above-mentioned various coupling relationships, an embodiment of the present invention constructs a time series knowledge graph between yield, influencing factors and time. Among them, the time series knowledge graph is a graph structure with time as the axis, which is used to record the relationship and evolution process of things in the time dimension. On the basis of the traditional knowledge graph, by adding time variables, the knowledge graph can better meet the actual application scenarios and mine more useful knowledge. The essence of the time series knowledge graph is a multi-version, multi-time dynamic graph that can display knowledge on the time axis. By adding the time variable to the triple, the relational model is expanded to become a four-tuple representation (head entity, relationship, tail entity, timestamp), thereby better capturing the dynamic changes of entities and relationships. In the embodiment of the present invention, the head entity is the output, the relationship is the influencing relationship, the tail entity is the influencing factor, and the timestamp is the time.
[0048] Step S103, using the time series knowledge graph as the input of the pre-built yield time series prediction model to perform yield prediction and obtain the estimated yield of crops within the target area.
[0049] Specifically, in an embodiment of the present invention, a production time series prediction model is pre-constructed, and the model includes a graph neural network (GNN) model, a sequence model (Transformer) and a multi-layer neural network model (Multi-Layer Perceptron, MLP) connected in sequence. Among them, GNN refers to the general term for algorithms that use neural networks to learn graph structure data, extract and explore features and patterns in graph structure data, and meet the needs of graph learning tasks such as clustering, classification, prediction, segmentation, and generation. Transformer is a sequence model based on the attention mechanism. Due to its powerful feature extraction ability and modeling ability for long sequences, it can better integrate time series information. MLP is a common artificial neural network model, which is widely used in various tasks, including classification, regression and feature learning. The structure of MLP includes an input layer, one or more hidden layers and an output layer. Each layer is composed of several neurons. The neurons in each layer are connected to all neurons in the previous layer in a fully connected manner, which enables the network to learn complex nonlinear relationships. In the embodiment of the present invention, spatial semantic information features can be extracted for each yield influencing factor in the time series knowledge graph. On this basis, Transformer is used to fuse the time series data (2004-2024) respectively, and finally the MLP multi-layer neural network is used to regress and predict the wheat yield to obtain the estimated wheat yield in 2025.
[0050] Step S104, based on the interpretable model, an interpretable analysis is performed on the prediction results of the yield time series prediction model to determine the degree of influence of the yield influencing factors on the estimated yield.
[0051] Specifically, in an embodiment of the present invention, considering the poor interpretability of the machine learning model, the prediction results are further explained and analyzed. The method used, LIME (Local Interpretable Model-agnostic Explanations), is a method for explaining machine learning models. Its basic principle is to generate disturbance samples similar to the original data and use these samples to train a simple explanation model to explain the prediction results of complex machine learning models. In an embodiment of the present invention, on the basis of the completion of the construction of the wheat yield time series prediction model, the explainable LIME model is integrated to realize the mining of the dynamic coupling relationship of multiple influencing factors of wheat yield and the construction of an explainable model, so that the influence effect and influence weight of different yield influencing factors on the current prediction of the estimated yield can be determined through interpretable analysis, so as to guide technicians to adjust the crop cultivation process according to the influence effect and influence weight, so that the crop yield exceeds the estimated yield as much as possible.
[0052] The yield prediction method provided by the embodiment of the present invention constructs a time series knowledge graph between yield, influencing factors, and time according to the historical time of the target area, the historical yield of crops within the historical time, and the historical parameters of yield influencing factors, and uses the time series knowledge graph as the input of the yield time series prediction model to perform yield prediction, thereby obtaining the estimated yield of crops within the target area, and determining the degree of influence of yield influencing factors on the estimated yield based on an interpretable model. By constructing a time series knowledge graph, and performing model yield prediction and interpretable analysis, the present invention can effectively associate and fuse the influencing factor data within a continuous time, reveal the dynamic coupling relationship between multiple influencing factors, improve the accuracy of yield prediction, and provide reliable data reference and scientific guidance for technical personnel.
[0053] In this embodiment, a method for predicting production is provided, which can be used in the above-mentioned mobile terminal, such as a computer, etc. Figure 3 is a flow chart of a method for predicting output according to an embodiment of the present invention. Figure 3 As shown, the process includes the following steps:
[0054] Step S301, obtaining the historical yield of crops and historical parameter values of yield influencing factors within the target area within a historical period.
[0055] Specifically, the above step S301 includes:
[0056] Step S3011, determining the yield influencing factors affecting the crop yield within the target area based on factor analysis, the yield influencing factors include: climate factors, soil factors, cultivation management factors and variety characteristics.
[0057] Step S3012, taking years as the unit of measurement, obtain historical parameter values of climate factors, soil factors, cultivation management factors and variety characteristics in multiple consecutive historical years, where the historical parameter value is at least one of a historical parameter value or a historical parameter intensity.
[0058] Specifically, in an embodiment of the present invention, taking wheat as an example, factor analysis can be performed based on big data to determine the yield influencing factors as follows: climate factors corresponding to light, temperature, and precipitation, soil factors corresponding to soil fertility and soil texture, cultivation management factors corresponding to irrigation and fertilization, and variety characteristics. In order to predict the wheat yield in the target area in 2025, the historical parameter values corresponding to each yield influencing factor in each year from 2004 to 2024 are obtained in units of years, and one year corresponds to a set of historical parameter values. Due to the essential characteristics of yield influencing factors, some yield influencing factors can be described by numerical values, such as precipitation, irrigation, temperature, soil fertility, and fertilization, while some yield influencing factors cannot be described by numerical values, but can only be described by intensity (strong or weak), such as light, soil texture, and variety. Therefore, different yield influencing factors are described according to historical parameter values or historical parameter intensity.
[0059] Step S302, construct a time series knowledge graph between output, influencing factors and time based on historical time, historical output and historical parameter values.
[0060] Specifically, the above step S302 includes:
[0061] Step S3021, taking the historical output as the first entity, the output influencing factors as the second entity, and the historical year as the third entity.
[0062] Step S3022, constructing a temporal knowledge graph with the first entity as the central entity based on the first entity, the second entity and the third entity.
[0063] Specifically, in an embodiment of the present invention, based on the collected historical output, historical output influencing factors, and historical time, the output, influencing factors, and time are respectively regarded as an entity, and the entity is the node in the time series knowledge graph. After the entity is determined, the relationship between the entities is extracted, and the relationship between the entities of the present invention is the influence relationship between output and output influencing factors. On this basis, the time dimension is added, so the triple of the ordinary knowledge graph (entity-relationship-entity, such as temperature-influence-output) is converted into a quadruple of the time series knowledge graph, and then through the processes of knowledge fusion, knowledge processing, model layer construction, data layer construction, etc., a time series knowledge graph with output as the central entity is obtained. The specific construction process of the time series knowledge graph is a conventional technical means in this field, which will not be repeated here, and the obtained time series knowledge graph is as follows Figure 4 As shown, but not limited to.
[0064] Step S303: Use the time series knowledge graph as the input of the pre-built yield time series prediction model to perform yield prediction and obtain the estimated yield of crops within the target area. Figure 1 Step S103 of the illustrated embodiment will not be described in detail here.
[0065] Step S304: Based on the interpretable model, an interpretable analysis is performed on the prediction results of the production time series prediction model to determine the degree of influence of the production influencing factors on the estimated production. Figure 1 Step S104 of the illustrated embodiment will not be described in detail here.
[0066] The yield prediction method provided by the embodiment of the present invention constructs a time series knowledge graph between yield, influencing factors, and time according to the historical time of the target area, the historical yield of crops within the historical time, and the historical parameters of yield influencing factors, and uses the time series knowledge graph as the input of the yield time series prediction model to perform yield prediction, thereby obtaining the estimated yield of crops within the target area, and determining the degree of influence of yield influencing factors on the estimated yield based on an interpretable model. By constructing a time series knowledge graph, and performing model yield prediction and interpretable analysis, the present invention can effectively associate and fuse the influencing factor data within a continuous time, reveal the dynamic coupling relationship between multiple influencing factors, improve the accuracy of yield prediction, and provide reliable data reference and scientific guidance for technical personnel.
[0067] In this embodiment, a method for predicting production is provided, which can be used in the above-mentioned mobile terminal, such as a computer, etc. Figure 5 is a flow chart of a method for predicting output according to an embodiment of the present invention. Figure 5 As shown, the process includes the following steps:
[0068] Step S501, obtaining the historical crop yields and historical parameter values of yield influencing factors within the target area within the historical period. Figure 3 Step S301 of the illustrated embodiment will not be described in detail here.
[0069] Step S502: construct a time series knowledge graph between output, influencing factors and time based on historical time, historical output and historical parameter values. Figure 3 Step S302 of the illustrated embodiment will not be described in detail here.
[0070] Step S503, using the time series knowledge graph as the input of the pre-built yield time series prediction model to perform yield prediction and obtain the estimated yield of crops within the target area.
[0071] Specifically, the above step S503 includes:
[0072] Step S5031, taking the year before the predicted year as the end year, and determining the start year according to the preset year interval.
[0073] Step S5032, filter data from the time series knowledge graph of the target area according to the time range corresponding to the start year to the end year to obtain an input data set.
[0074] Step S5033, input the input data set into the production time series forecasting model to obtain the estimated production in the forecast year.
[0075] Specifically, in an embodiment of the present invention, a yield time series prediction model is pre-constructed, and a sufficient number of training data sets need to be constructed at this time. The embodiment of the present invention pre-acquires the historical yields and historical parameter values of yield influencing factors between 2004 and 2024 in different regions, and constructs time series knowledge graphs corresponding to different regions. Because there are differences in climate factors, soil factors, cultivation management factors and crop variety characteristics in different regions, there are also differences in the constructed time series knowledge graphs, that is, there are differences in the dynamic coupling relationship between the various influencing factors. In order to further enrich the training data set, the embodiment of the present invention divides the data of each regional range, takes each historical year as the starting year, determines the ending year according to the preset year interval, and then screens the data from the time series knowledge graph according to the time range corresponding to the different starting years to the ending years, and obtains multiple training data sets. For example, a certain region contains relevant data between 2004 and 2024. First, 2004 is taken as the starting year, and the end year is determined to be 2008 according to the 5-year interval. Therefore, data is screened in the time series knowledge graph according to 2004, 2005, 2006, 2007 and 2008 to obtain a set of training data sets; then 2005 is taken as the starting year, and the end year is determined to be 2009 according to the 5-year interval. The relevant data from 2005 to 2009 is screened in the time series knowledge graph to obtain a set of training data sets. By analogy, 17 sets of training data sets can be obtained within this region, and 17 sets of training data sets can also be obtained in the corresponding other regions. This is only used as a distance and is not limited to this.
[0076] In some optional embodiments, after a sufficient amount of training data sets are obtained, a yield time series prediction model including a graph neural network model GNN, a sequence model Transforme and a multi-layer neural network model MLP connected in sequence is constructed, wherein the yield time series prediction model includes 3 Transforme, and different training data sets are input into the yield time series prediction model for model training. During the training process, the historical yield of the year following the termination year in the training data set is used as the actual yield, and the historical parameter values and historical yields corresponding to each year in the data set are used as input parameters of the yield time series prediction model. The output value of the yield time series prediction model is the predicted yield, and a suitable loss function (such as a mean square error loss function) and an optimization algorithm (such as a stochastic gradient descent algorithm) are used to gradually adjust the parameters of the yield time series prediction model until the error between the output value of the yield time series prediction model and the actual yield is less than a preset threshold, so that the yield time series prediction model outputs a more accurate wheat yield prediction value.
[0077] In some optional implementations, after the model training is completed, when the crop yield in a certain area is actually predicted, the input data set can be constructed according to the division method of the training data set. Taking the prediction of wheat yield in 2025 as an example, 2024 is used as the end year, and the starting year is determined as 2020 according to the 5-year year interval. The historical yield and historical parameter values corresponding to 2020-2024 are used as input data sets and input into the yield time series prediction model. Alternatively, all historical yields and historical parameter values corresponding to 2004-2024 can be directly input into the yield time series prediction model as input data sets, which is not limited here. The prediction process is as follows Figure 6 As shown in the figure, the graph neural network model GNN extracts spatial semantic information features of the interaction relationship between various influencing factors. On this basis, Transformer is used to fuse the time series data (2004-2024). The three Transformers respectively perform time series feature fusion extraction on the historical parameter intensity of categorical influencing factors (variety, soil texture, light), the historical parameter values of numerical influencing factors (soil fertility, temperature, precipitation, fertilizer application, irrigation amount) and historical yield. Finally, the MLP multi-layer neural network is used to regress and predict wheat yield to obtain the estimated yield in 2025.
[0078] In some optional embodiments, an embodiment of the present invention proposes a crop yield prediction method based on a time series graph neural network, which extracts spatial semantic features of multiple influencing factors in the crop yield time series graph through the graph neural network, and then fuses the time series features through Transformer, thereby improving the accuracy of wheat yield prediction results by fusing spatial semantic features with time series features.
[0079] Step S504, based on the interpretable model, an interpretable analysis is performed on the prediction results of the yield time series prediction model to determine the degree of influence of the yield influencing factors on the estimated yield.
[0080] Specifically, the above step S504 includes:
[0081] Step S5041, selecting historical parameter values corresponding to a preset year from the input data set as a set of sample data, and generating multiple sets of disturbance data based on the sample data.
[0082] Step S5042, inputting multiple groups of disturbance data into the production time series prediction model respectively to obtain multiple prediction values, and the number of multiple groups and multiple corresponding values is equal.
[0083] Step S5043, fitting multiple groups of disturbance data and multiple predicted values based on the interpretable model to obtain the impact effect and impact weight of the yield influencing factors on the yield.
[0084] Specifically, in an embodiment of the present invention, the construction method of the dynamic coupling relationship mining and interpretable model of multiple influencing factors of wheat yield is that on the time series graph of the coupling relationship of multiple influencing factors with yield as the central entity, the LIME algorithm fits a simple interpretable model of the wheat yield prediction result in a specific local area, and explains the prediction result of the complex model in an understandable way. In an embodiment of the present invention, the input is a set of instances (for example, the influencing factor data of a certain area), and a simple model (for example, a linear regression model) is fitted according to the LIME model principle and the yield time series prediction model. The weight of the linear regression model can show the influence of each influencing factor on the yield, and then the positive and negative influence of the dynamic coupling relationship of multiple factors of wheat yield is explained. A set of sample data of a certain year selected arbitrarily between 2020 and 2024 is used as a set of instances. The sample data includes 8 parameter values corresponding to 8 kinds of yield influencing factors. On this basis, neighboring sampling is performed to expand a set of sample data into multiple sets of disturbance data, for example, 1000 sets of disturbance data are generated. 1000 sets of disturbance data are respectively input into the yield time series prediction model, and 1000 prediction values can be obtained. 1000 sets of disturbance data and the corresponding 1000 output values can constitute 1000 sets of annotated data. By fitting the 1000 sets of annotated data with a linear regression model, we can obtain the effects and weights of various yield influencing factors on yield. The effects include positive and negative effects, and the weights are represented numerically, so that we can intuitively explain the specific examples in the wheat yield time series knowledge graph, such as Figure 7 As shown, 8 influencing factors have specific weight values on the edge, positive numbers represent positive influences, and negative numbers represent negative influences. It can be seen that in the embodiment of the present invention, temperature, irrigation amount, variety, etc. have positive influences, and temperature is the most critical; however, precipitation and organic matter have negative influences.
[0085] In some optional implementations, the embodiments of the present invention propose a method for mining and interpreting the dynamic relationship of multiple influencing factors of crop yield that integrates a time series graph with a LIME model, thereby revealing the dynamic coupling relationship of multiple influencing factors of wheat yield, and being able to intuitively display the degree of positive and negative effects of each influencing factor in the coupling dynamic relationship. The visualized root cause analysis of the wheat yield prediction results obtained according to the embodiments of the present invention can provide researchers in the agricultural field with a detailed explanation and scientific basis for the dynamic coupling relationship of multiple influencing factors, and automatically generate scientific guidance and suggestions for improving wheat yields, including recommending suitable wheat varieties based on soil and meteorological conditions, and determining the optimal sowing density, fertilizer type and amount, and irrigation frequency based on variety characteristics and soil fertility. Figure 7 The analysis results shown can suggest appropriate increases in irrigation and fertilization, thus providing well-reasoned and automated scientific guidance.
[0086] In some optional implementations, during the actual wheat planting process, a test field and a control field may be set up. Planting is carried out in the test field according to the guidance of the optimal coupling relationship, and traditional planting methods are used in the control field. The effectiveness of the optimal coupling relationship is verified by comparing the wheat yield, quality and other indicators of the test field and the control field. At the same time, new data is continuously collected, and the knowledge graph and model are updated and optimized based on the new data to adapt to the ever-changing wheat production environment.
[0087] The yield prediction method provided by the present invention constructs a time series knowledge graph between yield, influencing factors, and time according to the historical time of the target area, the historical yield of crops within the historical time, and the historical parameters of yield influencing factors, and uses the time series knowledge graph as the input of the yield time series prediction model to perform yield prediction, obtain the estimated yield of crops within the target area, and determine the degree of influence of yield influencing factors on the estimated yield based on an interpretable model. By constructing a time series knowledge graph, and performing model yield prediction and interpretable analysis, the present invention can effectively associate and fuse the influencing factor data within a continuous time, reveal the dynamic coupling relationship between multiple influencing factors, improve the accuracy of yield prediction, and provide reliable data reference and scientific guidance for technical personnel.
[0088] In the present embodiment, a prediction device for output is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions thereof will not be repeated. As used below, the term "module" may implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and contemplated.
[0089] This embodiment provides a production prediction device, such as Figure 8As shown, including:
[0090] The data acquisition module 801 is used to obtain the historical yield of crops and historical parameter values of yield influencing factors within the historical period of the target area.
[0091] The graph construction module 802 is used to construct a time series knowledge graph between output, influencing factors and time based on historical time, historical output and historical parameter values.
[0092] The yield prediction module 803 is used to perform yield prediction using the time series knowledge graph as the input of a pre-built yield time series prediction model to obtain the estimated yield of crops within the target area.
[0093] The model interpretation module 804 is used to perform an interpretable analysis on the prediction results of the yield time series prediction model based on the interpretable model to determine the degree of influence of the yield influencing factors on the estimated yield.
[0094] In some optional implementations, the data acquisition module 801 includes:
[0095] The factor analysis unit is used to determine the yield influencing factors affecting the crop yield within the target area based on factor analysis. The yield influencing factors include: climate factors, soil factors, cultivation management factors and variety characteristics.
[0096] The data collection unit is used to obtain historical parameter values of climate factors, soil factors, cultivation management factors and variety characteristics in multiple consecutive historical years with years as the unit of measurement. The historical parameter value is at least one of the historical parameter value or the historical parameter intensity.
[0097] In some optional implementations, the graph construction module 802 includes:
[0098] The entity determination unit is used to take the historical output as the first entity, the output influencing factors as the second entity, and the historical year as the third entity.
[0099] A graph determination unit is used to construct a temporal knowledge graph with the first entity as the central entity based on the first entity, the second entity and the third entity.
[0100] In some optional embodiments, the device further comprises: a model building module, comprising:
[0101] The training data set construction module is used to obtain multiple training data sets based on the pre-constructed time series knowledge graphs of different regional scopes, with each historical year as the starting year, and the ending year determined according to the preset year interval. The data is filtered from the time series knowledge graph according to the time range corresponding to different starting years to the ending years.
[0102] The model training module is used to construct a production time series prediction model that includes a graph neural network model, a sequence model, and a multi-layer neural network model connected in sequence, and uses the historical production corresponding to the year following the termination year in the training data set as the actual production, and the historical parameter values and historical production corresponding to each year in the training data set as input parameters to train the production time series prediction model until the error between the output value of the production time series prediction model and the actual production is less than a preset threshold.
[0103] In some optional implementations, the yield prediction module 803 includes:
[0104] The data partitioning unit is used to take the year before the predicted year as the end year and determine the start year according to the preset year interval.
[0105] The input data set construction unit is used to filter data from the time series knowledge graph of the target area according to the time range corresponding to the start year to the end year to obtain the input data set.
[0106] The first model prediction unit is used to input the input data set into the yield time series prediction model to obtain the estimated yield in the prediction year.
[0107] In some optional implementations, the model interpretation module 804 includes:
[0108] The data expansion unit is used to select historical parameter values corresponding to a preset year from an input data set as a set of sample data, and generate multiple sets of disturbance data based on the sample data.
[0109] The second model prediction unit is used to input multiple groups of disturbance data into the production time series prediction model respectively to obtain multiple prediction values, and the number of multiple groups and multiple corresponding values is equal.
[0110] The result interpretation unit is used to fit multiple groups of disturbance data and multiple predicted values based on an interpretable model to obtain the impact effect and impact weight of yield influencing factors on yield.
[0111] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0112] The output prediction device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0113] The embodiment of the present invention also provides a computer device having the above Figure 8The predicted means of production are shown.
[0114] See also Fig. 9 , Fig. 9 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Fig. 9 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Fig. 9 A processor 10 is taken as an example.
[0115] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.
[0116] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.
[0117] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0118] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.
[0119] The computer device further comprises a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0120] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.
[0121] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.
[0122] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A method for predicting yield, characterized in that: The method comprises: Determine the yield influencing factors affecting the crop yield within the target area according to factor analysis, wherein the yield influencing factors include: climate factors, soil factors, cultivation management factors and variety characteristics; obtain historical parameter values and historical yields of the climate factors, soil factors, cultivation management factors and variety characteristics in multiple consecutive historical years with years as the measurement unit, wherein the historical parameter value is at least one of the historical parameter value and the historical parameter intensity; Constructing a time series knowledge graph between output, influencing factors and time according to the historical years, the historical output and the historical parameter values; The time series knowledge graph is used as the input of a pre-constructed yield time series prediction model to perform yield prediction, and the estimated yield of the crop within the target area is obtained. The method for constructing the yield time series prediction model includes: based on the pre-constructed time series knowledge graphs of different regional ranges, each of the historical years is taken as the starting year, and the ending year is determined according to a preset year interval, and data is screened from the time series knowledge graph according to the time range corresponding to the different starting years to the ending years to obtain multiple training data sets; a yield time series prediction model including a graph neural network model, a sequence model and a multi-layer neural network model connected in sequence is constructed, and the historical yield corresponding to the year after the ending year in the training data set is taken as the actual yield, and the historical parameter values and historical yields corresponding to each year in the training data set are taken as input parameters, and the yield time series prediction model is trained until the error between the output value of the yield time series prediction model and the actual yield is less than a preset threshold; Based on the interpretable model, an interpretable analysis is performed on the prediction results of the production time series prediction model to determine the degree of influence of the production influencing factors on the estimated production.
2. The method according to claim 1, characterized in that The constructing of a time series knowledge graph between output, influencing factors and time according to the historical years, the historical output and the historical parameter values includes: The historical output is taken as the first entity, the output influencing factors are taken as the second entity, and the historical year is taken as the third entity; A temporal knowledge graph with the first entity as the central entity is constructed based on the first entity, the second entity and the third entity.
3. The method according to claim 1, characterized in that The method of using the time series knowledge graph as the input of a pre-built yield time series prediction model to perform yield prediction to obtain the estimated yield of the crop within the target area includes: The year before the predicted year is used as the end year, and the start year is determined according to the preset year interval; According to the time range corresponding to the starting year to the ending year, data is screened from the time series knowledge graph within the target area to obtain an input data set; The input data set is input into the production time series forecasting model to obtain the estimated production in the forecast year.
4. The method according to claim 3, characterized in that The method of performing an interpretable analysis on the prediction results of the production time series prediction model based on the interpretable model to determine the influence degree of the production influencing factors on the estimated production includes: Selecting historical parameter values corresponding to a preset year from the input data set as a set of sample data, and generating multiple sets of disturbance data based on the sample data; Inputting the plurality of groups of disturbance data into the production time series prediction model respectively to obtain a plurality of prediction values, wherein the plurality of groups and the plurality of corresponding values are equal in number; The multiple groups of disturbance data and the multiple predicted values are fitted based on an interpretable model to obtain the influence effect and influence weight of the yield influencing factor on the yield.
5. A yield prediction device, characterized in that: The device comprises: A data acquisition module is used to determine the yield influencing factors affecting the crop yield within the target area according to factor analysis, wherein the yield influencing factors include: climate factors, soil factors, cultivation management factors and variety characteristics; taking years as the measurement unit, obtaining historical parameter values and historical yields of the climate factors, soil factors, cultivation management factors and variety characteristics in multiple consecutive historical years, wherein the historical parameter value is at least one of a historical parameter value or a historical parameter intensity; A graph construction module, for constructing a time series knowledge graph between output, influencing factors and time according to the historical years, the historical output and the historical parameter values; A yield prediction module is used to use the time series knowledge graph as an input of a pre-constructed yield time series prediction model to perform yield prediction, and obtain the estimated yield of the crop within the target area. The method for constructing the yield time series prediction model includes: based on the pre-constructed time series knowledge graphs of different regional scopes, taking each of the historical years as the starting year, determining the ending year according to a preset year interval, and filtering data from the time series knowledge graph according to the time range corresponding to the different starting years to the ending years to obtain multiple training data sets; constructing a yield time series prediction model including a graph neural network model, a sequence model, and a multi-layer neural network model connected in sequence, and taking the historical yield corresponding to the year following the ending year in the training data set as the actual yield, and taking the historical parameter values and historical yields corresponding to each year in the training data set as input parameters, the yield time series prediction model is trained until the error between the output value of the yield time series prediction model and the actual yield is less than a preset threshold; The model interpretation module is used to perform an interpretable analysis on the prediction results of the production time series prediction model based on the interpretable model to determine the degree of influence of the production influencing factors on the estimated production.
6. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the yield prediction method according to any one of claims 1 to 4 by executing the computer instructions.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the yield prediction method according to any one of claims 1 to 4.
8. A computer program product, characterized in that The method comprises computer instructions for causing a computer to execute the yield prediction method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Interpretable recommendation modeltraining method and device
CN113360772A
Crop disaster loss and yield prediction system and method based on knowledge graph
CN115115126A