Crop identification method by simultaneously optimizing prediction network and confidence network
Through the reinforcement learning model of the Actor-Critic architecture and the use of time series data to optimize the prediction network and confidence network, the problem of insufficient recognition accuracy of traditional crop recognition methods in complex environments is solved, and accurate recognition and prediction of crop types and yields are achieved.
Patent Information
- Application Number
- CN202510786538.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-26
AI Technical Summary
Traditional crop identification methods have difficulty effectively capturing the nonlinear relationships between multidimensional factors in complex and dynamic agricultural environments, resulting in decreased recognition accuracy, and existing reinforcement learning models lack long-term prediction capabilities for complex environments.
The reinforcement learning model adopts the Actor-Critic architecture, optimizes the prediction network and the confidence network simultaneously, uses time series data to generate states, and combines the TD-error optimization model to dynamically adjust the error function to improve recognition accuracy.
It improves the accuracy of crop identification and the ability to adapt to complex environments, and can accurately identify crop types and predict yields in different growth stages and regions.
Smart Images

Figure CN120708056A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to remote sensing technology, and in particular, to a crop identification method and information processing device thereof by simultaneously optimizing a prediction network and a confidence network. Background Art
[0002] Crop identification is a crucial issue in agricultural production, directly impacting market supply, agricultural decision-making, and food security. Remote sensing imagery data is inherently limited by insufficient temporal and spatial resolution. Furthermore, the spectral characteristics of crops of different species and at different growth stages vary significantly, and the growing environments of the same crop vary across different regions, complicating remote sensing identification. Traditional crop identification methods rely primarily on statistical and machine learning models and remote sensing data. Common models include linear regression, support vector machines (SVMs), and neural networks. These models often rely on a direct mapping between input features and outputs, lacking the ability to provide long-term predictions in complex environments.
[0003] In recent years, reinforcement learning (RL) methods have seen increasing application in various prediction tasks, particularly in decision-making and prediction in complex dynamic environments. The actor-critic architecture, a common approach in reinforcement learning, divides the model into two relatively independent sub-models: a policy network (actor) that outputs the optimal action strategy based on the current state, and a value network (critic) that evaluates various strategies and assigns values. Both networks are optimized independently. This separation of policy and value training strategy adapts to complex and diverse environments, enabling the model to perform better in these environments. The actor-critic (AC) architecture typically uses TD-error (temporal difference error) as a learning objective, allowing for rapid, small-step adjustments to the strategy during the learning process for more efficient model optimization. Typical applications include robotic control and recommender systems, but its application in agricultural species identification remains relatively limited.
[0004] Traditional crop species identification methods typically only optimize model parameters, while the error function, a fixed evaluation metric, fails to dynamically adjust. This results in insufficient predictive capabilities when dealing with complex and dynamic agricultural production environments. In complex agricultural environments, crop growth is influenced by multiple factors, including climate, soil, and moisture. Traditional prediction methods struggle to effectively capture the nonlinear relationships between these multidimensional factors, resulting in reduced identification accuracy. Summary of the Invention
[0005] The purpose of the present invention is to propose a crop identification method by simultaneously optimizing the prediction network and the confidence network, aiming to overcome the defects in the prior art.
[0006] According to a first aspect of the present application, a method for crop identification based on a reinforcement learning model is provided, wherein the reinforcement learning model includes a prediction network and a confidence network, and the method includes: preprocessing remote sensing data of a target area to generate time series data; generating a state STATE for the reinforcement learning model based on the time series data, and providing the state STATE to the prediction network; the prediction network generates a predicted type of the crop based on the state STATE; wherein the process of training the prediction network includes: at each time step t, generating the state STATE at time t based on the training sample t Provided to the prediction network; the prediction network outputs the prediction result; based on the prediction result output by the prediction network, a reward r is generated t ; Generate the state STATE at time t+1 t+1 ; The confidence network is based on the reward r t STATE t+1 and STATE t Calculate the TD-error, expressed as δ; where t represents the current time step;
[0007] In each training round, the following steps S1 to S4 are also performed: S1, using the δ of multiple time steps of the training round, calculate the objective function of the belief network JV(w) = ∑δ 2 , where w represents the parameters of the confidence network; S2, calculate the objective function J(θ) = ∑(log(prob)δ) of the prediction network, where prob represents the probability distribution corresponding to the prediction result output by the prediction network, and θ is the parameter of the prediction network; S3, update the feature extraction network, update the prediction network and update the confidence network.
[0008] According to the first aspect of the present application, the crop identification method based on reinforcement learning model, wherein:
[0009] The prediction results include the predicted crop type p; and the calculation of the reward r at time t t , where if the crop type is correctly judged, r t =Total number of identified crop types. If the crop type is wrong, r t =0.
[0010] According to the first aspect of the present application, a method for crop identification based on a reinforcement learning model is provided, wherein a state STATE for a reinforcement learning model is generated based on the time series data, including: generating time series data of a length N with a specified value as a placeholder after the time series data, and generating a state STATE for a reinforcement learning model using the time series data of a length N, wherein N represents the maximum number of remote sensing data that can be obtained during the crop growth cycle; so that even if the number of remote sensing data in the time series data changes, the generated state STATE has the same structure.
[0011] According to the first aspect of the present application, a method for crop identification based on a reinforcement learning model is provided, wherein a state STATE for a reinforcement learning model is generated based on the time series data, including: processing the time series data through a state embedding network to generate a state STATE for a reinforcement learning model; so that even if the amount of remote sensing data in the time series data changes, the generated state STATE has the same structure.
[0012] According to the first aspect of the present application, the method for crop identification based on a reinforcement learning model comprises the following steps: training the prediction network and the confidence network comprises multiple training rounds, each training round comprising multiple time steps; at each time step t, performing the following steps A1 to A5:
[0013] A1. Processing the time series data SS of training samples t Generate the state STATE at time t t ;
[0014] A2. Set the state to STATE t Provided to the prediction network, the prediction network outputs a prediction result;
[0015] A3. Generate time series data SS for the next time step t+1 t+1 , and generate a reward r based on the prediction result output by the prediction network t ; Among them, r t Represents the reward at time step t, which is obtained by the difference between the prediction result output by the prediction network and the actual historical data;
[0016] A4. Processing time series data SS t+1 Generate the state STATE at time t+1 t+1 ;
[0017] A5. According to the reward r t STATE t+1 and STATE t Calculate TD-error, expressed as δ, δ = r t +λ*Eval(STATEt+1 )-Eval(STATE t ), where Eval() represents the processing performed by the confidence network, and λ represents a hyperparameter; wherein the prediction result includes the predicted crop type p.
[0018] According to the first aspect of the present application, the crop identification method based on reinforcement learning model, wherein:
[0019] Time Series Data SS t =(S1, S2, ... S t ), represents the time series data at time t, which includes the remote sensing data from the earliest remote sensing data S1 to the remote sensing data S at time t t The time series data formed, in which the remote sensing data S t The subscript represents the time attribute of the remote sensing data, or the position or sequence number of the remote sensing data in the time series data.
[0020] According to the first aspect of the present application, the crop identification method based on the reinforcement learning model, wherein the remote sensing data of the target area is preprocessed to generate time series data, including: obtaining remote sensing data of multiple time points within the growth cycle of the crop, the remote sensing data of each time point including year-day attributes; normalizing the data of each band of the remote sensing data at each time point to between 0 and 1; using the remote sensing data at each time point to generate one or more remote sensing indices as features to expand the dimension of the remote sensing data; generating extended remote sensing data at each time point, the extended remote sensing data including the normalized band data of the remote sensing data, one or more remote sensing indices and year-day attributes; and combining the extended remote sensing data at each time point to obtain time series data.
[0021] According to the first aspect of the present application, the method for crop identification based on reinforcement learning model, wherein, in step S1, the objective function of the confidence network JV(w)=∑δ 2 The number of summation items is the number of time steps in the current training round; in step S2, the number of summation items of the objective function J(θ)=∑(log(prob)δ) of the prediction network is the number of time steps in the current training round.
[0022] According to the first aspect of the present application, the method for crop identification based on a reinforcement learning model, wherein the number of time steps in each training round is the number N of time points of remote sensing data within the growing season of the crop; in each training round, it also includes recording each state transition tuple (f t ,p,r t ,f t+1 );
[0023] The training samples include time series data generated based on historical real remote sensing data and corresponding crop type data;
[0024] The time series data used in each time step has a different end time point;
[0025] The time series data SS for the next time step is generated t+1 , including obtaining the time series data SS at time step t t The extended remote sensing data of the next time point of the end time point is appended to the time series data SS of time step t t The time series data SS obtained later t+1 .
[0026] According to a second aspect of the present application, a first method for crop identification and yield prediction based on a reinforcement learning model is provided, wherein the reinforcement learning model includes a prediction network and a confidence network, and the method includes: preprocessing remote sensing data of a target area to generate time series data; generating a state STATE for the reinforcement learning model based on the time series data, and providing the state STATE to the prediction network; the prediction network generates a predicted type and yield of the crop based on the state STATE;
[0027] The process of training the prediction network includes:
[0028] At each time step t, the state STATE at time t is generated based on the training sample t Provided to the prediction network; the prediction network outputs the prediction result; based on the prediction result output by the prediction network, a reward r is generated t ; Generate the state STATE at time t+1 t+1 ; The confidence network is based on the reward r t STATE t+1 and STATE t Calculate the TD-error, expressed as δ; where t represents the current time step;
[0029] In each training round, the following steps S1 to S4 are also performed: S1, using the δ of multiple time steps of the training round, calculate the objective function of the belief network JV(w) = ∑δ 2 , where w represents the parameters of the confidence network; S2, calculate the objective function J(θ) = ∑(log(prob)δ) of the prediction network, where prob represents the probability distribution corresponding to the prediction result output by the prediction network, and θ is the parameter of the prediction network; S3, update the feature extraction network, update the prediction network and update the confidence network.
[0030] According to the second aspect of the present application, the method for crop identification and yield prediction based on reinforcement learning model, wherein:
[0031] The prediction results include the predicted crop type p and yield a t ; and according to r t =r x +r y Calculate the reward r at time t t , where r x is the reward for crop type discrimination, r y is the output forecast reward, and the calculation formula is
[0032] If the crop type is correctly judged, r x =Total number of identified crop types. If the crop type is wrong, r x = 0, and
[0033] r y = Total yield class - ABS (forecast crop yield code - actual crop yield code),
[0034] Wherein, ABS() represents the absolute value; and wherein a relative vector is obtained by ratioing the crop yield to the average yield of multiple crop growing seasons of the crop yield, and the relative yield is mapped to the crop yield code according to the value size.
[0035] According to the second aspect of the present application, the method for crop identification and yield prediction based on a reinforcement learning model, wherein a state STATE for the reinforcement learning model is generated according to the time series data, comprises:
[0036] A time series data of length N is generated by using a specified value as a placeholder after the time series data, and a state STATE for a reinforcement learning model is generated using the time series data of length N, where N represents the maximum number of remote sensing data that can be obtained during the crop growth cycle; so that even if the number of remote sensing data in the time series data changes, the generated state STATE has the same structure.
[0037] According to the second aspect of the present application, the method for crop identification and yield prediction based on reinforcement learning model, wherein:
[0038] Generating a state STATE for a reinforcement learning model according to the time series data includes: processing the time series data through a state embedding network to generate a state STATE for the reinforcement learning model; so that even if the amount of remote sensing data in the time series data changes, the generated state STATE has the same structure.
[0039] According to the second aspect of the present application, the method for crop identification and yield prediction based on a reinforcement learning model, wherein the process of training the prediction network and the confidence network includes multiple training rounds, each training round including multiple time steps;
[0040] At each time step t, the following steps A1 to A5 are performed:
[0041] A1. Processing the time series data SS of training samples t Generate the state STATE at time t t ;
[0042] A2. Set the state to STATE t Provided to the prediction network, the prediction network outputs a prediction result;
[0043] A3. Generate time series data SS for the next time step t+1 t+1 , and generate a reward r based on the prediction result output by the prediction network t ; Among them, r t Represents the reward at time step t, which is obtained by the difference between the prediction result output by the prediction network and the actual historical data;
[0044] A4. Processing time series data SS t+1 Generate the state STATE at time t+1 t+1 ;
[0045] A5. According to the reward r t STATE t+1 and STATE t Calculate TD-error, expressed as δ, δ = r t +λ*Eval(STATE t+1 )-Eval(STATE t ), where Eval() represents the processing performed by the confidence network, and λ represents a hyperparameter; wherein the prediction results include the predicted crop type p and yield a t .
[0046] According to the second aspect of the present application, the method for crop identification and yield prediction based on reinforcement learning model, wherein:
[0047] Time Series Data SS t =(S1, S2, ... S t ), represents the time series data at time t, which includes the remote sensing data from the earliest remote sensing data S1 to the remote sensing data S at time t t The time series data formed, in which the remote sensing data S t The subscript represents the time attribute of the remote sensing data, or the position or sequence number of the remote sensing data in the time series data.
[0048] According to the second aspect of the present application, the method for crop identification and yield prediction based on a reinforcement learning model, wherein the remote sensing data of the target area is preprocessed to generate time series data, including: obtaining remote sensing data of multiple time points within the growth cycle of the crop, the remote sensing data of each time point including year-day attributes; normalizing the data of each band of the remote sensing data at each time point to between 0 and 1; using the remote sensing data of each time point to generate one or more remote sensing indices as features to expand the dimension of the remote sensing data; generating extended remote sensing data for each time point, the extended remote sensing data including the normalized band data of the remote sensing data, one or more remote sensing indices and year-day attributes; and combining the extended remote sensing data of each time point to obtain time series data.
[0049] According to the second aspect of the present application, the method for crop identification and yield prediction based on reinforcement learning model, wherein, in step S1, the objective function of the confidence network JV(w)=∑δ 2 The number of summation items is the number of time steps in the current training round; in step S2, the number of summation items of the objective function J(θ)=∑(log(prob)δ) of the prediction network is the number of time steps in the current training round.
[0050] According to the second aspect of the present application, the method for crop identification and yield prediction based on a reinforcement learning model, wherein the number of time steps in each training round is the number N of time points of remote sensing data within the growing season of the crop; in each training round, it also includes recording each state transition tuple (f t ,(a t ,p),r t ,f t+1 ); the training samples include time series data generated based on historical real remote sensing data and corresponding crop species and yield data; the time series data used in each time step has a different end time point; the generation of time series data SSt+1 for the next time step includes obtaining the time series data SS for time step t t The extended remote sensing data of the next time point of the end time point is appended to the time series data SS of time step t t The time series data SS obtained later t+1 .
[0051] According to the third aspect of the present application, an information processing device according to the third aspect of the present application is provided, comprising a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor implements one of the methods according to the first aspect or the second aspect of the present application when executing the program. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The present application, together with the preferred mode of use and further objects and advantages thereof, will be best understood by reference to the following detailed description of illustrative embodiments when read in conjunction with the accompanying drawings, in which:
[0053] Figure 1 The remote sensing data structure of an embodiment of the present application is shown.
[0054] Figure 2 The embodiment of the present application shows an Actor-Critic architecture reinforcement learning model for predicting crop species.
[0055] Figure 3 A flowchart showing how to predict crop species using a trained reinforcement learning model.
[0056] Figure 4 A reinforcement learning model according to another embodiment of the present application is presented.
[0057] Figure 5 The identification results of crop types according to the present invention are shown.
[0058] Figure 6 The change in prediction accuracy during the training process of the present invention is shown. DETAILED DESCRIPTION
[0059] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and are not to be construed as limiting the present application.
[0060] Remote sensing data is obtained from satellites, drones, and other means. Remote sensing data has multiple dimensions, such as data obtained in different bands (visible light, infrared, microwave, etc.) and from different sources (satellites, drones, balloons, etc.).
[0061] Figure 1 The remote sensing data structure of an embodiment of the present application is shown.
[0062] Figure 1In the remote sensing data, for example, 30 fields are included, and each field is distinguished by a "number". The fields numbered 1-13 are remote sensing data of different bands. Remote sensing data can be in a variety of forms, such as data series or data sets, and the remote sensing data of each band can also be presented in the form of remote sensing images. The fields numbered 14-20 are different indexes extracted from the remote sensing data. By analyzing and processing remote sensing data, a variety of vegetation indices (such as normalized difference preparation index, green normalized difference preparation index, wide range dynamic vegetation index, enhanced vegetation index, etc.) are also obtained. These vegetation indices can also be used as one of the dimensions of remote sensing data.
[0063] The remote sensing data may also include one or more dimensions representing crop species. The crop species may be identified from the remote sensing data or obtained based on external information. Figure 1 In this example, field 22 of the remote sensing data represents the crop category. Field 23 of the input data represents the yield per unit area, which is the actual crop yield of the plot corresponding to the input data obtained based on historical data.
[0064] Optionally, the remote sensing data further includes dimensions representing parameters such as latitude and longitude, altitude, average annual rainfall and / or soil (fields 24-30). The remote sensing data may further include dimensions representing the area of a plot of land.
[0065] Remote sensing data has a time attribute. For example, for a certain plot of land, a remote sensing satellite flies over the plot of land every 5 days, so that a copy of remote sensing data of the plot of land can be obtained every 5 days. Taking into account weather factors (cloud cover, etc.), a copy of high-quality remote sensing data of the plot of land can be obtained every 10-20 days. The day of year (DOY) (field 21) of the remote sensing data represents the time when the remote sensing data is generated, which serves as the time attribute of the remote sensing data. Thus, the remote sensing data can form time series data with several days as intervals. In this application, time series data is, for example, multiple copies of data at different time points. Figure 1 The length of the time series data is denoted as N, where N is a positive integer. Typically, N is the number of remote sensing data within a specified time period (e.g., 1 year). For example, a piece of remote sensing data for a plot of land is obtained every 10 days, and there are 36 pieces of remote sensing data in 1 year, so N = 36. The time period usually represents a growth cycle of a crop. In this application, time series data formed by multiple remote sensing data within a time period is used to identify crop species.
[0066] Optionally, meteorological data can also be added to the remote sensing data as a dimension of the remote sensing data.
[0067] Optionally, preprocess the remote sensing data, including:
[0068] Normalize each band of remote sensing data to between 0 and 1;
[0069] Use remote sensing data to generate a variety of remote sensing indices as features to expand the dimension of remote sensing data.
[0070] The present invention adopts a reinforcement learning model with an Actor-Critic architecture to identify crop species. The reinforcement learning model with an Actor-Critic architecture includes an Actor and a Critic. The Actor observes environmental features as states (such as time series data of remote sensing images), predicts the types of crops, and inputs the prediction results into the environment as actions (action=p), where p represents the identified crop species. The error between the actual crop species / yield and the predicted value is obtained from the environment as a reward (reward, r). The Critic is responsible for evaluating the actions (prediction results) generated by the Actor and outputting a value function (Eval()), which represents the expected value of future accumulated rewards.
[0071] Since the concept of "environment" includes the natural environment in which crops grow, it also includes the unit in the reinforcement learning model that generates environmental states and calculates rewards. To distinguish them, the former is called the "real environment" and the latter is called the "environmental agent." The environmental agent can be implemented by, for example, a computer program that generates environmental states and calculates rewards (r) from remote sensing data of the crop growth environment and provides them to the actor and critic. In this application, a prediction network is provided to implement the actor, and a confidence network is provided to implement the critic.
[0072] Will Figure 1 The displayed remote sensing data is denoted as S. The time series data SS formed by remote sensing data t (denoted as S1, S2, ... S t ), remote sensing data S t The subscript represents the time attribute of remote sensing data, or the position or sequence number of remote sensing data in time series data. t Represents the time series data at time t, which includes the remote sensing data from the earliest remote sensing data S1 to the remote sensing data S at time t t The generated time series data.
[0073] Figure 2 The embodiment of the present application shows an Actor-Critic architecture reinforcement learning model for predicting crop species.
[0074] In the present invention, the time series data SS is used t Generate the state required by the reinforcement learning model (denoted as STATE t In one example, the time series data SS t The remote sensing data at each moment (S1, S2, ... S t) as a combination or concatenation of STATE t In another example, using state embedding networks to process time series data SS t And generate the state STATE used by the reinforcement learning model t Thus, as time goes by, the state STATE t Generated from time series data with more time points, and carries more information about the crop types. The state STATE is obtained by training the reinforcement learning model of the present invention using time series data at different time points in the crop growth cycle. t , which enables the reinforcement learning model to make predictions based on the early state STATE containing a smaller amount of remote sensing data, and to make more accurate predictions based on the later state STATE containing a larger amount of remote sensing data.
[0075] Prediction network (Actor) based on the current state (STATE t ), generates a predicted value for the crop type (p), which acts on the environment as the action generated by the prediction network. During training, the action generated by the prediction network is provided to the environment agent, which calculates the reward using the received action and generates the state (STATE) at the next time point (t+1) based on the remote sensing data at the next time point. t+1 ).
[0076] The task of the Critic network is to determine the current state of the environment based on the state provided by the environment agent. t ) and the reward r given by the environment agent, calculate Eval(), which is used to evaluate the contribution of the prediction network's action strategy to future rewards. Eval() represents the confidence network's STATE at time t. t The value evaluation means that in the input state STATE t When , the expected cumulative reward is generated according to the current strategy. The confidence network is optimized using TD-error (temporal difference error), making the confidence network's estimate of value more accurate.
[0077] Optionally, the output of the prediction network is provided to both the environmental agent of the present invention and the real environment, for example, by publishing the crop species predicted by the policy network (Actor) to agricultural practitioners in the real environment to influence the real environment.
[0078] The process of predicting the network is represented as p←Pred(STATE t ), where Pred() represents the state STATE of the predicted network at time t for the input tProcessing is performed to obtain the result p, where p represents the identified crop type.
[0079] The process of the confidence network is represented as Eval(STATE t ), the confidence network calculates the TD error (TD-error)δ=r based on the error t +λ*Eval(STATE t+1 )-Eval(STATE t ), provides the prediction network as the confidence (δ), where r t Represents the reward at time t, and Eval() represents the belief network's STATE at time t t The value evaluation means that in the input state STATE t When , the expected cumulative reward is generated according to the current strategy. λ represents the degree of confidence in the predicted reward, ranging from 0 to 1. If λ is 1, it means that the predicted reward is absolutely trusted. If it is 0, it means that only the current reward is considered without considering future rewards. As an example, λ = 1 can be used. Reward r t It can be obtained from the crop type recognition error at time t. The prediction network uses the received confidence (δ) to update its own parameters.
[0080] In this application, the fully connected layer of the prediction network is fitted with Pred(), and the fully connected layer of the confidence network is fitted with Eval().
[0081] For the crop types, codes are provided for them to facilitate the calculation of rewards. Table 1 shows examples of crop type codes. In Table 1, different code values are set for each predicted crop type, and the code values are integers between 0 and 7, for example.
[0082] Optionally, the prediction network also outputs a prediction of crop yield. Both the crop type and yield predictions are used to calculate the reward, ensuring that the calculated reward is of the same or similar order of magnitude as the reward derived from the predicted yield.
[0083] Crop yields are converted to "relative yields" and then encoded. This encoding is used to calculate rewards and also eliminates the impact of larger actual yield values on reward calculations. Relative yields are generated by preprocessing the actual crop yields. Relative yield is the ratio of the yield (predicted or actual) of a certain crop in a certain location in a certain year (or a certain crop growing season) to the average yield of that location over the years (or multiple crop growing seasons). Using relative yields also allows the model to adapt to yield predictions for different types of crops and different regions. Table 2 provides an example of encoding relative yields. Through crop yield encoding, the relative yield representing the ratio is converted into a code value. The code value of the crop yield code is, for example, an integer between 0 and 6. The code value of the crop yield code has a similar value range to the code value of the crop type code, making it easier to combine the two prediction settlements to calculate rewards.
[0084] Table 1
[0085]
[0086] Table 2
[0087]
[0088] Rewards are given based on the difference between predicted output and actual output.
[0089] When only predicting the category, the reward r = r x = {Total number of identified crop species (correct judgment), 0 (wrong judgment)}. When predicting both crop species and yield, the specific formula is
[0090] r=r x +r y
[0091] Among them, r x is the reward for crop type discrimination, r y is the output forecast reward, and the calculation formula is
[0092] r x = {Total number of identified crop types (correct), 0 (incorrect)}
[0093] r y = Total yield class - ABS (forecast crop yield code - actual crop yield code)
[0094] This design allows two rewards r x and r y The output levels are roughly on the same order of magnitude to avoid bias during training. It also rewards close output levels, enhancing the robustness of the model.
[0095] In calculating the crop type discrimination reward r x If the prediction network predicts the correct crop types, the number of correctly predicted crop types is taken as r x For example, if there is only one type of crop in the data sample and the prediction network correctly predicts its type, then r x = 1. If there are three kinds of crops in the data sample, and the prediction network correctly predicts the types of two of them, but incorrectly predicts the type of the other, then r x =2.
[0096] In calculating the output prediction reward r y When , "Total Yield Grades" is the total number of possible values for the crop yield code. In the example in Table 2, "Total Yield Grades" is 7. ABS() represents a function that calculates absolute values. The predicted crop yield code comes from the prediction network's output, while the actual crop yield code comes from preprocessed data samples. This is calculated by mapping the actual yield of a particular crop in a particular location in a particular year (or crop growing season) to the average yield of that location over the past years (or multiple crop growing seasons) to the crop yield code.
[0097] Return to view Figure 1 , the prediction network (Actor) and confidence network (Critic) of the reinforcement learning model of this application obtain the state STATE from the environment agent at time t t The prediction network (Actor) is based on the received state STATE t Output the predicted crop type p as the action at time t and provide it to the environment agent. The environment agent generates a reward r at time t based on the received action t And provide the state STATE of the next moment (t+1) t+1 ,STATE t+1 The time series data SS at time t+1 t+1 (SS t+1 =(S1,S2,…S t+1 )) is obtained. As an example, the time series data SS at time t t Including the remote sensing data S from the first remote sensing data S1 of the crop growth cycle to the remote sensing data S at time t t+1 A sequence composed of multiple remote sensing data.
[0098] The environmental agent obtains the remote sensing data of the next moment from the data sample based on time t. The current moment t and the next moment t+1 are adjacent. For example, if the data sample is generated from the remote sensing data obtained from the satellite every 10 days, then the moment t+1 is 10 days later than the moment t. Therefore, the environmental agent of this application obtains the remote sensing data S from the data sample. tThe data sample of the next moment represented by the current moment t is used to generate remote sensing data S t+1 This is different from the reinforcement learning method of the prior art. In the reinforcement learning model of the prior art, the environment generates the state of the next moment based on the current state and the impact of the action at the current moment on the environment. For example, in a chess game, the action is a move, and the new situation of the chess game after the move forms the state of the next moment. The environmental agent of the present invention does not need to simulate the state change of the environment, but uses the data samples generated from the remote sensing data to generate the remote sensing data S at time t+1. t+1 , and then generate the time data series SS at time t+1 t+1 .
[0099] In this invention, the environment agent calculates the reward r at time t based on the difference between the predicted crop type and the actual crop type during the received action according to the aforementioned formula. In addition to using the predicted value output by the prediction network, the calculation of reward r also requires the actual crop type at time t+1.
[0100] Optionally, when the prediction value output by the prediction network also includes the predicted yield, in agricultural production, except for the time when the crops mature and are harvested, there is no actual yield corresponding to these times; in this application, the final yield of the crop in the year (or season) is used as the actual yield corresponding to each data sample in the data sample time series of the year (or season), and is used to calculate the reward r. Optionally, some predictions output by the prediction network include multiple crops, and each crop has its own predicted yield and actual yield. In calculating the reward r y When the yield of each crop is calculated, r y , and use its statistical values (sum, average, weighted average) to calculate the reward r.
[0101] The environment agent outputs the reward r corresponding to the output of the prediction network at time t t Provided to the confidence network. The confidence network is based on the state STATE at time t t Output value evaluation Eval(STATE t ). The confidence network also evaluates the value of the state of the two adjacent moments Eval (STATE t+1 ) and Eval(STATE t ) and the reward r at time t t Calculate TD-error, expressed as δ = r t +λ*Eval(f t+1 )-Eval(f t ). The TD-error calculated at each moment is used to update the parameters of the prediction network and the confidence network in subsequent training.
[0102] Next, the reinforcement learning model advances to the next moment. Using the next moment state STATE t+1 Provided to the prediction network and confidence network, and repeat the above process, the prediction network outputs the prediction result (crop type and / or yield) at the next moment, and the environment agent calculates the reward r at the next moment t+1 , and the confidence network based on the state STATE t+2 and STATE t+1 Eval(STATE t+2 ) and Eval(STATE t+1 ) calculates the TD-error at the next moment. This process is called a time step.
[0103] The training process of a reinforcement learning model consists of multiple epochs, each of which includes multiple time steps. The number of time steps is, for example, the number of data samples in a time series obtained from remote sensing data over a crop growing season. After one epoch of learning, for example, the feature extraction network, prediction network, and confidence network are updated. Further epochs of learning are also performed.
[0104] Optionally, the complete growth cycle of crops is used for training. The image features of crops at different stages contained in the complete growth cycle are the key features that distinguish the image of one crop from the image of another crop, and the crop yield data in the training data that directly corresponds to the end time point of the complete growth cycle can be utilized. It should be understood that during the training phase, the time series data of the complete growth cycle of crops can be used to generate STATE, but during the prediction phase, since the harvest season is often not yet reached when the prediction occurs, remote sensing data for the complete growth cycle has not yet been obtained. Therefore, the state STATE generated by the time series data of the incomplete growth cycle can be used to identify the crop type and / or predict the yield. Optionally, the training process can also use the state STATE generated by the time series data of the incomplete crop growth cycle.
[0105] Model training
[0106] Table 3 shows the training process of the reinforcement learning model according to the present invention.
[0107] One epoch of training consists of multiple time steps, where t represents the time step. At each time step, the environment agent obtains remote sensing data from the data sample and generates input time series data SS t According to the time series data SS t Generate the state STATE at time t t .
[0108] Temporal difference (TD-error) is used to optimize the prediction network and the confidence network. TD-error is calculated using rewards. The learning goal of the prediction network is to maximize the policy value, while the goal of the confidence network is to minimize the prediction error. The network parameters are updated based on TD-error and the backpropagation algorithm.
[0109] See Table 3, a t Represents the prediction results output by the policy network (Actor) at time step t, including the predicted crop type. t Represents the state at time step t, which is derived from remote sensing data and historical data on the actual types of crops during the growing season. P represents the probability density corresponding to the predicted result, provided by the policy network (Actor). p represents the predicted result output by the prediction network at time t. t Represents the state at time t, from the time series data SS t Generate. t Represents the reward at time t, which is obtained by the difference between the prediction network's prediction result and the actual historical data. Prob represents the probability distribution corresponding to the prediction result p output by the prediction network, provided by the prediction network.
[0110] At each time step in the training process, the state STATE t The prediction network outputs the prediction result p (step 4 of Table 3). In Table 3, Pred() represents the processing performed by the prediction network. The prediction result p output by the prediction network is provided to the environment agent, which then determines the prediction result p based on the state STATE. t Generate the next state STATE at time step t+1 t+1 and the reward r at time step t t (Step 5 of Table 3). As an example, the environment agent obtains the state S at time t+1 based on the time series of remote sensing data in the training data. t+1 , and calculate the reward r according to the reward calculation method provided by the present invention t In Table 3, EVN() represents the processing performed by the environment agent. Optionally, the environment agent outputs a time data series SS t+1 rather than STATE t+1 , and by additional processing according to SS t+1 Generate STATE t+1 The confidence network is based on the state STATE t+1 and STATE t Calculate the TD-error, denoted as δ (step 5 in Table 3). Eval() in Table 3 represents the processing performed by the belief network. Generate a state transition tuple (STATE t ,p,r t ,STATEt+1 ), which is used to subsequently update the parameters of the prediction network and the confidence network.
[0111] A training epoch includes multiple time steps. The number of time steps (denoted as N) is, for example, the number of remote sensing data in a remote sensing data time series within a crop growing season. Alternatively, the number of time steps in a training epoch can take other values.
[0112] For each time step of training, repeat steps 2 to 6 of Table 3. If the current time step is t, it is used to generate state STATE t The time series data is SS t , then the next time step is t+1, and the time series sequence used to generate the state for the next time step is SS t+1 Therefore, in the later time steps, the time series data used include more remote sensing data (and closer to the end of the crop growth cycle). In each growing season of crops, there is only one real crop species data as a data sample for training the reinforcement learning model. When calculating the reward at each time step of a training round (step 4 of Table 3), regardless of whether the time series sequence of the current time step is SS t Regardless of the number of remote sensing data points, rewards and errors are calculated using the same real-world crop data from the same growing season. Temporal difference (TD-error) is combined with the prediction network and confidence network to optimize the reinforcement learning model, enabling it to learn to achieve more accurate identification and prediction results based on more remote sensing data from the same growing season.
[0113] Using the N time steps of a crop growing season, calculate the objective function J of the confidence network V (w)=∑δ 2 , where w represents the confidence network parameter, and the squares of the N δ values obtained over N time steps are summed (step 7 of Table 3). The objective function of the prediction network is calculated as J(θ) = ∑(log(prob)δ) (step 8 of Table 3), with N summation terms. The prediction network is updated accordingly (step 9 of Table 3) and the confidence network is updated accordingly (step 10 of Table 3). The goal of learning the prediction network is to maximize the confidence, while the goal of learning the confidence network is to minimize the error in the future error estimate.
[0114] In steps 7 and 8 of Table 3, and the corresponding steps 9 and 10, the state transition tuples of a data batch are sampled to update the parameters of the prediction network and the confidence network. A data batch includes, for example, 256 data items. A data batch can include other amounts of data. The data in a data batch can include data from different plots, different crop growing seasons, and different types of crops, so that the trained model has greater versatility, eliminating the need to train a dedicated model for each plot or each crop.
[0115] The present invention has no mandatory requirements for the number of training epochs, batch size, or time step. These are adjusted based on the natural spacing of the data and the speed of model convergence. For example, since Sentinel satellites pass over a fixed ground point once every five days, one time step is used, with every five days of satellite remote sensing data as one training batch. The number of training epochs is 256 samples, the number of training epochs is 5000, and the learning rate is set to 1e-5, with the learning rate reduced to 10% every 2000 steps.
[0116] Alternatively, the number of cycles from step 2 to step 6 in Table 3 represents a complete growth cycle of a crop. For example, the growth cycle of wheat is about 240 days, and a remote sensing data set is obtained every 5 days. The complete time series data consists of 48 remote sensing data sets. 48 time steps are used to input the time series data SS with lengths of 1 to 48. t Generate the corresponding state STATE t For example, if we take the natural year as the growth cycle of crops and obtain a remote sensing data set every 5 days, the complete time series data has 74 remote sensing data sets. Accordingly, one round of training has 74 time steps. In some cases, the remote sensing data of some time points are missing (for example, blocked by clouds), and the input state STATE t This may reflect missing data.
[0117] The trained prediction network is used to identify the crop species based on the state generated by the time series data of the current remote sensing data. In this case, the confidence network can be omitted. The state transition tuple (STATE t ,p,r t ,STATE t+1 ) and further train and optimize the prediction network and confidence network according to the training process shown in Table 3.
[0118] Optionally, the prediction network of the embodiment of the present application also predicts crop yields. Accordingly, the processing Pred() performed by the prediction network not only outputs the crop type recognition result p, but also outputs the prediction a of the crop yield based on the state at time t. t, and according to p and a t Calculate reward r t =r x +r y (According to the previous formula) (See also step 3 of Table 3). And the environment agent also uses the state including p and a when outputting the state at time t+1 t The prediction result at time t.
[0119] Table 3
[0120]
[0121] In traditional reinforcement learning models based on the actor-critic architecture, the action A generated by an actor affects the environment, changing its state. The output from the environment serves as a reward to guide the actor in optimizing its predictions. The rewards corresponding to different actions can vary significantly. In agricultural applications, reinforcement learning is used to optimize irrigation and fertilization plans for crops. These actions directly affect the environment, and the state of the environment changes as a result of these actions. However, crop species identification is not typically considered an action that can affect the environment. Consequently, the process lacks a closed-loop "action-environment-reward" mechanism, which limits the application of reinforcement learning methods for crop species identification. Therefore, traditionally, reinforcement learning models based on the actor-critic architecture cannot be applied to crop species identification scenarios. Therefore, the crop recognition model of the present invention is not mathematically equivalent to the traditional method. By converting remote sensing data into time series data, it is suitable for the characteristics of continuous interaction between reinforcement learning and the environment. The design method of optimizing the prediction network and the confidence network separately is essentially an architecture that separates prediction and evaluation, so that the value function can measure the uncertainty of the current state, which is lacking in traditional regression methods. It can adapt to more complex scenarios without considering hidden variables (meteorological conditions, soil, altitude, time, etc.).
[0122] Moreover, in traditional reinforcement learning models, each state processed has the same structure (e.g., different chess situations). However, in the present invention, each time step in training provides the state STATE of the prediction network and the confidence network. t Corresponding to time data series SS of different lengths t , as time step t advances, longer time data series SS t Not only does it represent different states, but the amount of information carried in the state increases, which helps to obtain more accurate prediction results as time t progresses. When applying the trained prediction network, the state STATE input to the prediction network t Can also be generated from time data series SS of different lengths t, so that the present invention can be used to identify crop species regardless of when in the crop growth cycle.
[0123] Figure 3 A flowchart showing how to predict crop species using a trained reinforcement learning model.
[0124] First, a target area is selected and remote sensing data of the target area is acquired. The acquired remote sensing data may be multiple copies, each representing remote sensing data at multiple consecutive time points. The multiple copies of remote sensing data are, for example, all available remote sensing data of the target area from the beginning of the current crop growth cycle to the present. When the crop growth cycle is based on a natural year, the multiple copies of remote sensing data are all available remote sensing data from the beginning of the year or after sowing to the present. Optionally, the remote sensing data also includes day-of-year (DOY) information to indicate the temporal position of the generation time of the remote sensing data in the crop growth cycle.
[0125] Preprocess remote sensing data to generate time series data. This includes, for example, normalizing the data across bands, generating one or more remote sensing indices to expand the dimensionality of the input data, and embedding year-day information. Preprocessing also optionally addresses missing data and supplements it with weather and geographic location data.
[0126] Process time series data to generate states for reinforcement learning models.
[0127] The state STATE is provided to the prediction network, which generates a predicted crop type based on the state STATE.
[0128] In the early stages of the crop growth cycle, the input STATE includes earlier and smaller amounts of remote sensing data, and the prediction network may not be accurate. As time goes by, more remote sensing data for the target area is available. Using more remote sensing data to generate time series data and STATE to provide to the feature network, the prediction network can make more accurate predictions accordingly. Figure 3 The process shown can be executed multiple times, for example, monthly. Figure 3 The processing flow is used to obtain the predicted value of the target area. As more remote sensing data of the current growth cycle are collected, the prediction results will become more accurate.
[0129] When predicting the types of crops, it is not necessary to use a confidence network. Optionally, when the types of crops currently planted in the target area are obtained, the training process shown in Table 3 can also be used to optimize the prediction network and confidence network provided by the present invention. For example, see Figure 3, by comparing the predicted crop type with the actual value to generate a reward, the reward can further calculate the prediction error (TD-Error, denoted as δ), so that the confidence network's prediction of confidence can be corrected and the prediction network and confidence network can be updated.
[0130] Due to the different growth periods of crops, the time data series SS t There are different amounts of remote sensing data, and the state STATE t It is beneficial to the processing of the prediction network and the confidence network. Optionally, the structure of STATE is determined according to the maximum number of remote sensing data in the crop growth cycle (denoted as N), and the remote sensing data after time t that has not been obtained at the current time t can be occupied by a specified value and form a time data sequence of length N. Then, the state STATE at time t is generated according to the time data sequence of length N. t , so that the state STATE generated at different times t t With the same structure. Alternatively, a state embedding network is used to process time data sequences SS of different lengths. t , and generate a state STATE with the same structure t .
[0131] Figure 4 A reinforcement learning model according to another embodiment of the present application is presented.
[0132] Figure 4 The reinforcement learning model shown is the same Figure 2 The reinforcement learning model of is similar in principle, except that Figure 4 The model also includes a state embedding network to transform the time series data SS t Generate the state STATE required by the reinforcement learning model t .
[0133] The state embedding network transforms the time series data SS t Converted into vector form to facilitate the processing of prediction network and confidence network, and also make it possible for the input time series data SS t =(S1, S2, ... S t ) includes how many copies of data, and the state after processing STATE t It has the same structure to facilitate its use. It can also process time series data SS through state embedding network t Some data in the missing, making the generated state STATE t It has the same structure and can still be accepted by the prediction network and the confidence network.
[0134] Figure 5 The identification results of crop types according to the present invention are shown.
[0135] Figure 5 In the figure, the horizontal axis is time, two remote sensing images are used, and the proportion line represents the proportion of the pixels to be tested in each image, that is, 100% of the pixels have one remote sensing image, and 99% of the pixels have two remote sensing images. Figure 5 The upper left graph is the accuracy index, the upper right graph is the precision rate, the lower left graph is the recall rate, and the lower right graph is the F1 harmonic mean of the accuracy rate and the precision rate. Figure 4 network structure.
[0136] exist Figure 5 In the upper left figure, when only one image was available (early stage of crop growth), the wheat species recognition accuracy (blue line) was 88.6%. 15 days later, when the second remote sensing image was obtained, the wheat species recognition accuracy increased to over 96%.
[0137] exist Figure 5 In the upper right figure, when there was only one image (early stage of crop growth), the wheat recognition accuracy (yellow line) was 75%. When the second remote sensing image was obtained 15 days later, the wheat recognition accuracy increased to over 90%.
[0138] It can be seen that most indicators have improved over time, indicating that the method of this patent invention can identify crops in the early stages of crop growth and improve accuracy and recall over time.
[0139] Figure 6 The figure shows the change in crop type prediction accuracy during the training process of the present invention. The CTP reward (crop type reward) is calculated based on the error and gradually increases as the training progresses.
[0140] The present invention includes the following key points.
[0141] The separate design of the confidence network and the prediction network enables predictions in complex scenarios: By separating the prediction model into two relatively independent networks, the "confidence network" and the "prediction network", and training them using different optimization objectives, they can better adapt to the complex and changing agricultural environment, be applicable to various crop types in different regions, and improve prediction accuracy.
[0142] The design implementation outputs confidence by taking the expected error as the objective function: the estimation of the error is used as the estimation of the uncertainty of the current state, and the prediction confidence is evaluated while the prediction is achieved.
[0143] TD-error-based value function optimization: Species identification error is used as a long-term goal, and network parameters are adjusted based on the error at each time step, allowing the model to predict crop production in the early stages of planting and gradually improving prediction accuracy over time.
[0144] Universal prediction architecture: The design of separating the prediction network and the confidence network proposes a new universal prediction architecture that can be applied to other similar complex scenarios.
[0145] The crop variety prediction model based on the present invention can be widely used in agricultural production, food security prediction, agricultural resource allocation and other fields, and can improve the decision-making efficiency and accuracy of agricultural production.
[0146] In addition, the present invention proposes improvements to the design of crop type and yield labels to further support medium- and long-term yield predictions.
[0147] This method innovatively combines state and action with a reward design mechanism, proposing a general implementation architecture for prediction tasks.
[0148] Although the present application has been described with reference to examples, this is for illustrative purposes only and is not intended to limit the present application, and changes, additions and / or deletions to the embodiments may be made without departing from the scope of the present application.
[0149] Those skilled in the art to which these embodiments relate and who benefit from the teachings presented in the above description and the associated drawings will recognize many modifications and other embodiments of the present application described herein. Therefore, it should be understood that this application is not limited to the specific embodiments disclosed, and modifications and other embodiments are intended to be included within the scope of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.
Claims
1. A method for crop identification based on a reinforcement learning model, wherein the reinforcement learning model includes a prediction network and a confidence network, and the method comprises: Preprocessing remote sensing data of the target area to generate time series data; Generate a state STATE for the reinforcement learning model according to the time series data, and provide the state STATE to the prediction network; The prediction network generates a predicted type of the crop according to the state STATE; The process of training the prediction network includes: At each time step t, the state STATE at time t is generated based on the training sample t Provided to the prediction network; the prediction network outputs the prediction result; based on the prediction result output by the prediction network, a reward r is generated t ; Generate the state STATE at time t+1 t+1 ; The confidence network is based on the reward r t STATE t+1 and STATE t Calculate TD-error, denoted as δ; Where t represents the current time step; In each training round, the following steps S1 to S4 are also performed: S1. Using the δ of multiple time steps of the training round, calculate the objective function of the belief network JV(w) = ∑δ 2 , where w represents the parameters of the belief network; S2. Calculate the objective function J(θ)=∑(log(prob)δ) of the prediction network, where prob represents the probability distribution corresponding to the prediction result output by the prediction network, and θ is a parameter of the prediction network; S3. Update the feature extraction network, update the prediction network, and update the confidence network.
2. The method according to claim 1, wherein The prediction result includes the predicted crop type p; as well as Calculate the reward r at time t t , in, If the crop type is correctly judged, r t =Total number of identified crop types. If the crop type is wrong, r t =0.
3. The method according to claim 2, wherein: Generating a state STATE for a reinforcement learning model according to the time series data includes: Generating time series data of length N using a specified value as a placeholder after the time series data, and generating a state STATE for a reinforcement learning model using the time series data of length N, where N represents the maximum number of remote sensing data that can be obtained during a crop growth cycle; Even if the number of remote sensing data in the time series data changes, the generated state STATE has the same structure.
4. The method according to claim 2, wherein: Generating a state STATE for a reinforcement learning model according to the time series data includes: Process time series data through a state embedding network to generate the state STATE for the reinforcement learning model; Even if the number of remote sensing data in the time series data changes, the generated state STATE has the same structure.
5. The method according to claim 3 or 4, wherein: The process of training the prediction network and the confidence network includes a plurality of training rounds, each training round including a plurality of time steps; At each time step t, the following steps A1 to A5 are performed: A1. Processing the time series data SS of training samples t Generate the state STATE at time t t ; A2. Set the state to STATE t Provided to the prediction network, the prediction network outputs a prediction result; A3. Generate time series data SS for the next time step t+1 t+1 , and generate a reward r based on the prediction result output by the prediction network t ; Among them, r t Represents the reward at time step t, which is obtained by the difference between the prediction result output by the prediction network and the actual historical data; A4. Processing time series data SS t+1 Generate the state STATE at time t+1 t+1 ; A5. According to the reward r t STATE t+1 and STATE t Calculate TD-error, expressed as δ, δ = r t +λ*Eval(STATE t+1 )-Eval(STATE t ), where Eval() represents the processing performed by the confidence network, and λ represents a hyperparameter; wherein the prediction result includes the predicted crop type p.
6. The method according to claim 5, wherein: Time Series Data SS t =(S1, S2, ... S t ), represents the time series data at time t, which includes the remote sensing data from the earliest remote sensing data S1 to the remote sensing data S at time t t The time series data formed, in which the remote sensing data S t The subscript represents the time attribute of the remote sensing data, or the position or sequence number of the remote sensing data in the time series data.
7. The method according to claim 6, wherein: The preprocessing of the remote sensing data of the target area to generate time series data includes: Acquiring remote sensing data at multiple time points within the growth cycle of the crop, wherein the remote sensing data at each time point includes year and day attributes; Normalize the data of each band of remote sensing data at each time point to between 0 and 1; Using remote sensing data at each time point to generate a variety of remote sensing indices as features, expanding the dimension of remote sensing data; Generate extended remote sensing data at each time point, the extended remote sensing data including normalized band data of remote sensing data, one or more remote sensing indices, and annual and daily attributes; The extended remote sensing data at each time point are combined to obtain time series data.
8. The method according to claim 7, wherein: In step S1, the objective function of the confidence network JV(w)=∑δ 2 The number of summations in is the number of time steps in the current training epoch; In step S2, the number of summation terms of the objective function J(θ)=∑(log(prob)δ) of the prediction network is the number of time steps in the current training round.
9. The method according to claim 8, wherein The number of time steps in each training round is the number of time points N of remote sensing data within the growing season of the crop; In each training round, it also includes a state transition tuple (f t ,p,r t ,f t+1 ); The training samples include time series data generated based on historical real remote sensing data and corresponding crop type data; The time series data used in each time step has a different end time point; The time series data SS for the next time step is generated t+1 , including obtaining the time series data SS at time step t t The extended remote sensing data of the next time point at the end time point is appended to the time series data SS of time step t t The time series data SS obtained later t+1 .
10. An information processing device comprising a memory, a processor, and a program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 9 is implemented.