Crop fixed-point harvesting and putting path optimization system based on deep reinforcement learning
By introducing deep reinforcement learning technology into the crop harvesting system, real-time analysis of farmland status and harvester location, and generating optimized operation strategies, the problems of inefficiency and resource waste in traditional agricultural operations are solved, efficient and accurate harvesting and delivery are achieved, and production costs are reduced.
Patent Information
- Application Number
- CN202510295260.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-13
AI Technical Summary
The harvesting and delivery of traditional crops mainly relies on manual operations, resulting in inefficiency, waste of resources and increased production costs.
A fixed-point harvesting and delivery path optimization system for crops based on deep reinforcement learning is adopted. Through environmental modeling, state space, action space, reward mechanism and deep learning module, farmland status and harvester location are obtained and analyzed in real time, feasible operation strategies are generated, and the accuracy and timeliness of harvesting are improved by dynamically adjusting path planning and task scheduling.
It effectively overcomes the inefficiency and resource waste caused by manual intervention in traditional agricultural operations, improves the accuracy and timeliness of harvesting, reduces production costs, and can adapt to the complex and changeable agricultural environment to achieve intelligent decision-making and automated operations.
Smart Images

Figure CN120143833A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of agricultural intelligence, and specifically to a system for optimizing the path of fixed-point harvesting and placing of crops based on deep reinforcement learning. Background Art
[0002] In recent years, global agricultural production has faced increasingly severe challenges. With the continuous growth of the population and the acceleration of urbanization, the demand for food has risen sharply, prompting agricultural production to improve efficiency and increase yields. At the same time, environmental problems such as climate change, soil degradation, and resource shortages have also had a profound impact on agricultural production. Against this background, traditional agricultural production methods have gradually revealed their deficiencies of low efficiency, high cost, and large environmental impact, and there is an urgent need for transformation and upgrading.
[0003] Traditional crop harvesting and placing mainly rely on manual operations, which are not only inefficient but also easily affected by human factors, resulting in resource waste and increased production costs. Summary of the Invention
[0004] (1) Technical Problems to be Solved
[0005] Aiming at the deficiencies of the prior art, the present invention provides a system for optimizing the path of fixed-point harvesting and placing of crops based on deep reinforcement learning. By introducing intelligent environmental modeling, state space, and action space modules, it can obtain and analyze the state of the farmland and the position of the harvester in real time, thereby generating feasible operation strategies, effectively overcoming the problems of low efficiency and resource waste caused by manual intervention in traditional agricultural operations. At the same time, by dynamically adjusting path planning and task scheduling, it not only improves the accuracy and timeliness of harvesting but also significantly reduces production costs. In addition, combined with the reward mechanism and deep learning technology, the system can adapt to complex and changing agricultural environments during continuous learning and optimization, realizing intelligent decision-making and automated operations.
[0006] (2) Technical Solutions
[0007] To achieve the above object, the present invention provides the following technical solutions: A system for optimizing the path of fixed-point harvesting and placing of crops based on deep reinforcement learning, including an environmental modeling module, a state space module, an action space module, a reward mechanism module, a deep learning module, and a scheduling and optimization module;
[0008] The environmental modeling module is used to obtain the coordinates of the farmland, the position and state of the harvester, the coordinates of the placement point, and calculate the distance between farmlands, the initial position of the harvester, and the remaining carrying capacity of the harvester, and transmit the calculated values to the state space module;
[0009] The state space module calculates the harvester state vector based on the position of the harvester, the remaining carrying capacity of the harvester, and the coordinates of the current farmland to be transported for subsequent update of the harvester position state;
[0010] The action space module obtains a list of feasible target farmlands and calculates the index of the next farmland, thereby generating a list of feasible actions based on the current state;
[0011] The reward mechanism module calculates the negative reward for task completion and the negative value of the travel distance based on the list of feasible actions, and feeds it back to the deep learning module;
[0012] The deep learning module obtains the state, action, reward, next state (experience replay), calculates the Q value and updates the loss function, extracts samples from the memory pool using experience replay, updates the parameters of the Q network, and trains the model;
[0013] The scheduling and optimization module performs path planning according to the current harvesting task state, the states of each farmland, and the Q value strategy output by the trained deep learning model, and schedules the movement and tasks of the harvester.
[0014] Preferably, the calculation formula for the distance between farmlands is as follows:
[0015]
[0016] In the formula, D i,j represents the distance from farmland 1 to farmland 2, x i , y i represent the position coordinates of farmland 1, x j , y j represent the position coordinates of farmland 2.
[0017] Preferably, the calculation formula for the initial position of the harvester is as follows:
[0018] (x 0 , y 0 ) = (x farm , y farm )
[0019] In the formula, (x 0 , y 0 ) represents the initial position of the harvester, x farm represents the x coordinate of the selected initial farmland, y farm represents the y coordinate of the selected initial farmland.
[0020] Preferably, the calculation formula for the remaining carrying capacity of the harvester is as follows:
[0021] C r = C max-C current
[0022] In the formula, C r represents the remaining carrying capacity of the harvester, and C max represents the maximum carrying capacity of the harvester, and C current represents the total weight of the crops that have been transported currently.
[0023] Preferably, the calculation formula of the harvester state vector is as follows:
[0024] S = (x r , y r , C r )
[0025] In the formula, S represents the harvester state vector, x r , y r represent the current coordinate position of the harvester, and C r represents the remaining carrying capacity of the harvester.
[0026] Preferably, the calculation formula of the index of the next farmland is as follows:
[0027]
[0028] In the formula, n represents the index of the selected next target farmland, F j represents the j-th farmland in the list of feasible target farmlands, and Q(S, F j ) represents the Q value between the current state S and the target farmland F j .
[0029] Preferably, the calculation formula of the negative reward for completing the task is as follows:
[0030] R task = -ω t *I
[0031] In the formula, R task represents the negative reward for completing the task, ω represents the negative reward weight for completing the task, and I represents the indicator function, which is 1 if the task is completed and 0 if the task is not completed.
[0032] Preferably, the calculation formula of the negative value of the driving distance is as follows:
[0033] R disk = -ω d *D i,j
[0034] In the formula, R disk represents the negative value of the driving distance, ω d represents the negative reward weight for the driving distance, and D i,jRepresents the distance from farmland 1 to farmland 2.
[0035] Preferably, the calculation formula of the Q value is as follows:
[0036]
[0037] In the formula, Q(S,A) represents the Q value of taking action A in state S, α represents the learning rate, R represents the total reward obtained currently, γ represents the discount factor, S′ represents the next state, and a′ represents the next possible action.
[0038] Preferably, the formula for updating the loss function is as follows:
[0039]
[0040] In the formula, L(θ) represents the loss function, N represents the number of samples used to calculate the loss, R i represents the reward of the i-th sample, Q(S′,a′;θ) represents the Q value in the target network, and Q(S,A i ;θ) represents the Q value in the current network.
[0041] Compared with the prior art, the present invention provides a crop fixed-point harvesting and placement path optimization system based on deep reinforcement learning, which has the following beneficial effects:
[0042] By introducing intelligent environment modeling, state space and action space modules, the present invention can obtain and analyze the state of the farmland and the position of the harvester in real time, so as to generate feasible operation strategies, effectively overcoming the problems of low efficiency and resource waste caused by manual intervention in traditional agricultural operations. At the same time, by dynamically adjusting path planning and task scheduling, not only the accuracy and timeliness of harvesting are improved, but also the production cost is greatly reduced. In addition, combined with the reward mechanism and deep learning technology, the system can adapt to complex and changing agricultural environments during continuous learning and optimization, realizing intelligent decision-making and automated operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a schematic diagram of the system flow of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0045] In view of the problem that traditional crop harvesting and placement mainly rely on manual operation, which is not only inefficient but also vulnerable to human factors, resulting in resource waste and increased production costs, a system for optimizing the path of fixed-point crop harvesting and placement based on deep reinforcement learning is proposed. Please refer to Figure 1 , the system includes an environment modeling module, a state space module, an action space module, a reward mechanism module, a deep learning module, and a scheduling and optimization module;
[0046] The main function of the environment modeling module in the system is to obtain relevant information about the farmland and the status of the harvester, so as to provide necessary inputs for the subsequent decision-making process. Specifically, this module obtains the coordinates of each farmland and the coordinates of the placement points through sensors or databases. After obtaining these coordinates, the module calculates the distance D between the farmlands i,j , and this distance can be expressed by the following formula:
[0047]
[0048] where i and j represent different farmland indices, and x i , y i represent the position coordinates of farmland 1, and x j , y j represent the position coordinates of farmland 2;
[0049] In addition, the environment modeling module also needs to determine the initial position (x 0 , y 0 ) of the harvester, and this position is usually set as the coordinates of the first farmland selected at the start of the harvesting task:
[0050] (x 0 , y 0 ) = (x farm , y farm )
[0051] This module also needs to calculate the remaining carrying capacity C r of the harvester, and this capacity will change dynamically during the harvesting process. Assuming that the maximum carrying capacity of the harvester is C max and the weight of the crops already transported is C current , then the formula for calculating the remaining carrying capacity is:
[0052] C r = C max - C current
[0053] The results of all these calculations, including the farmland coordinates, the initial position of the harvester, the distances between farmlands, and the remaining carrying capacity of the harvester, will subsequently be integrated and transmitted to the state space module. The state space module uses this data to generate a state vector of the harvester. The calculation formula for this state vector is:
[0054] S = (x r , y r , C r )
[0055] where x r and y r are the current coordinates of the harvester, and C r is the remaining carrying capacity. In this way, the environmental modeling module provides a sound data basis for subsequent decision-making processes and path optimization, enabling the harvester to effectively execute the harvesting task;
[0056] The action space module plays a crucial role in the intelligent decision-making system of the agricultural autonomous harvester. Its main function is to obtain a list of feasible target farmlands based on the current state of the harvester and calculate the index of the next farmland, thereby generating a corresponding list of actionable actions. First, the module obtains the information of currently reachable farmlands through information interaction with the environmental modeling module and organizes it into a list F. The elements in this list are all the farmland coordinates (x farm , y farm ) that the harvester can go to.
[0057] To determine the index n of the next target farmland, the module compares the current state vector S with the Q-value Q(S, F j ) corresponding to each farmland and selects the next optimal target using the following formula:
[0058]
[0059] This formula indicates that the index of the farmland with the highest Q-value will be selected in the next step, thereby formulating the best subsequent action plan for the harvester. At the same time, the list of actionable actions generated in this process covers all actionable actions from the current position to each candidate target farmland, which is an important basis for the system to make decisions based on the current state;
[0060] On this basis, the reward mechanism module plays a role in evaluating and providing feedback on each possible action. The module calculates the negative reward for completing the task and the negative value of the travel distance depending on the list of actionable actions. Specifically, the negative reward R task for completing the task can be expressed by the following formula:
[0061] R task = -ω t *I
[0062] where I is an indicator function that takes the value 1 when the task is successfully completed and 0 when the task is not completed, and ω t is the negative reward weight set when the task is completed;
[0063] Meanwhile, the negative value R of the driving distance disk is calculated according to the following formula:
[0064] R disk = -ω d *D i,j
[0065] where D i,j represents the driving distance from the current farmland to the target farmland, and ω d is the negative reward weight corresponding to the driving distance, usually a negative number;
[0066] Finally, by integrating these two aspects of calculations, the module will give a complete feedback to the deep learning module, and the calculation result of the reward mechanism is:
[0067] R = R task + R disk
[0068] This reward feedback not only provides a necessary basis for the learning process of the harvester, but also helps the deep learning algorithm to adjust its strategy to optimize the subsequent driving path and harvesting efficiency. Through the coordinated operation of the action space module and the reward mechanism module, the entire system can make wise decisions autonomously in a complex dynamic environment, greatly improving the intelligent level of agricultural mechanization;
[0069] The deep learning module plays a core computational and learning role in the entire agricultural intelligent harvesting system. Its main responsibility is to obtain information such as the current state, action, reward, and next state, and calculate the Q-value and update the loss function of the deep Q-network based on these inputs. This module first obtains the current state vector S, the action A taken, the reward R obtained, and the new state S' reached after executing this action from the environment. These information are crucial for understanding the operation effect of the harvester;
[0070] During the process of updating the Q-value, the deep learning module uses the following Q-learning formula:
[0071]
[0072] Here, Q(S,A) is the current Q-value of the harvester taking action A in state S, and R is the immediate reward obtained by the harvester after executing this action. The learning rate α controls the sensitivity of the Q-value update, and the discount factor γ reflects the importance of future rewards. Theoretically, The maximum Q value for various possible subsequent actions;
[0073] To further optimize the learning process, this module adopts an experience replay mechanism, which samples from the historical state transitions stored in the experience pool. Experience replay not only helps with the independence of samples, reducing the correlation between samples, but also improves the stability and efficiency of training. Specifically, the module randomly selects a batch of samples (S, A, R, S′) from the memory pool, and these samples are calculated according to the following loss function:
[0074]
[0075] In this formula, L(θ) is the loss function used for optimization, N is the size of the sample batch. By minimizing this loss function, the model gradually adjusts the parameters θ of the Q-network, thereby improving its evaluation accuracy for state-action pairs;
[0076] During the training process, to ensure the stability of the model, the deep learning module may also utilize a target network. The parameters θ of the target network are updated periodically to generate more stable Q-value estimates. This design can effectively eliminate the oscillation phenomenon that may occur during the training process;
[0077] In summary, the deep learning module obtains and integrates information on states, actions, rewards, and the next state, optimizes the learning process using the experience replay mechanism, effectively updates the parameters of the Q-network, and then continuously improves the decision-making ability and execution efficiency of the deep learning model in complex agricultural environments. The application of this technical means enables the system to quickly adapt in a dynamically changing environment, achieving better harvesting results and resource utilization efficiency;
[0078] The scheduling and optimization module is a key component in the agricultural intelligent harvesting system. Its main function is to perform intelligent path planning and scheduling decisions based on the current task status of the harvester, the status information of each farmland, and the Q-value strategy output by the trained deep learning model. First, this module continuously receives and integrates real-time data from the environmental modeling module, including the crop maturity of the current farmland, the position information of the harvester, the current list of farmlands to be processed, and the distance and status information between each farmland, to ensure a comprehensive understanding of the entire operation environment;
[0079] During path planning, the scheduling and optimization module relies on the Q-value strategy provided by the trained deep learning model and uses these Q-values to evaluate each possible action and the corresponding path selection. Specifically, the module calculates the expected return of each possible path through the following formula:
[0080]
[0081] Through this formula, the system can evaluate the total rewards that can be obtained from different actions, so as to select the optimal path planning strategy. The path selection not only takes into account the rewards after task completion, but also includes the cost of travel distance, enabling the harvester to achieve a balance between efficiency and effectiveness during operation;
[0082] Once the best path is determined, the scheduling and optimization module is also responsible for sending instructions to the harvester to schedule its movement and execute the required tasks. This process involves dynamically calculating the movement speed, steering angle, and estimated time to reach each farmland on the path, so as to achieve real-time adjustment and optimization. For example, if the system finds that the harvesting efficiency of a certain farmland is lower than expected, or the environmental conditions change, the scheduling module will immediately re-evaluate the path and strategy and make corresponding adjustments using the updated Q-value strategy;
[0083] In addition, the scheduling and optimization module can also consider other factors, such as weather changes, crop types, machine maintenance status, etc. These factors may all affect the harvesting efficiency and decision-making rationality. To ensure the flexibility and responsiveness of the system operation, the module will regularly update its decision-making basis and path planning strategy to ensure that the harvester can adaptively adjust in a dynamically changing operating environment, thereby maximizing the harvesting efficiency and resource utilization rate;
[0084] The scheduling and optimization module realizes efficient path planning and scheduling strategies by combining real-time status information and the output of the deep learning model, providing technical support for the implementation of intelligent agriculture. This enables the harvester to complete tasks more accurately when moving between multiple farmlands, effectively improving the automation and intelligence level of the entire agricultural production process.
[0085] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A crop fixed-point harvesting and delivery path optimization system based on deep reinforcement learning, characterized by: It includes environment modeling module, state space module, action space module, reward mechanism module, deep learning module and scheduling and optimization module; The environmental modeling module is used to obtain the coordinates of the farmland, the position and state of the harvester, the coordinates of the delivery point, and calculate the distance between the farmlands, the initial position of the harvester and the remaining carrying capacity of the harvester, and transmit the calculated values to the state space module; The state space module calculates the harvester state vector for later harvester position state update according to the harvester position, the harvester remaining carrying capacity and the current farmland position coordinates to be transported; The action space module obtains a list of feasible target farmlands and calculates the index of the next farmland, thereby generating a list of feasible actions according to the current state; The reward mechanism module calculates the negative reward for completing the task and the negative value of the driving distance according to the feasible action list, and feeds it back to the deep learning module; The deep learning module obtains the state, action, reward, next state (experience replay), calculates the Q value and updates the loss function, extracts samples from the memory pool using experience replay, updates the parameters of the Q network and trains the model; The scheduling and optimization module performs path planning and schedules the movement and tasks of the harvester according to the current harvesting task status, the status of each farmland, and the Q value strategy output by the trained deep learning model.
2. According to claim 1, a crop fixed-point harvesting and delivery path optimization system based on deep reinforcement learning is characterized in that: The distance calculation formula between the farmlands is as follows: In the formula, D i,j Indicates the distance from farmland 1 to farmland 2, x i ,y i Indicates the location coordinates of farmland 1, x j ,y j Indicates the location coordinates of farmland 2.
3. The crop fixed-point harvesting and delivery path optimization system based on deep reinforcement learning according to claim 2, characterized in that: The calculation formula of the initial position of the harvester is as follows: (x0,y0)=(x farm ,y farm ) In the formula, (x0, y0) represents the initial position of the harvester, x farm Indicates the x-coordinate and y-coordinate of the selected initial farmland. farm Indicates the y coordinate of the selected initial farmland.
4. The crop fixed-point harvesting and delivery path optimization system based on deep reinforcement learning according to claim 3 is characterized in that: The calculation formula of the remaining carrying capacity of the harvester is as follows: C r =C max -C current In the formula, C r represents the remaining carrying capacity of the harvester, C max Indicates the maximum carrying capacity of the harvester, C current Indicates the total weight of crops delivered so far.
5. The crop fixed-point harvesting and delivery path optimization system based on deep reinforcement learning according to claim 4 is characterized in that: The calculation formula of the harvester state vector is as follows: S=(x r ,y r ,C r ) In the formula, S represents the harvester state vector, x r ,y r Indicates the current coordinate position of the harvester, C r Indicates the remaining carrying capacity of the harvester.
6. The crop fixed-point harvesting and delivery path optimization system based on deep reinforcement learning according to claim 5, characterized in that: The index calculation formula of the next farmland is as follows: In the formula, n represents the index of the next target farmland to be selected, and F j represents the jth farmland in the list of feasible target farmlands, Q(S, F j ) represents the current state S and the target farmland F j The Q value between .
7. The crop fixed-point harvesting and delivery path optimization system based on deep reinforcement learning according to claim 6, characterized in that: The negative reward calculation formula for completing the task is as follows: R task =-ω t *I In the formula, R task represents the negative reward for completing the task, ω represents the negative reward weight for completing the task, and I represents the indicator function, which is 1 if the task is completed and 0 if it is not completed.
8. The crop fixed-point harvesting and delivery path optimization system based on deep reinforcement learning according to claim 7, characterized in that: The negative value calculation formula of the driving distance is as follows: R dist =-ω d *D i,j In the formula, R disk Represents a negative value of the travel distance, ω d Denotes the negative reward weight of the driving distance, D i,j Indicates the distance from farmland 1 to farmland 2.
9. The crop fixed-point harvesting and delivery path optimization system based on deep reinforcement learning according to claim 8, characterized in that: The calculation formula of the Q value is as follows: In the formula, Q(S,A) represents the Q value of state S taking action A, α represents the learning rate, R represents the total reward currently obtained, γ represents the discount factor, S ′ Indicates the next state, A ′ Indicates the next possible action.
10. The crop fixed-point harvesting and delivery path optimization system based on deep reinforcement learning according to claim 9, characterized in that: The formula for updating the loss function is as follows: In the formula, L(θ) represents the loss function, N represents the number of samples used to calculate the loss, and R i represents the reward of the i-th sample, Q(S′,a′;θ) represents the Q value in the target network, Q(S,A i ; θ) represents the Q value in the current network.