Visual perception digital twin cotton weeding robot identification obstacle avoidance method
By establishing a digital cotton and weed database and twin virtual scenes, combined with the MADPPO model and the yolov8 weed recognition model, the problem of inaccurate weeding strategies of weeding robots in dynamic time-varying environments was solved, and efficient autonomous obstacle avoidance and accurate weeding were achieved.
Patent Information
- Application Number
- CN202510919102.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-10-14
AI Technical Summary
Existing weeding robots find it difficult to quickly adapt and make correct and accurate weeding strategies in dynamic, unstructured, and time-varying environments, especially due to the complexity and randomness of image information caused by changes in the direction and intensity of sunlight.
By establishing a digital cotton and weed database, building a twin cotton and weed virtual scene, using cloud servers and virtual simulation platforms for real-time data communication, combining machine learning algorithms and agronomic knowledge base, constructing a time-varying digital cotton field scene, using the MADPPO model for multi-agent strategy autonomous obstacle avoidance, combining the yolov8 weed recognition model and channel attention mechanism for image processing, and optimizing the weeding path.
It improves the consistency between virtual scenes and real environments, enhances the weeding accuracy and autonomous obstacle avoidance capabilities of the weeding robot, and enables it to adapt to complex environments with dynamic lighting changes.
Smart Images

Figure CN120779946A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of agricultural robots, and specifically provides a method for identifying and avoiding obstacles of a visually perceived digital twin cotton weeding robot. Background Art
[0002] Weed growth is one of the main obstacles to crop production, seriously reducing crop yields. Weeding robots can effectively replace manual labor to complete the complicated weeding work, playing a role in ensuring production and increasing income in the field of agricultural production.
[0003] How to enable weeding robots to quickly adapt and make correct and accurate weeding strategies in dynamic, unstructured, time-varying environments is a key issue in the autonomous operation of weeding robots. At this stage, related research mainly focuses on the decision-making and path planning of single agents, or the decision-making and path planning of multiple agents in known, time-invariant simple scenarios. In real weeding working conditions, the direction and intensity of sunlight will change regularly over time, which will make the image information in the actual weeding robot working environment contain more randomness and complexity. Summary of the Invention
[0004] The purpose of the present invention is to provide a visual perception digital twin cotton weeding robot recognition and obstacle avoidance method to solve the problems raised in the above background technology.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a visually perceived digital twin cotton weeding robot identification and obstacle avoidance method, the specific steps of which are as follows:
[0006] Step 1: Establish a digital cotton and weed database
[0007] The data acquisition system collects cotton field image information and cotton field environmental information under varying lighting conditions at different times of the day in real-world scenes in the same area. The timestamp of the acquisition and the collected real-time data are transmitted to the cloud server. The cloud server uses machine learning algorithms and a pre-set agronomic knowledge base to perform intelligent matching and analysis to establish a time-varying digital cotton field database.
[0008] Step 2: Build a virtual scene of twin cotton and weeds
[0009] Establish real-time data communication between the cloud server and the local virtual simulation platform. The virtual simulation platform calls the real-time data in the time-varying digital cotton field database through the cloud server to build a virtual time-varying cotton field scene and a virtual time-varying cotton field information environment.
[0010] Step 3: Build a virtual simulation platform
[0011] The virtual simulation platform includes a simulation runtime environment and an interactive control environment. When initializing and creating the simulation runtime environment, the 3D model of the weeding robot is imported, joint coordinate information is bound, and motion constraints are established to create an intelligent agent. Environmental modeling and simulation environment are built to address the weeding robot's visual recognition and autonomous collision avoidance problems in time-varying scenarios. A weed recognition and localization algorithm is designed. The decision-making process of the virtual weeding robot agent's interaction with the virtual scene is described based on MADPPO, solving the virtual weeding robot agent's scheduling problem.
[0012] Step 4: Set parameters and train
[0013] The system sets the operating strategy parameters, performs iterative optimization based on the information collected by the virtual weeding robot agent in the virtual scene, plans an optimal weeding path, issues control instructions to the virtual weeding robot agent based on the calculated path, and synchronously controls the weeding robot in the real world to complete weeding and autonomous obstacle avoidance tasks in a complex real-world environment.
[0014] As a preferred technical solution of the present invention, the data acquisition system described in step 1 includes a high-performance embedded computer, a binocular camera, a 5G communication module and a sensor module. The image information D collected by the data acquisition system for the i-th time img i , weed coordinates D pos i 、Weed Posture D ori i , light intensity D bri i , air temperature D tmp i , air humidity D hum i , soil moisture D wat i Composition of sample data D i ={D img i ,D pos i ,D ori i ,D bri i ,D tmp i ,D hum i ,D wat i ,…}, the data acquisition system generates a UNIX sample timestamp T corresponding to the i-th sampling moment by a high-performance embedded computer i Composition sample information S i ={D i ,Ti}.
[0015] As a preferred technical solution of the present invention, the virtual simulation platform described in step 2 and step 3 includes a global virtual timeline T(t)=t0+kt, where t0 is the start time, k is the time change ratio, and t is the virtual time. The timeline uses UNIX time to describe the time flow of the virtual scene and reflect the time-varying characteristics of the virtual scene. When the virtual simulation platform described in step 2 calls the real-time data in the time-varying digital cotton field database, it first updates the time T on the global timeline. t , and calculate the corresponding expected timestamp T exp Then, the expected timestamp is used as the request information to request data from the time-varying digital cotton field database. The time-varying digital cotton field database retrieves the sample timestamp T with the smallest interval based on the expected timestamp. i , and use the sample timestamp to retrieve the scene information S i Finally, the retrieved scene information is sent back to the virtual simulation platform, and the virtual simulation platform updates the virtual information environment and virtual cotton field scene based on the returned data.
[0016] As a preferred technical solution of the present invention, the virtual timeline is timed by starting the simulation operation. When the scene is updated and the simulation is scheduled at time t, the current time stamp T is first calculated. t =T(t t )=t0+kt t , then pass the timestamp T t As the expected timestamp T exp Load the latest scene information from the server and send a call request to the cloud server through the http protocol. After receiving the call request, the cloud server queries the time-varying database according to the expected timestamp in the request information to obtain the corresponding scene information, and drives the virtual simulation platform to update the data, so that the virtual environment and the collected and archived real environment data are always kept in the same state. After the virtual simulation platform loads the new scene data and completes the rendering, each intelligent agent makes a decision and scheduling. After this decision, a behavior A of the robot arm will be generated. t , the weeding robot arm will be based on A t Generates the corresponding motion, which will produce an estimated running time under the given constraints
[0017] As a preferred technical solution of the present invention, the weed identification and positioning algorithm described in step three adopts a weed identification model based on yolov8, matches positive and negative samples through a dynamic label allocation strategy and an alignment distributor, sets the classification loss of the yolov8 weed identification model to VFL-Loss, and sets the regression loss to CIoU-Loss. Feature extraction is performed through the backbone network (Backbone), and then feature fusion, target detection, and classification are performed through the head network (Head).
[0018] As a preferred technical solution of the present invention, the yolov8 weed recognition model is improved by adopting a multi-branch feature extraction strategy and introducing a channel attention (CA) mechanism. The multi-branch feature extraction strategy sets three branches with different convolution parameters to perform convolution operations on the input image respectively. Each branch uses a different convolution kernel size or step size to extract richer feature information. The calculation process is:
[0019]
[0020] in Indicates the number of channels is c in Input, Represents a convolution operation with the number of channels c1+c2+c3, Indicates the number of input channels is c m , the number of output channels is c n branches, Represents a convolution operation with n channels. After the convolution operation, the feature maps obtained by each branch will be spliced into a high-dimensional feature representation. Then, an additional convolution layer is used to fuse the spliced feature maps to further fuse the features extracted by each branch to obtain more comprehensive and effective image features.
[0021] The channel attention (CA) mechanism enhances the small target features at specific locations in the image through weighted operations. The weight matrix M c ∈R C×1×1 The calculation method is:
[0022]
[0023] Where MLP is a shared multi-layer perceptron, is the global average pooling of the c channel, It is the global maximum pooling of the c channel. This attention mechanism enables the network to capture the detailed information of small-sized targets more effectively.
[0024] As a preferred technical solution of the present invention, the MADPPO described in step 3 performs environmental modeling on the problem of multi-agent strategy autonomous obstacle avoidance and weeding in a time-varying virtual scene. The decision-making process is defined as a five-tuple (t t ,S t ,A t ,R t ,S t+1 ), the five-tuple describes the parameters of the agent's state transition from time t to time t+1, specifically including: t : The moment of this decision-making process on the virtual timeline; S t :The agent is at t t The state at each moment constitutes the state space; A t :The agent is at t t The actions taken at every moment constitute the behavioral space; R t :The agent is at t t The reward value obtained at each moment constitutes the reward function; S t+1 :The agent is at t t+1 The state at the moment, from t0 to t t During the time period, the weeding robot collects the five-tuple of decision-making process and constructs the set (t, S, A, R, S*). Through iterative optimization, it obtains an optimal strategy π(a|s) so that it runs from t0 to t t The cumulative reward R within the time period is maximized.
[0025] As a preferred technical solution of the present invention, the reward function is optimized based on the artificial potential field method, and the cumulative reward function expression is:
[0026]
[0027] Where γ∈(0,1] represents the time decay factor, which is the penalty coefficient of cumulative running time, R t The reward value obtained from the decision result at time t;
[0028] The reward value is calculated as R t =R gui -R obs -R imp -R time , where R gui is the weeding path guidance reward, R obs is the obstacle avoidance penalty term, R imp is the collision avoidance penalty for the weeding robot, R time is the runtime penalty:
[0029] R gui The calculation method is: according to the position of each robot end effector at time t With weed position Pgrs The distance between (x0,y0,z0) Where i = 1, 2, 3…n represents the number of agents and calculates the minimum distance at that moment If during operation If the distance is reduced, appropriate rewards will be given, otherwise appropriate penalties will be imposed. The degree of rewards and penalties is determined by the reward coefficient k1. When it is 0, the end effector completely overlaps with the weeds, giving the maximum reward value k2. The specific process is:
[0030]
[0031]
[0032] R obs The calculation method is as follows: the position and size of the obstacle are described by the circumscribed circle of the nearest obstacle, and the radius of the circumscribed circle is r. A concentric circle with a radius of R is set as the warning area. The horizontal distance between the axes of the robot arm is L, and the distance from the horizontal axis of the robot arm to the obstacle is when If the end effector does not enter the warning area, no penalty will be charged. When the end effector enters the warning area without colliding with the obstacle, a low penalty is imposed, and the degree of penalty is determined by the coefficient k3; When , the end effector will collide with the obstacle, and the maximum penalty will be imposed at this time. The degree of penalty is determined by the coefficient k4. The specific process is:
[0033]
[0034] R imp The calculation method is: detect the motion process of two adjacent weeding robot arms and Is there an intersection? If there is no intersection, no penalty is imposed. If there is an intersection, the maximum penalty is imposed. The degree of penalty is determined by the coefficient k5. The specific process is:
[0035]
[0036] R time The calculation method is: based on the current robot arm movement process O t Estimating running time based on robot arm motion parameters when Less than the expected general running time t tol No penalty is imposed when the running time exceeds t tol The penalty value is calculated based on the timeout period, and the degree of penalty is determined by the coefficient k6. The specific process is:
[0037]
[0038] As a preferred technical solution of the present invention, the MADPPO described in step three is an improved proximal policy optimization (PPO) model. Multiple agents share a virtual scene and collaboratively perform tasks. The multi-agent scheduling problem is solved by learning the optimal strategy. Centralized training and distributed execution are adopted. Each network consists of a policy network and a value network. The policy network is used to generate the control strategy to be adopted in the current state, and the value network is used to evaluate and estimate the reward value of the current control strategy, and improve the control strategy to maximize the reward value.
[0039] As a preferred technical solution of the present invention, the specific process of initializing and creating a simulation operating environment described in step three is as follows: first, the initial cotton field scene collected by the data acquisition system is loaded, the weeding robot STEP model file is imported and motion constraints are added; a perspective camera is used to perform static batch rendering of visible objects and LOD optimization to reduce GPU overhead, the ML-Agents plug-in is used to perform reinforcement learning of intelligent agents in the virtual scene, and the training process parameters are read and set from the JSON configuration file; a multi-process method is used to run the neural network on the GPU using the Pytorch framework in Python for efficient visual task processing, and each intelligent agent of the virtual simulation platform is connected through Socket communication. The interactive control environment described in step three includes a command control area, an operating parameter area, and a scene display area.
[0040] The beneficial effects of the present invention are as follows:
[0041] The present invention establishes a virtual scene of cotton and weeds through the mutual cooperation of a data acquisition system, a cloud server and a virtual simulation platform. By simulating the cotton field and the weeding robot on the simulation platform, the construction of the time-varying digital scene is improved, and the problem of the existing digital scene's incomplete restoration of the time-varying process of the influence of light on the cotton field in the real environment is solved. A rich, comprehensive and complete cotton field time-varying light scene is established, thereby further improving the consistency between the virtual scene and the real environment and the weeding accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 This is a framework diagram of the present invention;
[0043] Figure 2 Constructing a structural diagram for the present invention scenario;
[0044] Figure 3 This is a block diagram of the feature extraction method of the present invention;
[0045] Figure 4 Improved yolov8 structure diagram for the present invention;
[0046] Figure 5 Schematic diagram of the scheduling process of the present invention;
[0047] Figure 6 This is a diagram showing the structure of the reinforcement learning reward function based on the artificial potential field method of the present invention;
[0048] Figure 7 is the constraint parameter diagram of the present invention;
[0049] Figure 8 Update the process diagram for the virtual timeline of the present invention;
[0050] Figure 9 This is a diagram of the MADPPO decision-making process of the present invention;
[0051] Figure 10 This is the interface diagram of the simulation software of the present invention. DETAILED DESCRIPTION
[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0053] like Figures 1 to 10 As shown, the embodiment of the present invention provides a visual perception digital twin cotton weeding robot recognition and obstacle avoidance method, the specific steps are as follows:
[0054] Step 1: Establish a digital cotton and weed database
[0055] The data acquisition system collects cotton field image information and cotton field environmental information under varying lighting conditions at different times of the day in real-world scenes in the same area. The timestamp of the acquisition and the collected real-time data are transmitted to the cloud server. The cloud server uses machine learning algorithms and a pre-set agronomic knowledge base to perform intelligent matching and analysis to establish a time-varying digital cotton field database.
[0056] Step 2: Build a virtual scene of twin cotton and weeds
[0057] Establish real-time data communication between the cloud server and the local virtual simulation platform. The virtual simulation platform calls the real-time data in the time-varying digital cotton field database through the cloud server to build a virtual time-varying cotton field scene and a virtual time-varying cotton field information environment.
[0058] Step 3: Build a virtual simulation platform
[0059] The virtual simulation platform includes a simulation runtime environment and an interactive control environment. When initializing and creating the simulation runtime environment, the 3D model of the weeding robot is imported, joint coordinate information is bound, and motion constraints are established to create an intelligent agent. Environmental modeling and simulation environment are built to address the weeding robot's visual recognition and autonomous collision avoidance problems in time-varying scenarios. A weed recognition and localization algorithm is designed. The decision-making process of the virtual weeding robot agent's interaction with the virtual scene is described based on MADPPO, solving the virtual weeding robot agent's scheduling problem.
[0060] Step 4: Set parameters and train
[0061] The system sets the operating strategy parameters, performs iterative optimization based on the information collected by the virtual weeding robot agent in the virtual scene, plans an optimal weeding path, issues control instructions to the virtual weeding robot agent based on the calculated path, and synchronously controls the weeding robot in the real world to complete weeding and autonomous obstacle avoidance tasks in a complex real-world environment.
[0062] Cotton field image information and cotton field environmental information include information such as cotton field topography, weed appearance, weed species, weed growth status, weed coordinates and weed posture. The data acquisition system adopts a fixed-interval acquisition strategy to periodically collect data on the same scene at different times of the day (5:00 to 23:00, with an interval of 15 minutes). Through long-term continuous data acquisition, a time-varying database with time series characteristics is constructed to more accurately reflect the changing scene information of cotton fields at different times in the real environment.
[0063] The data acquisition system in step 1 includes a high-performance embedded computer, a binocular camera, a 5G communication module and a sensor module. The image information D collected by the data acquisition system for the i-th time img i , weed coordinates D pos i 、Weed Posture D ori i , light intensity D bri i , air temperature D tmp i , air humidity D hum i , soil moisture D wat i Composition of sample data D i ={D img i ,D pos i ,D ori i ,D bri i ,D tmpi ,D hum i ,D wat i ,…}, the data acquisition system generates the UNIX sample timestamp T corresponding to the i-th sampling moment by a high-performance embedded computer i Composition sample information S i ={D i ,T i}.
[0064] The sensor module includes a light intensity sensor module, an air temperature and humidity sensor module, a soil moisture sensor module, a soil pH sensor module and a carbon dioxide concentration sensor module; a high-performance embedded computer converts the sample information S i The universally unique identifier (UUID) of the corresponding scene and the sample sequence number are formatted in Json text format to facilitate server archiving of data. The image information Dimgi is encoded in Base64 format, the UUID is represented in string form, and other information is represented using floating-point numbers. The high-performance embedded computer then sends the data packet to the server through the network via the 5G communication module. To avoid data loss due to network failure, all data packets are backed up in local storage space and can be resent after network communication is restored. In addition, for each real scene, the data acquisition system collects a total of 2,000-5,000 images of different growth stages, angles, and lighting. After annotation, they are standardized, including size normalization and pixel size adjustment to 640x640. Data enhancement is performed using methods such as AutoAugment, RandAugment, and Mixup, and the weed dataset is expanded to 6,000-8,000 images to construct.
[0065] Among them, the virtual simulation platform in step 2 and step 3 includes a global virtual timeline T(t)=t0+kt, where t0 is the starting time, k is the time change ratio, and t is the virtual time. The timeline uses UNIX time, describes the time flow of the virtual scene, and reflects the time-varying characteristics of the virtual scene. When the virtual simulation platform in step 2 calls the real-time data in the time-varying digital cotton field database, it first updates the time T on the global timeline. t , and calculate the corresponding expected timestamp T exp , then the expected timestamp is used as the request information to request data from the time-varying digital cotton field database, and the time-varying digital cotton field database retrieves the sample timestamp T with the smallest interval based on the expected timestamp. i , and use the sample timestamp to retrieve the scene information S i Finally, the retrieved scene information is sent back to the virtual simulation platform, and the virtual simulation platform updates the virtual information environment and virtual cotton field scene based on the returned data.
[0066] Before the virtual simulation platform starts running the simulation, it performs time alignment, selects the scene information to be loaded from the time-varying database, and uses the sample timestamp T of the first sample information in the scene. 0 As the starting moment of the global timeline, the time starting point of the virtual simulation platform is synchronized with the time starting point of the time-varying scene.
[0067] Among them, the virtual timeline is timed by the start of the simulation run. When the scene is updated and the simulation is scheduled at time t, the current time stamp T is first calculated. t =T(t t )=t0+kt t , then pass the timestamp T t As the expected timestamp T exp Load the latest scene information from the server and send a call request to the cloud server through the http protocol. After receiving the call request, the cloud server queries the time-varying database according to the expected timestamp in the request information to obtain the corresponding scene information, and drives the virtual simulation platform to update the data, so that the virtual environment and the collected and archived real environment data are always kept in the same state. After the virtual simulation platform loads the new scene data and completes the rendering, each intelligent agent makes a decision and scheduling. After this decision, a behavior A of the robot arm will be generated. t , the weeding robot arm will be based on A t Generates the corresponding motion, which will produce an estimated running time under the given constraints
[0068] Time will be accumulated on the virtual timeline and made t t Time transfer is t t+1 At this moment, the transfer process is:
[0069]
[0070] Among them, the weed recognition and positioning algorithm in step three adopts a weed recognition model based on yolov8, matches positive and negative samples through a dynamic label allocation strategy and an alignment allocator, sets the classification loss of the yolov8 weed recognition model to VFL-Loss, and the regression loss to CIoU-Loss. Feature extraction is performed through the backbone network (Backbone), and then feature fusion, target detection and classification are performed through the head network (Head).
[0071] In the virtual simulation platform, when the virtual weeding robot collects images in the virtual scene through the camera, it calls the method with the scene data as described above, requests image data that matches the current moment of the virtual timeline from the time-varying database, and uses the scene data such as the weed position and weed posture that are synchronously transmitted back with the above image data to update the virtual scene environment around the virtual weeding robot.
[0072] Among them, the yolov8 weed recognition model is improved by adopting a multi-branch feature extraction strategy and introducing a channel attention (CA) mechanism. The multi-branch feature extraction strategy sets three branches with different convolution parameters to perform convolution operations on the input image respectively. Each branch uses a different convolution kernel size or step size to extract richer feature information. The calculation process is:
[0073]
[0074]
[0075] in Indicates the number of channels is c in Input, Represents a convolution operation with the number of channels c1+c2+c3, Indicates the number of input channels is c m , the number of output channels is c n branches, Represents a convolution operation with n channels. After the convolution operation, the feature maps obtained by each branch will be spliced into a high-dimensional feature representation. Then, an additional convolution layer is used to fuse the spliced feature maps to further fuse the features extracted by each branch to obtain more comprehensive and effective image features.
[0076] The channel attention (CA) mechanism enhances the small target features at specific locations in the image through weighted operations. The weight matrix M c ∈R C×1×1 The calculation method is:
[0077]
[0078] Where MLP is a shared multi-layer perceptron, is the global average pooling of the c channel, It is the global maximum pooling of the c channel. This attention mechanism enables the network to capture the detailed information of small-sized targets more effectively.
[0079] The images of weeds in the picture are small, and the features such as the edge texture of weeds during the weeding stage are similar to those of cotton. Therefore, the yolov8 weed recognition model is improved by adopting a multi-branch feature extraction strategy and introducing a channel attention (CA) mechanism. In addition, through the channel attention (CA) mechanism, the network can dynamically adjust the weights of rows and columns at different positions, so as to focus on important areas in the image, especially those containing small targets. This mechanism automatically assigns different attention weights to different areas according to the feature importance of each spatial position, thereby improving the overall image processing accuracy and the ability to recognize small targets.
[0080] Among them, the MADPPO in step 3 models the environment for the multi-agent strategy autonomous obstacle avoidance and weeding strategy problem in the time-varying virtual scene, and the decision process is defined as a five-tuple (t t ,S t ,A t ,R t ,S t+1 ), the five-tuple describes the parameters of the agent's state transition from time t to time t+1, specifically including: t : The moment of this decision-making process on the virtual timeline; S t :The agent is at t t The state at each moment constitutes the state space; A t :The agent is at t t The actions taken at every moment constitute the behavioral space; R t :The agent is at t t The reward value obtained at each moment constitutes the reward function; S t+1 :The agent is at t t+1 The state at the moment, from t0 to t t During the time period, the weeding robot collects the five-tuple of decision-making process and constructs the set (t, S, A, R, S*). Through iterative optimization, it obtains an optimal strategy π(a|s) so that it runs from t0 to t t The cumulative reward R within the time period is maximized.
[0081] Based on the Markov decision process (MADPPO) to describe the decision-making process of the virtual weeding robot's interaction with the virtual scene, after environmental modeling of the multi-agent strategy autonomous collision avoidance weeding strategy problem in the time-varying virtual scene, the virtual weeding robot scheduling problem can be effectively solved.
[0082] Among them, the reward function is optimized based on the artificial potential field method, and the cumulative reward function expression is:
[0083]
[0084] Where γ∈(0,1] represents the time decay factor, which is the penalty coefficient of cumulative running time, Rt The reward value obtained from the decision result at time t;
[0085] The reward value is calculated as R t =R gui -R obs -R imp -R time , where R gui is the weeding path guidance reward, R obs is the obstacle avoidance penalty term, R imp is the collision avoidance penalty for the weeding robot, R time is the runtime penalty:
[0086] R gui The calculation method is: according to the position of each robot end effector at time t With weed position P grs The distance between (x0,y0,z0) Where i = 1, 2, 3…n represents the number of agents and calculates the minimum distance at that moment If during operation If the distance is reduced, appropriate rewards will be given, otherwise appropriate penalties will be imposed. The degree of rewards and penalties is determined by the reward coefficient k1. When it is 0, the end effector completely overlaps with the weeds, giving the maximum reward value k2. The specific process is:
[0087]
[0088]
[0089] R obs The calculation method is as follows: the position and size of the obstacle are described by the circumscribed circle of the nearest obstacle, and the radius of the circumscribed circle is r. A concentric circle with a radius of R is set as the warning area. The horizontal distance between the axes of the robot arm is L, and the distance from the horizontal axis of the robot arm to the obstacle is when If the end effector does not enter the warning area, no penalty will be charged. When the end effector enters the warning area without colliding with the obstacle, a low penalty is imposed, and the degree of penalty is determined by the coefficient k3; When , the end effector will collide with the obstacle, and the maximum penalty will be imposed at this time. The degree of penalty is determined by the coefficient k4. The specific process is:
[0090]
[0091] R imp The calculation method is: detect the motion process of two adjacent weeding robot arms and Is there an intersection? If there is no intersection, no penalty is imposed. If there is an intersection, the maximum penalty is imposed. The degree of penalty is determined by the coefficient k5. The specific process is:
[0092]
[0093] R time The calculation method is: based on the current robot arm movement process O t Estimating running time based on robot arm motion parameters when Less than the expected general running time t tol No penalty is imposed when the running time exceeds t tol The penalty value is calculated based on the timeout period, and the degree of penalty is determined by the coefficient k6. The specific process is:
[0094]
[0095] Each reward coefficient and penalty coefficient are given independently. Generally, in order to make the reward value positive, the reward coefficient is set greater than the penalty coefficient.
[0096] Among them, MADPPO in step three is an improved proximal policy optimization (PPO) model. Multiple agents share a virtual scene and collaboratively perform tasks. It solves the multi-agent scheduling problem by learning the optimal strategy. It adopts centralized training and distributed execution. Each network consists of two parts: a policy network and a value network. The policy network is used to generate the control strategy to be adopted in the current state, and the value network is used to evaluate and estimate the reward value of the current control strategy, and improve the control strategy to maximize the reward value.
[0097] First, the local network running on each intelligent agent interacts with the surrounding environment independently and simultaneously to collect the required environmental data. This data is combined with the current operating status of the intelligent agent, including environmental image recognition information, end effector position, obstacle position, target picking point position, and the closest distance between the end effector and the target picking point, the distance between the obstacle and each rotation axis of the robot, and the distance between each rotation axis of the robot, etc., and is input into the policy network. The control strategy under the current state is obtained according to the normal distribution sampling, and is updated and optimized through the back propagation strategy, and the actions of each rotation axis of the robot are output. The value network will perform value evaluation based on the collected environmental information, and continuously update the parameters through back propagation to maximize the reward value.
[0098] The specific process of initializing and creating the simulation running environment in step three is as follows: first, load the initial cotton field scene collected by the data acquisition system, import the weeding robot STEP model file and add motion constraints; use the perspective camera to perform static batch rendering of visible objects and perform LOD optimization to reduce GPU overhead, use the ML-Agents plug-in to perform reinforcement learning of intelligent agents in the virtual scene, and read and set training process parameters from the JSON configuration file; use a multi-process method to run the neural network on the GPU using the Pytorch framework in Python for efficient visual task processing, and connect to various intelligent agents in the virtual simulation platform through Socket communication. The interactive control environment in step three includes a command control area, an operating parameter area, and a scene display area.
[0099] During training, TensorBoard can be used to view the training progress. After the training is completed, the system will export the parameters and results of each stage of the simulation as a Pytorch learning strategy model and store them in a pre-set working directory; the command control area is designed with start, pause, stop, and reset buttons for users to control the operation of the simulation process; the operation parameter area can set parameters such as obstacle density and position, number of agents, and virtual timeline starting time, and can also display information such as the working status, behavioral decisions, and coordinate positions of each agent; the scene display area uses a perspective camera whose viewing angle is controlled by the mouse for uniform rendering, and then adds the agent running path to the rendered image and displays it in the scene display area, so that users can observe the planned path of each agent under different set working conditions, providing a reference for the operation of the real weeding robot.
[0100] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0101] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A visually aware digital twin cotton weeding robot obstacle avoidance method, characterized in that: The specific steps are as follows: Step 1: Establish a digital cotton and weed database The data acquisition system collects cotton field image information and cotton field environmental information under varying lighting conditions at different times of the day in real-world scenes in the same area. The timestamp of the acquisition and the collected real-time data are transmitted to the cloud server. The cloud server uses machine learning algorithms and a pre-set agronomic knowledge base to perform intelligent matching and analysis to establish a time-varying digital cotton field database. Step 2: Build a virtual scene of twin cotton and weeds Establish real-time data communication between the cloud server and the local virtual simulation platform. The virtual simulation platform calls the real-time data in the time-varying digital cotton field database through the cloud server to build a virtual time-varying cotton field scene and a virtual time-varying cotton field information environment. Step 3: Build a virtual simulation platform The virtual simulation platform includes a simulation runtime environment and an interactive control environment. When initializing and creating the simulation runtime environment, the 3D model of the weeding robot is imported, joint coordinate information is bound, and motion constraints are established to create an intelligent agent. Environmental modeling and simulation environment are built to address the weeding robot's visual recognition and autonomous collision avoidance problems in time-varying scenarios. A weed recognition and localization algorithm is designed. The decision-making process of the virtual weeding robot agent's interaction with the virtual scene is described based on MADPPO, solving the virtual weeding robot agent's scheduling problem. Step 4: Set parameters and train The system sets the operating strategy parameters, performs iterative optimization based on the information collected by the virtual weeding robot agent in the virtual scene, plans an optimal weeding path, issues control instructions to the virtual weeding robot agent based on the calculated path, and synchronously controls the weeding robot in the real world to complete weeding and autonomous obstacle avoidance tasks in a complex real-world environment.
2. The visually-aware digital twin cotton weeding robot obstacle avoidance method according to claim 1, characterized in that: The data acquisition system described in step 1 includes a high-performance embedded computer, a binocular camera, a 5G communication module and a sensor module. The image information D collected by the data acquisition system for the i-th time img i , weed coordinates D pos i 、Weed Posture D ori i , light intensity D bri i , air temperature D tmp i , air humidity D hum i , soil moisture D wat i Composition of sample data D i ={D img i ,D pos i ,D ori i ,D bri i ,D tmp i ,D hum i ,D wat i ,…}, the data acquisition system generates a UNIX sample timestamp T corresponding to the i-th sampling moment by a high-performance embedded computer i Composition sample information S i ={D i ,T i }.
3. The visually-aware digital twin cotton weeding robot obstacle avoidance method according to claim 1, characterized in that: The virtual simulation platform described in step 2 and step 3 includes a global virtual timeline T(t)=t0+kt, where t0 is the start time, k is the time change ratio, and t is the virtual time. The timeline uses UNIX time to describe the time flow of the virtual scene and reflect the time-varying characteristics of the virtual scene. When the virtual simulation platform described in step 2 calls the real-time data in the time-varying digital cotton field database, it first updates the time T on the global timeline. t , and calculate the corresponding expected timestamp T exp Then, the expected timestamp is used as the request information to request data from the time-varying digital cotton field database. The time-varying digital cotton field database retrieves the sample timestamp T with the smallest interval based on the expected timestamp. i , and use the sample timestamp to retrieve the scene information S i Finally, the retrieved scene information is sent back to the virtual simulation platform, and the virtual simulation platform updates the virtual information environment and virtual cotton field scene based on the returned data.
4. The visually-aware digital twin cotton weeding robot obstacle avoidance method according to claim 3 is characterized by: The virtual timeline is timed by the start of the simulation run. When the scene is updated and the simulation is scheduled at time t, the current time stamp T is first calculated. t =T(t t )=t0+kt t , then pass the timestamp T t As the expected timestamp T exp Load the latest scene information from the server and send a call request to the cloud server through the http protocol. After receiving the call request, the cloud server queries the time-varying database according to the expected timestamp in the request information to obtain the corresponding scene information, and drives the virtual simulation platform to update the data, so that the virtual environment and the collected and archived real environment data are always kept in the same state. After the virtual simulation platform loads the new scene data and completes the rendering, each intelligent agent makes a decision and scheduling. After this decision, a behavior A of the robot arm will be generated. t , the weeding robot arm will be based on A t Generates the corresponding motion, which will produce an estimated running time under the given constraints 5. The visually-aware digital twin cotton weeding robot obstacle avoidance method according to claim 1, characterized in that: The weed recognition and localization algorithm described in step 3 adopts a weed recognition model based on yolov8. It matches positive and negative samples through a dynamic label allocation strategy and an alignment allocator. The classification loss of the yolov8 weed recognition model is set to VFL-Loss, and the regression loss is set to CIoU-Loss. Feature extraction is performed through the backbone network (Backbone), and then feature fusion, target detection and classification are performed through the head network (Head).
6. The visually-aware digital twin cotton weeding robot obstacle avoidance method according to claim 5, characterized in that: The yolov8 weed recognition model is improved by adopting a multi-branch feature extraction strategy and introducing a channel attention (CA) mechanism. The multi-branch feature extraction strategy sets three branches with different convolution parameters to perform convolution operations on the input image respectively. Each branch uses a different convolution kernel size or step size to extract richer feature information. The calculation process is: in Indicates the number of channels is c in Input, Represents a convolution operation with the number of channels c1+c2+c3, Indicates the number of input channels is c m , the number of output channels is c n branches, Represents a convolution operation with n channels. After the convolution operation, the feature maps obtained by each branch will be spliced into a high-dimensional feature representation. Then, an additional convolution layer is used to fuse the spliced feature maps to further fuse the features extracted by each branch to obtain more comprehensive and effective image features. The channel attention (CA) mechanism enhances the small target features at specific locations in the image through weighted operations. The weight matrix M c ∈R C×1×1 The calculation method is: Where MLP is a shared multi-layer perceptron, is the global average pooling of the c channel, It is the global maximum pooling of the c channel. This attention mechanism enables the network to capture the detailed information of small-sized targets more effectively.
7. The visually-aware digital twin cotton weeding robot obstacle avoidance method according to claim 1, characterized in that: The MADPPO described in step 3 models the environment for the multi-agent strategy of autonomous obstacle avoidance and weeding in a time-varying virtual scene. The decision process is defined as a five-tuple (t t ,S t ,A t ,R t ,S t+1 ), the five-tuple describes the parameters of the agent's state transition from time t to time t+1, specifically including: t : The moment of this decision-making process on the virtual timeline; S t :The agent is at t t The state at each moment constitutes the state space; A t :The agent is at t t The actions taken at every moment constitute the behavioral space; R t :The agent is at t t The reward value obtained at each moment constitutes the reward function; S t+1 :The agent is at t t+1 The state at the moment, from t0 to t t During the time period, the weeding robot collects the five-tuple of decision-making process and constructs the set (t, S, A, R, S*). Through iterative optimization, it obtains an optimal strategy π(a|s) so that it runs from t0 to t t The cumulative reward R within the time period is maximized.
8. The visually-aware digital twin cotton weeding robot obstacle avoidance method according to claim 7, characterized in that: The reward function is optimized based on the artificial potential field method, and the cumulative reward function expression is: Where γ∈(0,1] represents the time decay factor, which is the penalty coefficient of cumulative running time, R t The reward value obtained from the decision result at time t; The reward value is calculated as R t =R gui -R obs -R imp -R time , where R gui is the weeding path guidance reward, R obs is the obstacle avoidance penalty term, R imp is the collision avoidance penalty for the weeding robot, R time is the runtime penalty: R gui The calculation method is: according to the position of each robot end effector at time t With weed position P grs The distance between (x0,y0,z0) Where i = 1, 2, 3…n represents the number of agents and calculates the minimum distance at that moment If during operation If the distance is reduced, appropriate rewards will be given, otherwise appropriate penalties will be imposed. The degree of rewards and penalties is determined by the reward coefficient k1. When it is 0, the end effector completely overlaps with the weeds, giving the maximum reward value k2. The specific process is: R obs The calculation method is as follows: the position and size of the obstacle are described by the circumscribed circle of the nearest obstacle, and the radius of the circumscribed circle is r. A concentric circle with a radius of R is set as the warning area. The horizontal distance between the axes of the robot arm is L, and the distance from the horizontal axis of the robot arm to the obstacle is when If the end effector does not enter the warning area, no penalty will be charged. When the end effector enters the warning area without colliding with the obstacle, a low penalty is imposed, and the degree of penalty is determined by the coefficient k3; When , the end effector will collide with the obstacle, and the maximum penalty will be imposed at this time. The degree of penalty is determined by the coefficient k4. The specific process is: R imp The calculation method is: detect the motion process of two adjacent weeding robot arms and Is there an intersection? If there is no intersection, no penalty is imposed. If there is an intersection, the maximum penalty is imposed. The degree of penalty is determined by the coefficient k5. The specific process is: R time The calculation method is: based on the current robot arm movement process O t Estimating running time based on robot arm motion parameters when Less than the expected general running time t tol No penalty is imposed when the running time exceeds t tol The penalty value is calculated based on the timeout period, and the degree of penalty is determined by the coefficient k6. The specific process is:
9. The visual perception digital twin cotton weeding robot obstacle avoidance method according to claim 1, characterized in that: The MADPPO described in step three is an improved proximal policy optimization (PPO) model. Multiple agents share a virtual scene and collaboratively perform tasks. It solves the multi-agent scheduling problem by learning the optimal strategy. It adopts centralized training and distributed execution. Each network consists of two parts: a policy network and a value network. The policy network is used to generate the control strategy to be adopted in the current state, and the value network is used to evaluate and estimate the reward value of the current control strategy, and improve the control strategy to maximize the reward value.
10. The visual perception digital twin cotton weeding robot recognition and obstacle avoidance method according to claim 1, characterized in that: The specific process for initializing and creating the simulation runtime environment described in step 3 is as follows: first, load the initial cotton field scene collected by the data acquisition system, import the weeding robot STEP model file and add motion constraints; use a perspective camera to perform static batch rendering of visible objects and perform LOD optimization to reduce GPU overhead; use the ML-Agents plug-in to perform reinforcement learning of the intelligent agent in the virtual scene, and read and set the training process parameters from the JSON configuration file; Using a multi-process approach, the Pytorch framework is used in Python to run the neural network on the GPU for efficient visual task processing, and each intelligent agent of the virtual simulation platform is connected through socket communication. The interactive control environment described in step 3 includes a command control area, an operating parameter area, and a scene display area.
Citation Information
Cited By
Unmanned aerial vehicle indoor three-dimensional reconstruction autonomous acquisition method and system based on reinforcement learning
CN120997407A
Field robot adaptive navigation path planning system based on convolutional neural network
CN121655538A
Modularized construction dynamic hoisting scheduling decision-making system
CN122264339A