Discharging and feeding robot and path dynamic optimization method thereof
By configuring independent algorithms and deep reinforcement learning models for each unloading and feeding robot, combining image recognition and real-time communication to optimize the unloading path, the problems of low efficiency and path conflict in traditional unloading operations are solved, and efficient and low-cost multi-robot collaborative operations are achieved.
Patent Information
- Application Number
- CN202510771827.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-11
AI Technical Summary
Traditional unloading operations rely on low manual operation efficiency and are susceptible to fatigue. The central controller mode has high algorithm complexity and poor scalability in multi-robot systems, and path planning is difficult to adapt to complex environments, so path conflicts are prone to occur between robots.
Each unloading and feeding robot is configured with an independent algorithm, optimized the path through image recognition and deep reinforcement learning model, communicate and negotiate with adjacent robots in real time, select the optimal unloading path and avoid collisions.
Improves the efficiency and fluency of unloading operations, reduces maintenance costs, reduces communication bandwidth requirements and energy consumption, adapts to complex environments and reduces robot path conflicts.
Smart Images

Figure CN120269581A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent robots, and more specifically, to a discharging and feeding robot and a method for dynamically optimizing its path. Background Art
[0002] Traditional discharging operations mostly rely on manual labor or unified scheduling by a central controller. Manual operation has low efficiency and is easily affected by fatigue and human errors; while in the central controller mode, when the number of robots increases, the algorithm complexity and computational volume increase exponentially, with poor scalability and high maintenance costs.
[0003] In terms of path planning, traditional methods often rely on simple environmental models and are difficult to adapt to the complex and changeable environment of the discharging area. At the same time, when dealing with multi-robot collaborative operations, these algorithms lack sufficient consideration of the mutual influence between robots, easily causing path conflicts between robots and further reducing the fluency of operations. Therefore, there are deficiencies in the existing technology. Summary of the Invention
[0004] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a discharging and feeding robot and a method for dynamically optimizing its path. By configuring an independently running algorithm for each robot, the bottleneck of traditional centralized control is overcome, and through communication and negotiation with adjacent robots, the optimal discharging path is selected.
[0005] To achieve the above purpose, the present invention provides the following technical solutions: The present invention provides a method for dynamically optimizing the path of a discharging and feeding robot. The discharging and feeding robots are multiple and located in the same storage area. The storage area includes multiple discharging areas, and each discharging area corresponds to multiple nodes. The method for dynamically optimizing the path of the discharging and feeding robot is executed by one of the discharging and feeding robots. The method for dynamically optimizing the path of the discharging and feeding robot includes: Obtain images of multiple discharging areas within a first preset range, and select a first area and a first node according to the images; Determine whether there are adjacent robots. If so, obtain a second area and a second node, and update the first area and the first node according to the second area and the second node to obtain the current discharging area and the current discharging node; the second area and the second node are the first area and the first node of the adjacent robot, and the adjacent robot is a discharging and feeding robot located within a second preset range; Move to the current discharging node and complete the discharging work on the current discharging area according to the action sequence, and the action sequence is obtained through a deep reinforcement learning model.
[0006] As a further improvement of the present invention, the selecting a first area and a first node according to the images includes: Determine the depth of the material in the multiple discharging areas according to the image; Obtain the driving distances of the multiple discharging areas to themselves; According to the driving distances, the depth of the material, and a preset evaluation index, obtain the evaluation value corresponding to each discharging area; Take the discharging area with the highest evaluation value as the first area, and take the node closest to its own position among the nodes corresponding to the first area as the first node.
[0007] As a further improvement of the present invention, the updating the first area and the first node according to the second area and the second node to obtain the current discharging area and the current discharging node includes: Judge whether the second area and the second node are the same as the first area and the first node; If the second area is the same as the first area, update the first area and the first node according to the evaluation value corresponding to each discharging area to obtain the current discharging area and the current discharging node; If the second area is different from the first area and the second node is the same as the first node, take the first area as the current discharging area and update the first node to obtain the current discharging node.
[0008] As a further improvement of the present invention, the updating the first area and the first node according to the evaluation value corresponding to each discharging area to obtain the current discharging area and the current discharging node includes: Obtain the evaluation value corresponding to the second area; If the evaluation value corresponding to the second area is greater than the evaluation value corresponding to the first area, update the first area and the first node to obtain the current discharging area and the current discharging node.
[0009] As a further improvement of the present invention, the updating the first area and the first node to obtain the current discharging area and the current discharging node includes: Exclude the evaluation value corresponding to the first area from the evaluation values corresponding to each discharging area to obtain the remaining evaluation values; Take the discharging area corresponding to the highest evaluation value among the remaining evaluation values as the current discharging area, and take the node closest to its own position among the nodes corresponding to the current discharging area as the current discharging node.
[0010] As a further improvement of the present invention, the action sequence is obtained through a deep reinforcement learning model, including: Obtain the pose information after moving to the current discharging node, where the pose information includes the position coordinates of the robotic arm, the pose angle of the robotic arm, and the position coordinates of the material to be discharged; Perform iterative operations, where the iterative operations include: inputting the pose information into the deep reinforcement learning model to obtain the quality values of each action in a preset action space, selecting the action corresponding to the highest quality value, putting it into an action sequence, updating the pose information until the action of putting down the material is put into the action sequence, and outputting the current action sequence.
[0011] As a further improvement of the present invention, the deep reinforcement learning model is obtained through training and includes: Initializing the parameters of the deep reinforcement learning model; Generating multiple samples in sequence based on environmental interaction and storing each sample in an experience replay pool in sequence; Selecting multiple samples from the experience replay pool and updating the parameters of the deep reinforcement learning model according to the multiple samples until the deep reinforcement learning model converges.
[0012] As a further improvement of the present invention, the selecting multiple samples from the experience replay pool includes: Calculating the current quality value of each sample according to a preset main quality network and calculating the target quality value of each sample according to a preset target quality network; Calculating the temporal difference error of each sample according to the current quality value and the target quality value; Determining the priority of each sample according to the temporal difference error; Selecting multiple samples from the experience replay pool according to the priority.
[0013] As a further improvement of the present invention, the updating the parameters of the deep reinforcement learning model according to the multiple samples includes: Obtaining the current quality value and the target quality value of each sample; Calculating a loss value according to the current quality value, the target quality value, and a preset loss function; Updating the parameters of the deep reinforcement learning model according to the loss value and the backpropagation algorithm.
[0014] The present invention provides a discharging and feeding robot, which is applied to the above-mentioned path dynamic optimization method for a discharging and feeding robot. The discharging and feeding robot includes: a binocular camera, an image processing module, a computing module, a communication module, and a robotic arm; The binocular camera is used to obtain images of multiple discharging areas within a first preset range; The image processing module is used to select a first area and a first node according to the images; The communication module is used to obtain a second area and a second node; The calculation module is used to update the first area and the first node according to the second area and the second node, so as to obtain the current discharging area and the current discharging node; The communication module is used to obtain the second area and the second node; The robotic arm is used to complete the discharging work on the current discharging area according to the action sequence.
[0015] First, the present invention selects the first area and the first node through image recognition technology, and determines whether there are adjacent robots. Then, it interacts with adjacent robots in real time through a communication protocol to avoid collisions. Moreover, the method provided by the present invention can be executed independently by each robot, enabling each robot to make independent decisions according to the information interacted in real time, thus overcoming the problem of high algorithm complexity in traditional centralized control. Description of the Drawings
[0016] Figure 1 It is a schematic diagram of the steps of a method for dynamically optimizing the path of a discharging and feeding robot according to the present invention; Figure 2 It is a schematic diagram of a storage area; Figure 3 It is a schematic diagram of a first preset range; Figure 4 It is a schematic diagram of a second preset range; Figure 5 It is a schematic diagram of a first node; Figure 6 It is a schematic diagram of the scenario when the first area needs to be updated; Figure 7 It is a schematic diagram of the scenario when the first node needs to be updated.
[0017] Reference Signs: 1, discharging area; 2, conveyor belt; 3, moving passage; 4, node; 5, projection area; 6, first preset range; 7, second preset range. Detailed Embodiments
[0018] The technical solution of the present invention will be described in detail below through the drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention.
[0019] Among them, the same components are denoted by the same reference signs. It should be noted that the terms "front", "rear", "left", "right", "up" and "down" used in the following description refer to the directions in the drawings, and the terms "bottom surface" and "top surface", "inner" and "outer" refer to the directions towards or away from the geometric center of a specific component, respectively.
[0020] The term "and / or" in the following text is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, the character " / " generally indicates that the associated objects before and after are in an "or" relationship.
[0021] As Figure 1 shown, the present invention provides a method for dynamically optimizing the path of a discharging and feeding robot, including: Obtaining images of multiple discharging areas within a first preset range, and selecting a first area and a first node according to the images; Judging whether there is an adjacent robot. If so, obtaining a second area and a second node, and updating the first area and the first node according to the second area and the second node to obtain the current discharging area and the current discharging node; the second area and the second node are the first area and the first node of the adjacent robot, and the adjacent robot is a discharging and feeding robot located within a second preset range; Moving to the current discharging node and completing the discharging work for the current discharging area according to the action sequence, where the action sequence is obtained through a deep reinforcement learning model.
[0022] Among them, there are multiple discharging and feeding robots located in the same storage area. The method for dynamically optimizing the path of the discharging and feeding robot is executed separately by each discharging and feeding robot, and each discharging area is only unloaded by one robot, and each discharging node parks one robot.
[0023] As Figure 2 shown, the storage area includes multiple square discharging areas 1 arranged in a grid pattern. Between the discharging areas are conveyors 2 for placing the unloaded materials. Above the conveyors is the movement path 3 of the robot, and the structure of the movement path is the same as that of the conveyor. Figure 2 The virtual wireframe in [[ ]] is obtained by projecting upward from the discharging area and is called the projection area, which can be used to illustrate the corresponding relationship between the conveyor and the movement path. Each discharging area corresponds to multiple nodes 4, and the nodes are located on the movement path. If the nodes are projected onto the conveyor, the projected nodes are located on the midlines of each side of the discharging area. Figure 2 Only the nodes corresponding to two discharging areas are shown in [[ ]], and it can be seen that there is an overlap between the nodes of adjacent two discharging areas.
[0024] As Figure 3 shown, the first preset range of each discharging and feeding robot can be set according to actual working needs, usually set as a circular area centered on itself (node A), and the radius of the circular area is 2 times the distance between two diagonally adjacent nodes. At this time, the multiple discharging areas within the first preset range are the discharging areas corresponding to the first to eighth projection areas.
[0025] As shown Figure 4 in the figure, assume that the discharging area corresponding to the seventh projection area is the first area, and the second preset range is a circular range centered on the first area with the distance between itself and the first area as the radius. Here, the distance between itself and the first area is specifically the distance between itself and the center of the projection area corresponding to the first area, that is, the distance between node A in the figure and the center of the seventh projection area. Robots located within the second preset range are more likely to select the same first area and first node as themselves, which may lead to collision accidents. Therefore, they are used as adjacent robots for subsequent analysis.
[0026] If there are no adjacent robots within the second preset range, then the first area is used as the current discharging area, and the first node is used as the current discharging node. Then, the path planning algorithm is used to determine the path to the current discharging node, thereby completing the discharging work.
[0027] In this embodiment, an algorithm for independent operation is configured for each robot to overcome the bottleneck of traditional centralized control. Through communication with adjacent robots, the current discharging area and the current discharging node are updated and determined, and finally the discharging work is completed.
[0028] Furthermore, the embodiment of the present application provides a step of selecting the first area and the first node according to an image, including: Determine the depth of the materials in multiple discharging areas according to the image; Obtain the driving distance between multiple discharging areas and itself; According to the driving distance, the depth of the materials, and a preset evaluation index, obtain the evaluation value corresponding to each discharging area;
[0029] Use the discharging area with the highest evaluation value as the first area, and use the node closest to its own position among the nodes corresponding to the first area as the first node.
[0030] Among them, the above image is collected by a binocular camera. Then, according to the principle of similar triangles and the geometric model of the camera, the depth of the materials in the discharging area can be determined by using the parallax. Since each discharging area corresponds to multiple nodes, the driving distance between each discharging area and itself is the average value of the driving distances between each node corresponding to the discharging area and itself.
[0031] Specifically, the evaluation index is: ; Among them, represents the evaluation value corresponding to the th discharging area, represents the depth of the materials in the th discharging area, represents the The travel distance of a discharge area to itself, and are the weight coefficients of depth and travel distance respectively, is the maximum value of the material depth among all discharge areas, is the maximum value of the travel distance among all discharge areas. By calculating the ratio to the maximum value, the effect of normalization processing is achieved, enabling the depth and travel distance to be calculated on the same scale. The weight coefficients are determined according to actual work needs. If the current work focuses more on the discharge volume, a higher depth weight is selected. If the current work focuses more on transportation and time costs, a higher travel distance weight is selected.
[0032] After that, the discharge area with the highest evaluation value is taken as the first area, and the node closest to its own position in the first area is taken as the first node. However, as Figure 5 shown, since the robot can only travel in a straight line along the movement channel, there are usually multiple nodes closest to its own position among the nodes corresponding to the first area (such as Figure 5 node C and node D in
[0033] At this time, a node is randomly selected from multiple nodes as the first node.
[0034] Furthermore, this embodiment provides a step of updating the first area and the first node according to the second area and the second node to obtain the current discharge area and the current discharge node, including: Judging whether the second area and the second node are the same as the first area and the first node; If the second area is the same as the first area, update the first area and the first node according to the evaluation value corresponding to each discharge area to obtain the current discharge area and the current discharge node; If the second area is different from the first area and the second node is the same as the first node, take the first area as the current discharge area and update the first node to obtain the current discharge node.
[0035] Among them, there may be multiple adjacent robots, so there may be multiple second areas and second nodes. For each second area and second node, the above judgment steps need to be performed.
[0036] Furthermore, this embodiment of the present application provides a step of updating the first area and the first node according to the evaluation value corresponding to each discharge area to obtain the current discharge area and the current discharge node, including: Obtain the evaluation value corresponding to the second area; If the evaluation value corresponding to the second region is greater than the evaluation value corresponding to the first region, update the first region and the first node to obtain the current unloading region and the current unloading node.
[0037] Further, an embodiment of the present application provides a step of updating the first region and the first node to obtain the current unloading region and the current unloading node, including: Exclude the evaluation value corresponding to the first region from the evaluation values corresponding to each unloading region to obtain the remaining evaluation values; Take the unloading region corresponding to the highest evaluation value among the remaining evaluation values as the current unloading region, and take the node closest to its own position among the nodes corresponding to the current unloading region as the current unloading node.
[0038] Specifically, if there is a region in multiple second regions that is the same as the first region, as Figure 6 shown in, the adjacent robot is located at node B, the first region (second region) corresponding to the adjacent robot is the unloading region corresponding to the seventh projection region, and the second node is node E. At this time, it is necessary to interact and communicate with the adjacent robot corresponding to this second region to obtain the evaluation value corresponding to this second region, and send the evaluation value corresponding to the first region to the adjacent robot. If the evaluation value corresponding to this second region is greater than the evaluation value corresponding to the first region, it means that it is more efficient to select the adjacent robot to complete the unloading work in this region. Therefore, it is necessary to update the first region and the first node, that is, take the unloading region corresponding to the highest evaluation value among the remaining evaluation values as the current unloading region, and select the node closest to its own position among the nodes corresponding to the current unloading region as the current unloading node; if the evaluation value corresponding to this second region is less than the evaluation value corresponding to the first region, there is no need to update the first region and the first node, and the adjacent robot performs the update step; if the evaluation value corresponding to this second region is equal to the evaluation value corresponding to the first region, continue to communicate with the adjacent robot to obtain the highest evaluation value among the remaining evaluation values corresponding to the adjacent robot, and compare it with the highest evaluation value among its own remaining evaluation values. If the highest evaluation value corresponding to the adjacent robot is less than its own highest evaluation value, the adjacent robot performs the update step. If it is greater, it performs the update step itself. If they are equal, communicate again and repeat the above steps.
[0039] Further, if the second region is different from the first region and the second node and the first node are the same, it means that the two robots select different unloading regions but occupy the same node, as Figure 7As shown in the figure, the first area corresponding to itself is the unloading area corresponding to the seventh projection area, the first node is node C, the adjacent robot is located at node B, the second area is the unloading area corresponding to the eighth projection area, and the second node is node C. At this time, the robot closer to node C updates the first node, that is, selects the node D closest to its own position from the remaining nodes corresponding to the first area as the current unloading node.
[0040] After that, through the path planning algorithm, a driving path can be generated with the current position as the starting point and the current unloading node as the ending point. Among them, the path planning algorithm can use the A* algorithm, Dijkstra algorithm, etc. Specifically, in order to avoid collisions with adjacent robots during driving, different priorities can be set for different robots. Exemplarily, when both itself and the adjacent robot have completed the determination of the current unloading area and the current unloading node, each robot sends the evaluation value of its corresponding current unloading area to other robots. After each robot receives the evaluation value, it sorts all the robots including itself in descending order according to the evaluation value, and each robot generates its own driving path in turn according to the order in the sequence.
[0041] Moreover, each robot needs to send the generated driving path to the robots that need to generate paths subsequently. The robots that need to generate paths subsequently will set the nodes in this driving path as non-passable nodes to avoid overlapping of paths. Specifically, for the A* algorithm, the evaluation function values of these nodes can be set to a maximum value, and for the Dijkstra algorithm, these nodes can be marked as inaccessible nodes.
[0042] In this embodiment, after detecting an adjacent robot, the collision situation with the adjacent robot is considered from two aspects. On the one hand, the collision during the unloading operation is considered, that is, the situation where the current unloading area or the current unloading node is the same. On the other hand, the collision during the driving process is considered. For the first situation, through communication with the adjacent robot, the evaluation value of this area for the adjacent robot is obtained, and by comparing the evaluation values, it is judged whether the area needs to be replaced; for the second situation, through the setting of the algorithm and the generation order, the collision situation during the driving process is avoided.
[0043] Through the method of this embodiment, each robot can make independent decisions through real-time interactive information. When a new robot is added, it only needs to be connected to the communication network and made to have the ability to interact with other robots and make independent decisions. Compared with the method of unified scheduling by a central controller, this embodiment does not need to make large-scale modifications to the planning algorithm and architecture of the entire robot system, and the robots in this embodiment only interact with other robots when necessary, without frequently sending status information and receiving instructions to the central server like the unified planning mode of the central server, thus reducing the communication bandwidth requirements and energy consumption.
[0044] Further, an embodiment of the present application provides a step of obtaining an action sequence through a deep reinforcement learning model, including: Obtain the pose information after moving to the current discharging node, where the pose information includes the position coordinates of the robotic arm, the pose angle of the robotic arm, and the position coordinates of the material to be discharged; Perform iterative operations, where the iterative operations include: inputting the pose information into the deep reinforcement learning model to obtain the quality values of each action in the preset action space, selecting the action corresponding to the highest quality value, putting it into the action sequence, updating the pose information until the action of putting down the material is put into the action sequence, and outputting the current action sequence.
[0045] Among them, the position coordinates of the robotic arm are specifically the two-dimensional position coordinates where the end effector of the robotic arm is located; the pose angle of the robotic arm refers to the angle between the robotic arm and the horizontal direction; the preset action space includes actions such as the robotic arm moving forward 0.5 meters, the robotic arm moving backward 0.5 meters, the robotic arm translating left 0.2 meters, the robotic arm translating right 0.2 meters, the robotic arm rotating clockwise 15 degrees, the robotic arm rotating counterclockwise 15 degrees, grasping the material, and putting down the material.
[0046] Exemplarily, input the pose information into the main quality network of the deep reinforcement learning model, and the network will output the Q value (quality value) of each action. At this time, select the action corresponding to the highest quality value, for example, select "the robotic arm translates left 0.2 meters", and put it into the action sequence. Then calculate the pose information after executing this action, input it again and select the next action, repeating the above steps until the action of putting down the material is put into the action sequence. Among them, the actions selected each time in the action sequence are arranged in the order of being selected, and the current action sequence is output. For example, output the action sequence of "the robotic arm translates left 0.2 meters, the robotic arm moves backward 0.5 meters, grasps the material, the robotic arm rotates clockwise 15 degrees, puts down the material". Then the robot completes the discharging work according to this action sequence.
[0047] In this embodiment, the action sequence is obtained through the deep reinforcement learning model. This model directly uses the original pose information as input and finally outputs the discharging action sequence, without the need for manual design of intermediate feature extraction and processing links, reducing human intervention and errors, and improving the automation and intelligence level of the entire discharging system.
[0048] Further, an embodiment of the present application provides a step of obtaining a deep reinforcement learning model through training, including: Initialize the parameters of the deep reinforcement learning model; Generate multiple samples in sequence based on environmental interaction and store each sample in the experience replay pool in sequence; Select multiple samples from the experience replay pool and update the parameters of the deep reinforcement learning model according to the multiple samples until the deep reinforcement learning model converges.
[0049] Among them, the parameters of the deep reinforcement learning model include: Q-network parameters (quality network parameters), learning rate, discount factor, exploration rate, and the capacity of the experience replay pool. Specifically, the Q-network parameters include weights and biases, and their initial values are generated by a random number generator; the learning rate is set to 0.001 to control the step size during each parameter update; the discount factor is set to 0.9, reflecting the degree of emphasis on future rewards; the exploration rate is set to 0.8, and a higher exploration rate can avoid getting stuck in local optima; the capacity of the experience replay pool is set to 1000.
[0050] The specific steps for generating multiple samples sequentially based on environment interaction are as follows. First, set up a virtual environment. Randomly generate the pose information of a robot in the virtual environment. This pose information is called the old state. Then, select an action according to the exploration rate and calculate the pose information (this pose information is called the new state) and the obtained reward value after executing this action. For example, if the new state is closer to the material to be unloaded than the old state, the obtained reward value is +5; if it is farther, the obtained reward value is -5. Finally, take (old state, action, reward value, new state) as a sample and store it in the experience replay pool. Repeat the above steps until the number of samples in the experience replay pool reaches the preset capacity value.
[0051] Furthermore, this embodiment provides a step for selecting multiple samples from the experience replay pool, including: Calculate the current quality value of each sample according to the preset main quality network, and calculate the target quality value of each sample according to the preset target quality network; Calculate the temporal difference error of each sample according to the current quality value and the target quality value; Determine the priority of each sample according to the temporal difference error; Select multiple samples from the experience replay pool according to the priority.
[0052] Specifically, the Q-network in this embodiment includes a main Q-network (main quality network) and a target Q-network (target quality network). Both the main Q-network and the target Q-network are neural networks. For example, they can be set as multi-layer perceptrons or convolutional neural networks. Each sample can be represented as , where represents the old state, represents the action, represents the reward value, represents the new state.
[0053] Take Input into the main Q-network, through the multiplication operation of the weight matrix and the input data and the addition operation of the bias term, then through the activation function for non-linear transformation, and the result is passed to the next layer. Calculating layer by layer like this, finally the result is output by the output layer. Each neuron in the output layer corresponds to an action in the action space. Therefore, the final output result is a vector with the same length as the number of actions. Select the element corresponding to as the current quality value of the sample . Then input into the main Q-network, select the action corresponding to the maximum value from the elements of the output vector, denoted as , and input into the target Q-network, select the element corresponding to from the output vector, denoted as , and calculate the target quality value of the sample according to this element as: ; where represents the discount factor. Repeat the above steps for each sample, and the current quality value and target quality value of each sample can be obtained.
[0054] For each sample, take the difference between its corresponding current quality value and target quality value to obtain the temporal difference error of each sample. Obtain the priority of each sample according to the temporal difference error. where is a very small positive number, used to avoid the situation where the priority is zero when the error is zero, and ensure that each sample has a certain probability of being selected. Specifically, for the th sample, the probability of its being selected is , where N represents the total number of samples in the experience replay pool, represents the priority of the th sample, is a hyperparameter, used to control the influence degree of the priority. The closer
[0055] is to 1, the greater the probability of selecting high-priority samples. The number of selected samples is determined according to actual needs, and the way of selecting samples in this embodiment is sampling with replacement.In this embodiment, the traditional Q-network is improved. The Q-network in the prior art only includes a main Q-network. When using the same network for action selection and action evaluation, as the error accumulates, the problem of too high output value is likely to occur, affecting the accuracy of the result. In this embodiment, the action selection and action evaluation are decoupled and carried out using different networks, that is, the main Q-network is used to select actions, and the target Q-network is used to evaluate actions. Since the parameters in the target Q-network are updated slowly, compared with the main Q-network, the frequency of parameter update is greatly reduced, which can effectively avoid the transmission and accumulation of errors, and thus avoid the above problems.
[0056] Moreover, in this embodiment, samples are selected according to the priority of the samples. Compared with the traditional random selection method, the deep reinforcement learning model can learn more effectively from high-priority samples. High-priority samples usually contain more important and representative information. By preferentially learning these samples, the model can capture the key patterns and rules in the environment faster, so as to adjust its own parameters more effectively and accelerate the model convergence speed.
[0057] Further, an embodiment of the present application provides a step of updating the parameters of the deep reinforcement learning model according to multiple samples, including: Obtain the current quality value and the target quality value of each sample; Calculate the loss value according to the current quality value, the target quality value and a preset loss function; Update the parameters of the deep reinforcement learning model according to the loss value and the backpropagation algorithm.
[0058] Specifically, the mean square error of the current quality value and the target quality value can be used as the loss function. At this time, the loss value can be expressed as: ; where represents the number of selected samples. Then, the parameters in the main Q-network are updated through the loss value and the backpropagation algorithm. This step is the prior art and will not be elaborated in this embodiment.
[0059] After that, repeat the above steps of selecting samples and updating parameters, and set a preset number of repetitions. When the preset number of repetitions is reached, copy the parameters in the main Q-network to the target Q-network to ensure the relative stability of the parameters of the target Q-network, and at the same time can gradually follow the learning of the main Q-network for update. Then update the exploration rate, that is, multiply the current exploration rate by a preset decay factor. The preset decay factor is greater than zero and less than 1, and the specific value is determined according to the actual training situation. After that, repeat the above steps of selecting samples, updating parameters and updating the exploration rate until the model converges.
[0060] A method for dynamically optimizing the path of a discharging and feeding robot provided by an embodiment of the present application. Each robot independently runs a path decision algorithm, and through the interactive communication between robots, the discharging area and discharging nodes are updated in case of collision to improve the operation efficiency. Then, through the method of deep reinforcement learning, the discharging action sequence is further determined to ensure the smooth completion of the discharging work.
[0061] Furthermore, an embodiment of the present application provides a discharging and feeding robot, which is applied to the above method for dynamically optimizing the path of a discharging and feeding robot. The discharging and feeding robot includes: a binocular camera, an image processing module, a calculation module, a communication module, and a robotic arm; The binocular camera is used to acquire images of multiple discharging areas within a first preset range; The image processing module is used to select a first area and a first node according to the images; The calculation module is used to update the first area and the first node according to a second area and a second node to obtain the current discharging area and the current discharging node; The communication module is used to acquire the second area and the second node; The robotic arm is used to complete the discharging work on the current discharging area according to the action sequence.
[0062] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0063] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0064] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in the function.
[0065] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A method for dynamically optimizing the path of a unloading and feeding robot, characterized in that, There are multiple unloading and feeding robots located in the same storage area. The storage area includes multiple unloading areas, and each unloading area corresponds to multiple nodes. The method for dynamically optimizing the path of the unloading and feeding robot is executed by one of the unloading and feeding robots. The method for dynamically optimizing the path of the unloading and feeding robot includes: Obtain images of multiple unloading areas within a first preset range, and select a first area and a first node according to the images; Determine whether there are adjacent robots. If so, obtain a second area and a second node, and update the first area and the first node according to the second area and the second node to obtain the current unloading area and the current unloading node. The second area and the second node are the first area and the first node of the adjacent robot, and the adjacent robot is an unloading and feeding robot located within a second preset range; Move to the current unloading node and complete the unloading work for the current unloading area according to the action sequence, where the action sequence is obtained through a deep reinforcement learning model.
2. The path dynamic optimization method of a discharging and feeding robot according to claim 1, wherein The selecting the first area and the first node according to the images includes: Determine the depth of the materials in the multiple unloading areas according to the images; Obtain the driving distances between the multiple unloading areas and itself; Obtain the evaluation value corresponding to each unloading area according to the driving distance, the depth of the materials, and a preset evaluation index; Take the unloading area with the highest evaluation value as the first area, and take the node closest to its own position among the nodes corresponding to the first area as the first node.
3. A path dynamic optimization method for a discharging and feeding robot according to claim 2, characterized in that The updating the first area and the first node according to the second area and the second node to obtain the current unloading area and the current unloading node includes: Determine whether the second area and the second node are the same as the first area and the first node; If the second area is the same as the first area, update the first area and the first node according to the evaluation value corresponding to each unloading area to obtain the current unloading area and the current unloading node; If the second area is different from the first area and the second node is the same as the first node, take the first area as the current unloading area and update the first node to obtain the current unloading node.
4. A method for dynamically optimizing the path of a discharging and feeding robot according to claim 3, characterized in that, The updating the first area and the first node according to the evaluation value corresponding to each unloading area to obtain the current unloading area and the current unloading node includes: Obtain the evaluation value corresponding to the second area; If the evaluation value corresponding to the second area is greater than the evaluation value corresponding to the first area, update the first area and the first node to obtain the current unloading area and the current unloading node.
5. A path dynamic optimization method for a discharging and feeding robot according to claim 4, characterized in that The updating the first area and the first node to obtain the current unloading area and the current unloading node includes: Exclude the evaluation value corresponding to the first area from the evaluation values corresponding to each unloading area to obtain the remaining evaluation values; Take the unloading area corresponding to the highest evaluation value among the remaining evaluation values as the current unloading area, and take the node closest to its own position among the nodes corresponding to the current unloading area as the current unloading node.
6. The path dynamic optimization method of a discharging and feeding robot according to claim 1, characterized in that The action sequence is obtained through a deep reinforcement learning model, including: Obtain the pose information after moving to the current unloading node, where the pose information includes the position coordinates of the robotic arm, the pose angle of the robotic arm, and the position coordinates of the material to be unloaded; Execute an iterative operation, where the iterative operation includes: inputting the pose information into the deep reinforcement learning model to obtain the quality values of each action in the preset action space, selecting the action corresponding to the highest quality value, putting it into the action sequence, updating the pose information until the action of putting down the material is put into the action sequence, and outputting the current action sequence.
7. A path dynamic optimization method for a discharging and feeding robot according to claim 6, characterized in that, The deep reinforcement learning model is obtained through training and includes: Initialize the parameters of the deep reinforcement learning model; Generate multiple samples in sequence based on environment interaction and store each sample in the experience replay pool in sequence; Select multiple samples from the experience replay pool and update the parameters of the deep reinforcement learning model according to the multiple samples until the deep reinforcement learning model converges.
8. A path dynamic optimization method for a discharging and feeding robot according to claim 7, characterized in that The selecting multiple samples from the experience replay pool includes: Calculate the current quality value of each sample according to the preset main quality network, and calculate the target quality value of each sample according to the preset target quality network; Calculate the temporal difference error of each sample according to the current quality value and the target quality value; Determine the priority of each sample according to the temporal difference error; Select multiple samples from the experience replay pool according to the priority.
9. A method for dynamically optimizing the path of a discharging and feeding robot according to claim 8, characterized in that, The updating the parameters of the deep reinforcement learning model according to the multiple samples includes: Obtain the current quality value and the target quality value of each sample; Calculate the loss value according to the current quality value, the target quality value, and the preset loss function; Update the parameters of the deep reinforcement learning model according to the loss value and the backpropagation algorithm.
10. A discharging and feeding robot, which is applied to a method for dynamically optimizing the path of a discharging and feeding robot according to any one of claims 1-9, and is characterized in that, The unloading and feeding robot includes: a binocular camera, an image processing module, a calculation module, a communication module, and a robotic arm; The binocular camera is used to obtain images of multiple unloading areas within a first preset range; The image processing module is used to select a first area and a first node according to the image; The communication module is used to obtain a second area and a second node; The calculation module is used to update the first area and the first node according to the second area and the second node to obtain the current unloading area and the current unloading node; The robotic arm is used to complete the unloading work of the current unloading area according to the action sequence.
Citation Information
Patent Citations
Robot path planning method for target searching
CN110135644A
Feeding robot path collaborative optimization method
CN112859885A
Intelligent robot path planning method and system, computer equipment and medium
CN118518104A
Mobile robot path planning method and system in man-machine co-fusion environment
CN118567362A
Animal house terrain modeling and path planning navigation system and method using laser scanning
CN118706131A