Method for training action detection network, and control method and device of robot
By cropping topic data to generate the first cropped image and training a lightweight and efficient motion detection network, the problem of high complexity in existing models is solved, and efficient real-time control of robots delivering items to customers is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI MOJIA ZHICHUANG ROBOT TECHNOLOGY CO LTD
- Filing Date
- 2026-02-11
- Publication Date
- 2026-05-29
AI Technical Summary
Existing robots have high model complexity in detecting hand posture or hand bounding boxes, resulting in long recognition times and failing to meet the real-time requirements of reaching out to catch objects, thus affecting the efficiency of robots in delivering items to customers.
By cropping topic data during the robot's object-grabbing process, a first cropped image is generated and input into a lightweight and efficient action detection network for training and recognition. The training recognition results are compared with the reference action states in the topic data, and the network parameters are adjusted to improve training efficiency. When the robot grasps the target object, it uses target cropping parameters and weights to crop and recognize the real-time image, controlling the robot to release the object.
This improved the training speed and efficiency of the motion detection network, reduced recognition time, increased the efficiency of robots delivering items to customers, and met real-time requirements.
Smart Images

Figure CN122115960A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and more specifically, to a training method for an action detection network, a control method for a robot, and a device. Background Technology
[0002] With the rapid development of technology, using robots to perform simple service tasks has become a trend. Currently, when using robots to deliver items to customers, the robot typically uses neural networks to output bounding boxes and classification labels for hands and items in images. The customer's receiving status is determined based on the spatial relationship between the hand and item bounding boxes and coherent inter-frame tracking, or by observing hand posture. However, the models used to detect hand posture or bounding boxes are highly complex, and the recognition process on the robot's edge devices is time-consuming, causing inference delays and failing to meet the real-time requirements for reaching out to receive items. Therefore, improving the efficiency of robots delivering items to customers has become an urgent problem to be solved. Summary of the Invention
[0003] In view of this, embodiments of this application propose a training method for an action detection network, a control method for a robot, and an apparatus to improve the above-mentioned problems.
[0004] According to a first aspect of the embodiments of this application, a method for training an action detection network is provided. The method includes: acquiring topic data during a robot grasping an object, wherein the topic data includes images collected by the robot that include hand features and images that do not include hand features; cropping the original images in the topic data to obtain a first cropped image, wherein the first cropped image includes a portion of the image region of the original image; inputting the first cropped image into the action detection network for training and recognition to obtain a training and recognition result; and training the action detection network based on the recognized action state in the training and recognition result and a reference action state corresponding to the topic data.
[0005] According to a second aspect of the embodiments of this application, a robot control method is provided, wherein the robot is equipped with a motion detection network. The method includes: responding to a detection request, acquiring a real-time image of the robot grasping a target object and target cropping parameters and target weights of the motion detection network; cropping the real-time image according to the target cropping parameters to obtain a target cropped image; inputting the target cropped image into the motion detection network for recognition to obtain a target recognition result, wherein the motion detection network is a network configured according to the target weights; and controlling the robot to release the target object according to the target recognition result.
[0006] According to a third aspect of the embodiments of this application, a processing apparatus for an action detection network is provided. The apparatus includes a transceiver module and a processing module. The transceiver module is used to execute receiving / transmitting actions in a training method for the action detection network, and the processing module is used to execute other actions in the training method for the action detection network besides receiving / transmitting actions. Alternatively, the transceiver module is used to execute receiving / transmitting actions in a robot control method, and the processing module is used to execute other actions in the robot control method besides receiving / transmitting actions.
[0007] According to a fourth aspect of the embodiments of this application, a robot is provided, comprising: a processor; and a memory storing computer-readable instructions, wherein when the computer-readable instructions are executed by the processor, the training method of the action detection network or the control method of the robot as described above is implemented.
[0008] According to a fifth aspect of the present application, a computer-readable storage medium is provided, on which computer-readable instructions are stored, which, when executed by a processor, implement the training method of the action detection network or the control method of the robot as described above.
[0009] In the proposed solution, the training method for the action detection network first involves cropping the original image from the topic data acquired during the robot's object-grabbing process. This results in a first cropped image that includes a portion of the original image. This first cropped image is then input into the action detection network for training and recognition, yielding a training recognition result. The training network can then be trained based on the recognized action states identified in the training result and the corresponding reference action states in the topic data. This approach trains the action detection network to recognize cropped images, ensuring both training speed and efficiency. Furthermore, by training a lightweight and efficient action detection network, it avoids the problems of high complexity, large computational load, and long inference time associated with existing detection networks or pose recognition models. This training method also improves the efficiency of robots using action detection networks in delivering items to customers.
[0010] In this application's solution, the robot control method allows the robot's control unit to respond to detection requests and acquire real-time images and target cropping parameters and weights from the motion detection network during the robot's grasping of a target item. This enables the real-time image to be cropped according to the target cropping parameters, resulting in a cropped image. This cropped image is then input into the motion detection network configured according to the target weights for recognition, yielding a target recognition result. Finally, the robot can be controlled to release the target item based on the target recognition result. This solution uses a trained, lightweight, and efficient motion detection network to identify whether a reaching intention exists in the cropped image obtained after cropping based on the target cropping parameters, thereby controlling the robot to release the target item. This avoids the problem of excessively long recognition time caused by the motion detection network recognizing real-time images, improving the efficiency of the robot delivering items to customers.
[0011] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit the embodiments of this application. Attached Figure Description
[0012] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0013] Figure 1 This is a schematic diagram of a robot according to an embodiment of this application.
[0014] Figure 2 This is a flowchart illustrating a training method for an action detection network according to an embodiment of this application.
[0015] Figure 3 This is a flowchart illustrating a training method for an action detection network according to another embodiment of this application.
[0016] Figure 4 This is a flowchart illustrating a robot control method according to an embodiment of this application.
[0017] Figure 5 This is a schematic flowchart of a robot control method according to another embodiment of this application.
[0018] Figure 6 This is a schematic flowchart of a robot control method according to another embodiment of this application.
[0019] Figure 7This is a block diagram of a training apparatus for an action detection network according to an embodiment of this application.
[0020] Figure 8 This is a block diagram of a robot control device according to an embodiment of this application.
[0021] Figure 9 This is a hardware structure diagram of a robot according to another embodiment of this application.
[0022] The accompanying drawings have illustrated specific embodiments of the present application. More detailed descriptions will follow. These drawings and descriptions are not intended to limit the scope of the present application's embodiments in any way, but rather to illustrate the concepts of the present application's embodiments to those skilled in the art through specific embodiments. Detailed Implementation
[0023] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.
[0024] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.
[0025] Please see Figure 1 , Figure 1 This application illustrates a robot provided in one embodiment, such as... Figure 1 As shown below, the training method for implementing the motion detection network and the control method for the robot are illustrated by example.
[0026] In one optional implementation, the robot 100 includes a transceiver module 110 and a processing module 120, wherein the transceiver module 110 and the processing module 120 are software units or modules. Specifically, the transceiver module 110 executes the receiving / transmitting actions in the training method of the motion detection network, and the processing module 120 executes other actions in the training method of the motion detection network besides the receiving / transmitting actions. Alternatively, the transceiver module 110 executes the receiving / transmitting actions in the robot's control method, and the processing module 120 executes other actions in the robot's control method besides the receiving / transmitting actions.
[0027] In another optional implementation, the transceiver module 110 and the processing module 120 refer to hardware devices. For example, the robot 100 may further include a sensing unit, a control unit, and a controlled unit. The sensing unit acquires topic data during the robot's grasping process and identifies the topic data to accurately determine the reaching intention in a certain frame of an image. The control unit receives the sensing result of the reaching intention in a certain frame of an image from the sensing unit. The execution unit successfully controls the robot's arm movement and ensures that its motors are fault-free and that dexterous grasping is normal. Optionally, the parameters in Table 1 below may be included during the robot's execution of the motion detection network training method and the robot control method.
[0028] Table 1. Training methods for robot motion detection networks and parameters for robot control methods.
[0029]
[0030] Figure 1 The robots in the document can be used to achieve the following Figure 2 For the training method of the described action detection network, please refer to [link / reference]. Figure 2 , Figure 2 This application illustrates a method for training an action detection network according to an embodiment of the present application. In a specific embodiment, this method for training the action detection network can be applied to, for example... Figure 7 The training device 700 for the motion detection network shown, and the robot 100 equipped with the training device 700 for the motion detection network. Figure 1 or Figure 9 The specific process of this embodiment will be described below. Of course, it is understood that this method can be executed by an electronic device with computing power, such as a vehicle-mounted server, a cloud server, or other processors. The following will focus on... Figure 2 The process shown is described in detail. The training method of the action detection network may specifically include the following steps 210-240.
[0031] Step 210: Obtain topic data during the robot's object-grabbing process, wherein the topic data includes images collected by the robot that include hand features and images that do not include hand features.
[0032] As an alternative approach, image acquisition devices or other sensors can be pre-installed in the robot to acquire topic data during the robot's object-grabbing process. This topic data can be stored in a bag file within the robot. A bag file is a special data format used in the Robot Operating System (ROS) to record and replay topic data from the ROS system, comprehensively recording all sensor data, control commands, and status information during robot operation.
[0033] In one optional scenario, after the robot is activated, raw topic data from an RGB camera can be recorded using the robot's image acquisition device. During the recording process, a human hand can be used at different times to simulate various realistic states of reaching out, thereby obtaining rich topic data. Optionally, the topic data can be cleaned and processed after acquisition to ensure its accuracy.
[0034] Optionally, to improve the training speed of the action detection network, the collected topic data can be pre-classified, thereby classifying the topic data into images with hand features and images without hand features, and then labeling whether hand features exist in the images after classification. Hand features can include human hand features and robotic hand features.
[0035] Step 220: Crop the original image in the topic data to obtain a first cropped image, wherein the first cropped image includes a portion of the image region of the original image.
[0036] As an alternative approach, since the original images acquired by the robot are relatively large, accurately identifying hand features requires comprehensive recognition of the entire original image, which is time-consuming. To accelerate image recognition, the original image can be cropped to reduce its size. This reduces unnecessary feature recognition and speeds up the process. Optionally, cropping the original image reduces the number of pixels, significantly increasing the processing speed compared to recognizing the first cropped image with its larger pixel count. For example, if the original image is 640×480 pixels, direct recognition of that single image takes approximately 8ms, while recognizing the first cropped image takes approximately 5ms.
[0037] Step 230: Input the first cropped image into the action detection network for training and recognition, and obtain the training and recognition results.
[0038] As an alternative approach, after obtaining the first cropped image, an action detection network can be used to train and recognize the first cropped image, thereby obtaining the training recognition result corresponding to each first cropped image. Optionally, the training recognition result includes the presence of human hand features and robot hand features in the first cropped image, the presence of only robot hand features in the first cropped image, and the absence of any hand features in the first cropped image.
[0039] In one alternative scenario, the action detection network can be the s version of the YOLOv11s classification model. Compared to the 9.4M size of the YOLOv11-s detection model, the s version of the YOLOv11s classification model is 5.9M in size, which reduces the model size by 3.5M. This reduces the model complexity to some extent and improves the inference speed of a single image.
[0040] Optionally, during the training and recognition process, the action detection network only needs to classify the first cropped input image, that is, classify whether there are human hand features, robot hand features, or no hand features in the first cropped image, without needing to detect the position of the human hand in the image or the detection box of the hand image and the object.
[0041] Step 240: Train the action detection network based on the identified action states in the training recognition results and the reference action states corresponding to the topic data.
[0042] As an alternative approach, after obtaining the training recognition results, the loss value of the action detection network can be determined to meet the preset conditions and the cropping process of the first cropped image can be determined by checking whether the recognition action state corresponding to the training recognition results is consistent with the reference action state corresponding to the first cropped image in the original image of the topic data. This can be used to train the action detection network.
[0043] Optionally, the reference action state corresponding to the topic data is the reference action state of whether the original image in the topic data includes hand features (human hand features and / or robot hand features). When creating the training dataset, the reference action state can be labeled in advance according to the content of the original image, so that the reference action state corresponding to the topic data can be determined according to the labeling information of the original image.
[0044] In one optional scenario, if the loss value of the action detection network does not meet the preset conditions, the network parameters and weights in the action detection network are adjusted according to the loss value. The adjusted action detection network is then trained again based on the collected topic data until the loss value meets the preset conditions or the number of training iterations is greater than or equal to the threshold number of iterations, at which point the training of the action detection network ends.
[0045] As an alternative approach, the original images in the topic data can be pre-classified and labeled with categories. After obtaining the training recognition results, the loss value of the action detection network can be determined based on the recognition classification corresponding to the first cropped image in the training recognition results and the category label in the original image corresponding to the first cropped image. The network weights of the action detection network can then be determined based on the loss value.
[0046] In the embodiments of this application, the original image in the topic data acquired during the robot's object-grabbing process is first cropped to obtain a first cropped image including a portion of the original image. This first cropped image is then input into an action detection network for training and recognition, thereby obtaining training and recognition results. The action detection network can then be trained based on the recognized action states from the training results and the corresponding reference action states from the topic data. This solution trains the action detection network to recognize cropped images, ensuring the training speed and efficiency of the action detection network. Furthermore, by training a lightweight and efficient action detection network, it avoids the problems of high complexity, large computational load, and high inference time associated with existing detection networks or pose recognition models. This training method also improves the efficiency of robots using action detection networks in delivering items to customers.
[0047] Please see Figure 3 , Figure 3 This paper illustrates a training method for an action detection network provided in an embodiment of this application. The following will focus on... Figure 3 The process is described in detail. The action detection network determines the recognition action state, which includes a first state where there is robot grasping an object and human hand features, a second state where there is robot grasping an object but no human hand features, and a third state where there is neither robot grasping an object nor human hand features. The training method of the action detection network may specifically include the following steps 310-370.
[0048] Step 310: Obtain topic data during the robot's object-grabbing process, wherein the topic data includes images collected by the robot that include hand features and images that do not include hand features.
[0049] Step 320: Crop the original image in the topic data to obtain a first cropped image, wherein the first cropped image includes a portion of the image region of the original image.
[0050] Step 330: Input the first cropped image into the action detection network for training and recognition, and obtain the training and recognition results.
[0051] The specific steps of steps 310-330 can be found in steps 210-230, and will not be repeated here.
[0052] Step 340: Determine the reference action state of the first cropped image based on the topic data; the reference action state is the first state, the second state, or the third state.
[0053] As an alternative approach, to ensure the training speed and accuracy of the action detection network, the reference action state of the first cropped image can be determined and labeled in advance. This allows the decision on whether to end the training of the action detection network to be made based on the determined and labeled reference action state combined with the recognized action state in the training results.
[0054] In one alternative scenario, to ensure that the action detection network can accurately identify whether someone is reaching out when applied offline, the reference action state of the first cropped image can be determined based on the content of the original image in the topic data.
[0055] Optionally, when collecting topic data, a training dataset can be created in advance based on the topic data. That is, the original image is classified according to the specific content of the original image in the collected topic data. After obtaining the first cropped image, the first cropped image is directly classified and labeled according to the classification result corresponding to the original image based on the correspondence between the first cropped image and the original image, so as to determine the reference action state of the first cropped image.
[0056] In one alternative scenario, since the topic data is collected during the process of the robot grasping the item, the role of the action detection network is to identify whether the robot has grasped the item and whether someone has reached out to take the item. Therefore, the original image in the topic data can be divided into three states: a first state where there are robot grasping the item and human hand features, a second state where there are robot grasping the item but no human hand features, and a third state where there are no robot grasping the item and no human hand features.
[0057] Step 350: Determine the loss value based on the identified action state and the reference action state.
[0058] As an alternative approach, after determining the recognized action state in the training recognition result and the reference action state corresponding to the first cropped image, the loss value can be determined based on the difference between the recognized action state and the reference action state.
[0059] In one optional scenario, since multiple original images can be input each time during the training of the action detection network, multiple action states are obtained. Each original image corresponds to a first cropped image, and each first cropped image has its corresponding action state. Each action state indicates whether the corresponding first cropped image is in a first, second, or third state. Therefore, the loss value can be determined by determining the number of times the action state corresponding to each first cropped image is identical to the reference action state corresponding to each original image. Optionally, the loss value can be determined by the percentage of the first number of action states corresponding to first cropped images that are identical to the reference action states corresponding to original images out of the second number of all first cropped images.
[0060] Step 360: If the loss value is less than or equal to the loss threshold, then output the trained action detection network.
[0061] As an optional approach, a loss threshold can be preset, allowing the determination of whether to terminate the training of the action detection network by comparing a predetermined loss value with this threshold. Optionally, when the loss value is determined to be less than the loss threshold, it can be determined that the accuracy of the action detection network in cropping the original image and recognizing the action on the cropped image meets a preset condition. Therefore, the action detection network can stop training to avoid overfitting due to overtraining, which would lead to a decrease in the accuracy of the action detection network. Thus, when the loss value is determined to be less than or equal to the loss threshold, the training of the action detection network is terminated, and the trained action detection network is output. Optionally, the trained action detection network includes the network structure, network weights, and various network parameters.
[0062] In some embodiments, step 360 includes: obtaining a trained action detection network; obtaining target clipping parameters and target weights corresponding to the trained action detection network; and storing a combination of the target clipping parameters and the target weights.
[0063] As an alternative approach, after outputting the action detection network, in order to enable its offline application, the target cropping parameters and target weights corresponding to the trained action detection network can be obtained and stored after the network is trained. This ensures that when the action detection network is used offline in the robot, it can be configured directly based on the stored target weights and target cropping parameters. This avoids errors in the action detection network caused by resetting the robot each time it is used, which could lead to inaccurate recognition of the images collected by the robot.
[0064] Step 370: If the loss value is greater than the loss threshold, then update the parameters of the action detection network according to the loss value, and train the updated action detection network.
[0065] As an alternative approach, if the loss value is determined to be greater than the loss threshold, it can be determined that the error rate of the action detection network in recognizing the first cropped image is too high. Therefore, it is necessary to update the action detection network by adjusting its parameters, and continue training based on the updated action detection network. The loss value is then determined again based on the training recognition results until the loss value is less than or equal to the loss threshold, or the number of training iterations is equal to or greater than the number of iterations threshold, at which point the trained action detection network is output.
[0066] Optionally, when the loss value is determined to be greater than the loss threshold, it may be because the parameters in the action detection network are inappropriate, or it may be because the method or content of cropping the original image is inappropriate. Therefore, while updating the parameters of the action detection network according to the loss value, the cropping parameters of the original image can be reset to update the region in the original image included in the first cropped image.
[0067] In some embodiments, step 370 includes: adjusting a first cropping parameter to obtain a second cropping parameter; the first cropping parameter being a parameter for cropping the original image to obtain a first cropped image; cropping the original image according to the second cropping data to obtain a second cropped image; and training an action detection network updated based on the second cropped image.
[0068] As an alternative approach, since the original image was cropped before recognition, the loss value exceeding the loss threshold may be due to inappropriate parameters of the action detection network or inappropriate cropping parameters used to crop the original image. Therefore, when the loss value is determined to be greater than the loss threshold, the first cropping parameter can be adjusted to adaptively re-crop the original image. This ensures that the resulting cropped image accurately shows whether hand features were detected during the robot's grasping process, thus preserving the robot's grasping of the object and the human's hand in the cropped image.
[0069] In one optional scenario, the first cropping parameters may include the aiming point for cropping the original image, the image coordinates corresponding to the vertices of the first cropped image, the image width of the first cropped image, and the image height of the first cropped image. Optionally, adjusting the first cropping parameters may involve changing the coordinates corresponding to the vertices of the cropped image, decreasing or increasing the image width and image height, etc.
[0070] Optionally, after adjusting the first cropping parameters to obtain the second cropping parameters, in order to ensure the training efficiency and accuracy of the action detection network, the original image can be cropped again according to the second cropping parameters to obtain the second cropped image, and then the second cropped image can be input into the updated action detection network for training.
[0071] As an alternative approach, to ensure the accuracy of the cropping parameters, the original image can be pre-trained in an action detection network. Based on the training and recognition results of the original image and the reference action state within it, a first loss value is determined. This first loss value is then used to update the action detection network until it equals or falls below a first loss threshold, resulting in a well-trained network. Next, the original image is cropped to obtain a first cropped image, which is then input into the trained action detection network for recognition. Based on the recognition results and the reference action state of the original image corresponding to the first cropped image, a second loss value is determined. If the second loss value exceeds the second loss threshold, the first cropping parameters of the first cropped image are adjusted to obtain second cropping parameters. The original image is then cropped according to these second cropping parameters to obtain a second cropped image. This second cropped image is then input into the trained action detection network for recognition, and so on, until the second loss value equals or falls below the second loss threshold. The corresponding target cropping parameters and target weights of the trained action detection network are then determined, and finally, the target cropping parameters and target weights are associated and stored.
[0072] In this embodiment, a reference action state for the first cropped image is first determined based on topic data. This allows for the determination of a loss value based on the reference action state of the first cropped image and the recognized action state in the training results. When the loss value is less than or equal to a loss threshold, the trained action detection network is output; otherwise, when the loss value is greater than the loss threshold, the action detection network is updated and retrained. The first cropping parameters of the first cropped image are also updated, allowing for cropping of the original image based on the updated second cropping parameters. Furthermore, the action detection network is updated based on the second cropped image corresponding to the second cropping parameters. This approach adjusts the cropping parameters while training the action detection network, ensuring that the cropped image retains the image region containing the hand features from the original image. This allows the action detection network to be trained on cropped images that include the image region containing the hand features from the original image, thus ensuring the training speed of the action detection network.
[0073] Figure 1 The robots in the text can also be used to achieve the following Figure 4For the described robot control method, please refer to [link / reference]. Figure 4 , Figure 4 This application illustrates a robot control method according to an embodiment of the present application. In a specific embodiment, the robot control method can be applied to, for example... Figure 8 The robot control device 800 and the robot 100 equipped with the robot control device 800 are shown. Figure 1 or Figure 9 The specific process of this embodiment will be described below. Of course, it is understood that this method can be executed by an electronic device with computing power, such as a smartphone, tablet, smart wearable device, cloud server, or other processor. The following will focus on... Figure 4 The process shown is described in detail. The robot is equipped with a motion detection network, and the control method of the robot may specifically include the following steps 410-440.
[0074] Step 410: In response to the detection request, obtain real-time images of the robot grasping the target object and the target clipping parameters and target weights of the motion detection network.
[0075] As an alternative approach, after the robot is started and the task of grasping the target object is set for the robot, the robot's control unit generates a detection request to control whether the robot releases the target object after grasping it. Therefore, real-time images of the robot grasping the target object and the target weights and target clipping parameters of the stored motion detection network can be obtained first.
[0076] Optionally, the target weights and target clipping parameters of the trained action detection network can be pre-stored in a cloud server or the robot's control unit. Upon receiving a detection request, the target clipping parameters and target weights can be obtained from the cloud server or control unit, and the action detection network in the robot can be configured according to the target weights. This allows for direct detection and recognition based on the action detection network configured with the target weights.
[0077] Step 420: Crop the real-time image according to the target cropping parameters to obtain the target cropped image.
[0078] As an alternative approach, after obtaining the target cropping parameters, the real-time image can be cropped using these parameters to obtain a target cropped image. This target cropped image can then be input into the robot's motion detection network, enabling the motion detection network to quickly recognize target cropped images that are smaller in size and contain less content, thereby improving the robot's recognition speed of real-time images.
[0079] Step 430: Input the target cropped image into the action detection network for recognition to obtain the target recognition result, wherein the action detection network is a network configured according to the target weights.
[0080] As an alternative approach, after obtaining the target cropped image, the target cropped image is input into an action detection network for recognition to determine whether hand features exist in the real-time image. Based on the presence of hand features in the real-time image, the robot is controlled to immediately release the object. Optionally, the target recognition result may include three cases: recognition of both human and robot hand features, recognition of only robot hand features, and no recognition of any hand features.
[0081] Step 440: Control the robot to release the target item based on the target recognition result.
[0082] As an alternative approach, after determining the target recognition result, the robot can be controlled to release the target item based specifically on the target recognition result. Optionally, if the target recognition result indicates that a human hand feature has been detected, the robot is controlled to release the target item; if the target recognition result indicates that no human hand feature has been detected but the robot's hand feature has been detected, the robot can be controlled to hold the target item for a period of time, continuously acquiring real-time images and detecting whether a human hand feature has been detected during this period, and if no human hand feature has been detected after a period of time, the robot is controlled to release the target item; or, if no hand feature is detected, the robot is restarted to respond to the detection request, thus avoiding the robot constantly grabbing the target item and ignoring other customers, resulting in a poor customer experience.
[0083] In the embodiments of this application, the control unit in the robot can respond to a detection request to acquire real-time images of the robot grasping a target item, as well as the target cropping parameters and target weights of the motion detection network. This allows the real-time image to be cropped according to the target cropping parameters to obtain a cropped image. The cropped image is then input into the motion detection network configured according to the target weights for recognition, resulting in a target recognition result. Finally, the robot can be controlled to release the target item based on the target recognition result. This solution can identify whether there is a reaching intention by using a trained, lightweight, and efficient motion detection network to crop the target image based on the target cropping parameters, thereby controlling the robot to release the target item. This avoids the problem of excessively long recognition time caused by the motion detection network recognizing real-time images, improving the efficiency of the robot delivering items to customers.
[0084] Please see Figure 5 , Figure 5 This application illustrates a robot control method according to an embodiment of the present application. The following will focus on... Figure 5The process shown is described in detail, and the robot control method may specifically include the following steps 510-550.
[0085] Step 510: In response to the detection request, obtain real-time images of the robot grasping the target object and the target clipping parameters and target weights of the motion detection network.
[0086] Step 520: Crop the real-time image according to the target cropping parameters to obtain the target cropped image.
[0087] Step 530: Input the target cropped image into the action detection network for recognition to obtain the target recognition result, wherein the action detection network is a network configured according to the target weights.
[0088] The specific steps of steps 510-530 can be found in steps 410-430, and will not be repeated here.
[0089] Step 540: If the target recognition result is received within the first time period, and it is determined that the target recognition result indicates that a human hand feature has been recognized, then a first instruction is generated, and the robot is controlled to release the target item according to the first instruction.
[0090] As an alternative approach, to ensure the accuracy of robot control, a first time period can be set to determine whether a target recognition result is received within that first time period. This allows for robot control based on the recognition result received within the first time period. Alternatively, after the robot's motion detection network recognizes the real-time image, it can send the target recognition result to the robot's corresponding control unit. The robot can then control itself to retrieve or release the corresponding item based on the received target recognition result.
[0091] In one optional scenario, if a target recognition result is received within the first time period, it is determined whether hand features have been detected, as indicated by the target recognition result, to control whether the robot releases the target item. Optionally, if a target recognition result is received within the first time period, and the target recognition result indicates that hand features have been detected, it can be determined that a customer is reaching out to take an item. Therefore, a first instruction can be generated, and the robot can be controlled to release the corresponding target item based on the first instruction, thereby ensuring that the target item can be delivered to the customer.
[0092] Step 550: If the target recognition result is not received within the first time period or the target recognition result indicates that no human hand feature is recognized, a second instruction is generated, and the robot is controlled to release the target item according to the second instruction.
[0093] As an optional approach, if no target recognition result is received within the first time period, or if the target recognition result indicates that no hand features have been recognized, it can be determined that the robot may have an information transmission error or has temporarily failed to detect a customer reaching out to take an object. To avoid the robot continuously grabbing target items, which could hinder rapid customer service in situations with many customers, a second instruction can be generated. This second instruction can then control the robot to release the target item based on a second mass. Optionally, the first and second instructions can be the same instruction or different instructions, but both the first and second instructions aim to control the robot to release the target item. Optionally, the second instruction may also include timeout information or alarm information. Simultaneously, the second instruction is fed back to an electronic device connected to the robot for display. Engineers can then view the robot's operating information on the electronic device and perform maintenance or control operations based on this information.
[0094] In this embodiment, if a target recognition result is received within a first time period and it is determined that the target recognition result indicates that a human hand feature has been recognized, a first instruction is generated, and the robot is controlled to release the target item according to the first instruction; or if a target recognition result is not received within the first time period or the target recognition result indicates that a human hand feature has not been recognized, a second instruction is generated, and the robot is controlled to release the target item according to the second instruction, thereby ensuring that the robot can transfer the target item at any time.
[0095] Please see Figure 6 , Figure 6 This application illustrates a robot control method according to an embodiment of the present application. The following will focus on... Figure 6 The process shown is described in detail, and the robot control method may specifically include the following steps 610-660.
[0096] Step 610: If the robot executes a grasping command to grasp the target item, then determine the first position information of the target robotic arm of the robot executing the grasping command.
[0097] As one approach, since robots can be used to serve customers and deliver target items to them, the robot needs to move the target item from its placement location to the customer's location and release it after detecting the customer's hand features at that location, thus ensuring the customer can receive the item promptly. Therefore, after the robot executes the grasping command for the target item, the first position information of the target robotic arm executing the grasping command is determined in real time. This first position information is used to determine whether the robot has moved the target item to the customer's corresponding location. Optionally, the first position information can be determined by determining the robot's position, or it can be determined by determining the direction and length of movement of the target robotic arm.
[0098] Step 620: If the first location information is the same as the target location information, then generate a detection request.
[0099] As an alternative approach, the target location information can be the pre-set location of the customer's entrance or exit. When the first location information and the target location information are the same, the robot can be determined to move the target robotic arm that executes the grasping command to the customer's location. At this time, it can detect whether the customer reaches out to take the object, and then control the robot to release the target item.
[0100] In one optional scenario, the robot can be configured with an intermediate layer, a perception layer, and an execution layer. The intermediate layer can be a module that controls the robot, the perception layer contains an action detection network, and the execution layer can be the robot's robotic arm or displacement module. When the robot grasps a target object and continuously moves it to the corresponding position, the intermediate layer sends a detection request to the perception layer via the ROS service. Upon receiving the request, the perception layer initiates the action detection network, including the YOLOv11 classification model, and begins recognizing each frame of the real-time image. If the recognition result is "human_grasp" within the timeout period, it returns a final_status command with success set to true to the intermediate layer. After receiving the data, the intermediate layer sends a command to the execution layer to release the object, ensuring the customer's hand catches it. If "human_grasp" is not detected within the timeout period, the intermediate layer automatically sends a success set to true command and controls the execution layer to release the object.
[0101] Step 630: In response to the detection request, obtain real-time images of the robot grasping the target object and the target clipping parameters and target weights of the motion detection network.
[0102] Step 640: Crop the real-time image according to the target cropping parameters to obtain the target cropped image.
[0103] Step 650: Input the target cropped image into the action detection network for recognition to obtain the target recognition result, wherein the action detection network is a network configured according to the target weights.
[0104] Step 660: Control the robot to release the target item based on the target recognition result.
[0105] In this embodiment, when the robot executes the grasping command to grab the target item, the first position information of the target robotic arm is determined. If the first position information is the same as the target position information, a detection request is generated. The detection request is used to detect whether a user is reaching out to grab the item, thereby ensuring the accuracy of the robot in delivering the target item.
[0106] The above embodiments describe in detail the training method of the action detection network provided in this application. In other embodiments, this application also provides a training apparatus for the action detection network. Figure 7 This is a block diagram of a training device for an action detection network according to an embodiment of this application, such as... Figure 7 As shown, the training device 700 for the action detection network includes: a topic data acquisition module 710, a first cropping module 720, a first recognition module 730, and a training module 740.
[0107] The topic data acquisition module 710 is used to acquire topic data during the robot's object grasping process, wherein the topic data includes images collected by the robot that include hand features and images that do not include hand features; the first cropping module 720 is used to crop the original image in the topic data to obtain a first cropped image, wherein the first cropped image includes a portion of the image region of the original image; the first recognition module 730 is used to input the first cropped image into an action detection network for training and recognition to obtain training and recognition results; the training module 740 is used to train the action detection network based on the recognized action state in the training and recognition results and the reference action state corresponding to the topic data.
[0108] In some embodiments, the action detection network determines the recognition action state as follows: a first state where there is a robot grasping an object and human hand features, a second state where there is a robot grasping an object but no human hand features, and a third state where there is no robot grasping an object and no human hand features; the training module 740 includes: a reference action state determination submodule, used to determine a reference action state of the first cropped image based on the topic data; the reference action state is the first state, the second state, or the third state; a loss value determination submodule, used to determine the loss value based on the recognition action state and the reference action state; an output submodule, used to output the trained action detection network if the loss value is less than or equal to a loss threshold; and an update submodule, used to update the parameters of the action detection network based on the loss value if the loss value is greater than the loss threshold, and train the updated action detection network.
[0109] In some embodiments, the update submodule includes: an adjustment unit, configured to adjust a first cropping parameter to obtain a second cropping parameter; the first cropping parameter being a parameter for cropping the original image to obtain a first cropped image; a cropping unit, configured to crop the original image according to the second cropping data to obtain a second cropped image; and a training unit, configured to train an action detection network updated based on the second cropped image.
[0110] In some embodiments, the output submodule includes: a first acquisition unit for acquiring the trained action detection network; a second acquisition unit for acquiring the target pruning parameters and target weights corresponding to the trained action detection network; and a storage unit for storing the combination of the target pruning parameters and the target weights.
[0111] The above embodiments describe in detail the robot control method provided in this application. In other embodiments, this application also provides a robot control device. Figure 8 This is a block diagram of a robot control device according to an embodiment of this application, such as... Figure 8 As shown, the robot's control device 800 includes: a response module 810, a second cutting module 820, a second recognition module 830, and a control module 840.
[0112] A response module 810 is used to respond to a detection request and acquire real-time images of the robot grasping the target object, as well as the target cropping parameters and target weights of the motion detection network; a second cropping module 820 is used to crop the real-time image according to the target cropping parameters to obtain a target cropped image; a second recognition module 830 is used to input the target cropped image into the motion detection network for recognition to obtain a target recognition result, wherein the motion detection network is a network configured according to the target weights; and a control module 840 is used to control the robot to release the target object according to the target recognition result.
[0113] In some embodiments, the control module 840 includes: a first control subunit, configured to generate a first instruction and control the robot to release the target item according to the first instruction if the target recognition result is received within a first time period and it is determined that the target recognition result indicates that a human hand feature has been recognized; and a second control subunit, configured to generate a second instruction and control the robot to release the target item according to the second instruction if the target recognition result is not received within the first time period or the target recognition result indicates that a human hand feature has not been recognized.
[0114] In some embodiments, the robot control device 800 further includes: a position information determination module, configured to determine first position information of the target robotic arm that the robot executes the grasping instruction if the robot executes a grasping instruction to grasp the target object; and a request generation module, configured to generate a detection request if the first position information is the same as the target position information.
[0115] According to one aspect of the embodiments of this application, a robot is also provided, such as Figure 9 As shown, the robot 100 also includes a processor 130 and one or more memories 140. The one or more memories 140 are used to store program instructions executed by the processor 130. When the processor 130 executes the program instructions, it implements the above-mentioned training method of the motion detection network and / or the control method of the robot.
[0116] Furthermore, the processor 130 may include one or more processing cores. The processor 130 runs or executes instructions, programs, code sets, or instruction sets stored in the memory 140, and retrieves data stored in the memory 140. Optionally, the processor 130 may be implemented using at least one hardware form selected from Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). The processor 130 may integrate one or a combination of several of the following: a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), and a modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor and may be implemented using a separate communication chip.
[0117] According to one aspect of this application, a computer-readable storage medium is also provided, which may be included in the cloud server described in the above embodiments; or it may exist independently and not assembled into the cloud server. The aforementioned computer-readable storage medium carries computer-readable instructions that, when executed by a processor, implement the methods in any of the above embodiments.
[0118] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. Computer-readable storage media can be, for example, but not limited to: electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such transmitted data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.
[0119] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.
[0120] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.
[0121] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A training method for an action detection network, characterized in that, The method includes: The topic data is obtained during the process of the robot grasping an object, wherein the topic data includes images collected by the robot that include hand features and images that do not include hand features; The original image in the topic data is cropped to obtain a first cropped image, wherein the first cropped image includes a portion of the image region of the original image; The first cropped image is input into the action detection network for training and recognition, and the training and recognition results are obtained. The action detection network is trained based on the identified action states in the training results and the reference action states corresponding to the topic data.
2. The method according to claim 1, characterized in that, The action detection network determines the recognition action states, including a first state where a robot grasps an object and human hand features are present, a second state where the robot grasps an object but no human hand features are present, and a third state where neither robot grasps an object nor human hand features are present. Training the action detection network based on the recognition action states identified in the training results and the reference action states corresponding to the topic data includes: The reference action state of the first cropped image is determined based on the topic data; the reference action state is the first state, the second state, or the third state. The loss value is determined based on the identified action state and the reference action state; If the loss value is less than or equal to the loss threshold, then the trained action detection network is output. If the loss value is greater than the loss threshold, the parameters of the action detection network are updated according to the loss value, and the updated action detection network is trained.
3. The method according to claim 2, characterized in that, The step of updating the parameters of the action detection network based on the loss value and training the updated action detection network includes: The first cropping parameter is adjusted to obtain the second cropping parameter; the first cropping parameter is the parameter used to crop the original image to obtain the first cropped image; The original image is cropped according to the second cropping data to obtain a second cropped image; The action detection network is trained based on the updated action detection network of the second cropped image.
4. The method according to claim 2, characterized in that, The output trained action detection network includes: Obtain the trained action detection network; Obtain the target cropping parameters and target weights corresponding to the trained action detection network; Store the combination of the target pruning parameters and the target weights.
5. A method for controlling a robot, characterized in that, The robot is equipped with an action detection network, and the method includes: In response to a detection request, real-time images of the robot grasping the target object and the target cropping parameters and target weights of the motion detection network are obtained. The real-time image is cropped according to the target cropping parameters to obtain the target cropped image; The target cropped image is input into the action detection network for recognition to obtain the target recognition result, wherein the action detection network is a network configured according to the target weights; Based on the target recognition result, the robot is controlled to release the target item.
6. The method according to claim 5, characterized in that, The step of controlling the robot to release the target item based on the target recognition result includes: If the target recognition result is received within the first time period, and it is determined that the target recognition result indicates that a human hand feature has been identified, then a first instruction is generated, and the robot is controlled to release the target item according to the first instruction; If the target recognition result is not received within the first time period or the target recognition result indicates that no human hand feature was recognized, a second instruction is generated, and the robot is controlled to release the target item according to the second instruction.
7. The method according to claim 5, characterized in that, Before responding to the detection request and acquiring real-time images of the robot grasping the target object and the target cropping parameters and target weights of the motion detection network, the method further includes: If the robot executes a grasping command to grasp the target item, then the first position information of the target robotic arm of the robot executing the grasping command is determined; If the first location information is the same as the target location information, a detection request is generated.
8. A processing device for an action detection network, characterized in that, The processing device includes a transceiver module and a processing module; the transceiver module is used to perform the receiving / transmitting action in the method of any one of claims 1-4, and the processing module is used to perform other actions in the method of any one of claims 1-4 besides the receiving / transmitting action; Alternatively, the transceiver module is used to perform the receiving / transmitting action in the method of any one of claims 5-7, and the processing module is used to perform other actions in the method of any one of claims 5-7 besides the receiving / transmitting action.
9. A robot, characterized in that, The robot includes: processor; A memory, wherein computer-readable instructions are stored thereon, which, when executed by the processor, implement the method as described in any one of claims 1-4 or 5-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains program code that can be invoked by a processor to execute the method as described in any one of claims 1-4 or 5-7.