Mechanical arm training data generation method and training method, medium and mechanical arm system
By generating training data through multi-view video analysis and single-cycle video cutting, the problems of high cost of generating training data for the robotic arm and limited coverage of scenarios are solved, and the operational capability and robustness of the robotic arm in complex environments are improved.
Patent Information
- Application Number
- CN202510881530.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-06-27
AI Technical Summary
The existing method of generating training data for robotic arms is costly and the dataset coverage is limited, making it difficult to meet the robustness and generalization requirements of dexterous hand robots in complex environments.
The robotic arm’s grasping process is captured by multiple cameras to form a multi-view video, identify the start and end positions of the grasping cycle, cut the single-cycle video, analyze whether the grasping is successful or not, and generate training data.
The quality and accuracy of training data generation are improved, and the operational capability and robustness of the robotic arm in complex environments are enhanced.
Smart Images

Figure CN120735007A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a method for generating training data for a robotic arm, a training method, a medium, and a robotic arm system. Background Art
[0002] In recent years, demand for dexterous hand robots has been growing in applications across industrial automation, service robotics, and specialized operations. A key challenge lies in achieving the precise, robust, and adaptive manipulation capabilities of a human hand. Deep learning-based control strategies (such as reinforcement learning and imitation learning) have become a mainstream approach to improving the performance of dexterous hands. However, the effectiveness of these approaches is highly dependent on high-quality, large-scale, and diverse training data.
[0003] There are two existing methods for generating training data. One is to use real-machine data. However, real-machine experiments (such as grasping and manipulating objects) are often time-consuming, labor-intensive, and hardware-intensive (e.g., wear and tear, potential damage), resulting in extremely high data acquisition costs. Furthermore, the complexity of the physical world and experimental constraints (e.g., time and resources) mean that datasets obtained through real-machine experiments often have limited scene coverage, sparse state spaces, and uneven motion distribution. This directly limits the generalization ability of the trained model and its robustness in complex and unknown environments. The second option is to use simulated data. However, the authenticity of simulated data still fails to meet the training requirements for dexterous hand robots. Summary of the Invention
[0004] The purpose of this application is to provide a method, device, medium and equipment for processing on-board information of an environmentally friendly vehicle to solve at least one of the above-mentioned technical problems.
[0005] In a first aspect of the present application, a method for generating training data for a robotic arm is provided, the method comprising: Use multiple cameras to shoot the process of the robotic arm grasping the target object in the environment to form a multi-view video; Locating the starting position and ending position of each complete grasping cycle based on the grasping motion trajectory of the robotic arm; Cutting the multi-view captured video based on the starting position and the ending position to form one or more single-cycle captured videos, each single-cycle captured video corresponding to a complete capture cycle; Performing target object grasping analysis on the single-cycle grasping video to identify whether the robotic arm in the single-cycle grasping video successfully completes the target object grasping cycle; Training data for training the robotic arm is generated based on a single-cycle grasping video in which a target object grasping cycle is successfully completed.
[0006] Optionally, the cutting of the multi-view video based on the starting position and the ending position to form one or more single-cycle captured videos includes: Obtaining a start time corresponding to the start position and an end time corresponding to the end position; A video segment between a start time and an end time of a capture cycle is extracted from the multi-view video, and the video segment is used as the single-cycle capture video.
[0007] Optionally, performing target object grasping analysis on the single-cycle grasping video to identify whether the robotic arm in the single-cycle grasping video successfully completes the target object grasping cycle includes: Extract frames from each single-cycle captured video to form data slices; Identifying whether each frame of image in the data slice represents a successful grasping of a target object, and assigning a grasping value to each frame of image according to a grasping success recognition result of each frame of image; The weighted summation is performed based on the grasping assignment of each frame of the image to obtain the grasping success confidence of the data slice of a complete grasping cycle; Based on the grasping success confidence level, it is determined whether the robotic arm successfully completes a target object grasping cycle.
[0008] Optionally, identifying whether each frame of the image in the data slice indicates that the target object is successfully grasped includes: Identify the positional relationship between the robotic arm and the target object based on the sub-image at each shooting angle in each frame of the image; The image to be analyzed is used as the current frame image. When it is recognized based on the positional relationship that the robotic arm contacts the target object, whether there is a change between a first position of the target object appearing in the current frame image and a second position appearing in an adjacent frame image is analyzed. If there is a change, it is determined that the current frame image indicates that the target object has been successfully grasped. If there is no change, it is determined that the current frame image indicates that the object has failed to be grasped. When it is identified based on the position relationship that the robotic arm has not contacted the target object, it is identified whether the first position of the target object in the current frame image matches the preset position, and based on the matching between the first position and the preset position, it is determined whether the current frame image indicates that the object has been successfully grasped. The preset position is set based on the position of the current frame image in the data slice, sorted in chronological order according to timestamps.
[0009] Optionally, identifying the positional relationship between the robotic arm and the target object based on the sub-image at each shooting angle in each frame of the image includes: Identifying a first sub-region where the target object is located in a sub-image at a first shooting angle of view of the current frame image, and forming a mask image at the first shooting angle of view based on the first sub-region; Identifying a second sub-region where the robotic arm is located in the sub-image at the first shooting angle of view of the current frame image; When the first sub-region and the second sub-region are in contact with each other, identifying a third sub-region where the target object is located in the sub-image at the second shooting angle of view of the current frame image; The position of the third sub-area is transformed based on the perspective relationship between the first shooting perspective and the second shooting perspective to obtain the fourth sub-area corresponding to the third sub-area in the mask image, and the fourth sub-area is marked in the mask image. The first sub-area and the fourth sub-area are used as the first position of the target object in the current frame image.
[0010] Optionally, assigning a value to each frame of image according to a successful capture recognition result of each frame of image includes: assigning a value of 1 to an image indicating successful capture, and assigning a value of 0 to an image indicating failed capture.
[0011] Optionally, the weighted summation based on the captured values of each frame of image includes: setting a weight for each frame of image according to the timestamp of each frame of image in the data slice, and among the images in the data slice sorted in chronological order according to the timestamp, the later the order, the greater the corresponding weight.
[0012] Optionally, when a target item grasping cycle fails to complete, the type of the target item that failed to be grasped is identified; the type of the target item that failed to be grasped is used as the type to be trained, and the target grasping number of target items of the type to be trained is determined based on the number of samples of the existing training data, and the target grasping number of target items is re-performed for the target item.
[0013] In a second aspect of the present application, a method for training a robotic arm to grasp is provided, the method comprising: Obtain a training data set, where the training data in the training data set is training data generated by the robotic arm training data generation method described in any embodiment of the present application; The preset robotic arm grasping model is iteratively trained according to the training data set to form a trained robotic arm grasping model.
[0014] In a third aspect of the present application, a computer-readable storage medium is provided, on which executable instructions are stored. When the executable instructions are executed by a processor, the processor executes the robot arm training data generation method and / or the robot arm grasping training method as described in any embodiment of the present application.
[0015] In a fourth aspect of the present application, a robotic arm system is provided, which includes: a robot body; a robotic arm; a camera, configured on the robot body or the robotic arm, for capturing an image of the environment; one or more processors, for controlling the robotic arm to grasp an object according to a trained robotic arm grasping model generated by the robotic arm grasping training method in an embodiment of the present application, and / or A memory for storing one or more programs, which, when executed by the one or more processors, enables the one or more processors to execute the robotic arm training data generation method and / or the robotic arm grasping training method described in any embodiment of the present application.
[0016] The robotic arm training data generation method and training method, medium and robotic arm system in the present application control the robotic arm to grasp the target object, collect videos of the robotic arm's grasping process from multiple angles, and determine the starting position and ending position of a complete grasping cycle based on the robotic arm's grasping motion trajectory, and then cut each single-cycle grasping video from the multi-angle shooting video according to the starting position and ending position, thereby improving the accuracy of single-cycle grasping video cutting; a target object grasping success analysis is carried out on the single-cycle grasping video, and training data is formed based on the single-cycle grasping video in which the target object is successfully grasped, which can improve the generation quality of training data of the robotic arm's grasping model. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope of the present application.
[0018] Figure 1 1 is a flow chart of a method for generating training data for a robotic arm according to an embodiment; Figure 2 Schematic diagram of a robotic arm grasping a target object in one embodiment; Figure 3 A schematic diagram of capturing images from multiple perspectives according to an embodiment; Figure 4 is a schematic diagram of a multi-view image including a mask image in one embodiment; Figure 5 is a schematic diagram of a multi-view image including a mask image according to another embodiment; Figure 6 A flowchart of an embodiment of performing target object grasping analysis on a single-cycle grasping video to identify whether the robotic arm in the single-cycle grasping video successfully completes the target object grasping cycle; Figure 7The figure is a flow chart of identifying the positional relationship between the robot arm and the target object based on the sub-image at each shooting angle in each frame of image in one embodiment. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0020] All terms (including technical and scientific terms) used in this application have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0021] For example, the terms "first," "second," etc. used in this application may be used herein to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish a first element from another element.
[0022] For example, the terms "include", "comprising", etc. used in this application indicate the existence of features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.
[0023] This application provides a method for generating training data for a robotic arm, such as Figure 1 As shown, the method includes: Step 110 , using multiple cameras to shoot the process of the robotic arm grasping the target object in the environment to form a multi-view video.
[0024] In this embodiment, combined with Figure 2 As shown in FIG, a grabbing scene for a robot arm to grab an object is pre-arranged. The scene includes one or more target objects that need to be grabbed by the robot arm, and a placement area for placing the target objects. Figure 2 As shown, for example, multiple objects can be placed on an object placement table as target objects, and one or more placement areas can be set up on the placement table, such as placement area A and placement area B. The robotic system can control the robotic arm to grasp the target object and move the grasped object to placement area A or placement area B. Specifically, for example, one object can be grasped from placement area A as the target object and moved to placement area B. Then, another target object can be grasped from placement area B and moved to placement area A. By periodically grasping objects back and forth between the two placement areas, the efficiency of grasping the target objects can be improved.
[0025] Optionally, there can be multiple ways to arrange the grabbing scene for the robotic arm to grab the target object. For example, only one placement area can be set up to place the grabbed object. The robotic arm can be controlled to grab the target object from the starting point and move it to a unified placement area.
[0026] The robot system is also equipped with multiple cameras, each fixed relative to the other. These cameras can be fixed at different locations on a rigid component, such as a rigid fixture on the robot body, to provide the same viewing angle. Each camera simultaneously captures the robotic arm's grasping process, generating a multi-view video. This multi-view video includes sub-videos captured by each camera.
[0027] Since the positions of the cameras are different, the captured gripping angles of the robotic arm are also different, thus forming videos shot from multiple angles.
[0028] For example, the robot system is equipped with a hand camera and a head camera. The relative positions of the two cameras are fixed. During the movement of the robot arm, the two cameras record the movement of the robot arm in real time to form a multi-angle video. Figure 3 As shown, this is an example of one frame of an image in the multi-view video. The left image is a sub-image captured by the hand camera, and the right image is a sub-image captured by the head camera. These two sub-images constitute one frame of the multi-view video. It is understood that any other suitable method can also be used to set the number and position of cameras.
[0029] Step 120 : Locate the starting position and the ending position of each complete grasping cycle based on the grasping motion trajectory of the robotic arm.
[0030] Step 130 : cutting the multi-view video based on the starting position and the ending position to form one or more single-cycle captured videos.
[0031] In this embodiment, each single-cycle grabbing video corresponds to a complete grabbing cycle. A multi-view video may include one or more complete grabbing cycles of the robot arm. One complete grabbing cycle represents the process of the robot arm grabbing the target object from the starting position and placing it at the end position. For example Figure 2 As shown in , the process of grabbing a target object from placement area A and placing it in placement area B is a complete grabbing cycle. Similarly, the process of grabbing a target object from placement area B and placing it in placement area A is also a complete grabbing cycle.
[0032] The robot system can shoot videos from multiple perspectives and identify the positions of the robotic arm and / or target objects. Based on the changes in the moving position of the robotic arm and the position of the target object, the starting and ending times of a complete grasping cycle can be determined. The position of the robotic arm at the starting time is the starting position, and the position of the robotic arm at the ending time is the ending position.
[0033] The multi-view video is cut based on the identified start and end times (starting position and ending position) to form one or more single-cycle captured videos. The first frame image in the single-cycle captured video is the image at the corresponding start time (starting position), and the last frame image is the image at the corresponding ending time (ending position).
[0034] In one embodiment, step 130 includes: obtaining a start time corresponding to the start position and an end time corresponding to the end position; extracting a video segment between the start time and the end time in a capture cycle from the multi-view video, and using the video segment as a single-cycle capture video.
[0035] Among them, the starting position and the ending position can be pre-set fixed positions. For example, the starting position and the ending position of the robot arm during a complete grasping cycle can be pre-set, and the starting time when the robot arm is at the starting position and the ending time when the robot arm is at the ending position can be recorded.
[0036] Furthermore, the starting and ending positions do not need to be pre-set fixed positions. The robotic arm system can adaptively determine the starting and ending positions based on the robotic arm's grasping motion trajectory. When the robotic arm moves to the starting position, it indicates the beginning of a corresponding grasping process; when the robotic arm moves to the preset ending position, it indicates the end of the corresponding grasping process.
[0037] For example, a complete grasping cycle involves picking up a target object from placement area A and placing it in placement area B. The corresponding start time is when the target object is within placement area A and the robotic arm is at the preset starting position, before grasping the target object. The end time is when the target object is within placement area B and the robotic arm is at the preset ending position, after placing the object. Conversely, the start and end positions of a complete grasping cycle can also be set based on the process of picking up a target object from placement area B and placing it in placement area A.
[0038] For the multi-view video, identify the image frames shot at the starting time and the ending time, and cut the multi-view video according to the starting time and the ending time, so that the first frame image in the cut video segment is the image frame at the starting time, and the last frame image is the image at the ending time. Then the image segment is a single-cycle captured video.
[0039] In one embodiment, when the robotic arm is in the starting position, the robotic arm is in a corresponding specific first area in the corresponding image frame in the multi-view image. Similarly, when the robotic arm is in the ending position, it is also in a corresponding specific second area in the corresponding image frame. By identifying the area where the robotic arm is located for each frame of the multi-view video, the image frames in the first area and the second area can be located respectively, and cutting is performed based on the located image frames to form a corresponding single-cycle captured video.
[0040] The image of the robot arm being identified in the first area is used as the first frame image of the single-cycle captured video, and the image of the robot arm being identified in the second area for the first time after the frame image is used as the second frame image of the single-cycle captured video.
[0041] Some image frames in the multi-view video that do not contain a complete capture cycle are discarded.
[0042] Step 140 : performing target object grasping analysis on the single-cycle grasping video to identify whether the robotic arm in the single-cycle grasping video successfully completes the target object grasping cycle.
[0043] Step 150 : Generate training data for robot arm training based on the single-cycle grasping video of a target object grasping cycle successfully completed.
[0044] The robotic arm may not always be able to successfully complete the grabbing of the target object. Image recognition is performed on the single-cycle grabbing video to detect whether the image frames therein can reflect the successful grabbing and successful placement of the target object by the robotic arm. Based on this, it is determined whether the grabbing cycle is successfully completed. For example, when it is determined that the target object falls during the grabbing process, or is not correctly placed in the placement area, the grabbing is determined to have failed. When the object is successfully grabbed and successfully placed in the preset placement area, it is determined to have been completed successfully. Optionally, it is determined whether the target object is damaged during the grabbing process. If the target object is damaged, the grabbing is determined to have failed.
[0045] Among them, the position of the robotic arm, the placement area, and the damage / integrity of the item can all be comprehensively judged in combination with multi-angle images to improve the accuracy of the placement area judgment.
[0046] If the capture fails, the single-cycle captured video is discarded. If the capture succeeds, the single-cycle captured video can be used directly as training data, or it can be processed accordingly and the processed data used as training data. The processing process for the single-cycle captured video can include extracting frames from the video and using the resulting data slices as training data; or performing other processing such as robot arm posture analysis based on the video and combined with other relevant data, and using the resulting posture data as training data.
[0047] The robotic arm training data generation method in this application controls the robotic arm to grasp the target object, collects videos of the robotic arm's grasping process from multiple angles, and determines the starting position and ending position of a complete grasping cycle based on the robotic arm's grasping motion trajectory. Then, each single-cycle grasping video is cut out from the multi-view shooting video based on the starting position and ending position, thereby improving the accuracy of single-cycle grasping video cutting.
[0048] A successful target object grasping analysis is performed on the single-cycle grasping video, and training data is formed based on the single-cycle grasping video in which the target object is successfully grasped, which can improve the generation quality of training data of the robotic arm's grasping model.
[0049] In one embodiment, Figure 6 As shown, step 140 includes: Step 610: extract frames of each captured video in a single cycle to form data slices.
[0050] In this embodiment, each camera in the robotic arm system can capture video at a preset frame rate. The frame rate of each camera can be the same, for example, any suitable frame rate such as 5 FPS, 7.5 FPS, 10 FPS, 15 FPS, or 20 FPS. For a single-cycle captured video, images are uniformly extracted from the video based on the number of image frames required for capture and analysis. An appropriate number of image frames are extracted to form data slices, so that each data slice contains a consistent number of image frames.
[0051] For example, the shooting frame rate of a single-cycle captured video is 10FPS, and the time of each video cycle is 10 seconds. The robot system extracts images according to the preset extraction interval, such as extracting one frame every 10 frames, and extracts 10 frames of images as data slices.
[0052] Step 620 , identifying whether each frame of image in the data slice indicates that the target object is successfully grasped, and assigning a grasping value to each frame of image according to the grasping success recognition result of each frame of image.
[0053] In this embodiment, target object grasping and identification is performed on each frame image in the data slice to determine the target object grasping result presented by each frame image, to obtain whether each frame image successfully grasps the target object, and to assign a grasping value to it according to the identification result.
[0054] Optionally, the grasp success recognition result can be expressed as a grasp success or a grasp failure, or as a probability of grasp success. A larger value indicates a higher probability of grasp success. For example, an image indicating a grasp success can be assigned a value of 1, while an image indicating a grasp failure can be assigned a value of 0. Alternatively, the value can be assigned based on the probability of grasp success, with a higher value assigned for a higher probability.
[0055] Specifically, for each frame, the system can identify one or more of the following: whether the target object has been successfully grasped by the robotic arm, whether the target object has been grasped to the correct target location, and whether the target object remains intact before and after being grasped by the robotic arm. This results in a target object grasping and recognition result for the corresponding frame. This analysis can be performed independently on each frame based on images from multiple perspectives, or by combining the current frame (i.e., the image to be analyzed) with its preceding and following frames for correlation analysis.
[0056] Step 630 : Perform weighted summation based on the capture values of each frame of image to obtain the capture success confidence level of the data slice of a complete capture cycle.
[0057] Step 640: Determine whether the robotic arm successfully completes the target object grasping cycle based on the grasping success confidence level.
[0058] In this embodiment, the robotic arm system further sets the weight of each frame image in the data slice, and calculates the value, adds it to the corresponding weight, and the obtained value is used as the corresponding grasping success confidence.
[0059] The calculated confidence level is compared with a preset confidence threshold. When the confidence level exceeds the preset confidence threshold, it is determined that the robotic arm in the data slice has successfully completed the target object grasping cycle.
[0060] Furthermore, data slices identifying successful completion of a grasping cycle can be labeled, including labels for one or more dimensions, including target item information, grasping environment, and robotic arm motion characteristics. Target item information includes one or more of the following: target item identification, item category, item material, item geometry, and item state and position; grasping environment includes one or more of the following: illumination level and grasping background (e.g., black background, white background, or complex background lighting) in the current grasping environment; and robotic arm motion characteristics may include one or more of the following: picking, transforming, placing, and moving.
[0061] In one embodiment, identifying whether each frame image in the data slice indicates successful grasping of the target object includes: identifying the positional relationship between the robotic arm and the target object based on the sub-image at each shooting angle in each frame image, and judging whether each frame image indicates successful grasping of the target object based on the positional relationship.
[0062] In this embodiment, the positional relationship between the robotic arm and the target object can indicate whether the robotic arm is in contact with the target object. When the robotic arm is in contact with the target object, it indicates that the robotic arm needs to move the target object. When the robotic arm is not in contact with the target object, it indicates that the robotic arm is moving toward the target object to grasp the target object, or that the robotic arm has completed moving the target object to the placement area and placed the target object.
[0063] Optionally, under different shooting angles, the positions / areas of the robotic arm and the target object in the corresponding sub-images are different. The robotic arm system can perform image recognition on each sub-image in each frame of image to identify the robotic arm and the target object in each sub-image, and obtain the areas where the robotic arm and the target object are located in each sub-image. When the areas where the robotic arm and the target object are located intersect in each sub-image of a frame of image, or when it is recognized that the target object and the robotic arm are mutually blocked, it means that the corresponding robotic arm has been in contact with the target object. Figure 3 and Figure 4 / Figure 5 As shown, it can be identified Figure 3 In the two sub-images, the robotic arm and the target object are mutually occluded, which means Figure 3 The represented robot arm is in contact with the target object. Figure 4 / Figure 5 In the two sub-images, the robotic arm with fingers in the left sub-image does not touch the target object, while in the middle sub-image, the robotic arm occludes the target object (the area where the robotic arm is located in the sub-image and the area where the target object is located in the sub-image intersect), which means that the robotic arm does not touch the target object.
[0064] In one embodiment, Figure 7 As shown, the positional relationship between the robotic arm and the target object is identified based on the sub-image at each shooting angle in each frame of the image, including: Step 710 : Identify a first sub-region where the target object is located in the sub-image at the first shooting angle of view of the current frame image, and form a mask image at the first shooting angle of view based on the first sub-region.
[0065] In this embodiment, image conversion is performed on the sub-image under the first shooting angle to form a mask image of the sub-image. The mask image is used to reflect the area (i.e., the first sub-area) where the target object is located under the first shooting angle. The target object identified in the sub-image under the first shooting angle is displayed in the mask image with a first pixel value, and the non-target object is displayed in the mask image with a second pixel value. Figure 4 As shown, for example, for the target object in the identified middle sub-image, in the corresponding mask image ( Figure 4 The same location in the right sub-image (shown in the figure) is displayed in green, while other non-target areas are displayed in black. By performing mask conversion, the location of the target object can be intuitively determined.
[0066] Step 720 : Identify the second sub-region where the robotic arm is located in the sub-image at the first shooting angle of view of the current frame image.
[0067] Similar to the recognition of the target object, the robotic arm is also recognized in the sub-image to obtain the area where the robotic arm is located in the sub-image (ie, the second sub-area).
[0068] Step 730 : When the first sub-region and the second sub-region are in contact with each other, a third sub-region where the target object is located in the sub-image at the second shooting angle of view of the current frame image is identified.
[0069] After obtaining the first and second sub-areas, a check can be performed to determine whether they are in contact. If contact is present, this indicates that the robotic arm has touched the target object, or that the robotic arm has obstructed the target object. Further confirmation can be obtained by combining the sub-image from the second viewing angle. The location of the target object in the sub-image from the second viewing angle can then be identified, and based on this location, the third sub-area within the sub-image can be determined.
[0070] Specifically, it may be detected whether there is overlap between the edges of the first sub-region and the second sub-region. If there is overlap, it is determined that the two sub-regions are in contact with each other.
[0071] Step 740: Perform position transformation on the third sub-region based on the perspective relationship between the first shooting perspective and the second shooting perspective to obtain the fourth sub-region corresponding to the third sub-region in the mask image, mark the fourth sub-region in the mask image, and use the first sub-region and the fourth sub-region as the first position of the target object in the current frame image.
[0072] Combined with the perspective relationship between the first shooting perspective and the second shooting perspective, the position of the pixel points at each position in the sub-image in the second shooting perspective in the sub-image in the first shooting perspective can be obtained. Based on this positional relationship, the third sub-region is mapped to the mask image to obtain the fourth sub-region corresponding to the third sub-region in the mask image. The fourth sub-region and the first sub-region are merged, and the region obtained after the merger is the first position of the target object in the current frame image, and the first position is the position under the first shooting perspective. Figure 5 As shown in the figure, the green area in the right image is the first sub-region of the target object, and the orange area is the third sub-region of the target object identified in the left sub-image. This third sub-region is converted to the mask image and then merged with the first sub-region to form the supplementary area. The green and orange areas together constitute the first position of the target object.
[0073] By combining the sub-images from two perspectives to identify the location of the target object, the accuracy of the target object location recognition can be improved.
[0074] In one embodiment, when the first sub-region and the second sub-region are not in contact, the position of the first sub-region is directly used as the first position of the target object. In this case, there is no need to perform target object recognition on the sub-image under the second shooting angle.
[0075] In one embodiment, the image to be analyzed is used as the current frame image. When it is identified that the robotic arm contacts the target object based on the position relationship, it is analyzed whether there is a change between the first position of the target object appearing in the current frame image and the second position appearing in the adjacent frame image. When there is a change, it is determined that the current frame image indicates that the target object has been successfully grasped. When there is no change, it is determined that the current frame image indicates that the object has failed to be grasped.
[0076] In this embodiment, the adjacent frame image can be any one of the previous and subsequent frames of the current frame image. The position of the target object identified in the current frame image is used as the first position, and the position of the target object identified in the adjacent frame image of the current frame image is used as the second position. The first position and the second position are positions at the same shooting angle. By comparing the two positions, it is determined whether the positions have changed. If a change has occurred, it indicates that the robot arm has moved the target object, and thus it can be determined that the object in the current frame image has been successfully grasped. If no change has occurred, it means that the robot arm has contacted the target object but failed to move it, and the object grasping has failed.
[0077] Furthermore, when it is identified based on the position relationship that the robotic arm has not touched the target object, it is also analyzed whether there is a change between the first position of the target object appearing in the current frame image and the second position appearing in the adjacent frame image. When there is a change, it is determined that the current frame image indicates that the target object has been successfully grasped. When there is no change, it is determined that the current frame image indicates that the object has failed to be grasped.
[0078] In this embodiment, position change recognition is performed for all images in the data slice. If there is a position change, it is determined that the capture is successful. If there is no position change, it is determined that the capture fails, which simplifies the method of successful capture recognition.
[0079] In one embodiment, when it is identified based on the position relationship that the robotic arm has not contacted the target object, it is identified whether the first position of the target object in the current frame image matches the preset position, and based on the matching between the first position and the preset position, it is determined whether the current frame image indicates that the object has been successfully grasped. The preset position is set based on the position of the current frame image in the data slice, sorted in chronological order according to timestamps.
[0080] In this embodiment, the preset position may include the initial placement position and the target placement position of the target object. The initial placement position is the position of the target object before the robotic arm grasps the target object; the target placement position is the target position to which the robotic arm will move the target object. The robotic arm system may determine whether the preset position is the initial placement position or the target placement position based on the position of the current frame image in the data slice. For example, the preset position corresponding to the first frame image in the data slice is the initial placement position, and the preset position corresponding to the last frame image in the data slice is the target placement position.
[0081] By comparing the identified first position with the preset position, if the two are consistent, it means that the robotic arm has successfully grasped the target object; if they are inconsistent, it means that the target object has failed to be grasped.
[0082] In one embodiment, weighted summation is performed based on the captured values of each frame of image, including: setting a weight for each frame of image according to the timestamp of each frame of image in the data slice, and among the images in the data slice that are sorted in order according to the timestamp, the later the order, the greater the corresponding weight.
[0083] In this embodiment, the later an image is ranked, the more it can reflect the capture success of the corresponding capture cycle, and therefore a larger value is assigned to it.
[0084] Optionally, you can set the weight of the i-th frame image in the data slice , where α is the weight attenuation factor, and its value can be between 0 and 1, and is set according to the actual situation. The assignment of the i-th frame image , when the i-th frame image indicates that the capture is successful, then ri=1, when it indicates that the capture fails, then r i =0.
[0085] Based on the assignment and weights set above, the corresponding confidence K= , n represents the total number of frames in the data slice.
[0086] For example, with n=5 and α=0.8, the values of each frame image form the value vector r=[0, 0, 1, 1,1], and the weight vector of each frame image is = [0.4096,0.512,0.64,0.8,1], the corresponding confidence level K=72.6% can be calculated.
[0087] In one embodiment, when a target item grasping cycle fails to complete, the type of the target item that failed to be grasped is identified; the type of the target item that failed to be grasped is used as the type to be trained, and the target grasping number of target items of the type to be trained is determined based on the number of samples of the existing training data, and the target grasping number of target items is re-performed.
[0088] In this embodiment, for grasping cycles in which grasping fails, the type of target item corresponding to the failed grasp can be recorded. Based on the pre-planned training data volume required for each item type, the number of grasps for the failed target item is determined. The robotic arm is then controlled to grasp the corresponding target item according to the required number of grasps, thereby generating the required sample size. By allowing the robotic arm to autonomously collect grasping data for target items with poor performance, and then autonomously allocating training data according to pre-set label screening logic, and determining the number of grasps for target items with each label, the quality of the samples used for subsequent model training can be improved.
[0089] In one embodiment, a robotic arm grasping training method is also proposed, the method comprising: obtaining a training data set; iteratively training a preset robotic arm grasping model according to the training data set to form a trained robotic arm grasping model.
[0090] The training data set is a training data set generated by the method for generating training data for a robotic arm according to the present application. The training data in the training data set are provided with corresponding data labels, and the robotic arm grasping model can be a VLA (visual language action) model.
[0091] The training data generated by the above-mentioned method for generating training data for the robotic arm is labeled to form a training dataset. The pre-trained robotic arm grasping model is iteratively trained a preset number of times using this training dataset. When the number of iterative training reaches a pre-determined threshold, or the accuracy of the trained model reaches a preset accuracy, the iterations are terminated, resulting in a trained robotic arm grasping model.
[0092] In one embodiment, a computer-readable storage medium is provided, on which executable instructions are stored. When the instructions are executed by a processor, the processor executes the steps in the above-mentioned method embodiments.
[0093] In one embodiment, a robotic arm system is also provided, which includes: a robot body; a robotic arm; a camera, configured on the robot body or the robotic arm, for capturing environmental images; and one or more processors, for controlling the robotic arm to grasp objects based on a trained robotic arm grasping model generated by the robotic arm grasping training method in this application.
[0094] In one embodiment, the robotic arm system further includes a memory for storing one or more programs, which, when executed by the one or more processors, enables the one or more processors to execute the method described in any one of the embodiments of the present application.
[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.
[0096] Furthermore, those skilled in the art will appreciate that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are intended to be within the scope of this application and to form different embodiments. For example, all of the above embodiments may be used in any combination. The information disclosed in this background section is intended solely to enhance understanding of the overall background of this application and should not be construed as an admission or any form of implication that such information constitutes prior art known to those skilled in the art.
Claims
1. A method for generating training data for a robotic arm, characterized in that: The method comprises: Use multiple cameras to shoot the process of the robotic arm grasping the target object in the environment to form a multi-view video; Locating the starting position and ending position of each complete grasping cycle based on the grasping motion trajectory of the robotic arm; Cutting the multi-view captured video based on the starting position and the ending position to form one or more single-cycle captured videos, each single-cycle captured video corresponding to a complete capture cycle; Performing target object grasping analysis on the single-cycle grasping video to identify whether the robotic arm in the single-cycle grasping video successfully completes the target object grasping cycle; Training data for training the robotic arm is generated based on a single-cycle grasping video in which a target object grasping cycle is successfully completed.
2. The method according to claim 1, characterized in that The step of cutting the multi-view shooting video based on the starting position and the ending position to form one or more single-cycle captured videos includes: Obtaining a start time corresponding to the start position and an end time corresponding to the end position; A video segment between a start time and an end time of a capture cycle is extracted from the multi-view video, and the video segment is used as the single-cycle capture video.
3. The method according to claim 1, characterized in that The performing target object grasping analysis on the single-cycle grasping video to identify whether the robotic arm in the single-cycle grasping video successfully completes the target object grasping cycle includes: Extract frames from each single-cycle captured video to form data slices; Identifying whether each frame of image in the data slice represents a successful grasping of a target object, and assigning a grasping value to each frame of image according to a grasping success recognition result of each frame of image; The weighted summation is performed based on the grasping assignment of each frame of the image to obtain the grasping success confidence of the data slice of a complete grasping cycle; Based on the grasping success confidence level, it is determined whether the robotic arm successfully completes a target object grasping cycle.
4. The method according to claim 3, characterized in that The identifying whether each frame of the image in the data slice indicates that the target object is successfully grasped includes: Identify the positional relationship between the robotic arm and the target object based on the sub-image at each shooting angle in each frame of the image; The image to be analyzed is used as the current frame image. When it is recognized based on the positional relationship that the robotic arm contacts the target object, whether there is a change between a first position of the target object appearing in the current frame image and a second position appearing in an adjacent frame image is analyzed. If there is a change, it is determined that the current frame image indicates that the target object has been successfully grasped. If there is no change, it is determined that the current frame image indicates that the object has failed to be grasped. When it is identified based on the position relationship that the robotic arm has not contacted the target object, it is identified whether the first position of the target object in the current frame image matches the preset position, and based on the matching between the first position and the preset position, it is determined whether the current frame image indicates that the object has been successfully grasped. The preset position is set based on the position of the current frame image in the data slice, sorted in chronological order according to timestamps.
5. The method according to claim 4, characterized in that The identifying the positional relationship between the robotic arm and the target object based on the sub-image at each shooting angle in each frame of the image includes: Identifying a first sub-region where the target object is located in a sub-image at a first shooting angle of view of the current frame image, and forming a mask image at the first shooting angle of view based on the first sub-region; Identifying a second sub-region where the robotic arm is located in the sub-image at the first shooting angle of view of the current frame image; When the first sub-region and the second sub-region are in contact with each other, identifying a third sub-region where the target object is located in the sub-image at the second shooting angle of view of the current frame image; The position of the third sub-area is transformed based on the perspective relationship between the first shooting perspective and the second shooting perspective to obtain the fourth sub-area corresponding to the third sub-area in the mask image, and the fourth sub-area is marked in the mask image. The first sub-area and the fourth sub-area are used as the first position of the target object in the current frame image.
6. The method according to claim 3, characterized in that The grabbing and assigning values to each frame of image according to the successful grabbing recognition result of each frame of image includes: assigning a value of 1 to an image indicating successful grabbing and assigning a value of 0 to an image indicating failed grabbing; The weighted summation based on the captured values of each frame of image includes: setting a weight for each frame of image according to the timestamp of each frame of image in the data slice, and among the images in the data slice sorted according to the timestamp, the later the order is, the greater the corresponding weight is.
7. The method according to any one of claims 1 to 6, characterized in that When the target object grasping cycle fails to complete, identifying the type of the target object that failed to be grasped; The type of target object that failed to be grasped is used as the type to be trained, and the target number of grasping times of the target object of the type to be trained is determined based on the number of samples of the existing training data, and the target object is grasped again for the target number of grasping times.
8. A robotic arm grasping training method, characterized in that: The method comprises: Acquire a training data set, wherein the training data in the training data set is training data generated by the method according to any one of claims 1 to 7; The preset robotic arm grasping model is iteratively trained according to the training data set to form a trained robotic arm grasping model.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores executable instructions, which, when executed by a processor, enable the processor to perform the method according to any one of claims 1 to 8.
10. A robotic arm system, characterized in that: include: Robot body; robotic arm; A camera, configured on the robot body or the robotic arm, for capturing images of the environment; One or more processors, configured to control the robotic arm to grasp an object according to the trained robotic arm grasping model generated according to claim 8, and / or A memory for storing one or more programs, which, when executed by the one or more processors, causes the one or more processors to perform the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Training method of model for cross-view gait feature extraction
CN118506454A
Robot data generation method and device, equipment and storage medium
CN119217370A
Bin picking performance evaluation device and bin picking performance evaluation method
JP2014240110A
Robotic grasping prediction method based on triplet contrastive network
WO2024087331A1