Grabbing pose determination method and device and pose estimation model training method and device

Through the two-stage network model processing image and depth images, the robotic arm grab position is determined, which solves the problems of large amount of point cloud data and information loss, and realizes efficient pose estimation on hardware with limited computing power.

CN120287296APending Publication Date: 2025-07-11SHENZHEN SWEET POTATO ROBOT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510559868.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the grab position estimation method of the robot arm is based on point cloud data, with large calculation volume and low inference efficiency, making it difficult to operate efficiently on hardware with limited computing power, and information is seriously lost after point cloud data is quantized, affecting the effect.

Method used

The pose estimation model composed of a two-stage network is adopted, and the image and depth image are processed through the first network to determine the position information of the grab center point, and the grabable area is obtained using the sampling layer. Finally, the grab pose is determined by the second network to avoid the use of point cloud data, reduce the calculation amount and improve the hardware resource utilization rate.

Benefits of technology

Efficient operation on hardware with limited computing power improves the computing efficiency and real-time performance of grabbing pose estimation, reduces computing resources, and improves the hardware resource utilization rate and accuracy of pose estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120287296A_ABST
    Figure CN120287296A_ABST
Patent Text Reader

Abstract

The invention discloses a grabbing pose determination method and device and a pose estimation model training method and device, and relates to the field of mechanical arm control. The method comprises the following steps: determining a first image comprising an object to be grabbed and a depth image corresponding to the first image; processing the first image and the depth image based on a first network in a pose estimation model to obtain position information of a grabbing center point of the to-be-grabbed object in the first image; on the basis of a sampling layer in the pose estimation model, coordinates in the position information of the grabbing center point are processed, and a grabbable area comprising the grabbing center point is obtained; and the position information of the grabbable area and the grabbing center point is processed based on a second network in the pose estimation model, and the grabbing pose for grabbing the to-be-grabbed object is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of intelligent robotic arm grasping, and particularly to a method and device for determining a grasping pose, and a method and device for training a pose estimation model. Background Art

[0002] With the development of industrial automation and intelligence, robotic arms have been widely used in various fields such as manufacturing, logistics, and the service industry. In embodied intelligence applications, in order to enable a robotic arm to accurately complete a grasping task, it is necessary to predict / estimate the grasping pose (e.g., 6-DOF (6 degrees of freedom) pose) corresponding to the grasping task in real time when the robotic arm is working.

[0003] In related technologies, the estimation method of the grasping pose of a robotic arm usually processes point cloud data based on algorithms of the type of pointnet++, etc., so as to complete the estimation of the grasping pose. However, due to the fact that point cloud data is sparse, irregular, and unordered data, and the point cloud data has a large amount of noise, the algorithm for pose estimation based on point cloud data has a large computational amount and low inference efficiency, and it is difficult to operate efficiently when deployed on hardware with limited computing power. Summary of the Invention

[0004] In order to solve the above technical problems, the present disclosure provides a method and device for determining a grasping pose, and a method and device for training a pose estimation model, so as to improve the computational efficiency of grasping pose estimation, improve the inference efficiency, and enable efficient operation when deployed on hardware with limited computing power.

[0005] In the first aspect of the present disclosure, a method for determining a grasping pose is provided, including: determining a first image including an object to be grasped and a depth image corresponding to the first image; processing the first image and the depth image based on a first network in a pose estimation model to obtain position information of a grasping center point of the object to be grasped in the first image; processing coordinates in the position information of the grasping center point based on a sampling layer in the pose estimation model to obtain a graspable area including the grasping center point; processing the graspable area and the position information of the grasping center point based on a second network in the pose estimation model to obtain a grasping pose for grasping the object to be grasped.

[0006] In the second aspect of the present disclosure, a method for training a pose estimation model is provided, including:

[0007] determining multiple groups of sample image data and sample grasping center point data corresponding to the sample image data; the sample image data includes a sample image and a sample depth image corresponding to the sample image; the sample grasping center point data includes position information of a sample grasping center point of a sample grasping object in the sample image and a sample grasping pose;

[0008] Process the sample image and the sample depth image based on the first network in the initial pose estimation model to obtain the position information of the training grasping center point of the sample grasping object in the sample image;

[0009] Based on the sampling layer in the initial pose estimation model, process the coordinates in the position information of the training grasping center point to obtain the training graspable area including the training grasping center point in the sample image;

[0010] Process the training graspable area and the position information of the training grasping center point based on the second network in the initial pose estimation model to obtain the training grasping pose for grasping the sample grasping object; use the training grasping pose as the initial training output of the initial pose estimation model, and use the sample grasping center point data as the supervision information to iteratively train the initial pose estimation model to obtain the trained pose estimation model.

[0011] An embodiment of the third aspect of the present disclosure provides a grasping pose determination device, including: an acquisition module for determining a first image including an object to be grasped and a depth image corresponding to the first image; a processing module for processing the first image and the depth image acquired by the acquisition module based on the first network in the pose estimation model to obtain the position information of the grasping center point of the object to be grasped in the first image; the processing module is further configured to process the coordinates in the position information of the grasping center point based on the sampling layer in the pose estimation model to obtain the graspable area including the grasping center point in the first image; the processing module is further configured to process the graspable area and the position information of the grasping center point based on the second network in the pose estimation model to obtain the grasping pose for grasping the object to be grasped.

[0012] An embodiment of the fourth aspect of the present disclosure provides a pose estimation model training device, including: an acquisition module for determining multiple groups of sample image data and sample grasping center point data corresponding to the sample image data; the sample image data includes a sample image and a sample depth image corresponding to the sample image; the sample grasping center point data includes the position information of the sample grasping center point of the sample grasping object in the sample image and the sample grasping pose;

[0013] A training module for processing the sample image and the sample depth image acquired by the acquisition module based on the first network in the initial pose estimation model to obtain the position information of the training grasping center point of the sample grasping object in the sample image;

[0014] The training module is further configured to process the coordinates in the position information of the training grasping center point based on the sampling layer in the initial pose estimation model to obtain the training graspable area including the training grasping center point in the sample image;

[0015] The training module is further configured to process the position information of the training graspable region and the training grasp center point based on the second network in the initial pose estimation model to obtain the training grasp pose for the grasping sample to grasp the object.

[0016] The training module is further configured to use the training grasp pose as the initial training output of the initial pose estimation model, and the sample grasp center point data as the supervision information, and iteratively train the initial pose estimation model to obtain the trained pose estimation model.

[0017] An embodiment of the fourth aspect of the present disclosure provides a computer-readable storage medium storing a computer program for executing the method provided in any of the above aspects.

[0018] An embodiment of the fifth aspect of the present disclosure provides an electronic device, which includes: a processor; a memory for storing executable instructions executable by the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the method provided in any of the above aspects.

[0019] Based on the grasp pose determination method provided in the embodiments of the present disclosure, the image (or color space information) and depth image (such as RGB-D information) of the object to be grasped are processed in two stages by a pose estimation model composed of a two-stage network to obtain the grasp pose of the object to be grasped. First, the first network in the pose estimation model processes the image and depth image including the object to be grasped to determine the position information of the grasp center point on the object to be grasped. Then, the sampling layer in the pose estimation model processes the center point position information to obtain a graspable region including the center of the grasp region. Finally, the second network in the pose estimation model processes the position information of the grasp center point and the graspable region to determine the grasp pose for grasping the object to be grasped. It can be seen that in the present disclosure, by performing pose estimation on the two-dimensional image and the depth information corresponding to the two-dimensional image, the use of three-dimensional data of the point cloud is avoided, thereby reducing the model calculation amount and providing the inference efficiency, so that it can run efficiently when deployed on hardware with limited computing power; secondly, through the staged processing of pose estimation by the first network and the second network, the computing resources are further reduced and it can be efficiently deployed on hardware with limited computing power; in addition, the first network and the second network can be implemented by models with independent structures, so that the first network and the second network are set on different hardware, improving the utilization rate of hardware resources; in addition, the steps executed by the first network and the second network for grasping pose estimation are independent of each other. For example, when the first network processes the current frame data, the second network processes the previous frame data, thereby further improving the processing efficiency of the pose estimation model and the real-time performance of the grasp pose estimation. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a schematic diagram of the principle of the grasping pose determination method provided by an exemplary embodiment of the present disclosure.

[0021] Figure 2 It is a schematic diagram of the structure of the pose estimation system provided by an exemplary embodiment of the present disclosure.

[0022] Figure 3 It is a schematic diagram of the structure of the electronic device provided by an exemplary embodiment of the present disclosure.

[0023] Figure 4 It is a schematic diagram of the structure of the training device provided by an exemplary embodiment of the present disclosure.

[0024] Figure 5 It is a flowchart of the grasping pose determination method provided by an exemplary embodiment of the present disclosure Figure 1 。

[0025] Figure 6 It is a schematic diagram of the principle of the grasping pose determination method provided by an exemplary embodiment of the present disclosure Figure 2 。

[0026] Figure 7 It is a flowchart of the grasping pose determination method provided by an exemplary embodiment of the present disclosure Figure 2 。

[0027] Figure 8 It is a schematic diagram of the principle of the first network in the pose estimation model provided by an exemplary embodiment of the present disclosure Figure 1 。

[0028] Figure 9 It is a flowchart of the grasping pose determination method provided by an exemplary embodiment of the present disclosure Figure 3 。

[0029] Figure 10 It is a schematic diagram of the principle of the first network in the pose estimation model provided by an exemplary embodiment of the present disclosure Figure 2 。

[0030] Figure 11 It is a flowchart of the grasping pose determination method provided by an exemplary embodiment of the present disclosure Figure 4 。

[0031] Figure 12 It is a schematic diagram of the principle of the first network in the pose estimation model provided by an exemplary embodiment of the present disclosure Figure 3 。

[0032] Figure 13 It is a flowchart of the grasping pose determination method provided by an exemplary embodiment of the present disclosure Figure 5 。

[0033] Figure 14It is a schematic diagram of the principle of the second network in the pose estimation model provided by an exemplary embodiment of the present disclosure.

[0034] Figure 15 It is a schematic flowchart of the pose estimation model training method provided by an exemplary embodiment of the present disclosure Figure 1 。

[0035] Figure 16 It is a schematic flowchart of the pose estimation model training method provided by an exemplary embodiment of the present disclosure Figure 2 。

[0036] Figure 17 It is a schematic flowchart of the pose estimation model training method provided by an exemplary embodiment of the present disclosure Figure 3 。

[0037] Figure 18 It is a schematic flowchart of the pose estimation model training method provided by an exemplary embodiment of the present disclosure Figure 4 。

[0038] Figure 19 It is a schematic diagram of the structure of the grasping pose determination device provided by an exemplary embodiment of the present disclosure Figure 1 。

[0039] Figure 20 It is a schematic diagram of the structure of the grasping pose determination device provided by an exemplary embodiment of the present disclosure Figure 2 。

[0040] Figure 21 It is a schematic diagram of the structure of the grasping pose determination device provided by an exemplary embodiment of the present disclosure Figure 3 。

[0041] Figure 22 It is a schematic diagram of the structure of the pose estimation model training device provided by an exemplary embodiment of the present disclosure Figure 1 。

[0042] Figure 23 It is a schematic diagram of the structure of the pose estimation model training device provided by an exemplary embodiment of the present disclosure Figure 2 。 Detailed implementation manners

[0043] To explain the present disclosure, exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. It should be understood that the present disclosure is not limited by the exemplary embodiments.

[0044] Hereinafter, the terms "first" and "second" are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this application, unless otherwise specified, the meaning of "a plurality" is two or more. "A and / or B" includes the following three combinations: only A, only B, and the combination of A and B.

[0045] It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present disclosure.

[0046] Application Overview

[0047] In embodied intelligence applications, robotic arms often need to perform grasping tasks, and grasping tasks require the prior completion of grasping detection, that is, the determination / prediction / estimation of the grasping pose (such as 6-DOF (6 degrees of freedom) pose) of the object to be grasped.

[0048] In related technologies, the estimation method of the grasping pose of a robotic arm usually processes the point cloud data obtained by sensors on the main body of the robotic arm based on algorithms of types such as the pointnet++ framework to complete the estimation of the grasping pose.

[0049] However, due to the unstructured and sparse data structure of the point cloud, the calculation is relatively complex, the calculation efficiency is not high, the real-time performance is low, and the requirement for computing power resources is relatively high. Therefore, the calculation amount of the grasping pose estimation algorithm in related technologies is large, the inference efficiency is low, and it is difficult to run in real time and efficiently on edge chips (such as neural network processing units (NPUs)) with high real-time requirements and limited computing power. In addition, when deploying the algorithm on the edge chip, in order to improve the calculation efficiency, it is usually necessary to quantize the algorithm model; however, since the point cloud data is of floating-point type (float) and each point carries a lot of information (such as coordinate and color space information, etc.), when quantizing this algorithm to the edge chip, the point cloud data needs to be converted from floating-point type to integer type, which will cause a lot of information loss, which will make the effect of the algorithm model after quantization deployment worse or even difficult to achieve the expected purpose. Therefore, when the grasping pose estimation algorithm in related technologies is quantized and deployed on the edge chip, the calculation efficiency is low and the effect is poor, which is not conducive to engineering deployment and use.

[0050] In view of the above technical problems, embodiments of the present disclosure provide a method for determining a grasping pose. Refer to Figure 1 , in the present disclosure, a pose estimation model composed of a two-stage network is used to perform two-stage processing on the image (or color space information) of the object to be grasped and the depth image (such as RGB-D information), so as to obtain the grasping pose of the object to be grasped. Specifically, first, the first network in the pose estimation model processes the image and depth image including the object to be grasped to determine the position information of the grasping center point on the object to be grasped. Then, the sampling layer in the pose estimation model processes the center point position information to obtain a graspable area including the center of the grasping area. Finally, the second network in the pose estimation model determines the grasping pose of the object to be grasped based on the position information of the grasping center point and the graspable area.

[0051] It can be seen that in the present disclosure, by performing pose estimation on the two-dimensional image and the depth information corresponding to the two-dimensional image, the three-dimensional data of the point cloud is avoided, thereby reducing the model calculation amount and providing the inference efficiency, so that it can run efficiently when deployed on hardware with limited computing power; secondly, through the staged processing of pose estimation by the first network and the second network, the computing resources are further reduced and it can be efficiently deployed on hardware with limited computing power; in addition, the first network and the second network can be implemented by models with independent structures, so that the first network and the second network are set on different hardware, improving the utilization rate of hardware resources; in addition, the steps executed by the first network and the second network during the grasping pose estimation are independent of each other. For example, when the first network processes the current frame data, the second network processes the previous frame data, thereby further improving the processing efficiency of the pose estimation model and the real-time performance of the grasping pose estimation.

[0052] Exemplary system

[0053] The method for determining a grasping pose and the method for training a pose estimation model provided by embodiments of the present disclosure can be applied to a Figure 2 pose estimation system as shown. Refer to Figure 2 As shown, in this pose estimation system, there may be an electronic device 01 and a training device 02. Among them, the training device 02 is mainly used to obtain sample data and train a pose estimation model, that is, to implement the method for generating a pose estimation model provided by embodiments of the present disclosure. The electronic device 01 is mainly used to process the image and depth image including the object to be grasped based on the pose estimation model trained by the training device, so as to determine the grasping pose of the robotic arm to grasp the object to be grasped, that is, to implement the method for determining a grasping pose provided by embodiments of the present disclosure. In addition, if the processing resources of the electronic device are sufficient, the electronic device can also implement the method for generating a pose estimation model provided by embodiments of the present disclosure, and the present disclosure does not make specific limitations on this.

[0054] It can be understood that the above-mentioned electronic device 02 and training device 01 can be two separate devices or the same device. The present application does not make specific restrictions on this.

[0055] Exemplarily, with reference to Figure 2 as shown, the electronic device 01 may include one or more processors 201 and a memory 202.

[0056] The processor 201 may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities (such as NPU, etc.), and may control other components in the electronic device 200 to perform desired functions.

[0057] The memory 202 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 201 may run the program instructions to implement the grasping pose determination method of each embodiment provided in the present disclosure and / or other desired functions.

[0058] In one example, the electronic device 01 may further include: an input device 203 and an output device 204, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0059] In the embodiments of the present disclosure, the input device may be used to input or acquire an image of an object to be grasped and the corresponding depth information. For example, the input device may be an RGB-D vision sensor (or called a depth vision sensor).

[0060] In the embodiments of the present disclosure, the output device may be used to output the grasping pose obtained after the electronic device executes the grasping pose determination method to the robotic arm. The grasping pose may be a 6-degree-of-freedom pose including three-dimensional translation data and three-dimensional rotation data. In some embodiments, the output device may be a control device in the robotic arm configured on the electronic device, and in this case, the output device may control the robotic arm to complete the grasping of the object to be grasped based on the grasping pose determined by the electronic device.

[0061] Of course, for simplicity, Figure 2Only some of the components related to this application in the electronic device 200 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device 200 may further include any other appropriate components.

[0062] In an embodiment of the present disclosure, the training device 02 may be a server, which may be a single server, a server cluster composed of multiple servers, or a cloud computing service center, and the present disclosure does not make specific limitations thereto.

[0063] Exemplarily, taking the training device as a server as an example, Figure 3 A schematic structural diagram of a server is shown. Referring to Figure 3 As shown, the server includes one or more processors 301, a communication line 302, and at least one communication interface ( Figure 3 only exemplified by including a communication interface 306 and one processor 301 for illustration). Optionally, the server may further include a memory 304.

[0064] The processor 301 may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the solution of the present application.

[0065] The communication line 302 may include a communication bus for communication between different components.

[0066] The communication interface 306 may be a transceiver module for communicating with other devices (such as the electronic device 01) or a communication network, such as Ethernet, RAN, wireless local area networks (WLAN), etc. For example, the transceiver module may be a device such as a transceiver or a transceiver. Optionally, the communication interface 306 may also be a transceiver circuit located within the processor 301 for realizing the signal input and signal output of the processor.

[0067] The memory 304 can be a device with storage functions. The memory 304 can include one or more computer program products, and the computer program products can include various forms of computer-readable storage media. For example, it can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or it can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but not limited to this.

[0068] The memory 304 can exist independently and be connected to the processor through the communication line 302. The memory can also be integrated with the processor.

[0069] Among them, one or more computer program instructions can be stored on the computer-readable storage media included in the memory 304, and the processor 301 or 307 can run the computer program instructions to implement the pose estimation model training method of each embodiment provided by the present disclosure and / or other desired functions. Optionally, the computer execution instructions in the embodiments of the present application can also be referred to as application code, and the embodiments of the present application do not make specific limitations on this.

[0070] Alternatively, optionally, in the embodiments of the present disclosure, it can also be the processor 301 that executes the functions related to the processing in the pose estimation model training method provided by the embodiments of the present disclosure.

[0071] In a specific implementation, as an embodiment, the processor 301 can include one or more CPUs, such as Figure 3 CPU0 and CPU1 in

[0072] In a specific implementation, as an embodiment, the server can include multiple processors, such as Figure 3The processors 301 and 307 therein. Each of these processors can be a single-core processor or a multi-core processor. The processors here can include but are not limited to at least one of the following: CPU, microprocessor, digital signal processing (DSP), microcontroller unit (MCU), NPU, or various computing devices running software such as artificial intelligence processors. Each computing device can include one or more cores for executing software instructions to perform operations or processing.

[0073] In a specific implementation, as an embodiment, the server may further include an output device 305 and an input device 306. The output device 305 communicates with the processor 301 and can display information in various ways. For example, the output device 305 can be a liquid crystal display (LCD), a light emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector, etc. The input device 306 communicates with the processor 301 and can receive user input in various ways, such as sample data and supervision information (or label data) required for training the pose estimation model. For example, the input device 306 can be a mouse, a keyboard, a touch screen device, or a sensing device, etc.

[0074] The above server can be a general-purpose device or a dedicated device. For example, the server can be a desktop computer, a laptop computer, a network server, a personal digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, an embedded device, the above terminal devices, the above network devices, or a device with a Figure 3 similar structure therein. The embodiments of the present application do not limit the type of the server.

[0075] Exemplary Method

[0076] The method for determining the grasping pose and the method for training the pose estimation model provided by the embodiments of the present disclosure will be introduced below with reference to the accompanying drawings.

[0077] Figure 5 is a schematic flowchart of the method for determining the grasping pose provided by an exemplary embodiment of the present disclosure. This method can be applied to the electronic device provided in the foregoing embodiment, such as Figure 5 shown, this method may include S501 - S504:

[0078] S501. Determine a first image including an object to be grasped and a depth image corresponding to the first image.

[0079] In the embodiments of the present disclosure, the object to be grasped can be any object that can be grasped or operated by a robotic arm, such as an express package, a device on an assembly line, an automobile charging interface, etc. In addition, the object to be grasped can be one or more according to different usage scenarios of the method. For example, in the assembly line scenario, the object to be grasped can be all the devices on the assembly line in the image captured by the electronic device that executes the grasping pose determination method. Another example is that in the vehicle charging scenario, the object to be grasped can be the charging interface on the vehicle charging device. The present disclosure does not make specific limitations in this regard.

[0080] In the embodiments of the present disclosure, the first image can be a two-dimensional color image, such as an RGB (red, green, blue color space) image, a YUV (luminance, chrominance) image, etc. The depth image is composed of depth information corresponding one-to-one to the pixels in the first image. The depth information refers to the distance between each pixel point in the first image and the image acquisition device (such as a camera).

[0081] In the embodiments of the present disclosure, the first image and the corresponding depth image can be acquired by a depth vision sensor configured on the electronic device; alternatively, the first image is acquired by a vision sensor configured on the electronic device, and the depth image is acquired by a depth sensor with the same shooting angle as the vision sensor configured on the electronic device. Exemplarily, the depth sensor can be any one of the following: a structured light sensor, a Time-of-Flight (TOF) sensor, an Active Stereo sensor, etc. The depth vision sensor is usually an integrated device of an ordinary camera and a depth sensor, or a sensor that can simultaneously complete depth acquisition and image acquisition, such as a Stereo Vision sensor, an RGB-D sensor, etc.

[0082] In a possible implementation, taking an example that an electronic device is configured with a depth vision sensor, the first image and the depth image can be two corresponding but independent images obtained after the depth vision sensor captures the scene within its shooting angle. For example, the depth vision sensor configured on the electronic device can be an RGB-D sensor, which can capture a depth image corresponding to the RGB image while capturing the RGB image. In another possible implementation, the target image captured by the depth vision sensor can include RGB information and depth information. For example, the depth vision sensor configured on the electronic device can be an active stereo vision sensor, which can calculate the depth while capturing color information, so as to obtain RGB information and the corresponding depth information, and the sensor can fuse the RGB information and the depth information into four-channel image data and then output it. In this implementation, the first image and the depth image mentioned in the embodiments of the present disclosure can be obtained based on the target image captured by the depth vision sensor. For example, the RGB information in the target image is extracted to generate the first image, and the depth information in the target image is extracted to generate the depth image corresponding to the first image.

[0083] Of course, the above-mentioned methods for obtaining the first image and the depth image are only examples, and in practice, other arbitrary possible implementation methods can also be used to obtain the first image of the object to be grasped and the corresponding depth image.

[0084] When the first image and the depth image are two corresponding images, after obtaining the first image and the depth image corresponding to the first image, both are input into the pose estimation model for processing. Generally, the first image usually includes three-channel two-dimensional data, and the depth image usually includes single-channel two-dimensional data. For the convenience of model processing, the first image and its corresponding depth image can be combined into a four-dimensional tensor and then input into the pose estimation model. The specific combination method can be any possible implementation method, and the present disclosure does not make specific limitations on this.

[0085] Due to the defects of the acquisition device for capturing the first image and the depth image on the electronic device or the influence of the environment, the following one or more problems may occur: there may be noise in both the first image and the depth image, there may be partial information loss in the depth image, there may be a spatial offset (partial pixels are not aligned) between the first image and the depth image, etc. For these problems, it is necessary to perform data preprocessing on the first image and the depth image before generating the four-dimensional tensor. For example, denoising the first image and the depth image, filling in missing values for the depth image, and aligning the pixels of the first image and the depth image through the calibration parameters of the acquisition device. The specific implementation method adopted for data preprocessing can be any possible implementation method, and the present disclosure does not make specific limitations on this.

[0086] In addition, to facilitate the processing of the finally generated four-dimensional tensor, the pixel values of the pixel points in the first image and the depth image can be normalized before merging to generate the four-dimensional tensor. For example, the pixel values of the pixel points in the first image and the depth image are both adjusted to values in the range [0, 1]. The specific normalization method can be any possible implementation method such as mean standard deviation normalization, range scaling normalization, maximum minimum normalization, global contrast normalization, etc., and the embodiments of the present disclosure do not make specific limitations thereto.

[0087] Of course, the data preprocessing before generating the four-dimensional tensor may include any other possible preprocessing methods in addition to the above, and the embodiments of the present disclosure do not make specific limitations thereto.

[0088] S502. Process the first image and the depth image based on the first network in the pose estimation model to obtain the position information of the grasping center point of the object to be grasped in the first image.

[0089] In the embodiments of the present disclosure, the pose estimation model is used to implement the estimation of the grasping pose, and the pose estimation model includes a first network, a second network, and a sampling layer located between the first network and the second network. Among them, the first network and the second network can be implemented by a neural network model.

[0090] In a possible implementation manner, to facilitate the processing of the first image and the depth image, the first network can be a two-dimensional convolutional neural network (2DCNN) model that can process two-dimensional images. In another possible implementation manner, the first network can be a network composed of a segmentation model, a grasping center point determination algorithm, and a depth information determination algorithm. Among them, the segmentation model can be used to determine the mask image of the object to be grasped, and the grasping center point determination algorithm and the depth information are algorithms jointly used to determine the position information of the grasping center point. The specific implementation of these two implementation manners is described in the subsequent embodiments, and will not be elaborated here.

[0091] In the embodiments of the present disclosure, the position information includes the coordinate information of the grasping center point and the depth information of the grasping center point. Among them, the coordinates of the grasping center point can be the coordinates of the grasping center point in the first image or the coordinates of the grasping center point in the camera coordinate system. It can be determined according to the actual situation, and the present disclosure does not make specific limitations thereto. The depth information of the grasping center can be obtained based on the depth image.

[0092] In a specific embodiment, the first network processes the first image and the depth image, including the first network performing feature extraction and feature fusion processing on the first image and the depth image, or the first network performing feature extraction on the first image and fusing the depth image and the result after feature extraction to obtain the position information of the grasping center point.

[0093] In the embodiments of the present disclosure, the grasping center point refers to the projection point of the center of the clamping device of the robotic arm on the two-dimensional image.

[0094] S503. Based on the sampling layer in the pose estimation model, process the coordinates in the position information of the grasping center point to obtain a graspable area including the grasping center point.

[0095] Combined Figure 1 , referring to Figure 6 As shown, between the first network and the second network in the pose estimation model, there may also be a sampling layer. After the first network obtains the position information of the grasping center point, the sampling layer can process the coordinates in the position information of the grasping center point in the first image to obtain the neighborhood of the grasping center point, that is, the graspable area including the grasping center point. In the embodiments of the present disclosure, the graspable area includes the coordinates of the graspable area and the depth information of the graspable area. Among them, the graspable area can be defined based on the shape of the object to be grasped. For example, the graspable area is circular, square, or diamond-shaped, etc. In a specific implementation manner, the range of the graspable area can be determined based on the coordinates of the center point of the object to be grasped and the shape of the object to be grasped, and further the coordinates and depth information of the graspable area can be determined.

[0096] In some embodiments, the sampling layer can also be regarded as a layer in the second network, and the present disclosure does not make specific limitations on this, which is specifically determined according to actual needs.

[0097] S504. Based on the second network in the pose estimation model, process the graspable area and the position information of the grasping center point to obtain the grasping pose for grasping the object to be grasped. Exemplarily, the second network can be a 2D CNN model that can process two-dimensional images.

[0098] After the graspable area is determined, the second network in the pose estimation model can accurately predict the pose of the grasping center point and the neighborhood of the grasping center point (i.e., the graspable area), so as to determine the grasping pose for grasping the object to be grasped. Exemplarily, the grasping pose can specifically be a 6-DOF pose, and the 6-DOF pose can specifically include three-dimensional translation and three-dimensional rotation.

[0099] In the embodiments of the present disclosure, the grasping center point of the object to be grasped can be one or more, and the grasping pose can also be one or more. When there are multiple grasping poses, to improve the accuracy of the object to be grasped, when the pose estimation model estimates the grasping pose, it can output the grasping pose and the confidence or accuracy of each grasping pose. When determining the accurate grasping pose, the electronic device determines different grasping poses with spatial overlap and pose overlap among at least one grasping pose of the same object to be grasped. Among them, spatial overlap specifically refers to the spatial distance (such as Euclidean distance) between two grasping poses being less than a preset distance. Pose overlap can refer to the difference in the rotation angle of two grasping poses rotating around the same axis being less than a preset angle value.

[0100] Then, according to the confidence level from high to low, different grasping poses are screened in turn until a preset number (such as two) of grasping poses are retained. Among them, the screening rule includes: eliminating other grasping poses that have spatial overlap and pose overlap with this grasping pose. For example, if the grasping poses include: pose 1, pose 2, pose 3, pose 4, and pose 5. The confidence levels of these five grasping poses from high to low are: pose 1 > pose 3 > pose 4 > pose 2 > pose 5. Among them, pose 2 has spatial overlap with pose 1, and pose 4 has pose overlap with pose 1. If the preset number is 2, then first eliminate pose 2 and pose 4, and the remaining poses are pose 1, pose 3, and pose 5. Then, retain the top two poses with the highest confidence levels, namely pose 1 and pose 3, and eliminate pose 5. Finally, use pose 1 as the main grasping pose for controlling the robotic arm, and pose 3 as the backup grasping pose for controlling the robotic arm. That is, when the electronic device controls the robotic arm, it preferentially uses pose 1 to grasp the object to be grasped, and then uses pose 3 to grasp the object to be grasped to ensure the accuracy of the object to be grasped. It can be seen that determining a more effective grasping pose from at least one grasping pose can ensure the effectiveness of the final grasping pose used to control the robotic arm and ensure that the robotic arm can effectively perform the grasping task.

[0101] The above method of screening grasping poses can be called non-maximum suppression (NMS). Its specific implementation method can also be any other possible implementation method, and the embodiments of the present disclosure do not make specific limitations on this.

[0102] In some embodiments, in order to ensure that the robotic arm does not collide with surrounding objects when performing the grasping pose, after obtaining at least one grasping pose, collision detection needs to be performed to eliminate the grasping poses that will cause the robotic arm to collide with surrounding objects among the at least one grasping pose. In a possible implementation, when performing collision detection, the grasping pose can be first converted into the robotic arm coordinate system, and then based on the motion rules of the robotic arm, the trajectories of the relevant components of the robotic arm (such as robotic arm links, grippers, etc.) from the current position to the grasping pose are determined. Then, based on the positions of the objects in the first image in the robotic arm coordinate system and the trajectories of the relevant components of the robotic arm, it is determined whether there will be a collision when controlling the robotic arm to perform the grasping task according to the grasping pose. If there is a collision, the grasping pose is eliminated. After performing this collision detection on all at least one grasping pose, a grasping pose with higher safety can be determined. Based on this collision detection, the electronic device can determine a safer grasping pose from at least one grasping pose, ensuring the safety of the grasping pose finally used to control the robotic arm, and improving the safety during the process of the electronic device controlling the robotic arm to perform the grasping task. Of course, the above specific implementation of collision detection is only an example, and in practice, it can also be any other possible implementation, and the present disclosure does not make specific limitations on this.

[0103] In practice, only one of the above NMS and collision detection can be executed, or both can be executed, depending on actual needs, and the present disclosure does not make specific limitations on this. Combining Figure 1 , referring to Figure 6As shown, the above-mentioned NMS and / or collision detection can also be referred to as post-processing. In addition to NMS and collision detection, post-processing in practice can also include any other possible processing flows that can improve the effectiveness and safety of the final output of the grasping pose. The present disclosure does not make specific limitations in this regard. Based on the technical solution provided by the embodiments of the present disclosure, the pose estimation model composed of a two-stage network is used to perform two-stage processing on the image (or color space information) and depth image (such as RGB-D information) of the object to be grasped, and the grasping pose of the object to be grasped is obtained. First, the first network in the pose estimation model processes the image and depth image including the object to be grasped to determine the position information of the grasping center point on the object to be grasped. Then, the sampling layer in the pose estimation model processes the center point position information to obtain a graspable area including the center of the grasping area. Finally, the second network in the pose estimation model processes the position information of the grasping center point and the graspable area to determine the grasping pose of grasping the object to be grasped. It can be seen that in the present disclosure, by performing pose estimation on the two-dimensional image and the depth information corresponding to the two-dimensional image, the use of three-dimensional data of point clouds is avoided, thereby reducing the model calculation amount and providing the inference efficiency, so that it can run efficiently when deployed on hardware with limited computing power; secondly, through the staged processing of pose estimation by the first network and the second network, the computing resources are further reduced and it is efficiently deployed on hardware with limited computing power; in addition, the first network and the second network can be implemented by models with independent structures, so that the first network and the second network are set on different hardware to improve the utilization rate of hardware resources; in addition, the first network and the second network are independent of each other when performing steps for grasping pose estimation. For example, when the first network processes the current frame data, the second network processes the previous frame data, thereby further improving the processing efficiency of the pose estimation model and the real-time performance of grasping pose estimation.

[0104] In some embodiments, the first network in the embodiments of the present disclosure can be a 2D CNN model with a small amount of computation suitable for processing two-dimensional images. Specifically, the 2D CNN model determines the grasping pose based on the first image and the corresponding depth image, and both of these images are two-dimensional images. Therefore, using the 2D CNN model for pose estimation can reduce the computation amount and improve the computing efficiency.

[0105] Combined Figure 5 , referring Figure 7 As shown, the above-mentioned S502 can specifically include S5021-S5023:

[0106] S5021. Based on the first feature processing layer in the first network, perform feature processing on the first image and the depth image to obtain a first feature map.

[0107] Exemplarily, referring Figure 8As shown, the first network may include a first feature processing layer for extracting features, an upsampling layer for determining a grasping point heatmap based on a feature map (i.e., the first feature map), and an algorithm layer for determining the position information of the grasping center point based on the grasping point heatmap. Among them, the first feature processing layer can be implemented by the convolutional layer and the pooling layer in the first network.

[0108] During the feature processing of the CNN model, generally, the underlying features -> middle-level features -> high-level features (semantic features) in the image are gradually extracted. The first feature map in the embodiments of the present disclosure can be obtained by fusing a low-level feature map for characterizing the underlying features and a high-level feature map for characterizing the high-level features, so as to obtain richer feature information.

[0109] S5022. Perform upsampling processing on the first feature map based on the upsampling layer in the first network to obtain a grasping point heatmap corresponding to the first image.

[0110] Since the first feature map is obtained through feature extraction including downsampling, its scale (or can be called the control resolution) is different from that of the original first image. Therefore, in order to obtain a grasping point heatmap with the same scale as the original first image, so as to more accurately select the grasping center point. Refer to Figure 8 As shown, it is necessary to perform downsampling processing on the first feature map so that the finally obtained grasping point heatmap has the same scale as the first image.

[0111] In the embodiments of the present disclosure, the upsampling processing can be specifically implemented by any possible processing operation methods such as deconvolution, bilinear interpolation, nearest neighbor interpolation, etc. The resolution of the feature map can be increased through upsampling. In addition, after the upsampling processing is performed on the first feature map in the present disclosure, probability prediction processing is performed on the upsampled feature map, so that the probability value of each pixel point in the feature map as the grasping center point can be obtained. Then, the upsampling layer will also perform pixel value conversion on the probability value of each pixel point as the grasping center point to obtain a grasping point heatmap.

[0112] In the embodiments of the present disclosure, the probability prediction can be specifically any possible processing operation such as probability normalization (completed through an activation function), and the present disclosure does not make specific limitations on this. The pixel points in the grasping point heatmap (heatmap) correspond one by one to the pixel points in the first image. The pixel value of the pixel point in the grasping point heatmap is used to represent the probability value of being the grasping center point. For example, if the pixel value of pixel point A in the grasping point heatmap is 255, it means that the probability of pixel point B corresponding to pixel point A in the first image being the grasping center point is 100%. If the pixel value of pixel point C in the grasping point heatmap is 0, it means that the probability of pixel point D corresponding to pixel point C in the first image being the grasping center point is 0%.

[0113] In some embodiments, while the upsampling layer of the first network processes the first feature map to obtain the grasping point heat map, it can also obtain partial grasping poses when the pixel points in the grasping point heat map are used as the grasping center points. Since the main purpose of the first network is to output a two-dimensional grasping point heat map and thus obtain the grasping center points, when the first network outputs the grasping poses, it can only output partial grasping poses obtained by projecting the six-degree-of-freedom poses onto a two-dimensional plane. Exemplarily, this partial grasping pose can be a two-degree-of-freedom pose, such as the translation amounts along the X-axis and Y-axis in three-dimensional space; or, this partial grasping pose can also be a three-degree-of-freedom pose (such as the translation amounts along the X-axis and Y-axis in three-dimensional space, and the rotation amount around the Z-axis).

[0114] S5023. Based on the grasping point heat map and the depth image, determine the coordinates and depth information in the position information of the grasping center point of the object to be grasped in the first image.

[0115] In the embodiments of the present disclosure, S5023 can specifically be completed by the algorithm layer in the first network mentioned in the foregoing embodiments.

[0116] After the first network obtains the grasping point heat map, based on the probability values represented by the pixel values of each pixel point in the grasping point heat map, select the pixel points whose probability values rank in the top preset order (or top preset percentage) among the pixel points corresponding to the object to be grasped in the first image as the grasping center points. Then, determine the depth information represented by the pixel value of the pixel point in the depth image that has the same position as the grasping center point in the first image as the depth information of the grasping center point. In this way, the position information of the grasping center point is obtained.

[0117] In the embodiments of the present disclosure, through the multi-layer structure (feature extraction layer, upsampling) of the first network in the pose estimation model, the first image is subjected to feature extraction to obtain a feature map, and further based on the feature map, a grasping point heat map is determined, and then in combination with the depth image, the coordinates and depth information in the position information of the grasping center point of the object to be grasped are determined. Since the first network can directly process the image through a two-dimensional convolutional neural network, and the two-dimensional convolutional neural network model has a simple structure, low computational complexity, and avoids using three-dimensional data of point clouds, when determining the position information of the grasping center point through the first network, the model calculation amount is reduced, the calculation efficiency and inference efficiency are improved, so that the model can run efficiently when deployed on hardware with limited computing power.

[0118] In some embodiments, in combination with Figure 8 , with reference to Figure 10As shown, the first feature processing layer in the first network may include a first feature extraction layer that extracts a low-level feature map (i.e., the third feature map) and a high-level feature map (i.e., the second feature map), and a feature fusion layer that fuses the low-level feature map and the high-level feature map. Based on this, combined with Figure 7 , referring to Figure 9 as shown, S5021 may specifically include S901 and S902:

[0119] S901. Based on the first feature extraction layer in the first feature processing layer, perform feature extraction processing on the first image and the depth image to obtain a second feature map and a third feature map with different scales.

[0120] In some embodiments, the first feature extraction layer may specifically be a backbone model or any other possible model.

[0121] In a possible implementation, the first feature extraction layer may include multiple sub-layers for extracting feature maps of different scales. Based on this, S901 may specifically include: based on the first sub-layer in the first feature extraction layer, perform feature extraction processing on the first image and the depth image to obtain a second feature map of the first scale; based on the second sub-layer in the first feature extraction layer, perform feature extraction processing on the second feature map to obtain a third feature map of the second scale. Among them, the first scale is greater than the second scale. For example, the first scale may be 200×200, and the second scale may be 100×100. The extraction size can be set based on actual needs.

[0122] In the embodiments of the present disclosure, by hierarchically and gradually extracting features of different levels in the image through the first feature extraction layer, second feature maps and third feature maps of different scales can be obtained, thereby improving the accuracy of feature extraction.

[0123] S902. Based on the feature fusion layer in the first feature processing layer, perform fusion processing on the second feature map and the third feature map to obtain a first feature map.

[0124] In some embodiments, the feature fusion layer may specifically be a feature pyramid network (FPN) or any other possible network architecture, and the present disclosure does not make specific limitations on this.

[0125] In an embodiment of the present disclosure, feature extraction at different levels is completed through a first feature extraction layer in a first feature processing layer in a first network to obtain feature maps at different levels. Then, through a feature fusion layer, fusion of the feature maps at different levels is completed to obtain a first feature map that can represent multi-level features. In this way, a more accurate heat map of the extracted feature points can be generated based on the first feature map subsequently, and then a more accurate grasping center point can be determined, and further an accurate grasping pose can be determined based on the accurate grasping center point. Thus, the electronic device can complete accurate control of the robotic arm.

[0126] In some other embodiments, the first network in the embodiments of the present disclosure may also be composed of a specific segmentation model, a grasping center point determination algorithm for determining a grasping center point based on a mask image, and a depth information determination algorithm. In this case, after determining the detection frame of the object to be grasped, the mask image of the object to be grasped can be determined through the first network, and then the position information of the grasping center point can be determined based on the mask image. Based on this, in combination with Figure 5 and referring to Figure 11 as shown, before S502, S503 may be included, and S502 may include S1101 and S1102.

[0127] S1101, determine the detection frame of the object to be grasped in the first image.

[0128] Exemplarily, the detection frame may be a bounding box that includes the object to be grasped.

[0129] In the embodiments of the present disclosure, there are at least two possible implementation manners for determining the detection frame of the object to be grasped in the first image:

[0130] The first implementation manner: In combination with Figure 1 and referring to Figure 12 as shown, after obtaining the first image and the corresponding depth image, the first image can be detected and processed based on a semantic detection model to obtain the detection frame of the object to be grasped in the first image. Exemplarily, the semantic detection model may be any detection model or detection algorithm that can detect a specific object, such as the YOLO (you only look once) model, the SSD (single shot multibox detector) model, the DETR (detection transformer) model, etc. The present disclosure does not make specific limitations thereto. Based on the first implementation manner, when the category of the object to be grasped is known, the detection frame of the object to be grasped in the first image can be accurately obtained by using the semantic detection model.

[0131] The second implementation manner, referring to Figure 12As shown, after obtaining the first image and the corresponding depth image, the detection frame of the object to be grasped can be determined by means of human-computer interaction box selection. Specifically, the first image can be displayed first, and then in response to the user's box selection operation on the object to be grasped in the first image, the detection frame of the object to be grasped in the first image can be obtained. Among them, the box selection operation can be any possible operation for determining the detection of the object to be grasped. For example, the box selection operation can be a sliding operation of drawing a box including the object to be grasped on the touch screen. For another example, the box selection operation can also be a click operation on the object to be grasped in the first image. In response to this click operation, the electronic device can generate a detection frame of the object to be grasped corresponding to the click operation through a specific algorithm. Based on the second implementation method, the method of human-computer interaction box selection can be adopted to allow the user to independently determine the object to be grasped that needs to be grasped by the robotic arm and its detection frame, avoiding selecting the object to be grasped that is not the grasping target, and improving the user experience at the same time.

[0132] Based on the above two implementation methods, the detection frame of the object to be grasped can be determined by using a semantic detection model, or the user can select the object to be grasped that needs to be grasped and the corresponding detection frame through human-computer interaction, and then the mask image of the object to be grasped can be determined based on the detection frame, and then the accurate grasping center point can be determined.

[0133] S1102. Segment the object to be grasped in the detection frame based on the segmentation model in the first network to obtain the mask image of the object to be grasped.

[0134] Since there will be parts that do not belong to the object to be grasped in addition to the object to be grasped in the detection frame, in order to determine a more accurate grasping center point. After obtaining the detection frame of the object to be grasped, the object to be grasped in the detection frame can be segmented based on the segmentation model in the first network, so as to obtain the mask image of the object to be grasped.

[0135] Exemplarily, the segmentation model in the first network can be any feasible image segmentation model, such as SAM (Segment Anything Model), Fast-SCNN (Fast Semantic Segmentation Network), etc. The embodiments of the present disclosure do not make specific limitations on this.

[0136] S1103. Based on the mask image of the object to be grasped, determine the coordinates in the position information of the grasping center point of the object to be grasped, and determine the depth information in the position information of the grasping center point based on the depth image.

[0137] In the embodiments of the present disclosure, after obtaining the mask image of the object to be grasped, the centroid of the object to be grasped or multiple points within a certain area where the centroid is located in the mask image of the object to be grasped can be determined as the grasping center point. Further, the depth information in the position information of the grasping center can be determined from the depth image according to the coordinates of the grasping center point.

[0138] In the embodiments of the present disclosure, the part of determining the coordinates in the position information of the grasping center point in S1103 can be implemented by the grasping center point determination algorithm and the depth information determination algorithm in the first network.

[0139] Based on the above-mentioned disclosed embodiments, the electronic device can first determine the detection frame, then determine the mask image of the object to be grasped through the segmentation model in the first network, and further obtain the coordinates and depth information of the grasping center point. When determining the grasping center point in this way, existing mature and stable models or algorithms that are convenient to deploy on edge devices can be used to obtain the detection frame and mask image of the object to be grasped, and then determine the accurate grasping center point and its position information. The second network can then determine the accurate grasping pose based on the position information of the accurate grasping center point. At the same time, since the methods of determining the detection frame, mask image, and position information of the grasping center point are calculated step by step based on the two-dimensional image, the amount of calculation required for each step is small and the calculation efficiency is high. Furthermore, the amount of calculation for pose estimation is reduced, and the calculation efficiency and inference efficiency of the pose estimation model are improved, so that the model can run efficiently when deployed on hardware with limited computing power.

[0140] In some embodiments, for the electronic device to determine the graspable area, it can specifically determine the image of the graspable area and the depth information of the pixel points in the image. Based on this, in combination with Figure 5 , referring to Figure 3 shown, S503 can specifically include S5031 and S5032:

[0141] S5031. Based on the sampling layer, using the coordinates in the position information of the grasping center point, determine the image of the graspable area including the grasping center point from the first image.

[0142] In the embodiments of the present disclosure, the coordinates in the position information of the grasping center point can specifically be the pixel coordinates of the grasping center point in the first image. After the first network obtains the position information of the grasping center point, through the sampling layer, based on the coordinates in the position information, a preset range including this coordinate in the first image is determined as the graspable area, and the image within this preset range is also the image of the graspable area. The image of the graspable area can be a local feature map in the first image.

[0143] Exemplarily, the preset range may be a square area, a circular area, or an area of other shapes with a preset side length centered on the coordinates of the grasping center point. For example, the preset range is a square area of 10×10 centered on the coordinates of the grasping center point.

[0144] S5032. Based on the sampling layer, determine the depth information of the pixel points in the image of the graspable area from the depth image.

[0145] After obtaining the graspable area, according to the corresponding relationship between the pixel positions in the first image and the depth image, determine the depth information of the pixel points in the image of the graspable area from the depth image.

[0146] In the embodiments of the present disclosure, the sampling layer can sample in the first image to obtain the relevant information of the graspable area that is the neighborhood of the grasping center point, so that the second network can obtain an accurate grasping pose based on the relevant information of the graspable area. Since the graspable area obtained by the sampling layer is closely related to the grasping center point, the grasping center point is closely related to the grasping pose, and the number of pixel points in the grasping area is very small, this can enable the second network to use less computational effort and higher computational efficiency to complete the determination of the grasping pose. Furthermore, it enables the model to run efficiently when deployed on hardware with limited computing power.

[0147] In some embodiments, in addition to determining the graspable area from the first image, it can also be selected from the target feature map extracted by the first feature extraction layer in the first network. The target feature map is a middle-level feature map that can represent both some low-level features and some high-level features. Since the features represented by the target feature map are rich in levels and the scale is smaller than that of the first image, the graspable area selected based on the target feature map can better reflect the position characteristics of the grasping center. To obtain the target feature map, in combination with the relevant description after S901 in the foregoing embodiments, the second sublayer in the first feature extraction layer may include a third sublayer and a fourth sublayer. Among them, the third sublayer can be used to perform feature extraction processing on the second feature map to obtain the target feature map, and the fourth sublayer can be used to perform feature extraction processing on the target feature map to obtain the third feature map. In this way, the first feature extraction layer can generate three feature maps with increasing levels, and the target feature map among them can be a middle-level feature map that can represent both some low-level features and some high-level features.

[0148] Based on this, when S503 is executed, the sampling layer can determine the image of the graspable area in the target feature map generated by the first network based on the coordinates of the grasp center point. Among them, since the coordinates in the position information of the grasp center point are the coordinates of the grasp center point in the first image, when determining the graspable area from the target feature map based on the coordinates of the grasp center point, the coordinates of the grasp center point can be converted into the coordinates in the target feature map based on the scale ratio between the first image and the target feature map, and then the graspable area is selected. For example, if the scale of the first image is 400×400, the scale of the target feature map is 100×100, and the coordinates of the grasp center point are (200, 200), the converted coordinates should be (50, 50). Then, based on the coordinates (50, 50), the graspable area centered on (50, 50) and its image can be determined from the target feature map.

[0149] After that, the sampling layer can determine the depth information of the pixel points in the image of the graspable area from the depth image based on the scale ratio between the target feature map and the depth image. For example, if the scale of the depth image is 400×400, the scale of the target feature map is 100×100, and the coordinates of pixel A in the image of the graspable area are (1, 2), the depth information of the pixel point with coordinates (4, 8) in the depth image can be used as the depth information of pixel A.

[0150] In the embodiments of the present disclosure, the sampling layer can sample in the target feature map to obtain the relevant information of the graspable area that is the neighborhood of the grasp center point. Since the feature levels of the target feature map are richer, the graspable area obtained based on the target feature map can better reflect the position characteristics of the grasp center, and thus the accuracy of the grasp pose obtained based on the graspable area can be improved.

[0151] In some embodiments, the second network in the embodiments of the present disclosure can be a 2D CNN model with a small amount of computation suitable for processing two-dimensional images. Specifically, when the second network determines the grasp pose, it is completed based on the graspable area, and the graspable area is a two-dimensional image. Therefore, using a 2D CNN model for pose estimation can reduce the amount of computation and improve the computational efficiency.

[0152] Based on this, in combination with Figure 5 , with reference to Figure 13 shown, S504 can specifically include S5041 and S5042:

[0153] S5041. Based on the second feature processing layer in the second network, perform feature extraction processing on the graspable area and the position information of the grasp center point to obtain fourth feature maps and fifth feature maps with different scales.

[0154] The second network may include a second feature processing layer for extracting features and a pose prediction layer for determining the grasping pose based on the feature map. Among them, the second feature processing layer can be implemented by the convolutional layer and the pooling layer in the second network.

[0155] In addition, in order to make the prediction of the grasping pose more accurate, the second network can determine the grasping pose based on the low-level feature map that can reflect the underlying features and the high-level feature map that can reflect the high-level features. Refer to Figure 14 As shown, through the second feature processing layer in the first network, the fourth feature map and the fifth feature map with different scales can be obtained. Of course, in practice, in order to make the subsequent determined grasping pose more accurate, the second feature processing layer can also obtain more feature maps with different scales, not limited to the fourth feature map and the fifth feature map shown here.

[0156] In some embodiments, the first feature extraction layer can specifically be a backbone model or any other possible model.

[0157] In a possible implementation manner, the second feature extraction layer can include multiple extraction layers for extracting feature maps with different scales. Based on this, S5041 can specifically include: based on the first extraction layer in the second feature processing layer, performing feature extraction processing on the position information of the graspable area and the grasp center point to obtain the fourth feature map of the third scale; based on the second extraction layer in the second feature processing layer, performing feature extraction processing on the fourth feature map to obtain the fifth feature map of the fourth scale.

[0158] Among them, the third scale is larger than the fourth scale. For example, the first scale can be 20×20, and the second scale can be 4×4.

[0159] In this way, through the hierarchical step-by-step extraction in the second feature extraction layer, the fourth feature map and the fifth feature map with different scales can be obtained, that is, the feature maps with different levels can be obtained. Subsequently, the accurate grasping pose can be predicted by using the feature maps with different scales, and then the electronic device can accurately control the robotic arm.

[0160] S5042: Based on the pose prediction layer in the second network, perform prediction processing on the fourth feature map and the fifth feature map to obtain the grasping pose for grasping the object to be grasped.

[0161] In the embodiments of the present disclosure, the pose prediction layer performs prediction processing on the fourth feature map and the fifth feature map, including: fusing the fourth feature map and the fifth feature map, and performing pose grasping prediction based on the fused image to obtain the grasping pose. Among them, the pose prediction layer includes an FPN (Feature Pyramid Network), which fuses feature information of different scales through the FPN, thereby avoiding the problem of information loss when processing multi-scale targets through a convolutional neural network. The pose prediction layer further includes a pose prediction model for predicting the grasping pose based on the fused features. The pose prediction model predicts the position information of the graspable region and the grasping center point in the fused image, and obtains the grasping pose for the robotic arm to perform grasping based on the position information from the current pose to the grasping center point. The current pose herein refers to the pose of the robotic arm before grasping based on the predicted grasping pose.

[0162] It should be noted that the network structure in the pose prediction layer in the present disclosure is not specifically limited, and the fusion of feature maps and pose prediction can be implemented based on existing algorithms or models.

[0163] In some embodiments, referring to the relevant description after the aforementioned S5022, if the first network outputs the grasping pose of each pixel point in the grasping point heat map as the grasping center point when outputting the grasping point heat map, then the partial grasping pose corresponding to the grasping center point finally determined by the first network can be obtained. Further, in order to make the grasping pose obtained by the pose prediction layer of the second network more accurate, when predicting the grasping pose, it can use the partial grasping pose corresponding to the grasping center point obtained by the first network as prior information, and jointly complete the prediction processing in combination with the fourth feature map and the fifth feature map.

[0164] Of course, in practice, any possible data can be used as prior information for the pose prediction layer of the second network to improve the accuracy of the predicted grasping pose. In this way, the second network can use the prior information as verification information or prediction basis in the process of predicting the grasping pose, so that the second network can predict a more accurate grasping pose.

[0165] Based on the above-mentioned disclosed embodiments, after obtaining the information of the relevant neighborhood (graspable area) based on the position information of the grasping center point obtained from the first network, feature extraction is completed through the second feature processing layer in the second network of the pose estimation model to obtain feature maps at multiple levels. Then, the pose prediction layer in the second network performs prediction processing on the feature maps at different levels to obtain the grasping pose corresponding to the grasping center point. Since this second network can directly process images through a two-dimensional convolutional neural network, and the two-dimensional convolutional neural network has a simple model structure, low computational complexity, and avoids using three-dimensional data of point clouds, when determining the grasping pose through the second network, the model calculation amount is reduced, the computational efficiency and inference efficiency are improved, so that the model can run efficiently when deployed on hardware with limited computing power.

[0166] To improve the prediction accuracy of the pose estimation model, the pose estimation model used in the above-mentioned grasping pose determination method can be pre-trained (at least before executing S502) to obtain the pose estimation model used in the foregoing embodiments. Based on this, the disclosed embodiments also provide a method for training a pose estimation model, and this method can be specifically applied to the training device mentioned in the foregoing embodiments. Referring to Figure 15 As shown, the method for training the pose estimation model may include S1501 - S1505:

[0167] S1501. Determine multiple sets of sample image data and sample grasping center point data corresponding to the sample image data.

[0168] Among them, the sample image data includes a sample image and a sample depth image corresponding to the sample image; the sample grasping center point data includes the position information of the sample grasping center point of the sample grasping object in the sample image and the sample grasping pose.

[0169] In some embodiments, the electronic device can obtain the sample image data and the corresponding sample grasping center point data in any possible implementation manner. For example, it can be obtained by actual measurement; or, it can be obtained from an existing training data set (such as GraspNet-1Billion). The disclosed embodiments do not make specific limitations on this.

[0170] After obtaining multiple groups of sample image data and the sample grasping center point data corresponding to the sample image data, the sample image and the corresponding sample depth image in each group of sample image data can be input into the initial pose estimation model to obtain the training grasping pose for grasping the object in the sample; then, based on the training grasping pose, using the sample grasping center point data corresponding to this group of sample images as the supervision information, the initial pose estimation model is iteratively trained to obtain the trained pose estimation model. This trained pose estimation model can predict the grasping pose more accurately. The training process of the pose estimation model is introduced below through S1502 - S1505.

[0171] S1502. Process the sample image and the sample depth image based on the first network in the initial pose estimation model to obtain the position information of the training grasping center point of the object to be grasped in the sample image.

[0172] In some embodiments, during initial training, the parameters in the initial pose estimation model can be set to 0 or any other possible initial values.

[0173] It can be understood that the specific implementation of S1502 can refer to the specific implementation of S502 in the foregoing embodiments, and will not be elaborated here.

[0174] In some embodiments, in combination with Figure 15 , with reference to Figure 16 shown, S1502 may specifically include S1601 - S1603:

[0175] S1601. Perform feature processing on the sample image and the sample depth image based on the first feature processing layer in the first network to obtain the first training feature map.

[0176] S1602. Perform upsampling processing on the first training feature map based on the upsampling layer in the first network to obtain the training grasping point heat map corresponding to the sample image.

[0177] S1603. Based on the training grasping point heat map and the sample depth image, determine the coordinate and depth information in the position information of the training grasping center point of the object to be grasped in the sample image.

[0178] It can be understood that the specific implementation of S1601 - S1603 can refer to the specific implementation of S5021 - S5023 in the foregoing embodiments, and will not be elaborated here.

[0179] S1503. Process the coordinates in the position information of the training grasping center point based on the sampling layer in the initial pose estimation model to obtain the training graspable area in the sample image that includes the training grasping center point.

[0180] S1504. Process the position information of the training graspable area and the training grasp center point based on the second network in the initial pose estimation model to obtain the training grasp pose for the grasp sample to grasp the object.

[0181] It can be understood that the specific implementation of S1503 and S1504 can refer to the specific implementation of S503 and S504 in the foregoing embodiments, and will not be elaborated here.

[0182] In some embodiments, in combination with Figure 15 , with reference to Figure 17 as shown, S1504 may specifically include S1701 and S1702:

[0183] S1701. Based on the second feature processing layer in the second network, perform feature extraction processing on the position information of the training graspable area and the training grasp center point to obtain the fourth training feature map and the fifth training feature map with different scales.

[0184] S1702. Based on the pose prediction layer in the second network, perform prediction processing on the fourth training feature map and the fifth training feature map to obtain the training grasp pose for the grasp sample to grasp the object.

[0185] The specific implementation of S1701 and S1702 can refer to the specific implementation of S5041 and S5042 in the foregoing embodiments, and will not be elaborated here.

[0186] S1505. Use the training grasp pose as the initial training output of the initial pose estimation model, and the sample grasp center point data as the supervision information, and iteratively train the initial pose estimation model to obtain the trained pose estimation model.

[0187] Based on the above embodiments, the training device can train the pose estimation model in a supervised learning manner. The pose estimation model has the ability to determine the grasping pose for grasping the object to be grasped in the two-dimensional image by using the corresponding two-dimensional image and depth image. The pose estimation model consists of a two-stage network. First, the first network is used to determine the position information of the grasping center point on the object to be grasped based on the image and depth image including the object to be grasped. Secondly, the sampling layer between the two-stage networks is used to determine the neighborhood information (graspable area) of the grasping center point based on the position information of the grasping center point in combination with the first image and depth image. Finally, the second network can determine the grasping pose for grasping the object to be grasped based on the neighborhood information of the grasping center point. It can be seen that in the present disclosure, by performing pose estimation on the two-dimensional image and the depth information corresponding to the two-dimensional image, the use of three-dimensional data of the point cloud is avoided, thereby reducing the model calculation amount and providing the inference efficiency, so that it can run efficiently when deployed on hardware with limited computing power; secondly, through the staged processing of pose estimation by the first network and the second network, the computing resources are further reduced and it can be efficiently deployed on hardware with limited computing power; in addition, the first network and the second network can be implemented by models with independent structures, so that the first network and the second network are set on different hardware, improving the utilization rate of hardware resources; in addition, the steps performed by the first network and the second network during the grasping pose estimation are independent of each other. For example, when the first network processes the current frame data, the second network processes the previous frame data, thereby further improving the processing efficiency of the pose estimation model and improving the real-time performance of the grasping pose estimation.

[0188] In some embodiments, in combination with Figure 15 , with reference to Figure 18 shown, S1505 may specifically include S1801 - S1803:

[0189] S1801. Determine a first loss value based on the training grasping pose and the sample grasping pose in the sample grasping center point data.

[0190] S1802. Determine a second loss value based on the position information of the training grasping center point and the position information of the sample grasping center point in the sample grasping center point data.

[0191] In the embodiments of the present disclosure, the calculation of the first loss value and the second loss value can be any possible calculation method, and the present disclosure does not make specific limitations thereto.

[0192] In the embodiments of the present disclosure, since the initial pose estimation model processes the input data by first processing it based on the first network and then based on the second network, the second loss value can be calculated after the first network finishes processing the data, and the first loss value can be calculated after the second network finishes processing. Based on this, usually, S1802 can be executed before S1801. Of course, in practice, all the loss values can also be calculated after the second network finishes processing. Therefore, S1801 can also be executed before S1802, or S1801 and S1802 can be executed simultaneously. The execution order of S1801 and S1802 can be specifically determined according to actual needs, and the present disclosure does not make specific limitations on this.

[0193] S1803. Iteratively update the initial pose estimation model based on the first loss value and the second loss value to obtain the trained pose estimation model.

[0194] It can be understood that based on the technical solutions provided in the above-mentioned disclosed embodiments, the training device can continuously iteratively optimize through the first loss value and the second loss value, and thus obtain a pose estimation model with small computational complexity, high computational / inference efficiency, which can operate efficiently when deployed on hardware with limited computing power and can estimate the grasping pose.

[0195] In addition, it should be noted that since the above-mentioned pose estimation model is composed of two networks, it can not only be jointly trained as in the above embodiments, but also the two networks can be trained separately. Among them, since the input of the second network needs to be obtained based on the output of the first network, the first network can be trained first. The sample data of the first network can be the multiple groups of sample image data in the above embodiments, and the supervision information can be the position information of the sample grasping center points in the sample grasping center point data corresponding to the sample image data. After the first network is trained, the input samples of the second network can be generated based on the first network, and thus the second network can be trained. The sample data of the second network can be the position information of the grasping center points obtained by the first network processing the sample image data, combined with the multiple groups of graspable regions and the corresponding position information of the grasping center points obtained from the sample image data. The supervision information of the second network can be the sample grasping poses in the sample grasping center point data corresponding to the sample image data.

[0196] The training methods of the first network and the second network can be any possible supervised learning methods, and the present disclosure does not make specific limitations on this.

[0197] Exemplary Device

[0198] Figure 19 A grasping pose determination device provided in an embodiment of the present disclosure. Referring to Figure 19 As shown, the grasping pose determination device includes an acquisition module 1901 and a processing module 1902.

[0199] Among them, an acquisition module 1901 is configured to determine a first image including an object to be grasped and a depth image corresponding to the first image; a processing module 1902 is configured to process the first image and the depth image acquired by the acquisition module 1901 based on a first network in a pose estimation model to obtain position information of a grasping center point of the object to be grasped in the first image; the processing module 1902 is further configured to process coordinates in the position information of the grasping center point based on a sampling layer in the pose estimation model to obtain a graspable area including the grasping center point in the first image; the processing module 1902 is further configured to process the graspable area and the position information of the grasping center point based on a second network in the pose estimation model to obtain a grasping pose for grasping the object to be grasped.

[0200] In some embodiments, in combination with Figure 19 , with reference to Figure 20 shown, the processing module 1902 may specifically include a feature sub-module 19021, an upsampling sub-module 19022, and a determination sub-module 19023. Among them, the feature sub-module 19021 is configured to perform feature processing on the first image and the depth image determined by the acquisition module 1901 based on a first feature processing layer in the first network to obtain a first feature map; the upsampling sub-module 19022 is configured to perform upsampling processing on the first feature map based on an upsampling layer in the first network to obtain a grasping point heat map corresponding to the first image; the determination sub-module 19023 is configured to determine coordinates and depth information in the position information of the grasping center point of the object to be grasped in the first image based on the grasping point heat map obtained by the upsampling sub-module 19022 and the depth image determined by the acquisition module 1901.

[0201] In some embodiments, the feature sub-module 19021 may specifically include a first feature unit and a second feature unit. Among them, the first feature unit is configured to perform feature extraction processing on the first image and the depth image based on a first feature extraction layer in the first feature processing layer to obtain second feature maps and third feature maps with different scales; the second feature unit is configured to perform fusion processing on the second feature map and the third feature map based on a feature fusion layer in the first feature processing layer to obtain a first feature map.

[0202] In some embodiments, the first feature unit is specifically configured to: perform feature extraction processing on the first image and the depth image based on a first sub-layer in the first feature extraction layer to obtain a second feature map at a first scale; perform feature extraction processing on the second feature map based on a second sub-layer in the first feature extraction layer to obtain a third feature map at a second scale.

[0203] In some embodiments, in combination with Figure 19 , with reference to Figure 20As shown, the device further includes a detection module 1903. Before the processing module 1902 processes the first image and the depth image determined by the acquisition module 1901 based on the first network in the pose estimation model to obtain the position information of the grasping center point of the object to be grasped in the first image, the detection module 1903 is further configured to determine the detection frame of the object to be grasped in the first image. The processing module 1902 further includes a segmentation sub-module 19024 and a calculation sub-module 19025.

[0204] Among them, the segmentation sub-module 19024 is configured to perform segmentation processing on the object to be grasped in the detection frame determined by the detection module based on the segmentation model in the first network to obtain a mask image of the object to be grasped; the calculation sub-module 19025 is configured to determine the coordinates in the position information of the grasping center point of the object to be grasped based on the mask image of the object to be grasped obtained by the segmentation sub-module 19024, and determine the depth information in the position information of the grasping center point based on the depth image determined by the acquisition module 1901.

[0205] In some embodiments, in combination with Figure 20 , with reference to Figure 21 As shown, the detection module 1903 specifically includes a display unit 19031 and a processing unit 19032. Among them, the processing unit 19032 is configured to perform object detection processing on the first image based on the semantic detection model to obtain the detection frame of the object to be grasped in the first image;

[0206] Or,

[0207] The display unit 19031 is configured to display the first image; the processing unit 19032 is configured to obtain the detection frame of the object to be grasped in the first image in response to the frame selection operation on the object to be grasped in the first image.

[0208] In some embodiments, in combination with Figure 19 , with reference to Figure 20 As shown, the processing module 1902 may further include a sampling sub-module 19026. Among them, the sampling sub-module 19026 is configured to determine, based on the sampling layer, an image of the graspable area including the grasping center point from the first image using the coordinates in the position information of the grasping center point; the determination sub-module 19023 is specifically configured to determine, based on the sampling layer, the depth information of the pixel points in the image of the graspable area determined by the sampling sub-module 19026 from the depth image.

[0209] In some embodiments, in combination with Figure 19 , with reference to Figure 20As shown, the processing module 1902 further includes a prediction sub-module 19027. The feature sub-module 19021 is further configured to perform feature extraction processing on the position information of the graspable region and the grasp center point based on the second feature processing layer in the second network, to obtain a fourth feature map and a fifth feature map with different scales; the prediction sub-module 19027 is configured to perform prediction processing on the fourth feature map and the fifth feature map based on the pose prediction layer in the second network, to obtain the grasp pose for grasping the object to be grasped.

[0210] In some embodiments, the feature sub-module 19021 further includes a third feature unit and a fourth feature unit. Among them, the third feature unit is configured to perform feature extraction processing on the position information of the graspable region and the grasp center point based on the first extraction layer in the second feature processing layer, to obtain a fourth feature map with a third scale; the fourth feature unit is configured to perform feature extraction processing on the fourth feature map based on the second extraction layer in the second feature processing layer, to obtain a fifth feature map with a fourth scale.

[0211] Regarding the grasp pose determination device in the above embodiments, the specific manners of operations performed by each module and the corresponding beneficial effects have been described in detail in the embodiments of the grasp pose determination method described above, and will not be elaborated herein.

[0212] Referring to Figure 22 As shown, the present disclosure further provides a pose estimation model training device, which includes an acquisition module 2201 and a training module 2202.

[0213] Among them, the acquisition module 2201 is configured to determine multiple groups of sample image data and sample grasp center point data corresponding to the sample image data; the sample image data includes a sample image and a sample depth image corresponding to the sample image; the sample grasp center point data includes the position information and the sample grasp pose of the sample grasp center point of the sample grasp object in the sample image; the training module 2202 is configured to process the sample image and the sample depth image acquired by the acquisition module 2201 based on the first network in the initial pose estimation model, to obtain the position information of the training grasp center point of the sample grasp object in the sample image; the training module 2202 is further configured to process the coordinates in the position information of the training grasp center point based on the sampling layer in the initial pose estimation model, to obtain the training graspable region including the training grasp center point in the sample image;

[0214] The training module 2202 is further configured to process the position information of the training graspable area and the training grasp center point based on the second network in the initial pose estimation model to obtain the training grasp pose for the grasping sample to grasp the object; the training module 2202 is further configured to use the training grasp pose as the initial training output of the initial pose estimation model, and the sample grasp center point data as the supervision information, and iteratively train the initial pose estimation model to obtain the trained pose estimation model.

[0215] In some embodiments, in combination with Figure 22 , with reference to Figure 23 as shown, the training module 2202 includes a feature sub-module 22021, a sampling sub-module 22022, and a determination sub-module 22023. Among them, the feature sub-module 22021 is configured to perform feature processing on the sample image and the sample depth image based on the first feature processing layer in the first network to obtain a first training feature map; the sampling sub-module 22022 is configured to perform upsampling processing on the first training feature map based on the upsampling layer in the first network to obtain a training grasp point heat map corresponding to the sample image; the determination sub-module 22023 is configured to determine the coordinate and depth information in the position information of the training grasp center point of the sample grasping object in the sample image based on the training grasp point heat map and the sample depth image.

[0216] In some embodiments, in combination with Figure 22 , with reference to Figure 23 as shown, the training module 2202 further includes a prediction sub-module 22024. The feature sub-module 22021 is further configured to perform feature extraction processing on the training graspable area determined by the determination sub-module 22023 and the position information of the training grasp center point based on the second feature processing layer in the second network to obtain a fourth training feature map and a fifth training feature map with different scales; the prediction sub-module 22024 is configured to perform prediction processing on the fourth training feature map and the fifth training feature map based on the pose prediction layer in the second network to obtain the training grasp pose for the grasping sample to grasp the object.

[0217] In some embodiments, in combination with Figure 22 , with reference to Figure 23As shown, the training module 2202 further includes a loss sub-module 22025 and an iteration sub-module 22026. Among them, the loss sub-module 22025 is used to determine a first loss value based on the training grasping pose obtained by the fourth sub-module 22024 and the sample grasping pose in the sample grasping center point data determined by the acquisition module 2201; the loss sub-module 22025 is further used to determine a second loss value based on the position information of the training grasping center point determined by the determination sub-module 22023 and the position information of the sample grasping center point in the sample grasping center point data determined by the acquisition module 2201; the iteration sub-module 22026 is used to iteratively update the initial pose estimation model based on the first loss value and the second loss value determined by the loss sub-module 22025 to obtain a trained pose estimation model.

[0218] Regarding the pose estimation model training device in the above embodiments, the specific manners in which each module performs operations and the corresponding beneficial effects have been described in detail in the embodiments of the pose estimation model training method described above, and will not be elaborated here.

[0219] Exemplary electronic device

[0220] An electronic device provided by an embodiment of the present disclosure includes at least one processor and a memory. Among them, the memory is used to store executable instructions for the processor; the processor is used to read the executable instructions from the memory and execute the instructions to implement the grasping pose determination method provided in the above-mentioned disclosed embodiments, the pose estimation model training method provided in the above-mentioned disclosed embodiments, and / or other desired functions.

[0221] The specific structure of this electronic device can refer to Figure 3 the electronic device shown in Figure 4 or the structure of the training device shown in

[0222] Exemplary Computer Program Product and Computer Readable Storage Medium

[0223] In addition to the above methods and devices, an embodiment of the present disclosure may also provide a computer program product, including computer program instructions, which, when run by a processor, cause the processor to execute the steps in the grasping pose determination method of various embodiments of the present disclosure described in the "Exemplary Method" section above, or the steps in the pose estimation model training method of various embodiments of the present disclosure.

[0224] A computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The programming code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0225] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium storing computer program instructions, which when run by a processor cause the processor to execute the steps of the method for determining the grasping pose of various embodiments of the present disclosure or the steps in the method for training the pose estimation model described in the "Exemplary Method" section above.

[0226] The computer-readable storage medium may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium, for example but not limited to, includes systems, devices or components of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0227] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that they are essential for each embodiment of the present disclosure. In addition, the above-mentioned specific details are only for the purpose of illustration and easy understanding, and not for limitation. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0228] Those skilled in the art can make various changes and modifications to the present disclosure without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present disclosure and their equivalent technologies, the present disclosure is also intended to include these changes and modifications.

Claims

1. A method for determining a grasping pose, comprising: Determining a first image including an object to be grasped and a depth image corresponding to the first image; Processing the first image and the depth image based on a first network in a pose estimation model to obtain position information of a grasping center point of the object to be grasped in the first image; Processing coordinates in the position information of the grasping center point based on a sampling layer in the pose estimation model to obtain a graspable region including the grasping center point; Processing the graspable region and the position information of the grasping center point based on a second network in the pose estimation model to obtain a grasping pose for grasping the object to be grasped.

2. The method according to claim 1, wherein, The processing the first image and the depth image based on a first network in a pose estimation model to obtain position information of a grasping center point of the object to be grasped in the first image includes: Performing feature processing on the first image and the depth image based on a first feature processing layer in the first network to obtain a first feature map; Performing upsampling processing on the first feature map based on an upsampling layer in the first network to obtain a grasping point heat map corresponding to the first image; Determining coordinate and depth information in the position information of the grasping center point of the object to be grasped in the first image based on the grasping point heat map and the depth image.

3. The method according to claim 2, wherein The performing feature processing on the first image and the depth image based on a first feature processing layer in the first network to obtain a first feature map includes: Performing feature extraction processing on the first image and the depth image based on a first feature extraction layer in the first feature processing layer to obtain a second feature map and a third feature map with different scales; Performing fusion processing on the second feature map and the third feature map based on a feature fusion layer in the first feature processing layer to obtain the first feature map.

4. The method according to claim 3, wherein The performing feature extraction processing on the first image and the depth image based on a first feature extraction layer in the first feature processing layer to obtain a second feature map and a third feature map with different scales includes: Performing feature extraction processing on the first image and the depth image based on a first sub-layer in the first feature extraction layer to obtain the second feature map at a first scale; Performing feature extraction processing on the second feature map based on a second sub-layer in the first feature extraction layer to obtain the third feature map at a second scale.

5. The method according to claim 1, wherein Before the processing the first image and the depth image based on a first network in a pose estimation model to obtain position information of a grasping center point of the object to be grasped in the first image, the method further includes: determining a detection frame of the object to be grasped in the first image; The processing the first image and the depth image based on a first network in a pose estimation model to obtain position information of a grasping center point of the object to be grasped in the first image includes: Performing segmentation processing on the object to be grasped in the detection frame based on a segmentation model in the first network to obtain a mask image of the object to be grasped. Based on the mask image of the object to be grasped, determine the coordinates in the position information of the grasping center point of the object to be grasped, and based on the depth image, determine the depth information in the position information of the grasping center point.

6. The method according to claim 5, wherein, The determining of the detection box of the object to be grasped in the first image includes: Performing object detection processing on the first image based on a semantic detection model to obtain the detection box of the object to be grasped in the first image; Or, Display the first image; in response to a box selection operation on the object to be grasped in the first image, obtain the detection box of the object to be grasped in the first image.

7. The method according to claim 1, wherein The processing of the coordinates in the position information of the grasping center point based on the sampling layer in the pose estimation model to obtain a graspable area including the grasping center point includes: Based on the sampling layer, using the coordinates in the position information of the grasping center point, determine an image of the graspable area including the grasping center point from the first image; Based on the sampling layer, determine the depth information of the pixel points in the image of the graspable area from the depth image.

8. The method according to claim 1, wherein The processing of the graspable area and the position information of the grasping center point based on the second network in the pose estimation model to obtain the grasping pose for grasping the object to be grasped includes: Based on the second feature processing layer in the second network, perform feature extraction processing on the graspable area and the position information of the grasping center point to obtain fourth and fifth feature maps with different scales; Based on the pose prediction layer in the second network, perform prediction processing on the fourth and fifth feature maps to obtain the grasping pose for grasping the object to be grasped.

9. The method according to claim 8, wherein The processing of the graspable area and the position information of the grasping center point based on the second feature processing layer in the second network to obtain fourth and fifth feature maps with different scales includes: Based on the first extraction layer in the second feature processing layer, perform feature extraction processing on the graspable area and the position information of the grasping center point to obtain a fourth feature map of the third scale; Based on the second extraction layer in the second feature processing layer, perform feature extraction processing on the fourth feature map to obtain the fifth feature map of the fourth scale.

10. A method for training a pose estimation model, including: Determine multiple groups of sample image data and sample grasping center point data corresponding to the sample image data; the sample image data includes a sample image and a sample depth image corresponding to the sample image; the sample grasping center point data includes the position information and sample grasping pose of the sample grasping center point of the sample grasping object in the sample image; Based on the first network in the initial pose estimation model, process the sample image and the sample depth image to obtain the position information of the training grasping center point of the sample grasping object in the sample image; Based on the sampling layer in the initial pose estimation model, process the coordinates in the position information of the training grasping center point to obtain the training graspable area including the training grasping center point in the sample image; Process the position information of the training grippable region and the training grasp center point based on the second network in the initial pose estimation model to obtain the training grasp pose for grasping the sample grasping object; Use the training grasp pose as the initial training output of the initial pose estimation model, and use the sample grasp center point data as supervision information to iteratively train the initial pose estimation model to obtain the trained pose estimation model.

11. The method according to claim 10, wherein, The process of using the first network in the initial pose estimation model to process the sample image and the sample depth image to obtain the position information of the training grasp center point of the sample grasping object in the sample image includes: Based on the first feature processing layer in the first network, perform feature processing on the sample image and the sample depth image to obtain the first training feature map; Based on the upsampling layer in the first network, perform upsampling processing on the first training feature map to obtain the training grasp point heat map corresponding to the sample image; Based on the training grasp point heat map and the sample depth image, determine the coordinate and depth information in the position information of the training grasp center point of the sample grasping object in the sample image.

12. The method according to claim 10, wherein The process of using the second network in the initial pose estimation model to process the position information of the training grippable region and the training grasp center point to obtain the training grasp pose for grasping the sample grasping object includes: Based on the second feature processing layer in the second network, perform feature extraction processing on the position information of the training grippable region and the training grasp center point to obtain the fourth training feature map and the fifth training feature map with different scales; Based on the pose prediction layer in the second network, perform prediction processing on the fourth training feature map and the fifth training feature map to obtain the training grasp pose for grasping the sample grasping object.

13. The method according to claim 10, wherein Using the training grasp pose as the initial training output of the initial pose estimation model, and using the sample grasp center point data as supervision information to iteratively train the initial pose estimation model to obtain the trained pose estimation model includes: Determine the first loss value based on the training grasp pose and the sample grasp pose in the sample grasp center point data; Determine the second loss value based on the position information of the training grasp center point and the position information of the sample grasp center point in the sample grasp center point data; Based on the first loss value and the second loss value, iteratively update the initial pose estimation model to obtain the trained pose estimation model.

14. A grasping pose determination device, comprising: An acquisition module for determining a first image including an object to be grasped and a depth image corresponding to the first image; A processing module for processing the first image and the depth image acquired by the acquisition module based on the first network in the pose estimation model to obtain the position information of the grasp center point of the object to be grasped in the first image; The processing module is further configured to process the coordinates in the position information of the grasping center point based on the sampling layer in the pose estimation model, so as to obtain a graspable area including the grasping center point in the first image; The processing module is further configured to process the graspable area and the position information of the grasping center point based on the second network in the pose estimation model, so as to obtain a grasping pose for grasping the object to be grasped.

15. A pose estimation model training device, comprising: An acquisition module, configured to determine multiple sets of sample image data and sample grasping center point data corresponding to the sample image data; the sample image data includes a sample image and a sample depth image corresponding to the sample image; the sample grasping center point data includes the position information and sample grasping pose of a sample grasping object in the sample image; A training module, configured to process the sample image and the sample depth image acquired by the acquisition module based on a first network in an initial pose estimation model, so as to obtain the position information of a training grasping center point of the sample grasping object in the sample image; The training module is further configured to process the coordinates in the position information of the training grasping center point based on the sampling layer in the initial pose estimation model, so as to obtain a training graspable area including the training grasping center point in the sample image; The training module is further configured to process the training graspable area and the position information of the training grasping center point based on the second network in the initial pose estimation model, so as to obtain a training grasping pose for grasping the sample grasping object; The training module is further configured to use the training grasping pose as an initial training output of the initial pose estimation model, and the sample grasping center point data as supervision information, and iteratively train the initial pose estimation model to obtain a trained pose estimation model.

16. A computer-readable storage medium, storing a computer program, where the computer program is used to execute the grasping pose determination method according to any one of claims 1-9 above, or execute the pose estimation model training method according to any one of claims 10-13 above.

17. An electronic device, the electronic device includes: A processor; A memory for storing executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the grasping pose determination method according to any one of claims 1-9 above, or execute the pose estimation model training method according to any one of claims 10-13 above.