Lightweight grabbing attitude prediction method and system
By introducing the CA attention mechanism and monocular depth estimation task into the GR-CNN model, the improved GR-CNN model achieves high-precision grasping pose prediction in scenarios that only rely on RGB images, solving the problems of insufficient grasping pose prediction accuracy and lack of depth information, and improving the target object positioning and background separation capabilities.
Patent Information
- Application Number
- CN202510624796.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-09-16
AI Technical Summary
The existing grasping posture prediction model has problems such as poor grasping effect in working scenarios that only rely on RGB images, insufficient grasping posture prediction accuracy, poor ability to separate target objects from background, and lack of depth information.
The CA attention mechanism and monocular depth estimation auxiliary task are introduced to improve the GR-CNN model. Through the dual-branch encoder-decoder architecture, the collaborative learning of grasping pose prediction and depth estimation is realized, thereby enhancing feature expression and target object positioning accuracy.
It improves the accuracy and robustness of grasping pose prediction, enhances the ability to separate target objects from backgrounds, solves the problem of missing depth information, and is suitable for high-precision grasping in complex scenes.
Smart Images

Figure CN120655706A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robot visual grasping, and in particular to a lightweight grasping posture prediction method and system. Background Art
[0002] In confined environments like surgical robots, spatial limitations prevent the deployment of depth sensors (such as RGB-D cameras). This results in a lack of depth information, which traditional grasping algorithms rely on, making high-precision grasping difficult. The current mainstream method, GR-ConvNet, uses a generative residual convolutional network to predict the grasp pose (quality score, angle, and width). However, it suffers from the following issues: positional bias: traditional convolution operations lose spatial position information, resulting in inaccurate predictions of the grasp box center and rotation angle; weak global perception: the model overly focuses on local features and ignores the overall structure of the object, resulting in low grasp quality scores in the center region; and depth dependency: the model relies on depth image input and is not adaptable to scenarios where only RGB images are available.
[0003] In summary, existing grasping posture prediction models have problems such as insufficient grasping posture prediction accuracy, poor ability to separate target objects from background, and lack of depth information in working scenarios that rely only on RGB images. Summary of the Invention
[0004] The present invention provides a lightweight grasping posture prediction method and system to solve the technical problems of insufficient grasping posture prediction accuracy, poor target object and background separation ability, and lack of depth information in existing grasping posture prediction models in working scenarios that only rely on RGB images.
[0005] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0006] In one aspect, the present invention provides a lightweight grasping posture prediction method, comprising:
[0007] Acquire a target image; wherein the target image is a single-view RGB image of the object to be grasped in the operation scene;
[0008] Preprocessing the acquired target image;
[0009] The preprocessed target image is input into the preset grasping posture prediction model to generate the corresponding grasping posture.
[0010] Furthermore, the preprocessing of the acquired target image includes:
[0011] The acquired target image is cropped, resized, and normalized.
[0012] Furthermore, the preset grasping posture prediction model is an improved GR-CNN model; the improved GR-CNN model introduces the CA attention mechanism into the GR-CNN model and adds a monocular depth estimation auxiliary task.
[0013] Furthermore, the improved GR-CNN model is a dual-branch encoder-decoder architecture, including a main branch and an auxiliary branch; the main branch and the auxiliary branch extract multi-task common features through a shared encoder, and use task-specific decoders to achieve collaborative learning of the main and auxiliary tasks of grasping pose prediction and depth estimation; wherein, the main branch focuses on grasping pose prediction, and the auxiliary branch introduces geometric prior constraints through the depth estimation task.
[0014] Furthermore, the process of generating the grasping pose by the improved GR-CNN model includes:
[0015] In the encoding stage, three convolutional layers are first used to construct a feature pyramid through a downsampling operation with a stride of 2, and batch normalization and ReLU activation function are combined to realize multi-scale spatial feature extraction. Then, based on the extracted multi-scale spatial features, the CA module is used to parallelly calculate the channel attention and spatial attention weights through a channel segmentation strategy to achieve adaptive calibration of the feature map along the height and width dimensions. Subsequently, the features processed by the CA module are input into five stacked residual layers for deep feature refinement to obtain deep features. Finally, the deep features processed by the CA module are used to obtain multi-task general features.
[0016] In the decoding stage, the network adopts a mirror-symmetric decoder design, which consists of two parallel and structurally identical decoder branches: a main task decoder branch and an auxiliary task decoder branch; both decoder branches gradually restore the spatial resolution through three transposed convolutional layers, and connect batch normalization layers and nonlinear activation functions after upsampling at each level, and finally each output pixel-level prediction maps; among them, the main branch outputs three maps: grasping quality score, grasping angle, and grasping width required by the end effector, and the auxiliary branch outputs a pixel-by-pixel depth estimation map.
[0017] On the other hand, the present invention also provides a lightweight grasping posture prediction system, comprising:
[0018] The target image acquisition module is used to acquire a target image; wherein the target image is a single-view RGB image of the object to be grasped in the operation scene;
[0019] A preprocessing module, used for preprocessing the acquired target image;
[0020] The grasping posture prediction module is used to input the preprocessed target image into the preset grasping posture prediction model to generate the corresponding grasping posture.
[0021] Furthermore, the preprocessing module is specifically used to:
[0022] The acquired target image is cropped, resized, and normalized.
[0023] Furthermore, the preset grasping posture prediction model is an improved GR-CNN model; the improved GR-CNN model introduces the CA attention mechanism into the GR-CNN model and adds a monocular depth estimation auxiliary task.
[0024] Furthermore, the improved GR-CNN model is a dual-branch encoder-decoder architecture, including a main branch and an auxiliary branch; the main branch and the auxiliary branch extract multi-task common features through a shared encoder, and use task-specific decoders to achieve collaborative learning of the main and auxiliary tasks of grasping pose prediction and depth estimation; wherein, the main branch focuses on grasping pose prediction, and the auxiliary branch introduces geometric prior constraints through the depth estimation task.
[0025] Furthermore, the process of generating the grasping pose by the improved GR-CNN model includes:
[0026] In the encoding stage, three convolutional layers are first used to construct a feature pyramid through a downsampling operation with a stride of 2, and batch normalization and ReLU activation function are combined to realize multi-scale spatial feature extraction. Then, based on the extracted multi-scale spatial features, the CA module is used to parallelly calculate the channel attention and spatial attention weights through a channel segmentation strategy to achieve adaptive calibration of the feature map along the height and width dimensions. Subsequently, the features processed by the CA module are input into five stacked residual layers for deep feature refinement to obtain deep features. Finally, the deep features processed by the CA module are used to obtain multi-task general features.
[0027] In the decoding stage, the network adopts a mirror-symmetric decoder design, which consists of two parallel and structurally identical decoder branches: a main task decoder branch and an auxiliary task decoder branch; both decoder branches gradually restore the spatial resolution through three transposed convolutional layers, and connect batch normalization layers and nonlinear activation functions after upsampling at each level, and finally each output pixel-level prediction maps; among them, the main branch outputs three maps: grasping quality score, grasping angle, and grasping width required by the end effector, and the auxiliary branch outputs a pixel-by-pixel depth estimation map.
[0028] On the other hand, the present invention further provides an electronic device, comprising a processor and a memory; wherein the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the above method.
[0029] In yet another aspect, the present invention further provides a computer-readable storage medium, wherein the storage medium stores at least one instruction, and the instruction is loaded and executed by a processor to implement the above method.
[0030] The beneficial effects brought about by the technical solution provided by the present invention include at least:
[0031] 1. The present invention improves the accuracy of grasping pose prediction: By introducing the CA attention mechanism, the present invention significantly improves the accuracy of predicting the grasping box position and rotation angle.
[0032] 2. The present invention enhances the ability to separate target objects from backgrounds: The CA attention mechanism dynamically selects key information, suppresses background interference, and improves the positioning accuracy of target objects.
[0033] 3. This invention solves the problem of missing depth information: through the auxiliary task of monocular depth estimation, the relative depth of the target object is inferred from the RGB image, thereby improving the accuracy of grasping pose inference.
[0034] 4. The present invention is applicable to complex scenarios: the improved model can better handle complex scenarios and significantly improves the accuracy and robustness of grasp prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0036] Figure 1 This is the GR-CNN model structure diagram;
[0037] Figure 2 This is a diagram of the CA network structure;
[0038] Figure 3 This is a structural diagram of the improved GR-CNN model provided by an embodiment of the present invention;
[0039] Figure 4 This is a system block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0040] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0041] First, it should be noted that in the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as an "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "exemplarily" is intended to present concepts in a concrete manner. In addition, in the embodiments of the present invention, the meaning of "and / or" can be both or either of the two.
[0042] First embodiment
[0043] This embodiment provides a lightweight grasping pose prediction method that improves the accuracy and robustness of grasping pose prediction by introducing an attention mechanism and a monocular depth estimation auxiliary task. This method is suitable for high-precision grasping tasks in constrained scenarios such as surgical robots, especially for scenarios that rely solely on monocular RGB image input. The method can be implemented by an electronic device. Specifically, the execution process of the method includes the following steps:
[0044] S1, acquire target image;
[0045] It should be noted that the target image is a single-view RGB image of the object to be grasped in the working scene.
[0046] S2, preprocessing the acquired target image;
[0047] The preprocessing may include cropping, resizing, and normalizing the acquired target image.
[0048] S3, inputting the preprocessed target image into a preset grasping posture prediction model to generate a grasping posture;
[0049] Among them, in order to clarify the implementation process of this solution, the grasping problem is first defined; specifically, this embodiment defines the grasping problem as predicting the corresponding grasping posture of an unknown object from a single-view RGB image obtained from the working scene by a visual sensor (monocular camera) and executing it on a robotic arm or surgical actuator. An improved grasping representation proposed by Morrison et al. is used. This representation clearly simulates the physical size of the end gripper and strictly limits the scope of feature calculation to solve the problem of mismatch between the support area and the physical space in the grasping point or point pair representation. The grasping posture in the robotic arm coordinate system is expressed as:
[0050] G r =(P,θ r ,W r ,Q) (1)
[0051] Where P = (x, y, z) is the center position of the gripper, θ ris the rotation angle of the gripper around the z-axis (unit: rad), W r is the required width of the gripper and Q is the grasp quality score.
[0052] From an RGB image I=R with height h and width w 3×h×w Detect grasp in , which can be expressed as:
[0053] G i =(x,y,θ i ,W i ,Q) (2)
[0054] Where (x, y) is the grasp center in the image coordinate system, θ i is the rotation angle around the z axis in the camera reference frame, W i is the required width of the gripper in the image coordinate system, and Q is the grasping quality score (same meaning as in formula (1)).
[0055] θ i Indicates the corresponding measurement value of the angular rotation required for each pixel when grasping the target object, expressed in the range of [-π / 2,π / 2]. i is the desired width, expressed as a measure of uniform depth, in [0,W max ]The value within the pixel range indicates that W max The maximum width of the antipodal gripper. Q is the grasp quality score for each grasped pixel in the image, expressed as a value in the range [0,1], where values closer to 1 indicate a greater chance of successful grasping, and vice versa.
[0056] Finally, the grasp acquired in the image coordinate system is performed on the robot arm, and the grasp pose representation in the image coordinate system is converted to the grasp pose representation in the robot arm coordinate system through the following transformation:
[0057] G r =T rc (T ci (G i )) (3)
[0058] Among them, T ci It is the transformation matrix that transforms the 2D image coordinate system to the 3D camera coordinate system using the camera’s intrinsic parameters (i.e., internal parameters). rc It is the transformation matrix that converts the camera coordinate system to the manipulator coordinate system using the camera attitude calibration value (i.e., external parameter).
[0059] The set of all grasps can be simplified as:
[0060] G=(θ,W,Q)∈R 3×h×w (4)
[0061] Among them, θ, W and Q represent the three images of Angle, Width and Quality in the form of grasping angle, grasping width and grasping quality score respectively, and are calculated at each grasped pixel point in the image using formula (2).
[0062] Based on the above, this embodiment uses an improved GR-CNN model to generate grasp poses. It's important to note that the GR-CNN model is a generative, pixel-by-pixel grasping model that accepts an n-channel input image and generates three pixel-level output images for grasp pose prediction. This model effectively decodes and extracts latent features, focusing on key features while suppressing irrelevant ones. This allows it to efficiently process multiple input modalities and generate accurate grasp pose information.
[0063] Before inputting into GR-CNN, the image needs to be preprocessed, including cropping, resizing, and normalization. If the input data contains depth images, it is also necessary to use low-gradient regularization methods to perform depth repair to obtain effective depth representation and ensure the integrity and consistency of the data. Figure 1 As shown in the figure, after preprocessing, a 224×224 n-channel input image is fed into the GR-CNN model. The model first passes through three convolutional layers, which extract local features of the image using local receptive fields. The feature maps are then passed to five residual layers. These layers optimize the learning of the identity mapping through skip connections, effectively alleviating the accuracy degradation caused by vanishing gradients and dimensionality errors in deep networks while improving the model's training stability and feature representation. After processing through the convolutional and residual layers, the image size is reduced to 56×56. To preserve the spatial features of the image and improve interpretability, the model uses three convolutional transpose layers to upsample the image back to the original input size. The model then generates three output images representing the grasp quality score, the grasp angle, and the required grasp width for the end-effector. The grasp angle is represented by two elements (sin2θ and cos2θ) to avoid the antipodal problem of angles near [-π / 2, π / 2] and ensure unambiguous and unambiguous angle representation. Based on these three output images, the model ultimately infers the precise grasping pose. However, GR-CNN suffers from insufficient grasping pose prediction accuracy and poor ability to separate the target object from the background when relying solely on RGB images.
[0064] To address the above problems, this embodiment improves the GR-CNN model by introducing the CA attention mechanism into the GR-CNN model and adding a monocular depth estimation auxiliary task to obtain an improved GR-CNN model.
[0065] Specifically, this embodiment introduces Figure 2The CA attention mechanism module shown here improves the quality score and localization accuracy of the grasped region by preserving positional information, enhancing feature representation, and dynamically selecting key information. A lightweight, plug-and-play attention module aims to simultaneously model inter-channel dependencies and precise spatial positional information. Using a bidirectional decomposition encoding strategy (horizontal and vertical feature aggregation), it preserves spatial structure while enhancing feature representation, making it suitable for visual grasp prediction tasks requiring precise localization. Its core concept consists of two steps: coordinate information embedding and coordinate attention generation. In the coordinate information embedding stage, traditional global pooling is decomposed into 1D feature encoding operations in the horizontal and vertical directions, which are then pooled separately on the input feature map to generate a direction-aware feature map. This operation not only preserves precise positional information but also captures long-range dependencies. In the coordinate attention generation stage, the horizontal and vertical feature maps are concatenated and transformed using a shared 1×1 convolution to generate an intermediate feature map. The intermediate feature map is then split into two independent tensors, each of which is subjected to a convolutional transformation to generate attention weights. These weights are then applied to the input feature map to dynamically adjust the spatial and channel-wise information of the features. Through the above two stages, CA can more accurately locate the target object and enhance the feature expression capability.
[0066] Furthermore, this embodiment uses the monocular depth estimation auxiliary task to infer the relative depth of the target object from the RGB image, solving the problem that the monocular camera cannot directly obtain depth information.
[0067] Based on the above-mentioned coordinate attention module model, this embodiment proposes a method such as Figure 3The architecture shown in Figure 1 is a multi-task architecture. This architecture extracts common features across multiple tasks through a shared encoder and utilizes task-specific decoders to collaboratively learn the primary and secondary tasks of grasp pose prediction and depth estimation. The network extracts cross-modal representations through a shared underlying feature encoding layer. The primary branch focuses on grasp pose prediction, while the secondary branch introduces geometric prior constraints through the depth estimation task. During the encoding phase, the network adopts a progressive feature extraction strategy: the low-level encoding module consists of three convolutional layers, constructing a feature pyramid through downsampling with a stride of 2. It combines batch normalization with ReLU activation functions to extract multi-scale spatial features. The mid-level feature enhancement module innovatively introduces a two-stage attention mechanism. A coordinate attention (CA) module is deployed before the stacking of five residual layers. Using a channel-splitting strategy, channel-wise and spatial attention weights are computed in parallel, achieving adaptive calibration of feature maps along the height and width dimensions. Deep features are then refined through residual layers (including skip connections), effectively alleviating the gradient decay problem. To further strengthen the spatial and channel-wise dependencies of features and ensure that deep features fully capture global context during refinement, a coordinate attention (CA) module is also deployed after the stacking of residual layers. During the decoding phase, the network employs a mirror-symmetric decoder design consisting of two parallel, identical decoder branches: a primary task decoder branch and an auxiliary task decoder branch. Both decoder branches progressively restore spatial resolution through three transposed convolutional layers. After upsampling at each level, batch normalization layers and nonlinear activation functions are applied. Each decoder branch ultimately outputs a pixel-level prediction map. The primary task outputs three maps: grasp quality score, grasp angle, and the required grasp width for the end effector. The auxiliary task outputs a pixel-by-pixel depth estimate.
[0068] Furthermore, after constructing the improved GR-CNN model, this embodiment uses the rectangle metric method proposed by Jiang et al. to measure the proximity of the predicted grab box position to the theoretical grab position to verify the performance of the network model. According to the rectangle metric, a grab is considered valid when the following two conditions are met:
[0069] 1. The angular deviation between the predicted grasping rectangle and the theoretical grasping direction is less than 30 degrees.
[0070] 2. The intersection over union (IoU) between the theoretical grasped rectangle and the predicted grasped rectangle exceeds 25%.
[0071]
[0072] Among them, G t ∩G p is the overlapping area between the theoretical grasping rectangle and the predicted grasping rectangle, G t ∪G p is the area of the union of the two.
[0073] In summary, this embodiment proposes an improved GR-CNN model, and based on this, implements a lightweight grasping posture prediction method. By introducing the CA attention mechanism, the accuracy of predicting the grasping box position and rotation angle is significantly improved; and the CA attention mechanism suppresses background interference by dynamically selecting key information, thereby improving the positioning accuracy and robustness of the target object. In addition, this embodiment infers the relative depth of the target object from the RGB image through the auxiliary task of monocular depth estimation, thereby improving the accuracy of grasping posture inference. The improved model can achieve high-precision grasping posture prediction in working scenarios that rely only on RGB images, so it can better handle complex scenes and significantly improve the accuracy and robustness of grasping prediction.
[0074] Second embodiment
[0075] This embodiment provides a lightweight grasping posture prediction system, which includes the following modules:
[0076] The target image acquisition module is used to acquire a target image; wherein the target image is a single-view RGB image of the object to be grasped in the operation scene;
[0077] A preprocessing module, used for preprocessing the acquired target image;
[0078] The grasping posture prediction module is used to input the preprocessed target image into the preset grasping posture prediction model to generate the corresponding grasping posture.
[0079] It should be noted that the lightweight grasping posture prediction system of this embodiment corresponds to the lightweight grasping posture prediction method of the above-mentioned first embodiment; the functions implemented by each functional module in the lightweight grasping posture prediction system of this embodiment correspond one-to-one to each process step in the lightweight grasping posture prediction method of the above-mentioned first embodiment; therefore, they will not be repeated here.
[0080] Third embodiment
[0081] This embodiment provides an electronic device, such as Figure 4 As shown, the electronic device includes: a processor and a memory; wherein the processor and the memory can be connected via a communication bus; the memory stores at least one instruction, which is loaded and executed by the processor to implement the method of the first embodiment described above. In addition, the electronic device may also include a transceiver; the processor and the transceiver can be connected via a communication bus, and the transceiver is used to communicate with other devices.
[0082] Next, combine Figure 4 A detailed introduction to the various components of the electronic device is given below:
[0083] Among them, the processor is the control center of the electronic device, and the electronic device may include multiple processors, each of which may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here can be a processor or a general term for multiple processing elements. For example, the processor is one or more central processing units (CPUs), or other general-purpose processors, application specific integrated circuits (ASICs), or one or more integrated circuits configured to implement an embodiment of the present invention, such as one or more microprocessors (digital signal processors, DSPs), or one or more field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor can perform various functions of the electronic device by running or executing software programs stored in the memory and calling data stored in the memory.
[0084] In a specific implementation, as an embodiment, the processor may include one or more CPUs, such as Figure 4 The CPU0 and CPU1 shown in FIG are, of course, only exemplary.
[0085] The memory is used to store the software program for executing the solution of the present invention, and the execution is controlled by the processor. The specific implementation method can refer to the above method embodiment and will not be repeated here.
[0086] Optionally, the memory may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory may be integrated with the processor or exist independently and accessed through the interface circuit ( Figure 4 (not shown) is coupled to the processor, which is not specifically limited in this embodiment of the present invention.
[0087] The transceiver may include a receiver and a transmitter ( Figure 4 The receiver is used to implement the receiving function, and the transmitter is used to implement the sending function. The transceiver can be integrated with the processor or exist independently and communicate with the electronic device through the interface circuit ( Figure 4 (not shown) is coupled to the processor, which is not specifically limited in this embodiment of the present invention.
[0088] In addition, it should be noted that Figure 4 The structure of the electronic device shown in the figure does not constitute a limitation on the device. The actual device may include more or fewer components than shown, or may combine certain components, or arrange the components differently. In addition, the technical effects achieved by the electronic device when executing the method of the first embodiment can refer to the technical effects described in the first embodiment above, and therefore will not be repeated here.
[0089] Fourth embodiment
[0090] This embodiment provides a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the method of the first embodiment described above. The computer-readable storage medium may be a ROM, random access memory, CD-ROM, magnetic tape, floppy disk, or optical data storage device. The instructions stored therein can be loaded by a processor in a terminal to execute the method described above.
[0091] Furthermore, it should be noted that the present invention may be provided as a method, apparatus, or computer program product. Therefore, embodiments of the present invention may take the form of a fully or partially hardware embodiment, a fully or partially software embodiment, or an embodiment combining software and hardware aspects. Furthermore, when implemented using software, embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The computer program product comprises one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the processes or functions described in accordance with the embodiments of the present invention are fully or partially generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired connection (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium accessible by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium. The semiconductor medium may be a solid state drive.
[0092] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the process in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0093] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for implementing the process in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0094] It should also be noted that, in this document, relational terms such as first and second are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any actual relationship or order between these entities or operations. The terms "include," "comprises," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or terminal device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. Without further limitation, an element defined by the phrase "comprising a..." does not preclude the presence of other identical elements in the process, method, article, or terminal device comprising the element. In addition, the term "and / or" is merely a description of an associative relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: the presence of A alone, the presence of A and B simultaneously, or the presence of B alone, where A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding. "At least one" means one or more, and "more" means two or more. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can be represented by: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.
[0095] In addition, it can be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0096] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0097] In the several embodiments provided herein, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of functional modules / units is merely a logical functional division. In actual implementation, other division methods may be used, such as multiple units or components being combined or integrated into another device, or some features being ignored or not implemented. Furthermore, the coupling or direct coupling or communication connection shown or discussed between each other may be through some interface, indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms. Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, i.e., they may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of the present embodiment according to actual needs. In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit.
[0098] If the method is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0099] Finally, it should be noted that the above description is merely a preferred embodiment of the present invention. It should be noted that, although preferred embodiments of the present invention have been described, those skilled in the art, once understanding the basic inventive concepts of the present invention, may make various improvements and modifications without departing from the principles of the present invention. Such improvements and modifications should also be considered within the scope of protection of the present invention. Therefore, the appended claims are intended to be interpreted as covering the preferred embodiments and all variations and modifications that fall within the scope of the embodiments of the present invention.
Claims
1. A lightweight grasping posture prediction method, characterized in that: include: Acquire a target image; wherein the target image is a single-view RGB image of the object to be grasped in the operation scene; Preprocessing the acquired target image; The preprocessed target image is input into the preset grasping posture prediction model to generate the corresponding grasping posture.
2. The lightweight grasping posture prediction method according to claim 1, wherein: The preprocessing of the acquired target image includes: The acquired target image is cropped, resized, and normalized.
3. The lightweight grasping posture prediction method according to claim 1, wherein: The preset grasping posture prediction model is an improved GR-CNN model; the improved GR-CNN model introduces the CA attention mechanism into the GR-CNN model and adds a monocular depth estimation auxiliary task.
4. The lightweight grasping posture prediction method according to claim 3, wherein: The improved GR-CNN model is a two-branch encoder-decoder architecture, including a main branch and an auxiliary branch. The main branch and the auxiliary branch extract multi-task common features through a shared encoder, and use task-specific decoders to achieve collaborative learning of the main and auxiliary tasks of grasping pose prediction and depth estimation. Among them, the main branch focuses on grasping pose prediction, while the auxiliary branch introduces geometric prior constraints through the depth estimation task.
5. The lightweight grasping posture prediction method according to claim 4, characterized in that: The process of generating grasping poses by the improved GR-CNN model includes: In the encoding stage, three convolutional layers are first used to construct a feature pyramid through a downsampling operation with a stride of 2, and batch normalization and ReLU activation function are combined to realize multi-scale spatial feature extraction. Then, based on the extracted multi-scale spatial features, the CA module is used to parallelly calculate the channel attention and spatial attention weights through a channel segmentation strategy to achieve adaptive calibration of the feature map along the height and width dimensions. Subsequently, the features processed by the CA module are input into five stacked residual layers for deep feature refinement to obtain deep features. Finally, the deep features processed by the CA module are used to obtain multi-task general features. In the decoding stage, the network adopts a mirror-symmetric decoder design, which consists of two parallel and structurally identical decoder branches: a main task decoder branch and an auxiliary task decoder branch; both decoder branches gradually restore the spatial resolution through three transposed convolutional layers, and connect batch normalization layers and nonlinear activation functions after upsampling at each level, and finally each output pixel-level prediction maps; among them, the main branch outputs three maps: grasping quality score, grasping angle, and grasping width required by the end effector, and the auxiliary branch outputs a pixel-by-pixel depth estimation map.
6. A lightweight grasping posture prediction system, characterized in that: include: The target image acquisition module is used to acquire a target image; wherein the target image is a single-view RGB image of the object to be grasped in the operation scene; A preprocessing module, used for preprocessing the acquired target image; The grasping posture prediction module is used to input the preprocessed target image into the preset grasping posture prediction model to generate the corresponding grasping posture.
7. The lightweight grasping posture prediction system according to claim 6, wherein: The preprocessing module is specifically used for: The acquired target image is cropped, resized, and normalized.
8. The lightweight grasping posture prediction system according to claim 6, wherein: The preset grasping posture prediction model is an improved GR-CNN model; the improved GR-CNN model introduces the CA attention mechanism into the GR-CNN model and adds a monocular depth estimation auxiliary task.
9. The lightweight grasping posture prediction system according to claim 8, wherein: The improved GR-CNN model is a two-branch encoder-decoder architecture, including a main branch and an auxiliary branch. The main branch and the auxiliary branch extract multi-task common features through a shared encoder, and use task-specific decoders to achieve collaborative learning of the main and auxiliary tasks of grasping pose prediction and depth estimation. Among them, the main branch focuses on grasping pose prediction, while the auxiliary branch introduces geometric prior constraints through the depth estimation task.
10. The lightweight grasping posture prediction system according to claim 9, wherein: The process of generating grasping poses by the improved GR-CNN model includes: In the encoding stage, three convolutional layers are first used to construct a feature pyramid through a downsampling operation with a stride of 2, and batch normalization and ReLU activation function are combined to realize multi-scale spatial feature extraction. Then, based on the extracted multi-scale spatial features, the CA module is used to parallelly calculate the channel attention and spatial attention weights through a channel segmentation strategy to achieve adaptive calibration of the feature map along the height and width dimensions. Subsequently, the features processed by the CA module are input into five stacked residual layers for deep feature refinement to obtain deep features. Finally, the deep features processed by the CA module are used to obtain multi-task general features. In the decoding stage, the network adopts a mirror-symmetric decoder design, which consists of two parallel and structurally identical decoder branches: a main task decoder branch and an auxiliary task decoder branch; both decoder branches gradually restore the spatial resolution through three transposed convolutional layers, and connect batch normalization layers and nonlinear activation functions after upsampling at each level, and finally each output pixel-level prediction maps; among them, the main branch outputs three maps: grasping quality score, grasping angle, and grasping width required by the end effector, and the auxiliary branch outputs a pixel-by-pixel depth estimation map.