Method and device for reconstructing robot grasping scene, and electronic device

Through multi-view image feature extraction and grid-based three-dimensional workspace methods, the robot capture scene is reconstructed, which solves the problems of large amount of computing and poor generalization capabilities in the existing technology, and realizes high-precision transparent and mirror object scene reconstruction, which improves the success rate of the capture task.

CN119445029BActive Publication Date: 2025-05-09启元实验室
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510026516.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-09
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

In the six-degree of freedom grasping task guided by robot vision, the existing technology has a large amount of calculation and poor generalization ability in three-dimensional scene reconstruction, making it difficult to adapt to different objects and scenes, affecting the success rate of grasping.

Method used

By extracting multi-view image features of the captured image, a gridded three-dimensional workspace including multiple voxels is established, and the captured image volume features are generated, and the robot capture scene is reconstructed using these features, and the scene volume features and background volume features are fused to improve the reconstruction accuracy.

Benefits of technology

It realizes clear geometric reconstruction under sparse and narrow viewing angle conditions, improves the scene reconstruction accuracy of transparent and mirror objects, and enhances the accuracy and efficiency of robot grasping.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119445029B_ABST
    Figure CN119445029B_ABST
Patent Text Reader

Abstract

The present application proposes a method, device and electronic device for reconstructing a robot grasping scene, the method comprising: grasping a scene image according to preset camera parameters to obtain a grasped image; performing multi-view image feature extraction on the grasped image to obtain grasped image features; using the grasped image features to establish a gridded three-dimensional workspace including multiple voxels to generate grasped image volume features; using the grasped image volume features to reconstruct the robot grasping scene to obtain a reconstructed scene. According to an embodiment of the present application, hierarchical aggregation of grasping scenes is performed using multiple feature methods, achieving clear geometric reconstruction under sparse and narrow viewing conditions, providing a basis for subsequent accurate grasping detection. And by generating grasped image volume features, surface reconstruction is achieved using background priors, improving the scene reconstruction accuracy of transparent and mirror objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robots, and in particular to a method and device for reconstructing a robot grasping scene, an electronic device, and a non-transient computer-readable storage medium. Background Art

[0002] In robot vision-guided six-degree-of-freedom (6-DoF) grasping tasks, accurate object recognition and grasping are always difficult due to the challenges brought by the diversity of object shapes, textures, and material properties. For effective grasping operations, the 3D scene geometry must be accurately reconstructed. However, since depth sensors often have unreliable depth measurements of transparent and mirrored objects, the generated geometry information is inaccurate, which in turn affects the success rate of grasping.

[0003] In recent years, researchers have tried to use neural radiance fields to address the challenges of grasping transparent and mirrored objects. Although it can capture complex light interactions and is suitable for scene reconstruction of non-Lambertian objects, it requires dense input images and long training before each grasp, which limits its application in real scenes. Some methods rely on 360-degree full-view acquisition and truncated signed distance function true value supervision, which limits their practical application. In addition, although robust image input grasping is achieved by integrating pre-trained depth estimation and grasp detection models, it still relies on dense perspective input and requires frequent retraining to adapt to different scenes. It is time-consuming and computationally intensive, and has insufficient generalization ability. It is difficult to adapt to different objects and scenes and cannot meet the efficiency requirements in real scenes. Summary of the invention

[0004] The present application aims to propose a method and device for reconstructing a robot grasping scene, an electronic device, and a non-transitory computer-readable storage medium to solve the current problems of large computational complexity and poor generalization ability in three-dimensional scene reconstruction.

[0005] According to one aspect of the present application, a method for reconstructing a robot grasping scene is proposed, comprising: grasping a scene image according to preset camera parameters to obtain a grasped image; performing multi-view image feature extraction on the grasped image to obtain grasped image features; using the grasped image features to establish a gridded three-dimensional workspace including multiple voxels to generate grasped image volume features; and using the grasped image volume features to reconstruct the robot grasping scene to obtain a reconstructed scene.

[0006] According to some embodiments, the captured image includes a scene image and a background image, wherein multi-perspective image feature extraction is performed on the captured image to obtain captured image features, including: multi-perspective image feature extraction is performed on the scene image and the background image respectively to obtain the captured image features, and the captured image features include scene image features and background image features.

[0007] According to some embodiments, a gridded three-dimensional workspace including a plurality of voxels is established using the captured image features to generate captured image volume features, including: generating the captured image volume features using the scene image features and the background image features, wherein the captured image volume features include scene volume features and background volume features.

[0008] According to some embodiments, the robot grasping scene is reconstructed using the grasped image volume feature to obtain a reconstructed scene, including: fusing the scene volume feature and the background volume feature to obtain a fused perspective feature; and reconstructing the robot grasping scene using the fused perspective feature to obtain the reconstructed scene.

[0009] According to some embodiments, the method further includes: obtaining residual image features using the scene image features and the background image features; and fusing the scene volume features, the background volume features, and the residual image features to obtain residual space information.

[0010] According to some embodiments, reconstructing the robot grasping scene using the fused perspective feature to obtain the reconstructed scene includes: reconstructing the robot grasping scene using the fused perspective feature and the residual space information to obtain the reconstructed scene.

[0011] According to some embodiments, the method further includes: generating grasping posture parameters of the robot using the reconstructed scene, so that the robot generates an arm control signal according to the grasping posture parameters to perform a grasping action.

[0012] According to one aspect of the present application, a device for reconstructing a robot grasping scene is proposed, comprising: a grasped image acquisition unit, used to grasp a scene image according to preset camera parameters to obtain a grasped image; a grasped image feature extraction unit, used to perform multi-view image feature extraction on the grasped image to obtain a grasped image feature; a grasped image volume feature generation unit, used to use the grasped image feature to establish a gridded three-dimensional workspace including multiple voxels to generate a grasped image volume feature; and a scene reconstruction unit, used to use the grasped image volume feature to reconstruct the robot grasping scene to obtain a reconstructed scene.

[0013] According to one aspect of the present application, an electronic device is provided, comprising: a processor; and a memory storing a computer program, wherein when the computer program is executed by the processor, the processor executes the method described in any of the preceding embodiments.

[0014] According to one aspect of the present application, a non-transitory computer-readable storage medium is provided, on which computer-readable instructions are stored. When the instructions are executed by a processor, the processor executes the method described in any of the previous embodiments.

[0015] According to the embodiments of the present application, the captured scene is hierarchically aggregated using multiple features, achieving clear geometric reconstruction under sparse and narrow viewing conditions, providing a basis for subsequent accurate capture detection. And by generating the volume features of the captured image, the surface reconstruction is completed using the background prior, improving the scene reconstruction accuracy of transparent and mirror objects.

[0016] It should be understood that the foregoing general description and the following detailed description are exemplary only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. By describing the exemplary embodiments in detail with reference to the drawings, the above and other objectives, features and advantages of the present invention will become more apparent.

[0018] Figure 1 A flow chart of a method for reconstructing a robot grasping scene according to an example embodiment of the present application is shown.

[0019] Figure 2 A flow chart of another method for reconstructing a robot grasping scene according to an exemplary embodiment of the present application is shown.

[0020] Figure 3 A block diagram of a device for reconstructing a robot grasping scene according to an example embodiment of the present application is shown.

[0021] Figure 4 An electronic device according to an exemplary embodiment of the present application is shown. DETAILED DESCRIPTION

[0022] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that the present invention will be comprehensive and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The same figures in the drawings represent the same or similar parts, and thus their repeated description will be omitted.

[0023] The described features, structures or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced without one or more of these specific details, or other modes, components, materials, devices or operations may be adopted. In these cases, known structures, methods, devices, implementations, materials or operations will not be shown or described in detail.

[0024] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.

[0025] The terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products or devices.

[0026] The specific embodiments according to the present invention are described in detail below with reference to the accompanying drawings.

[0027] Figure 1 A flow chart of a method for reconstructing a robot grasping scene according to an exemplary embodiment of the present application is shown. Figure 1 The method shown includes steps S101 , S103 , S105 and S107 .

[0028] like Figure 1 As shown, in step S101, a scene image is captured according to preset camera parameters to obtain a captured image.

[0029] The camera parameters record the spatial position and orientation information of each image. In a specific embodiment, the RGB camera is installed at the end of the robot's mechanical arm to ensure the consistency and repeatability of the image acquisition angle.

[0030] According to an embodiment of the present application, the preset camera parameters include a radius, a polar angle and / or an azimuth angle.

[0031] In a specific embodiment, the radius , polar angle , azimuth .

[0032] In some embodiments, the captured images are collected according to a fixed trajectory. For example, a spiral trajectory is used to collect the captured scene images, covering one-third of the spherical surface, achieving a balance between geometric reconstruction and capture prediction under sparse viewing angle conditions, and improving the flexibility of the system, so that this embodiment not only balances the scene coverage and image sparsity, but also can obtain clear geometric reconstruction under sparse viewing angle and narrow viewing angle conditions, so as to improve the accuracy of subsequent capture detection using the reconstruction model.

[0033] According to some other embodiments, after image acquisition is completed, data preprocessing is required for the captured image. The preprocessing method includes but is not limited to image denoising, illumination correction and resolution adjustment to ensure the quality of the input image. The preprocessed image is transmitted to the encoder for feature extraction, and step S103 is executed.

[0034] It should be noted here that the embodiments of the present application can be applied to a variety of capture scenarios, including but not limited to transparent and mirror object scenes.

[0035] In step S103, multi-view image feature extraction is performed on the captured image to obtain captured image features.

[0036] In the embodiment of the present application, the captured image includes a scene image and a background image, wherein the background image is used to generate prior information to help enhance the reconstruction effect of transparent and mirror objects.

[0037] In step S103, multi-view image feature extraction is performed on the scene image and the background image respectively to obtain captured image features, wherein the captured image features include scene image features and background image features.

[0038] In a specific embodiment, when extracting features from captured images, the scene image and the background image are passed through an encoder (e.g., Res U-net) to extract multi-view image features. Each image is processed through a convolutional neural network to generate spatial features of the image blocks.

[0039] To improve the reconstruction accuracy of transparent and mirror objects, in some embodiments, a lightweight convolutional network is introduced into the encoder structure to sample K points from the camera center to the pixel for each ray. At each sampling point, multi-view features are aggregated through the geometric network and the weight network, and the SDF (Signed Distance Function) value and mixing weight are estimated. Finally, the radiance of each point is calculated to ensure efficient feature extraction with limited computing resources.

[0040] In step S105 , a gridded three-dimensional workspace including a plurality of voxels is established using the captured image features to generate captured image volume features.

[0041] According to an embodiment of the present application, a captured image volume feature is generated using scene image features and background image features, wherein the captured image volume feature includes a scene volume feature and a background volume feature.

[0042] In a specific embodiment, in step S105, a corresponding three-dimensional workspace needs to be established according to the captured scene, and gridded to obtain voxels corresponding to each grid in the workspace. At the center of each voxel, by calculating the mean and variance of the multi-view image features obtained in step S103, that is, the captured image features, 3D U-Net is used to infer the shape details, and a high-level shape prior volume, that is, the captured image volume feature, is generated to provide a global shape prior for the reconstruction of the captured scene.

[0043] In step S107, the robot grasping scene is reconstructed using the grasping image volume features to obtain a reconstructed scene.

[0044] According to an embodiment of the present application, step S107 includes:

[0045] Step S1071, fusing the scene volume feature and the background volume feature to obtain a fused viewing angle feature; and

[0046] Step S1073, reconstructing the robot grasping scene using the fused view features to obtain a reconstructed scene.

[0047] In some embodiments, when the scene volume features and the background volume features are fused, the multi-view image features (scene image features and background image features) obtained in step S103 and the volume features obtained from the shape prior volume (including the scene volume features and the background volume features) are used to generate a unified view feature through a self-attention mechanism to obtain a fused view feature. In a specific embodiment, in the process of generating a unified view feature through a self-attention mechanism, each image feature is projected into the three-dimensional voxel space of the scene, and a weighted feature is generated through trilinear interpolation to improve the reconstruction effect of transparent and mirror objects.

[0048] According to other embodiments, when the robot grasping scene is reconstructed using the fused view feature, the spatial information is fused using the ray marching technique to generate geometric features, and the geometric features and the spatial coding are combined to generate the reconstructed scene. This embodiment uses the non-local characteristics of the ray marching technique and the fixed position coding embedding to provide occlusion information and improve the surface reconstruction effect of transparent objects.

[0049] In some embodiments, the fused view features are embedded in the voxel space, and the spatial information is optimized by combining position encoding and sampling order through a multi-layer perceptron (eg, MLP) to improve the accuracy of transparent object surface reconstruction.

[0050] according to Figure 1 The embodiment shown uses multiple features to hierarchically aggregate the captured scene, achieving clear geometric reconstruction under sparse and narrow viewing conditions, providing a basis for subsequent accurate capture detection. And by generating volume features of the captured image, surface reconstruction is achieved using background priors, improving the scene reconstruction accuracy of transparent and mirror objects.

[0051] Figure 2 A flowchart of another method for reconstructing a robot grasping scene according to an exemplary embodiment of the present application is shown. Figure 2 The method shown includes Figure 1 In addition to steps S101, S103, S105 and S107 shown in FIG. 1 , steps S109 and S111 are also included. Step S107 includes step S1071 and step S1073. For the sake of simplicity, this embodiment will not be further described. Figure 1 The similarities between the two are omitted and only the differences between them are described.

[0052] like Figure 2 As shown, in step S109, the scene image features and the background image features are used to obtain the residual image features.

[0053] In order to improve the segmentation effect of foreground objects in the reconstruction of transparent and mirror objects, in some embodiments, feature subtraction is performed using scene image features and background image features to obtain residual image features.

[0054] In order to highlight the details of the foreground object and improve the accuracy of reconstruction, in some embodiments, the background prior is used to enhance the feature. For example, the foreground attention feature is obtained by performing background subtraction in the residual image feature.

[0055] In step S111, the scene volume features, the background volume features, and the residual image features are fused to obtain residual space information.

[0056] The residual features implicitly indicate the occupancy of the foreground object. In a specific embodiment, in step S111, the residual image features are expanded from 2D to 3D residual volume features, and spatial information is aggregated with the captured image volume features. For example, the 3D residual volume features and the captured image volume features are first feature fused through the query Transformer in the multi-head self-attention mechanism, and then linear fusion is performed to optimize the reconstruction effect of transparent and mirror objects. This method pays more attention to the difference between the foreground object and the background through the subtractive attention mechanism, so as to perform more accurate surface reconstruction of transparent objects and mirror objects during reconstruction.

[0057] According to an embodiment of the present application, in step S1073, the robot grasping scene is reconstructed using fused viewing angle features and residual space information to obtain a reconstructed scene.

[0058] For example, MLP is used to fuse the fused view features and residual space information to generate the final reconstructed scene.

[0059] according to Figure 2 The embodiment shown, without the need for dense perspective or geometric supervision, improves the reconstructed scene's focus on foreground objects by integrating the residual features of the captured image into the reconstructed scene, thereby improving the system's practicality and real-time performance. It overcomes the limitations of traditional depth sensors in processing complex optical surfaces, and provides a basis for the accuracy of subsequent capture detection; and by utilizing the global implicit occupancy information in the residual feature map, spatial perception can be enhanced, which is particularly suitable for reconstructing scenes that process transparent and mirror objects.

[0060] According to an embodiment of the present application, after the scene reconstruction is completed, the reconstructed scene is also used for grasping detection, for example, to generate grasping posture parameters of the robot so that the robot generates an arm control signal according to the grasping posture parameters and performs a grasping action.

[0061] For example, in a specific embodiment, a voxel grid is generated through volume query and input into a 3D CNN network to generate grasping posture parameters of each voxel center, including grasping center, grasping quality, grasping direction, and gripper width.

[0062] Since this embodiment does not rely on explicit foreground segmentation or collision detection, it is able to automatically generate a grasping pose based on the reconstructed geometry, ensuring the accuracy of collision-free grasping. In some embodiments, after grasp detection is completed, the system applies mask processing and non-maximum suppression to ensure the stability of the grasping result.

[0063] In some other embodiments, the grab detection results are also used for feedback adjustment.

[0064] For example, the system monitors the operation in real time during the grasping process and adjusts the grasping posture based on sensor feedback to ensure high stability. After a successful grasp, the system records the grasping performance indicators for subsequent optimization.

[0065] The above mainly introduces the embodiments of the present application from the perspective of the method. Those skilled in the art should easily appreciate that, in combination with the operations or steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Those skilled in the art can use different ways to implement the described functions for each specific operation or method, and such implementation should not be considered to be beyond the scope of the present application.

[0066] The following describes the device embodiments of the present application. For details not described in the device embodiments of the present application, reference can be made to the method embodiments of the present application.

[0067] Figure 3 A block diagram of a device for reconstructing a robot grasping scene according to an exemplary embodiment of the present application is shown, Figure 3 The device shown includes a captured image acquisition unit 301, a captured image feature extraction unit 303, a captured image volume feature generation unit 305 and a scene reconstruction unit 307. The captured image acquisition unit 301 is used to capture a scene image according to preset camera parameters to obtain a captured image; the captured image feature extraction unit 303 is used to perform multi-view image feature extraction on the captured image to obtain captured image features; the captured image volume feature generation unit 305 is used to establish a gridded three-dimensional workspace including multiple voxels using the captured image features to generate captured image volume features; the scene reconstruction unit 307 is used to reconstruct the robot capture scene using the captured image volume features to obtain a reconstructed scene.

[0068] According to the embodiments of the present application, Figure 3 After the device shown is constructed, it is necessary to Figure 3 The apparatus shown is used for training. According to some embodiments, 2304 rays are sampled per batch during training, the learning rate is 1e-4 and exponential decay is adopted. By directly converting the corresponding grasping posture parameters into the robot's arm control signal, an end-to-end training method is adopted to optimize the reconstruction and grasping tasks at the same time. The corresponding loss functions include but are not limited to photometric loss, grasping loss and regularization loss. Among them, the photometric loss is used to compare the pixel color of the rendered image with the real image, and the mean square error (MSE) is used to calculate the photometric loss; the grasping loss is used to supervise the grasping quality, direction and gripper width, and the loss is calculated only for successful grasping; the regularization loss includes the Eikonal loss to ensure that the SDF gradient is close to 1.

[0069] According to the embodiments of the present application, the grasping success rate and reconstruction accuracy are significantly better than the traditional methods, and collision-free six-degree-of-freedom grasping operations can be achieved under narrow and sparse viewing angles by relying on background priors and multi-view feature fusion.

[0070] In this embodiment, after the scene reconstruction is completed, the corresponding grasping posture parameters are directly converted into the robot's arm control signal, thereby achieving seamless integration of reconstruction and grasping, so that it can quickly adapt to complex scenes, and is particularly suitable for grasping operations of transparent and mirror objects.

[0071] Figure 4 An electronic device according to an exemplary embodiment of the present application is shown. Figure 4 The electronic device 200 according to this embodiment of the present application is described. Figure 4 The electronic device 200 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0072] like Figure 4 As shown, the electronic device 200 is in the form of a general computing device. The components of the electronic device 200 may include, but are not limited to: at least one processing unit 210, at least one storage unit 220, a bus 230 connecting different system components (including the storage unit 220 and the processing unit 210), a display unit 240, etc.

[0073] The storage unit stores program codes, which can be executed by the processing unit 210, so that the processing unit 210 executes the methods described in this specification according to various exemplary embodiments of the present application. For example, the processing unit 210 can execute the following: Figure 1 The method shown in .

[0074] The storage unit 220 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 2201 and / or a cache storage unit 2202 , and may further include a read-only storage unit (ROM) 2203 .

[0075] The storage unit 220 may also include a program / utility 2204 having a set (at least one) of program modules 2205, such program modules 2205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0076] Bus 230 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0077] The electronic device 200 may also communicate with one or more external devices 300 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 200, and / or may communicate with any device that enables the electronic device 200 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 250. In addition, the electronic device 200 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 260. The network adapter 260 may communicate with other modules of the electronic device 200 via a bus 230. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 200, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0078] Through the description of the above implementation methods, it is easy for those skilled in the art to understand that the example implementation methods described here can be implemented by software, or by combining software with necessary hardware. The technical solution according to the implementation method of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the above method according to the implementation method of the present application.

[0079] The software product can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0080] Computer readable storage media may include data signals propagated in baseband or as part of a carrier wave, wherein readable program codes are carried. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or device. The program codes contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.

[0081] Program code for performing the operations of the present application may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).

[0082] The computer-readable medium carries one or more programs. When the one or more programs are executed by a device, the computer-readable medium implements the aforementioned functions.

[0083] Those skilled in the art will appreciate that the above modules can be distributed in the device according to the description of the embodiment, or can be changed accordingly and only used in one or more devices different from the embodiment. The modules of the above embodiments can be combined into one module, or further divided into multiple sub-modules.

[0084] According to an embodiment of the present application, a computer program is provided, including a computer program or an instruction. When the computer program or the instruction is executed by a processor, the method described above can be executed.

[0085] The embodiments of the present application are described in detail above. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and its core idea of ​​the present application. At the same time, changes or deformations made by those skilled in the art based on the ideas of the present application, the specific implementation methods and the scope of application of the present application, all belong to the scope of protection of the present application. In summary, the content of this specification should not be construed as a limitation on the present application.

Claims

1. A method for reconstructing a robot grasping scene, characterized in that: include: Capturing a scene image according to preset camera parameters to obtain a captured image; Performing multi-view image feature extraction on the captured image to obtain captured image features; Using the captured image features, a gridded three-dimensional workspace including a plurality of voxels is established to generate captured image volume features; Reconstructing the robot grasping scene using the grasping image volume feature to obtain a reconstructed scene; in, The captured image includes a scene image and a background image, wherein multi-view image feature extraction is performed on the captured image to obtain captured image features, including: Performing multi-view image feature extraction on the scene image and the background image respectively to obtain the captured image features, wherein the captured image features include scene image features and background image features; Using the captured image features, a gridded three-dimensional workspace including a plurality of voxels is established to generate captured image volume features, including: Generating the captured image volume feature by using the scene image feature and the background image feature, wherein the captured image volume feature includes a scene volume feature and a background volume feature; Reconstructing the robot grasping scene by using the grasping image volume feature to obtain a reconstructed scene, including: Fusing the scene volume feature and the background volume feature to obtain a fused viewing angle feature; The robot grasping scene is reconstructed using the fused view feature to obtain the reconstructed scene.

2. The method according to claim 1, characterized in that: Also includes: Obtaining residual image features using the scene image features and the background image features; The scene volume feature, the background volume feature, and the residual image feature are fused to obtain residual space information.

3. The method according to claim 2, characterized in that Reconstructing the robot grasping scene by using the fused view feature to obtain the reconstructed scene includes: The robot grasping scene is reconstructed using the fused viewing angle feature and the residual space information to obtain the reconstructed scene.

4. The method according to claim 1, characterized in that: Also includes: The reconstructed scene is used to generate the grasping posture parameters of the robot, so that the robot generates an arm control signal according to the grasping posture parameters and performs a grasping action.

5. A device for reconstructing a robot grasping scene, characterized in that: For executing the method according to any one of claims 1 to 4, the device comprises: A captured image acquisition unit, used to capture a scene image according to preset camera parameters to obtain a captured image; A captured image feature extraction unit, used for performing multi-view image feature extraction on the captured image to obtain captured image features; A captured image volume feature generating unit, used for establishing a gridded three-dimensional workspace including a plurality of voxels using the captured image feature to generate the captured image volume feature; The scene reconstruction unit is used to reconstruct the robot grasping scene by using the grasping image volume feature to obtain a reconstructed scene.

6. An electronic device comprising: processor; as well as A memory storing a computer program, which, when executed by the processor, enables the processor to perform the method according to any one of claims 1 to 4.

7. A non-transitory computer-readable storage medium having computer-readable instructions stored thereon, which, when executed by a processor, causes the processor to execute the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Image processing method and device, computer equipment and storage medium

    CN112053285A

  • Heart three-dimensional modeling method and device based on point cloud generation

    CN118334238A

  • Mechanical arm grabbing method and equipment for mirror surface medical instrument, medium and product

    CN118453114A