Mechanical arm obstacle avoidance method and device based on structured light
By acquiring structured light depth images to reconstruct the three-dimensional model and planning the path, the technical gap in robotic arm obstacle avoidance is solved, and the effect of real-time dynamic obstacle avoidance is achieved.
Patent Information
- Application Number
- CN202510393511.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
AI Technical Summary
The structured light technology has not been applied to robotic robotic arm obstacle avoidance in the prior art, resulting in the inability to achieve real-time dynamic obstacle avoidance.
By obtaining structured light-based depth images of the surrounding environment of the robot arm, reconstructing the environmental 3D model, generating the planned path of the robot arm, and updating the three-dimensional model in real time to control the planned path of the robot arm execution.
Real-time dynamic obstacle avoidance of robotic robot arms is realized to ensure safe operation in a dynamic environment.
Smart Images

Figure CN120244955A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of medical image recognition, and more particularly, to a method and device for obstacle avoidance of a robotic arm based on structured light. Background Art
[0002] In orthopedic joint replacement surgery, the robot needs to operate precisely to avoid colliding with surrounding obstacles. Structured light technology can sense the dynamic environment within the field of view in real time, but currently there is no solution to apply structured light technology to the obstacle avoidance of robotic arms. Summary of the Invention
[0003] The problem solved by this application is that currently there is no solution to apply structured light technology to the obstacle avoidance of robotic arms.
[0004] To solve the above problems, the first aspect of this application provides a method for obstacle avoidance of a robotic arm based on structured light, including:
[0005] Obtaining a depth image based on structured light of the environment around the robotic arm;
[0006] Reconstructing the depth image to obtain a three-dimensional model of the environment;
[0007] Generating a planned path for the robotic arm based on the three-dimensional model of the environment;
[0008] Updating the three-dimensional model of the environment in real time and controlling the robotic arm to execute the planned path.
[0009] The second aspect of this application provides a device for obstacle avoidance of a robotic arm based on structured light, which includes:
[0010] An image acquisition module for obtaining a depth image based on structured light of the environment around the robotic arm;
[0011] A three-dimensional reconstruction module for reconstructing the depth image to obtain a three-dimensional model of the environment;
[0012] A path planning module for generating a planned path for the robotic arm based on the three-dimensional model of the environment;
[0013] A robotic arm control module for updating the three-dimensional model of the environment in real time and controlling the robotic arm to execute the planned path.
[0014] The third aspect of this application provides an electronic device, which includes: a memory and a processor;
[0015] The memory is used for storing programs;
[0016] The processor is coupled to the memory and is used for executing the program to:
[0017] Obtain a structured-light-based depth image of the environment around the robotic arm;
[0018] Reconstruct the depth image to obtain a three-dimensional model of the environment;
[0019] Generate a planned path for the robotic arm based on the three-dimensional model of the environment;
[0020] Update the three-dimensional model of the environment in real time and control the robotic arm to execute the planned path.
[0021] The fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to implement the above-mentioned structured-light-based robotic arm obstacle avoidance method.
[0022] In the present application, a three-dimensional model of the environment around the robotic arm is constructed through structured light, and then path planning and execution of the robotic arm are realized within the three-dimensional model of the environment, thereby applying the structured light technology to the obstacle avoidance of the robotic arm of the robot to achieve real-time dynamic obstacle avoidance. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 FIG. is a flowchart of a structured-light-based robotic arm obstacle avoidance method according to an embodiment of the present application;
[0024] Figure 2 FIG. is an architecture diagram of three-dimensional model construction of a structured-light-based robotic arm obstacle avoidance method according to an embodiment of the present application;
[0025] Figure 3 FIG. is an architecture diagram of multi-scale fusion of a structured-light-based robotic arm obstacle avoidance method according to an embodiment of the present application;
[0026] Figure 4 FIG. is an architecture diagram of graph convolution processing of a structured-light-based robotic arm obstacle avoidance method according to an embodiment of the present application;
[0027] Figure 5 FIG. is a structural block diagram of a structured-light-based robotic arm obstacle avoidance device according to an embodiment of the present application;
[0028] Figure 6 FIG. is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] To make the above objects, features, and advantages of the present application more obvious and understandable, the following detailed description of the specific embodiments of the present application is provided with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.
[0030] It should be noted that, unless otherwise specified, the technical terms or scientific terms used in this application shall have the ordinary meanings understood by those skilled in the art to which this application pertains.
[0031] In view of the above problems, this application provides a new obstacle avoidance solution for a robotic arm based on structured light, which can embed a pre-trained large model into an image fusion module to eliminate the problem that it is difficult to achieve high-precision registration and fusion of CT images and MRI images.
[0032] An embodiment of this application provides a method for obstacle avoidance of a robotic arm based on structured light. The specific solution of this method is Figures 1 - 4 as shown. This method can be executed by an obstacle avoidance device for a robotic arm based on structured light, and this obstacle avoidance device for a robotic arm based on structured light can be integrated in electronic devices such as a computer, a server, a computer, a server cluster, and a data center. Combining Figure 1 as shown, it is a flowchart of a method for obstacle avoidance of a robotic arm based on structured light according to an embodiment of this application; wherein, the method for obstacle avoidance of a robotic arm based on structured light includes:
[0033] S101, obtaining a depth image based on structured light of the environment around the robotic arm;
[0034] In this application, near-infrared structured light (wavelength 850 - 940 nm) is used to reduce environmental light interference (such as operating room lights).
[0035] In this application, the structured light in a dynamic scene is a single-shot projected pseudo-random dot matrix, and the structured light in a static scene is multi-frame phase-shifted fringes (Gray code + sine fringes) to improve depth accuracy.
[0036] S102, reconstructing the depth image to obtain an environmental three-dimensional model;
[0037] In this step, the depth map is converted into a three-dimensional Mesh model that can be used for path planning.
[0038] S103, generating a planned path for the robotic arm based on the environmental three-dimensional model;
[0039] In this step, a collision-free motion path is generated in the three-dimensional Mesh model.
[0040] Among them, the shortest path can be searched in the voxelized three-dimensional grid through the A* algorithm; or a planned path can be generated through the RRT (Rapidly-Exploring Random Tree) algorithm.
[0041] In this application, the specific process of the random tree algorithm is as follows: Initialize the random tree, generate a source point, a target point, and the radius of the target point, and generate an obstacle matrix; generate random points; find the point on the tree closest to the random point; extend from the closest point to the random point by the distance from the closest point to the random point to obtain a new point; if the straight line segment from the closest point to the new point does not pass through an obstacle, the new point is not empty, otherwise, the new point is empty; if the new point is empty, return to the step of generating random points; add the new point to the tree as a node, and use the closest point as the parent node of the new point and record it; if the distance from the new point to the target point is less than the preset distance, the algorithm ends, otherwise return to the step of generating random points.
[0042] S104. Update the three-dimensional environment model in real time, and control the robotic arm to execute the planned path.
[0043] In this step, dynamically respond to environmental changes to ensure real-time obstacle avoidance.
[0044] Preferably, control the robotic arm to execute the planned path, and update the three-dimensional environment model in real time to predict whether there is a collision in the movement trajectory of the robotic arm; in the case of a collision, regenerate the movement path of the robotic arm based on the current three-dimensional environment model, and control the robotic arm to execute the regenerated movement path. In this way, it can be ensured that the robotic arm executes the planned path and reaches the end point.
[0045] In this application, a three-dimensional environment model around the robotic arm is constructed by structured light, and then path planning and execution of the robotic arm are realized within this three-dimensional environment model, thereby applying the structured light technology to obstacle avoidance of the robotic arm to achieve real-time dynamic obstacle avoidance.
[0046] In this application, the structured light is projected onto the surgical area, and the deformed light spots are captured by a camera to generate a depth map of the surgical scene in real time (such as ToF or fringe projection technology), and a dynamic three-dimensional environment model is constructed.
[0047] In one implementation, the obtaining of the structured light-based depth image of the environment around the robotic arm includes:
[0048] Project a grid-structured light onto the environment around the robotic arm;
[0049] Obtain the reflected structured light data and synchronously obtain the visible light data;
[0050] Generate a depth image based on the structured light data;
[0051] Generate an RGB image based on the visible light data;
[0052] Convert the depth image and the RGB image into the same coordinate system.
[0053] In this application, a programmable grid pattern (such as a sine stripe or a pseudo-random dot matrix) can be projected by a DLP projector; or projected by laser structured light.
[0054] Preferably, in a static scene, the four-step phase-shifting method (phase stripes of 0°, 90°, 180°, 270°) is used to solve the absolute phase. In a dynamic scene, a single-frame pseudo-random grid (such as M-array coding) is used for coding.
[0055] In this application, the extrinsic calibration of the projector-camera (using a checkerboard or a dedicated calibration board) is performed to ensure that the geometric relationship between projection and imaging is known.
[0056] In this application, the structured light imaging is performed by capturing the reflected structured light data; the visible light imaging is performed by capturing the visible light data; and the synchronization acquisition is performed to ensure the consistency between the structured light imaging and the visible light imaging.
[0057] In this application, for the structured light data, the process of generating a depth image is as follows: the wrapped phase is calculated for the four-step phase-shifted images, the absolute phase is solved using the Gray code or the multi-frequency method, and the phase is converted into depth through the triangulation formula to obtain the depth image.
[0058] In this application, the color and texture information of the environment is obtained through the visible light data to assist semantic segmentation. The RGB image is directly generated through the visible light data.
[0059] In this application, the relative pose of the RGB camera and the infrared camera is obtained, and based on the relative pose, the depth map is mapped to the perspective of the RGB camera through perspective transformation.
[0060] In one implementation, as shown in combination with Figure 2 The reconstruction of the depth image to obtain the three-dimensional model of the environment includes:
[0061] Feature extraction is performed on the depth image to obtain a depth feature map;
[0062] Feature extraction is performed on the RGB image to obtain an RGB feature map;
[0063] The depth feature map and the RGB feature map are fused at multiple scales to obtain fused feature maps of multiple sizes;
[0064] The fused feature maps of multiple sizes are processed by graph convolution to obtain a three-dimensional graph;
[0065] The three-dimensional model of the environment is constructed based on the three-dimensional image.
[0066] In this application, through feature extraction, geometric and texture features are respectively extracted from the depth map and the RGB map.
[0067] In this application, multi-scale feature fusion combines geometric (depth) and semantic (RGB) information to improve the integrity of reconstruction.
[0068] In this application, through graph convolution processing, the 2D fusion features are mapped to the 3D space to construct a Mesh model.
[0069] Among them, the specific process of graph convolution processing includes:
[0070] Initial Mesh generation: Basic shape: ellipsoid (156 vertices, 462 edges), centered and aligned with the camera frustum; Vertex feature initialization: Coordinate feature: Initial 3D position (x, y, z) of the vertex. Image feature: Project the 2D feature onto the Mesh vertex through bilinear interpolation.
[0071] Convolutional network (GCN) architecture: Graph definition: Vertex (Node): Position + feature vector (such as 64 dimensions). Edge (Edge): Based on Delaunay triangulation or K-nearest neighbor (K = 6); Cascade deformation: 3 GCN modules gradually deform, and each module is followed by graph upsampling (subdivide the mesh to increase details).
[0072] Loss function: Chamfer Loss: Minimize the distance between the predicted vertices and the real point cloud; Normal consistency loss: Constrain the surface smoothness; Laplacian regularization: Prevent excessive distortion.
[0073] In this application, based on the three-dimensional image, the three-dimensional environmental model is constructed, and the optimized Mesh model is converted into a three-dimensional environmental representation that can be used for path planning. Specifically, it includes: converting the Mesh into an occupancy grid map (resolution 1mm 3 ), which is convenient for collision detection; classifying the Mesh patches according to the RGB features (such as "blood vessel", "bone"), and marking the prohibited areas.
[0074] In this application, the graph convolution directly processes the irregular Mesh data, which is suitable for complex topological structures.
[0075] In one implementation, combined with Figure 3 As shown, in the feature extraction of the depth image to obtain the depth feature map, the depth image is downsampled to obtain depth feature maps of multiple sizes.
[0076] In one implementation, combined with Figure 3 As shown, in the feature extraction of the RGB image to obtain the RGB feature map, the RGB image is downsampled to obtain RGB feature maps of multiple sizes, and the sizes of the RGB feature maps and the depth feature maps correspond one by one.
[0077] In one embodiment, the depth feature map and the RGB feature map are fused at multiple scales to obtain fused feature maps of multiple sizes, including:
[0078] After upsampling the RGB feature map with the smallest size, it is fused with the depth feature map and the RGB feature map of the previous size to obtain a fused map of the previous size;
[0079] Through cyclic processing, fused maps of the second size, the third size, the fourth size, and the largest size are obtained in sequence;
[0080] The fused maps of the second size, the third size, and the largest size are the fused feature maps of multiple sizes.
[0081] In one embodiment, as shown in Figure 3 the number of channels of the fused feature maps of multiple sizes is 64, 256, and 512.
[0082] As shown in Figure 3 assuming that the maximum size - minimum size is L1 - L5 respectively, then in the figure, the maximum size L1 has 64 channels, L2 has 128 channels, L3 has 256 channels, L4 has 512 channels, and the depth feature map has 1024 channels.
[0083] Among them, the depth image is a 64-channel image, and through two convolutions, an L1 depth feature map is obtained; the L1 depth feature map undergoes pooling (downsampling) and convolution to obtain an L2 depth feature map; the L2 depth feature map undergoes pooling (downsampling) and convolution to obtain an L3 depth feature map; the L3 depth feature map undergoes pooling (downsampling) and convolution to obtain an L4 depth feature map.
[0084] Among them, the RGB image is a 64-channel image, and through two convolutions, an L1 RGB feature map is obtained; the L1 RGB feature map undergoes pooling (downsampling) and convolution to obtain an L2 RGB feature map; the L2 RGB feature map undergoes pooling (downsampling) and convolution to obtain an L3 RGB feature map; the L3 RGB feature map undergoes pooling (downsampling) and convolution to obtain an L4 RGB feature map; the L4 RGB feature map undergoes pooling (downsampling) to obtain an L5 RGB feature map.
[0085] In this application, the sizes of the RGB feature map and the depth feature map correspond one by one, which means that the feature map sizes corresponding to L1 with 64 channels, L2 with 128 channels, L3 with 256 channels, and L4 with 512 channels correspond one by one. The depth feature map does not have an L5 size.
[0086] In this application, the depth feature map and the RGB feature map are multi-scale fused to obtain fused feature maps of multiple sizes. Among them, the RGB feature map of L5 is convolved and upsampled to obtain the upsampled feature map of L4; the RGB feature map of L4, the depth feature map of L4, and the upsampled feature map of L4 are fused to obtain the fused map of L4; the fused map of L4 is upsampled to obtain the upsampled feature map of L3; the RGB feature map of L3, the depth feature map of L3, and the upsampled feature map of L3 are fused to obtain the fused map of L3; the fused map of L3 is upsampled to obtain the upsampled feature map of L2; the upsampled feature map of L2 is convolved to obtain the fused map of L2; the fused map of L2 is upsampled to obtain the upsampled feature map of L1; the RGB feature map of L1, the depth feature map of L1, and the upsampled feature map of L1 are fused to obtain the fused map of L1.
[0087] Among them, the fused maps of L1, L3, and L4 are output as fused feature maps of multiple sizes and output to graph convolution.
[0088] In this application, it should be noted that the generation process of the fused map of L2 is different from that of the other fused maps (the fused maps of L1, L3, and L4).
[0089] In this application, the RGB feature map (256) of L3, the depth feature map (256) of L3, and the upsampled feature map (256) of L3 are fused to obtain the fused map (256) of L3; it can be: the RGB feature map of L3 and the depth feature map of L3 are added to obtain an added feature map; the added feature map and the upsampled feature map of L3 are combined to obtain a combined feature map (512); the combined feature map is convolved to obtain the fused map (256) of L3.
[0090] In this application, the RGB image passes through the feature extraction module for rough reconstruction of the contour shape, and the depth image passes through the feature extraction module for fine reconstruction of the overall details and three-dimensional dimensions.
[0091] In this application, by simultaneously using the depth image and the RGB image as inputs, the model can extract richer and complementary feature information from two different visual signals.
[0092] In this application, in combination with Figure 4 As shown, the fused feature maps of multiple sizes are subjected to graph convolution processing to obtain a three-dimensional graph. In the three-dimensional Mesh model, it is basically composed of a series of vertices, edges, and faces, and these elements together outline the shape of the three-dimensional object.
[0093] In combination with Figure 4As shown, the three-dimensional vertices of the Mesh model are first projected onto the image plane, and the perceptual features are pooled from the image feature extraction layer through bilinear interpolation, so that each vertex extracts relevant features according to its projection position on the image and combines them with the existing three-dimensional shape features of the vertex, and then inputs them into the graph-based residual network. Finally, the updated coordinates of the vertices and the three-dimensional shape features are obtained through the output of the residual network.
[0094] In this application, graph convolution processing can enable the Mesh reconstruction module to gradually and finely reconstruct the Mesh model according to the features of the input image, ensuring that the generated three-dimensional model is both accurate and has rich details.
[0095] Preferably, after obtaining the depth image / RGB image, preprocessing is performed on the depth image / RGB image through adaptive processing, which can greatly reduce the preprocessing time.
[0096] Preferably, the adaptive processing process includes:
[0097] Partition the depth image / RGB image to obtain independent blocks;
[0098] For each independent block, obtain the first neighborhood block and the second neighborhood block with different spacings;
[0099] Generate the first feature block based on the independent block and the first neighborhood block;
[0100] Generate the second feature block based on the independent block and the second neighborhood block;
[0101] Perform feature compression on the first feature block and the second feature block to obtain a compressed block;
[0102] Traverse all independent blocks and generate the preprocessed depth image / RGB image based on the obtained compressed blocks.
[0103] In this application, partitioning the depth image / RGB image means dividing the depth image / RGB image into corresponding image blocks through a checkerboard; where the image block can be at the pixel level (i.e., each pixel is an image block), or at other levels, and the specific partitioning depends on the actual processing situation.
[0104] In this application, the image is divided into blocks of the same size using a sliding window or a fixed step size.
[0105] It should be noted here that if the depth image / RGB image is a two-dimensional image, then the two-dimensional image is directly divided by a checkerboard, and each grid is an image block.
[0106] Preferably, in the present application, each image block has 100 - 1000 pixels, so that on the basis of ensuring the generation accuracy and reducing the calculation amount, more feature calculations are performed between local regions.
[0107] In the present application, an image block is selected as an independent block, and the adjacent image blocks above, below, to the left, and to the right of the independent block are the first neighborhood blocks; the image blocks one grid apart above, below, to the left, and to the right of the independent block are the second neighborhood blocks. The distances between the first neighborhood blocks and the second neighborhood blocks and the independent block are different.
[0108] In the present application, neighborhood information is extracted for each independent block to capture the local structure.
[0109] In the present application, a first feature block is generated by using the independent block and its first neighborhood blocks to generate a local feature representation. Specifically, it can be: the independent block and the first neighborhood blocks are processed through a convolutional layer and an attention layer to obtain the first feature block.
[0110] In the present application, the specific structure and specific parameters of the convolutional layer and the attention layer can be obtained according to the training data or determined according to the actual situation.
[0111] It should be noted that in the present application, there are four first neighborhood blocks and multiple first feature blocks.
[0112] In the present application, the independent block and the first neighborhood blocks are processed through a convolutional layer and an attention layer to obtain the first feature block. The specific process is: the independent block and the four neighborhood blocks are concatenated together to form a multi-channel input, and the convolutional layer is used to extract features from the concatenated block; the self-attention mechanism or the channel attention mechanism is used to enhance important features, calculate the attention weights, and weight the output of the convolutional layer to enhance important features; the output of the attention layer is split into multiple feature blocks, and each feature block corresponds to the processing result of the independent block and at least one neighborhood block.
[0113] In the present application, a second feature block is generated by using the independent block and its second neighborhood blocks to generate a more extensive local feature representation. The specific generation process is the same as that of the first feature block, except that the parameters of the convolutional layer and the attention layer are different.
[0114] In the present application, the generated feature blocks are compressed into a more compact representation to reduce the calculation amount and retain key information. Pooling operations (such as max pooling or average pooling) or fully connected layers are used for feature compression.
[0115] In this way, through compression, multiple first feature blocks and second feature blocks are compressed into a compressed block, which corresponds to the independent block in size and position and is used to replace the independent block. All image blocks are replaced by compressed blocks to obtain the processed depth image / RGB image.
[0116] In this application, by means of traversal, each image block of the depth image / RGB image is traversed to obtain the corresponding compressed block.
[0117] In this application, for the image blocks / independent blocks near the edge, their first neighborhood blocks and second neighborhood blocks are incomplete. At this time, the first neighborhood blocks and second neighborhood blocks at the relative positions are copied for completion. For example, if the first neighborhood block above the independent block does not exist, the first neighborhood block below is copied and used as the block above.
[0118] In this application, through completion, the processing accuracy of the edge-adjacent image blocks is greatly improved.
[0119] The embodiment of this application provides a manipulator obstacle avoidance device based on structured light, which is used to execute a manipulator obstacle avoidance method based on structured light described above in this application. The following describes the manipulator obstacle avoidance device based on structured light in detail.
[0120] As Figure 5 shown, the manipulator obstacle avoidance device based on structured light includes:
[0121] An image acquisition module 101, which is used to acquire a depth image based on structured light of the environment around the manipulator;
[0122] A three-dimensional reconstruction module 102, which is used to reconstruct the depth image to obtain an environmental three-dimensional model;
[0123] A path planning module 103, which is used to generate a planned path of the manipulator based on the environmental three-dimensional model;
[0124] A manipulator control module 104, which is used to update the environmental three-dimensional model in real time and control the manipulator to execute the planned path.
[0125] In one implementation, the image acquisition module 101 is further used for:
[0126] Projecting a grid-structured light onto the environment around the manipulator; acquiring the reflected structured light data and synchronously acquiring the visible light data; generating a depth image based on the structured light data; generating an RGB image based on the visible light data; converting the depth image and the RGB image into the same coordinate system.
[0127] In one implementation, the three-dimensional reconstruction module 102 is further used for:
[0128] Extract features from the depth image to obtain a depth feature map; extract features from the RGB image to obtain an RGB feature map; perform multi-scale fusion on the depth feature map and the RGB feature map to obtain fused feature maps of multiple sizes; perform graph convolution processing on the fused feature maps of multiple sizes to obtain a three-dimensional graph; construct the environmental three-dimensional model based on the three-dimensional image.
[0129] In one implementation, the three-dimensional reconstruction module 102 is further configured to:
[0130] Downsample the depth image to obtain depth feature maps of multiple sizes.
[0131] In one implementation, the three-dimensional reconstruction module 102 is further configured to:
[0132] Downsample the RGB image to obtain RGB feature maps of multiple sizes, and the sizes of the RGB feature maps correspond one-to-one with the sizes of the depth feature maps.
[0133] In one implementation, the three-dimensional reconstruction module 102 is further configured to:
[0134] After upsampling the RGB feature map with the smallest size, fuse it with the depth feature map and the RGB feature map of the previous size to obtain a fused map of the previous size; perform cyclic processing to sequentially obtain fused maps of the second size, the third size, the fourth size, and the largest size; the fused maps of the second size, the third size, and the largest size are the fused feature maps of multiple sizes.
[0135] In one implementation, the number of channels of the fused feature maps of multiple sizes is 64, 256, and 512.
[0136] An obstacle avoidance device for a robotic arm based on structured light provided in the above embodiments of the present application and a method for obstacle avoidance of a robotic arm based on structured light provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by the application programs stored therein.
[0137] The internal functions and structures of an obstacle avoidance device for a robotic arm based on structured light are described above. As Figure 6 shown, in practice, the obstacle avoidance device for a robotic arm based on structured light can be implemented as an electronic device, including: a memory 301 and a processor 303.
[0138] The memory 301 can be configured to store programs.
[0139] In addition, the memory 301 can also be configured to store various other data to support the operations on the electronic device. Examples of such data include instructions for any application or method for operating on the electronic device, contact data, phone book data, messages, pictures, videos, etc.
[0140] The memory 301 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0141] The processor 303, coupled to the memory 301, is configured to execute the programs in the memory 301 for:
[0142] Obtain a structured light-based depth image of the environment around the robotic arm;
[0143] Reconstruct the depth image to obtain a three-dimensional model of the environment;
[0144] Generate a planned path for the robotic arm based on the three-dimensional model of the environment;
[0145] Real-time update the three-dimensional model of the environment and control the robotic arm to execute the planned path.
[0146] In one embodiment, the processor 303 is further configured to:
[0147] Project a grid-structured light onto the environment around the robotic arm; obtain the reflected structured light data and synchronously obtain the visible light data; generate a depth image based on the structured light data; generate an RGB image based on the visible light data; convert the depth image and the RGB image into the same coordinate system.
[0148] In one embodiment, the processor 303 is further configured to:
[0149] Extract features from the depth image to obtain a depth feature map; extract features from the RGB image to obtain an RGB feature map; perform multi-scale fusion on the depth feature map and the RGB feature map to obtain fusion feature maps of multiple sizes; perform graph convolutional processing on the fusion feature maps of multiple sizes to obtain a three-dimensional graph; construct the three-dimensional model of the environment based on the three-dimensional image.
[0150] In one embodiment, the processor 303 is further configured to:
[0151] Perform downsampling on the depth image to obtain depth feature maps of multiple sizes.
[0152] In one embodiment, the processor 303 is further configured to:
[0153] Downsample the RGB image to obtain RGB feature maps of multiple sizes, and the sizes of the RGB feature maps correspond one by one to the sizes of the depth feature maps.
[0154] In one embodiment, the processor 303 is further configured to:
[0155] Upsample the RGB feature map with the smallest size, and then fuse it with the depth feature map and the RGB feature map of the previous size to obtain a fused map of the previous size; perform cyclic processing to sequentially obtain fused maps of the second size, the third size, the fourth size, and the fused map of the largest size; the fused maps of the second size, the third size, and the largest size are fused feature maps of multiple sizes.
[0156] In one embodiment, the number of channels of the fused feature maps of multiple sizes is 64, 256, and 512.
[0157] In this application, Figure 6 only some components are schematically shown, which does not mean that the electronic device only includes Figure 6 the components shown.
[0158] The electronic device provided in this embodiment and a method for avoiding obstacles of a robotic arm based on structured light provided in an embodiment of this application are based on the same inventive concept, and have the same beneficial effects as the methods adopted, run, or implemented by the application programs stored therein.
[0159] Those skilled in the art should understand that the embodiments of this application can be provided as a method, a system, or a computer program product. Therefore, this application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can adopt the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0160] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate for implementing in the process Figure 1 this process or multiple processes and / or blocks Figure 1means for the functions specified in one or more boxes.
[0161] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the functions specified in one Figure 1 process or multiple processes and / or boxes Figure 1 means for the functions specified in one or more boxes.
[0162] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 process or multiple processes and / or boxes Figure 1 means for the functions specified in one or more boxes.
[0163] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0164] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (Flash RAM). The memory is an example of computer-readable media.
[0165] This application also provides a computer-readable storage medium corresponding to a method for obstacle avoidance of a robotic arm based on structured light provided in the foregoing embodiment. A computer program (i.e., a program product) is stored thereon. When the computer program is run by a processor, it will execute a method for obstacle avoidance of a robotic arm based on structured light provided in any of the foregoing embodiments.
[0166] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0167] The computer-readable storage medium provided in the above-mentioned embodiments of the present application and the structured light-based robotic arm obstacle avoidance method provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.
[0168] It should be noted that in the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known structures and technologies are not shown in detail so as not to obscure the understanding of this description.
[0169] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0170] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
Claims
1. A method for a robotic arm to avoid obstacles based on structured light, characterized in that Including: Obtain a structured-light-based depth image of the environment around the robotic arm; Reconstruct the depth image to obtain a three-dimensional environmental model; Generate a planned path for the robotic arm based on the three-dimensional environmental model; Update the three-dimensional environmental model in real time and control the robotic arm to execute the planned path.
2. The method for obstacle avoidance of a robotic arm based on structured light according to claim 1, wherein The obtaining of the structured-light-based depth image of the environment around the robotic arm includes: Project a grid-structured light onto the environment around the robotic arm; Obtain the reflected structured-light data and synchronously obtain visible-light data; Generate a depth image based on the structured-light data; Generate an RGB image based on the visible-light data; Convert the depth image and the RGB image into the same coordinate system.
3. The method for obstacle avoidance of a robotic arm based on structured light according to claim 1 or 2, characterized in that, The reconstructing of the depth image to obtain a three-dimensional environmental model includes: Extract features from the depth image to obtain a depth feature map; Extract features from the RGB image to obtain an RGB feature map; Perform multi-scale fusion on the depth feature map and the RGB feature map to obtain fusion feature maps of multiple sizes; Perform graph convolution processing on the fusion feature maps of multiple sizes to obtain a three-dimensional graph; Construct the three-dimensional environmental model based on the three-dimensional image.
4. The method for obstacle avoidance of a robotic arm based on structured light according to claim 3, characterized in that, In the extracting of features from the depth image to obtain a depth feature map, downsample the depth image to obtain depth feature maps of multiple sizes.
5. The method for obstacle avoidance of a robotic arm based on structured light according to claim 4, wherein In the extracting of features from the RGB image to obtain an RGB feature map, downsample the RGB image to obtain RGB feature maps of multiple sizes, and the sizes of the RGB feature maps and the depth feature maps correspond one by one.
6. The method for obstacle avoidance of a robotic arm based on structured light according to claim 5, wherein The performing of multi-scale fusion on the depth feature map and the RGB feature map to obtain fusion feature maps of multiple sizes includes: After upsampling the RGB feature map of the smallest size, fuse it with the depth feature map of the previous size and the RGB feature map of the previous size to obtain a fusion map of the previous size; Perform cyclic processing to sequentially obtain fusion maps of the second size, the third size, the fourth size, and the largest size; The fusion maps of the second size, the third size, and the largest size are fusion feature maps of multiple sizes.
7. The method for obstacle avoidance of a robotic arm based on structured light according to claim 3, characterized in that, The number of channels of the fusion feature maps of multiple sizes is 64, 256, 512.
8. An obstacle avoidance device for a robotic arm based on structured light, characterized in that, Including: An image acquisition module for obtaining a structured-light-based depth image of the environment around the robotic arm; A three-dimensional reconstruction module for reconstructing the depth image to obtain a three-dimensional environmental model; A path planning module for generating a planned path for the robotic arm based on the three-dimensional environmental model; A robotic arm control module for updating the three-dimensional environmental model in real time and controlling the robotic arm to execute the planned path.
9. An electronic device, characterized in that, Including: A memory and a processor; The memory for storing programs; The processor, coupled to the memory, for executing the programs to: Obtain a structured-light-based depth image of the environment around the robotic arm; Reconstruct the depth image to obtain a three-dimensional environmental model; Generate a planned path for the robotic arm based on the three-dimensional environmental model; Update the three-dimensional environmental model in real time and control the robotic arm to execute the planned path.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement a structured-light-based robotic arm obstacle avoidance method according to any one of claims 1-7.
Citation Information
Patent Citations
Mechanical arm path planning method and device based on improved random tree
CN119407784A