Mechanical arm obstacle avoidance method and device based on view field generation
By identifying and cropping the ROI region around the robotic arm, a cropping depth map is generated, which solves the problem of high computational cost in robotic arm path planning and achieves more efficient path planning and control.
Patent Information
- Application Number
- CN202510762401.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-10-28
AI Technical Summary
In current surgical robot control, the path planning calculation of the robotic arm is computationally intensive, resulting in low efficiency.
By acquiring depth image video streams, the robotic arm is identified and its surrounding ROI regions are cropped to generate a cropped depth map. The cropped depth map is then used for path planning and control, and combined with an improved RRT algorithm, a shorter and smoother path is generated.
It significantly reduces the computational load of path planning, improves planning speed and efficiency, and adapts to different task requirements and environmental changes.
Smart Images

Figure CN120839770A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of medical image processing technology, and more specifically, to a method and apparatus for obstacle avoidance of a robotic arm based on field of view generation. Background Technology
[0002] In current surgical robot control, obstacle avoidance path planning for the robotic arm requires planning based on the entire three-dimensional environment model. Although this planning method is relatively accurate, it requires the entire environment model for each planning, resulting in a large amount of computation. Summary of the Invention
[0003] The problem this application addresses is that current path planning computations are very computationally intensive.
[0004] To address the aforementioned problems, the first aspect of this application provides a robotic arm obstacle avoidance method based on field-of-view generation, which includes:
[0005] Acquire depth image video stream;
[0006] Identify robotic arms in depth image video streams;
[0007] Based on the identified robotic arm, the ROI region of the image and video stream is cropped to obtain a cropping depth map;
[0008] The robotic arm is controlled based on the cutting depth map.
[0009] The second aspect of this application provides a manufacturing system for a robotic arm obstacle avoidance method based on field-of-view generation, comprising:
[0010] The video acquisition module is used to acquire depth image video streams;
[0011] A robotic arm recognition module is used to identify robotic arms in depth image video streams;
[0012] The depth cropping module is used to crop the ROI region of the image and video stream based on the identified robotic arm to obtain a cropping depth map;
[0013] The robotic arm control module is used to control the robotic arm based on the cutting depth map.
[0014] A third aspect of this application provides an electronic device, including: a memory and a processor; the memory being configurable to store a program, and the processor being coupled to the memory for executing the program in the memory for:
[0015] Acquire depth image video stream;
[0016] Identify robotic arms in depth image video streams;
[0017] Based on the identified robotic arm, the ROI region of the image and video stream is cropped to obtain a cropping depth map;
[0018] The robotic arm is controlled based on the cutting depth map.
[0019] The fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the aforementioned obstacle avoidance method for a robotic arm based on field-of-view generation.
[0020] In this application, path planning and control are performed by identifying the robotic arm and cropping the ROI region around it, thereby greatly reducing the amount of computation in the path planning process and improving the speed and efficiency of planning. Attached Figure Description
[0021] Figure 1 This is a flowchart of a robotic arm obstacle avoidance method based on field-of-view generation according to an embodiment of this application;
[0022] Figure 2 This is an architecture diagram of the obstacle recognition model of the robotic arm obstacle avoidance method based on field of view generation according to an embodiment of this application;
[0023] Figure 3 This is an architecture diagram of the semantic extraction of the robotic arm obstacle avoidance method based on field of view generation according to an embodiment of this application;
[0024] Figure 4 This is an architecture diagram of the obstacle avoidance method for robotic arms based on field-of-view generation according to an embodiment of this application;
[0025] Figure 5 This is an architectural diagram of a robotic arm obstacle avoidance device based on field-of-view generation according to an embodiment of this application;
[0026] Figure 6 This is an architectural diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0027] To make the above-mentioned objects, features, and advantages of this application more apparent and understandable, specific embodiments of this application will be described in detail below with reference to the accompanying drawings. Although exemplary embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of this application and to fully convey the scope of this application to those skilled in the art.
[0028] It should be noted that, unless otherwise stated, the technical or scientific terms used in this application shall have the ordinary meaning as understood by one of ordinary skill in the art.
[0029] This application provides the above-described obstacle avoidance method for robotic arms based on field-of-view generation. The specific solution of this method is as follows: Figure 1 As shown, this method can be executed by a vision-based obstacle avoidance device for robotic arms, which can be integrated into electronic devices such as computers, servers, computer clusters, and data centers.
[0030] Combination Figure 1 As shown, the obstacle avoidance method for the robotic arm based on field of view generation includes:
[0031] S101, acquire depth image video stream;
[0032] In this application, depth image video streams of the environment are acquired in real time using a depth camera (such as an RGB-D camera or LiDAR).
[0033] S102, Identify the robotic arm in the depth image video stream;
[0034] In this application, a robotic arm is separated from the background using depth thresholding or semantic segmentation algorithms (such as deep learning models).
[0035] S103, based on the identified robotic arm, crop the ROI region of the image / video stream to obtain a cropping depth map;
[0036] In this application, the region of interest (ROI) is extracted based on the workspace and task requirements of the robotic arm, thereby reducing the amount of irrelevant data to be processed.
[0037] S104 controls the robotic arm based on the cutting depth map.
[0038] In this application, the path trajectory of the robotic arm is generated using a clipping depth map, and the robotic arm is controlled to complete the task.
[0039] In this application, path planning and control are performed by identifying the robotic arm and cropping the ROI region around it, thereby greatly reducing the amount of computation in the path planning process and improving the speed and efficiency of planning.
[0040] In this application, the field of view is limited by cropping the depth map, thereby reducing the amount of irrelevant data to be processed.
[0041] In this application, an improved RRT algorithm is used to generate shorter and smoother paths.
[0042] In this application, the ROI boundary and step size are dynamically adjusted to adapt to different task requirements and environmental changes.
[0043] In one implementation, step S103, based on the identified robotic arm, cropes the ROI region of the image / video stream to obtain a cropping depth map, including:
[0044] Obtain the workspace of the robotic arm and determine the 3D ROI boundaries;
[0045] The 3D ROI boundary is projected onto the image plane of the depth camera to obtain the 2D ROI region;
[0046] By cropping the image / video stream using a two-dimensional Region of Interest (ROI), a cropping depth map is obtained.
[0047] In this application, a three-dimensional ROI boundary is defined based on the task requirements and working range of the robotic arm.
[0048] In this application, the reachable range of the robotic arm is calculated based on its kinematic model (such as forward kinematics and inverse kinematics); the range of motion of the robotic arm's end effector in the task is determined, thereby determining the workspace of the robotic arm.
[0049] In this application, the three-dimensional ROI boundary is defined using the base coordinate system of the robotic arm as a reference, and is a three-dimensional cube, spherical region, or cylinder containing the workspace of the robotic arm.
[0050] Preferably, for dynamic task scenarios, the ROI boundary can be dynamically adjusted based on the real-time status of the robotic arm. For example, its workspace can be recalculated when the robotic arm moves to a new position.
[0051] In this application, the perspective projection formula is used to project key points (such as vertices) in the 3D ROI boundary onto the image plane of the depth camera.
[0052] In this application, a minimum rectangular region containing all projected points is generated as a two-dimensional Region of Interest (ROI) based on the projected points. (Then, by limiting the depth dimension of this rectangle, a clipping depth map of a three-dimensional cube can be generated.)
[0053] In this application, depth image video streams are cropped based on two-dimensional ROI regions to extract regions of interest and reduce the amount of irrelevant data to be processed.
[0054] In this application, the cropped depth image retains only the pixel values within the ROI region and ignores other regions; and only the pixel values at the preset depth are retained, while pixel values at other depths are ignored.
[0055] Preferably, the cropped depth image is filtered (e.g., Gaussian filtering or bilateral filtering) to remove noise and improve the accuracy of subsequent processing.
[0056] In this application, the projection is two-dimensional before cropping, thereby improving efficiency and avoiding the introduction of additional computational overhead.
[0057] In one implementation, after cropping the ROI region of the image / video stream based on the identified robotic arm to obtain a cropping depth map, the method further includes:
[0058] Perform connectivity analysis on the robotic arm in the clipping depth map to confirm the non-occluded or occluded state of the robotic arm.
[0059] In the occlusion state, image frames from the image video stream are input into a pre-trained obstacle recognition model to obtain the occluded image and dynamic obstacle;
[0060] Occluded images and dynamic obstacles are mapped to a cropping depth map.
[0061] In this application, connectivity analysis is performed on the robotic arm in the clipping depth map to determine whether the robotic arm is occluded.
[0062] In this application, a connectivity analysis algorithm is used to check whether the robotic arm region is connected to other object regions.
[0063] In this application, the unoccluded state is defined as follows: if the robotic arm area exists independently and is not connected to other object areas, the robotic arm is considered to be in an unoccluded state.
[0064] In this application, the occlusion state is defined as follows: if the robotic arm area is connected to other object areas, the robotic arm is considered to be in an occlusion state.
[0065] In this application, a pre-trained obstacle recognition model is used to identify occluded images and dynamic obstacles from an image video stream.
[0066] In this application, the occlusion image refers to a static object or background area that occludes the robotic arm.
[0067] In this application, a dynamic obstacle is an object that moves within an image frame.
[0068] In this application, the identified occluded images and dynamic obstacle information are mapped onto the clipping depth map to provide complete environmental information for subsequent path planning.
[0069] In this application, information about occluded images and dynamic obstacles is superimposed onto the clipping depth map; for dynamic obstacles, their motion trajectory can be represented by dashed lines or color markers.
[0070] In one implementation, combined with Figure 2 As shown, the step of inputting image frames from the image video stream into a pre-trained obstacle recognition model under occlusion conditions to obtain the occluded image and dynamic obstacle includes:
[0071] Feature extraction is performed on image frames to obtain feature pyramids of different levels;
[0072] Semantic extraction is performed on each level of the feature pyramid to obtain semantic feature maps for different levels;
[0073] The semantic feature maps of each level are fused separately to obtain augmented pyramids of different levels;
[0074] The head of the enhanced pyramid was subjected to detection processing to obtain the detection results;
[0075] Based on the detection results of adjacent image frames, occlusion images and dynamic obstacles are obtained.
[0076] In this application, static occlusion images and dynamic obstacles are distinguished by analyzing the detection results in consecutive image frames.
[0077] Preferably, the motion vector between adjacent image frames is calculated using optical flow or inter-frame difference algorithm; the motion vector is then used to determine whether the target object has moved.
[0078] This application fully utilizes the advantages of feature pyramids and multi-scale fusion, combined with motion detection technology, to accurately identify occluded images and dynamic obstacles, providing crucial environmental information for robotic arm path planning.
[0079] In one implementation, combined with Figure 3 As shown, semantic extraction is performed on each level of the feature pyramid to obtain semantic feature maps at different levels, including:
[0080] Perform 1×N and N×1 convolutions on any level of feature sequentially to obtain the first branch map;
[0081] The feature at this level is convolved with N×1 and 1×N sequentially to obtain the second branch map;
[0082] Add the first branch graph and the second branch graph to obtain the sum graph;
[0083] The additive image is subjected to convolution, activation, and convolution processing to obtain a convolutional image.
[0084] The semantic feature map of this level is obtained by adding the features of this level, the summation map, and the convolution map.
[0085] It should be noted that the above describes the processing procedure for one level. The processing procedure for each level is similar and will not be repeated here.
[0086] In this application, through semantic extraction, a large-size convolutional kernel is applied and the receptive field is significantly expanded through bilateral scanning operations, enabling the model to capture rich contextual features.
[0087] In this application, a larger convolutional kernel is used to broaden the coverage of the effective receptive field. First, the large-kernel 2D convolution is split into two parallel branches, each consisting of 1D convolutions along two different directions (vertical and horizontal). That is, the standard N×N large convolution (where N represents the height and width of the convolutional kernel) is replaced with a combination of 1×N+N×1 and N×1+1×N convolutions, which not only reduces computational complexity but also enables the model to more effectively capture the features of the target in different directions.
[0088] In this application, complementary features from two branches are fused through element-wise addition to generate a more refined feature map. Furthermore, pooling layers and fully connected layers are removed to maximize the preservation of spatial information in the feature map. This type of symmetric linear 1D filter, derived from a large-size 2D filter, not only solves the problem of additional computational burden caused by the large convolutional kernel CNN architecture, but also enables the model to perceive high-response contexts and transmit discriminative information to adjacent low-response target regions (non-discriminative regions), thereby improving the model's overall perception of the target.
[0089] In one implementation, combined with Figure 4 As shown, the fusion processing of semantic feature maps at each level to obtain augmentation pyramids at different levels includes:
[0090] Upsampling is performed on the layer above the semantic feature map at any level to obtain an upsampled map;
[0091] The semantic feature map of this level, the features of this level of the feature pyramid, and the upsampled map are concatenated to obtain the concatenated image;
[0092] By iteratively performing 1×1 convolutions and 3×3 convolutions on the stitched image, the features of the enhanced pyramid at this level are obtained.
[0093] In this application, high-quality target localization maps of different scales are obtained through fusion processing to more efficiently fuse them.
[0094] In this application, a sophisticated strategy is employed to fuse localization maps at different scales. First, the FRM fusion process upsamples the feature map output from the (i+1)th layer of semantic extraction, doubling its size to maintain resolution consistent with the previous layer's feature map. Next, these feature maps are concatenated along the channel dimension with feature maps from corresponding layers of the backbone network to learn inter-channel dependencies, thereby generating a feature map that incorporates rich channel contextual information. This concatenation operation effectively fuses multi-scale feature information, compensating for the limitations of single-scale features in representing targets.
[0095] In this application, the spliced image is cyclically subjected to 1×1 convolution and 3×3 convolution to obtain the features of the enhanced pyramid at this level, that is, 1×1 convolution, 3×3 convolution, 1×1 convolution, 3×3 convolution, 1×1 convolution are performed.
[0096] In this application, in order to avoid introducing too many model parameters in this operation, two 1×1 convolutional layers are used in the fusion processing interval to sequentially reduce the channel dimension to 256 and 64. This not only realizes the dimensionality reduction of the feature map but also further fuses the information between different channels, making the feature representation more compact and efficient. Subsequently, a series of standard convolution operations are implemented to eliminate the aliasing effect caused by interpolation.
[0097] In this application, by mining the channel information, the preliminary localization map generated by the network is refined and fused. This localization map originally only covered the sparse target area, but under the effect of the fusion processing, it is transformed into an accurate and reliable target localization map, greatly improving the detection accuracy and enabling the model to better cope with the target detection challenges in complex scenarios.
[0098] In one implementation, in S104, the control of the robotic arm based on the cropped depth map includes:
[0099] Generating a three-dimensional model based on the cropped depth map, where the occluded image and moving obstacles are marked in the three-dimensional model;
[0100] Setting a target biasing strategy;
[0101] Constructing a spherical set constraint and an angular constraint;
[0102] Based on the target biasing strategy, the spherical set constraint and the angular constraint, generating the motion trajectory of the robotic arm through the random tree algorithm;
[0103] Controlling the robotic arm to execute the motion trajectory.
[0104] In this application, the target biasing strategy:
[0105] Set a threshold P∈(0,1). When the random tree selects a sampling point Xr, use a random number generator to generate a random number Pr between (0,1); if P < Pr, then use the target point Xgoal as the sampling point for target-biased sampling; otherwise, randomly sample within the sampling area.
[0106] In this application, the target biasing strategy not only enhances the target tendency of the algorithm but also retains the randomness of the RRT algorithm, avoiding the algorithm from losing the ability to explore the space.
[0107] In this application, constructing a spherical set constraint:
[0108] In the three-dimensional model, connect the target point Xgoal and the current parent node Xnear, and construct a spherical subset using the line segment Xgoal-Xnear; the radius of the spherical subset is dynamically adjusted according to the distance between the parent node and the target point, and as the parent node approaches the target point, the sampling area gradually shrinks.
[0109] Generate a new sampling point Xr in the spherical subset, and determine whether the new node Xnew is within the spherical subset; if the new node is not within the spherical subset, delete the node and resample within the spherical subset.
[0110] In this application, construct an angular constraint:
[0111] With the starting point Xinit as the vertex, connect the starting point to the parent node Xnear and the target point Xgoal to form an included angle θ; the area within the included angle θ is the angular constraint range of the sampling points.
[0112] Set a threshold Q∈(0,1). After generating the sampling point, generate a random number Qrand between (0,1). If Q>Qrand, perform an angular constraint optimization judgment on this node; otherwise, skip the angular constraint and directly use this node as the parent node for the next planning.
[0113] In this application, based on the target bias strategy, spherical set constraint and angular constraint, generate the motion trajectory of the robotic arm through the random tree algorithm:
[0114] Adaptive step size expansion: Dynamically adjust the expansion step size according to the distance S between the node and the obstacle:
[0115] When S>D2 (a type of safety distance), use a large step size Sl.
[0116] When D2<S<D1 (a second type of safety distance), use a normal step size Sm.
[0117] When S<D1, use a small step size Ss.
[0118] Path generation: Each time the random tree expands, check whether the new node is directly connected to the target point. If there is no obstacle blocking, directly connect the new node to the target point to complete the path planning.
[0119] Path smoothing optimization: Remove redundant points in the path; use a cubic B-spline curve to fit the path to generate a smoother motion trajectory.
[0120] In this application, convert the generated motion trajectory into the joint angle sequence of the robotic arm, and control the robotic arm to move along the trajectory:
[0121] Inverse kinematics solution: According to the generated path points, use the inverse kinematics algorithm to calculate the angles of each joint of the robotic arm.
[0122] Trajectory interpolation: Interpolating the joint angle sequence (such as cubic spline interpolation) to generate a smooth motion trajectory.
[0123] Speed control: The movement speed of the robotic arm is dynamically adjusted according to the distance to obstacles to ensure smooth and safe movement.
[0124] Real-time monitoring: During the robotic arm's movement, environmental changes are monitored in real time. If a new obstacle or unreachable path is detected, the path planning algorithm is invoked again to generate a new trajectory.
[0125] Preferably, connectivity analysis is performed on the robotic arm in the clipping depth map to confirm the non-occluded or occluded state of the robotic arm, including:
[0126] Perform connectivity analysis on the robotic arm in the cutting depth map to determine the connected components of the robotic arm;
[0127] Obtain the posture projection of the robotic arm;
[0128] If the connected domain of the robotic arm coincides with the pose projection of the robotic arm, then the robotic arm is confirmed to be in an unoccluded state.
[0129] If the connected domain of the robotic arm exceeds the pose projection of the robotic arm, then the robotic arm is confirmed to be in an occluded state.
[0130] If the robot arm's pose projection exceeds the robot arm's connected domain, then the robot arm is confirmed to be in an occluded state.
[0131] In one implementation, feature extraction is performed on the image frame to obtain feature pyramids of different levels; prior to this, adaptive updating of the image frame is also included; the specific process of the adaptive update includes:
[0132] The image frame is divided into blocks to obtain independent blocks;
[0133] For each independent block, obtain the first and second neighboring blocks with different spacings;
[0134] A first feature block is generated based on the independent block and the first neighboring block;
[0135] A second feature block is generated based on the independent block and the second neighboring block.
[0136] The first and second feature blocks are compressed to obtain a compressed block.
[0137] Iterate through all the individual blocks and generate updated image frames based on the resulting compressed blocks.
[0138] In this application, the image frame is divided into blocks, that is, the image frame is divided into corresponding image blocks by using a checkerboard pattern; wherein, the image block can be at the pixel level (that is, each pixel is an image block) or other levels, and the specific division depends on the actual processing situation.
[0139] In this application, a sliding window or a fixed step size is used to divide the image into blocks of the same size.
[0140] It should be noted that if the image frame is a 3D image, then a face is selected and divided into a checkerboard pattern. Each square is a strip with a depth (the depth is the depth of the 3D image), and this strip is an image block.
[0141] Preferably, in this application, each image block has 1,001,000 pixels, thereby enabling more feature calculations between local regions while ensuring generation accuracy and reducing computational load.
[0142] In this application, an image block is selected as an independent block. The image blocks above, below, to the left, and to the right of this independent block are the first neighboring blocks; the image blocks one grid away from the top, bottom, left, and right of this independent block are the second neighboring blocks. The spacing between the first and second neighboring blocks and the independent block is different.
[0143] In this application, neighborhood information is extracted for each independent block to capture local structure.
[0144] In this application, generating the first feature block is to generate a local feature representation using an independent block and its first neighboring block. Specifically, this can be done by processing the independent block and the first neighboring block with convolutional layers and attention layers to obtain the first feature block.
[0145] In this application, the specific structure and parameters of the convolutional layer and attention layer can be obtained from the training data or determined according to the actual situation.
[0146] It should be noted that in this application, there are four first neighboring blocks and multiple first feature blocks.
[0147] In this application, the independent block and the first neighboring block are processed by convolutional layers and attention layers to obtain the first feature block. The specific process is as follows: the independent block and four neighboring blocks are concatenated together to form a multi-channel input, and the convolutional layer is used to extract features from the concatenated block; an important feature is enhanced by using a self-attention mechanism or a channel attention mechanism, the attention weight is calculated, and the output of the convolutional layer is weighted to enhance the important feature; the output of the attention layer is split into multiple feature blocks, and each feature block corresponds to the processing result of the independent block and at least one neighboring block.
[0148] In this application, a second feature block is generated to generate a broader local feature representation using the independent block and its second neighboring block. The specific generation process is the same as that of the first feature block, except that the parameters of the convolutional layer and the attention layer are different.
[0149] In this application, the generated feature blocks are compressed into a more compact representation to reduce computational cost while retaining key information. Feature compression is performed using pooling operations (such as max pooling or average pooling) or fully connected layers.
[0150] In this way, multiple first feature blocks and second feature blocks are compressed into a single compressed block, which corresponds to the size and position of the independent blocks and is used to replace them. All image blocks are replaced by the compressed blocks, resulting in an updated image frame.
[0151] In this application, each image block of the image frame is traversed to obtain the corresponding compressed block.
[0152] In this application, for image blocks / independent blocks near the edge, their first and second neighboring blocks are incomplete. In this case, the incomplete blocks are completed by copying the first and second neighboring blocks in their relative positions. For example, if the first neighboring block above an independent block does not exist, the first neighboring block below it is copied and used as the block above it.
[0153] In this application, by completing the image blocks, the processing accuracy of adjacent image blocks is greatly improved.
[0154] In this application, an adaptive adjustment module is used to capture the similarity relationship between local regions, thereby enhancing the feature representation.
[0155] This application provides a robotic arm obstacle avoidance device based on field of view generation, used to execute the robotic arm obstacle avoidance method based on field of view generation described above. The following is a detailed description of the robotic arm obstacle avoidance device based on field of view generation.
[0156] like Figure 5 As shown, the obstacle avoidance device for the robotic arm based on field of view generation includes:
[0157] Video acquisition module 101 is used to acquire depth image video stream;
[0158] Robotic arm recognition module 102, which is used to recognize robotic arms in depth image video streams;
[0159] The depth cropping module 103 is used to crop the ROI region of the image and video stream based on the identified robotic arm to obtain a cropping depth map.
[0160] The robotic arm control module 104 is used to control the robotic arm based on the cutting depth map.
[0161] In one implementation, the depth clipping module 103 is further configured to:
[0162] The workspace of the robotic arm is obtained, and the 3D ROI boundary is determined. The 3D ROI boundary is projected onto the image plane of the depth camera to obtain the 2D ROI region. The image and video stream is cropped through the 2D ROI region to obtain the cropping depth map.
[0163] In one implementation, the depth clipping module 103 is further configured to:
[0164] Connectivity analysis is performed on the robotic arm in the clipping depth map to confirm its unoccluded or occluded state. In the occluded state, image frames from the image video stream are input into a pre-trained obstacle recognition model to obtain the occluded image and dynamic obstacle. The occluded image and dynamic obstacle are then mapped to the clipping depth map.
[0165] In one implementation, the depth clipping module 103 is further configured to:
[0166] Feature extraction is performed on image frames to obtain feature pyramids of different levels; semantic extraction is performed on each level of the feature pyramid to obtain semantic feature maps of different levels; the semantic feature maps of each level are fused to obtain augmented pyramids of different levels; head detection is performed on the augmented pyramids to obtain detection results; based on the detection results of adjacent image frames, occlusion images and dynamic obstacles are obtained.
[0167] In one implementation, the depth clipping module 103 is further configured to:
[0168] Perform 1×N and N×1 convolutions on any level feature sequentially to obtain the first branch map; perform N×1 and 1×N convolutions on the same level feature sequentially to obtain the second branch map; add the first branch map and the second branch map to obtain the sum map; perform convolution, activation, and convolution processing on the sum map to obtain the convolution map; add the level feature, the sum map, and the convolution map together to obtain the semantic feature map of that level.
[0169] In one implementation, the depth clipping module 103 is further configured to:
[0170] Upsample the layer above the semantic feature map of any level to obtain an upsampled map; concatenate the semantic feature map of that level, the features of that level of the feature pyramid, and the upsampled map to obtain a concatenated map; perform 1×1 convolution and 3×3 convolution on the concatenated map in a loop to obtain the features of the augmented pyramid of that level.
[0171] In one embodiment, the robotic arm control module 104 is further configured to:
[0172] A 3D model is generated based on the cropped depth map, and the 3D model is labeled with occluded images and moving obstacles; a target bias strategy is set; spherical set constraints and angular constraints are constructed; based on the target bias strategy, spherical set constraints and angular constraints, the motion trajectory of the robotic arm is generated by a random tree algorithm; and the robotic arm is controlled to execute the motion trajectory.
[0173] The obstacle avoidance device for robotic arms based on field of view generation provided in the above embodiments of this application corresponds to the obstacle avoidance method for robotic arms based on field of view generation provided in the embodiments of this application. Therefore, the specific content in this system corresponds to the obstacle avoidance method for robotic arms based on field of view generation. The specific content can be referred to the records in the obstacle avoidance method for robotic arms based on field of view generation, which will not be repeated in this application.
[0174] The obstacle avoidance device for robotic arms based on field of view generation provided in the above embodiments of this application and the obstacle avoidance method for robotic arms based on field of view generation provided in the embodiments of this application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the applications stored therein.
[0175] The above describes the internal functions and structure of the obstacle avoidance device for a robotic arm based on field-of-view generation, such as... Figure 6 As shown, in practice, the obstacle avoidance device of the robotic arm based on field of view generation can be implemented as an electronic device, including: a memory 301 and a processor 303.
[0176] Memory 301 can be configured to store a program.
[0177] Additionally, memory 301 can also be configured to store various other data to support operation on the electronic device. Examples of this data include instructions for any application or method used to operate on the electronic device, contact data, phonebook data, messages, pictures, videos, etc.
[0178] Memory 301 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. Processor 303, coupled to memory 301, is used to execute programs in memory 301 for:
[0179] Acquire depth image video stream;
[0180] Identify robotic arms in depth image video streams;
[0181] Based on the identified robotic arm, the ROI region of the image and video stream is cropped to obtain a cropping depth map;
[0182] The robotic arm is controlled based on the cutting depth map.
[0183] In one implementation, the processor 303 is further configured to:
[0184] The workspace of the robotic arm is obtained, and the 3D ROI boundary is determined. The 3D ROI boundary is projected onto the image plane of the depth camera to obtain the 2D ROI region. The image and video stream is cropped through the 2D ROI region to obtain the cropping depth map.
[0185] In one implementation, the processor 303 is further configured to:
[0186] Connectivity analysis is performed on the robotic arm in the clipping depth map to confirm its unoccluded or occluded state. In the occluded state, image frames from the image video stream are input into a pre-trained obstacle recognition model to obtain the occluded image and dynamic obstacle. The occluded image and dynamic obstacle are then mapped to the clipping depth map.
[0187] In one implementation, the processor 303 is further configured to:
[0188] Feature extraction is performed on image frames to obtain feature pyramids of different levels; semantic extraction is performed on each level of the feature pyramid to obtain semantic feature maps of different levels; the semantic feature maps of each level are fused to obtain augmented pyramids of different levels; head detection is performed on the augmented pyramids to obtain detection results; based on the detection results of adjacent image frames, occlusion images and dynamic obstacles are obtained.
[0189] In one implementation, the processor 303 is further configured to:
[0190] Perform 1×N and N×1 convolutions on any level feature sequentially to obtain the first branch map; perform N×1 and 1×N convolutions on the same level feature sequentially to obtain the second branch map; add the first branch map and the second branch map to obtain the sum map; perform convolution, activation, and convolution processing on the sum map to obtain the convolution map; add the level feature, the sum map, and the convolution map together to obtain the semantic feature map of that level.
[0191] In one implementation, the processor 303 is further configured to:
[0192] Upsample the layer above the semantic feature map of any level to obtain an upsampled map; concatenate the semantic feature map of that level, the features of that level of the feature pyramid, and the upsampled map to obtain a concatenated map; perform 1×1 convolution and 3×3 convolution on the concatenated map in a loop to obtain the features of the augmented pyramid of that level.
[0193] In one implementation, the processor 303 is further configured to:
[0194] A 3D model is generated based on the cropped depth map, and the 3D model is labeled with occluded images and moving obstacles; a target bias strategy is set; spherical set constraints and angular constraints are constructed; based on the target bias strategy, spherical set constraints and angular constraints, the motion trajectory of the robotic arm is generated by a random tree algorithm; and the robotic arm is controlled to execute the motion trajectory.
[0195] In this application, Figure 6 The diagram only shows some components and does not mean that the electronic device includes only these components. Figure 6 The components shown.
[0196] The electronic device provided in this embodiment is based on the same inventive concept as the obstacle avoidance method of the robotic arm based on field of view generation provided in this application embodiment, and has the same beneficial effects as the methods adopted, run or implemented by the application stored therein.
[0197] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0198] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0199] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0200] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory. Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0201] This application also provides a computer-readable storage medium corresponding to the obstacle avoidance method for robotic arms based on field of view generation provided in the foregoing embodiments, wherein a computer program (i.e., a program product) is stored thereon. When the computer program is run by a processor, it executes the interactive image analysis assistance method for 3D aerial imaging provided in any of the foregoing embodiments.
[0202] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0203] The computer-readable storage medium provided in the above embodiments of this application and the interactive image analysis assistance method for 3D aerial imaging provided in the embodiments of this application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the applications stored therein.
[0204] It should be noted that numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this application may be practiced without these specific details. In some instances, well-known structures and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0205] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0206] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for obstacle avoidance by a robotic arm based on field-of-view generation, characterized in that, include: Acquire depth image video stream; Identify robotic arms in depth image video streams; Based on the identified robotic arm, the ROI region of the image and video stream is cropped to obtain a cropping depth map; The robotic arm is controlled based on the cutting depth map.
2. The obstacle avoidance method for robotic arms based on field-of-view generation according to claim 1, characterized in that, The process of cropping the ROI region of the image / video stream based on the identified robotic arm to obtain a cropping depth map includes: Obtain the workspace of the robotic arm and determine the 3D ROI boundaries; The 3D ROI boundary is projected onto the image plane of the depth camera to obtain the 2D ROI region; By cropping the image / video stream using a two-dimensional Region of Interest (ROI), a cropping depth map is obtained.
3. The obstacle avoidance method for robotic arms based on field-of-view generation according to claim 2, characterized in that, After cropping the ROI region of the image / video stream based on the identified robotic arm to obtain the cropping depth map, the process further includes: Perform connectivity analysis on the robotic arm in the clipping depth map to confirm the non-occluded or occluded state of the robotic arm. In the occlusion state, image frames from the image video stream are input into a pre-trained obstacle recognition model to obtain the occluded image and dynamic obstacle; Occluded images and dynamic obstacles are mapped to a cropping depth map.
4. The obstacle avoidance method for robotic arms based on field-of-view generation according to claim 3, characterized in that, In the occlusion state, inputting image frames from the image video stream into a pre-trained obstacle recognition model to obtain the occlusion image and dynamic obstacle includes: Feature extraction is performed on image frames to obtain feature pyramids of different levels; Semantic extraction is performed on each level of the feature pyramid to obtain semantic feature maps for different levels; The semantic feature maps of each level are fused separately to obtain augmented pyramids of different levels; The head of the enhanced pyramid was subjected to detection processing to obtain the detection results; Based on the detection results of adjacent image frames, occlusion images and dynamic obstacles are obtained.
5. The obstacle avoidance method for robotic arms based on field-of-view generation according to claim 4, characterized in that, The semantic extraction is performed on each level of the feature pyramid to obtain semantic feature maps at different levels, including: Perform 1×N and N×1 convolutions on any level of feature sequentially to obtain the first branch map; The feature at this level is convolved with N×1 and 1×N sequentially to obtain the second branch map; Add the first branch graph and the second branch graph to obtain the sum graph; The additive image is subjected to convolution, activation, and convolution processing to obtain a convolutional image. The semantic feature map of this level is obtained by adding the features of this level, the summation map, and the convolution map.
6. The obstacle avoidance method for a robotic arm based on field-of-view generation according to claim 4, characterized in that, The process of fusing the semantic feature maps at each level to obtain augmentation pyramids at different levels includes: Upsampling is performed on the layer above the semantic feature map at any level to obtain an upsampled map; The semantic feature map of this level, the features of this level of the feature pyramid, and the upsampled map are concatenated to obtain the concatenated image; By iteratively performing 1×1 convolutions and 3×3 convolutions on the stitched image, the features of the enhanced pyramid at this level are obtained.
7. The obstacle avoidance method for a robotic arm based on field-of-view generation according to any one of claims 1-6, characterized in that, The control of the robotic arm based on the cutting depth map includes: A 3D model is generated based on the cropping depth map, and the 3D model is marked with occluded images and moving obstacles. Set the target bias strategy; Construct spherical set constraints and angular constraints; Based on the target bias strategy, spherical set constraint and angle constraint, the motion trajectory of the robotic arm is generated by the random tree algorithm. Control the robotic arm to execute the motion trajectory.
8. A robotic arm obstacle avoidance device based on field of view generation, characterized in that, include: The video acquisition module is used to acquire depth image video streams; A robotic arm recognition module is used to identify robotic arms in depth image video streams; The depth cropping module is used to crop the ROI region of the image and video stream based on the identified robotic arm to obtain a cropping depth map; The robotic arm control module is used to control the robotic arm based on the cutting depth map.
9. An electronic device, characterized in that, include: memory and processor; The memory is used to store programs; The processor, coupled to the memory, is used to execute the program for: Acquire depth image video stream; Identify robotic arms in depth image video streams; Based on the identified robotic arm, the ROI region of the image and video stream is cropped to obtain a cropping depth map; The robotic arm is controlled based on the cutting depth map.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by a processor to implement the obstacle avoidance method for robotic arms based on field of view generation as described in any one of claims 1-7.