Robot algorithm oriented to dynamic shielding scene, storage medium and equipment

Through improved segmentation algorithms and background repair technology, the problems of inaccurate evaluation of dynamic object motion states and poor static background repair in the prior art are solved, and higher feature information accuracy and performance improvement of SLAM algorithms in dynamic environments are achieved.

CN120014597AActive Publication Date: 2025-05-16ANHUI POLYTECHNIC UNIV

Patent Information

Application Number
CN202510090444.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

The prior art cannot evaluate the real-time motion state of potential dynamic objects very accurately, resulting in the removal of dynamic feature points that will cause the problem of reducing feature information. After dynamic area removal, the static background repair effect is poor, affecting the accuracy of feature information and pose estimation.

Method used

The improved segmentation algorithm is adopted to improve the recognition accuracy of dynamic objects through multi-directional feature enhancement convolution, dual-pooling feature enhancement module and multi-attention mechanism module. At the same time, through object motion estimation and background repair algorithms, dynamic areas are eliminated and static backgrounds are restored to ensure the accuracy of feature point extraction and pose calculation.

Benefits of technology

It improves the accuracy of segmentation and identification of dynamic objects, reduces the loss of feature information, and improves the quality and number of static feature points through effective background repair, and improves the performance of SLAM algorithm in dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014597A_ABST
    Figure CN120014597A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of synchronous positioning and map creation, and discloses a robot algorithm, a storage medium and equipment for a dynamic shielding scene, and the algorithm comprises the following steps: S1, collecting image information, and carrying out the complete segmentation of a shielded object through an improved segmentation algorithm; s2, analyzing the motion state of the object through an object motion estimation method, and removing the accurately recognized dynamic region; s3, static information in a plurality of adjacent reference frames is utilized to complete restoration of a dynamic region, feature point extraction is performed on a restored data set, and the pose is computed.The method helps a network to deduce information of a shielded part through features in other directions when an object is shielded, and under the condition that boundary semantic information is weak, the network can still be subjected to feature point extraction, and the pose is calculated. The boundary of the object can still be accurately identified, the segmentation accuracy is improved, the key features of the shielded object can be captured, and the segmentation robustness is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of simultaneous location and mapping (SLAM), and in particular relates to a robot algorithm, a storage medium and a device for dynamic occlusion scenes. Background Art

[0002] Simultaneous Location And Mapping (SLAM) technology focuses on enabling mobile robots to achieve autonomous navigation and exploration in unknown environments. During the process of moving, this technology can accurately build models and maps of the environment, which is crucial for robots to understand and adapt to various complex environments, and has far-reaching significance for improving task execution efficiency and enhancing environmental interactivity. When a camera is used as a sensor, this technology is called visual SLAM. If a laser sensor is equipped, it is called laser SLAM. However, laser sensors have limitations in information acquisition and volume, which to a certain extent restrict the widespread application of laser SLAM technology. In contrast, visual SLAM has gradually become a research hotspot in the field of SLAM due to its small size, low cost and richer information acquisition capabilities. However, in actual scenes, due to the occlusion of objects, it is difficult to completely mark or accurately segment the occluded potential dynamic objects, and in the segmentation stage of potential dynamic objects, although possible dynamic elements can be preliminarily identified, there is a lack of effective verification mechanism to confirm whether these candidate objects have actually shifted, and the motion consistency of feature points between multiple consecutive frames cannot be guaranteed, such as the existing Oneformer segmentation network. Therefore, the existing technology cannot accurately evaluate the real-time motion state of potential dynamic objects, which leads to the problem that the traditional algorithm will reduce feature information when removing dynamic feature points. In addition, the existing technology is also insufficient in repairing the static background of the current frame RGB image after the dynamic area is removed. It fails to screen the most suitable image frame, and the static information in the adjacent frames may not be the best choice for static background repair, resulting in poor repair effect of the image static background, affecting the accuracy of feature information and pose estimation. Summary of the invention

[0003] The purpose of the present invention is to provide a robot algorithm for dynamic occlusion scenarios, which is used to solve the technical problems in the prior art that the real-time motion state of potential dynamic objects cannot be evaluated very accurately and the removal of dynamic feature points will result in a reduction in feature information.

[0004] The robot algorithm for dynamic occlusion scenes includes the following steps.

[0005] Step S1: collect image information and use an improved segmentation algorithm to completely segment the blocked object.

[0006] Step S2: Analyze the motion state of the object by using the object motion estimation method, and remove the accurately identified dynamic area.

[0007] Step S3: Use the static information in several adjacent reference frames to complete the repair of the dynamic area, extract feature points from the repaired data set, and calculate the pose.

[0008] Among them, the step S1 includes: 1) integrating multi-directional feature enhancement convolution into the backbone network, which consists of four strip convolutions in different directions and a dilated convolution, which significantly expands the receptive field of the convolution kernel and can flexibly control the range of the receptive field; 2) adding a double-pooling feature enhancement module to the Neck part of the network to perform feature enhancement processing on the input feature layer, aiming to strengthen the weak semantic information in these feature maps; 3) sending the high-dimensional feature map extracted by the backbone network to the multi-attention mechanism module, which can enhance the pixel weights of the dynamic object area, thereby improving the recognition accuracy of dynamically occluded objects.

[0009] Preferably, the calculation method of multi-directional feature enhancement convolution is as follows:

[0010]

[0011] Among them, M in Indicates input, represents the depth of convolution, 3×3 represents the size of the convolution kernel, Direction i Represents the convolution operation of multi-directional convolution in different directions, M dfec Represents the result of multi-directional feature enhancement convolution; the result M dfec First, batch normalization is performed, then it is processed by a multi-layer perceptron classifier, and then it is fused element by element with the input image for output.

[0012] Preferably, in the dual-pooling feature enhancement module, when the feature map D is input, the module first uses pooling kernels of various sizes to capture the multi-scale features of the image; for each size of pooling operation, two parallel branches are designed: one branch performs average pooling to fuse the information of the pooling area; the other branch performs maximum pooling to accurately capture the salient features in the pooling area; next, the feature maps generated by the two branches are merged in the channel dimension, and sequentially passed through a batch normalization (BN) layer and a ReLU activation function to enhance the stability and nonlinear expression ability of the features; then, a 1×1 convolution, a BN layer, and a ReLU activation function are applied. The merged feature maps are further processed with bilinear interpolation to extract features and reduce their dimensions. The feature maps are then upsampled to the same size as the module input feature maps using a 1×1 convolution, a BN layer, and a ReLU activation function, and the number of channels of the module input feature maps is compressed to one-fourth of the number of original input channels. The processed feature maps are then merged with the feature maps obtained by multi-scale pooling in the channel dimension. Finally, a 1×1 convolution, a BN layer, and a ReLU activation function are applied again to further fuse these multi-scale features, thereby generating the output feature map of the dual-pooling feature enhancement module.

[0013] Preferably, in step S2, for a feature point p in space, the current frame Y, the key frame Y o , where key frame Y o After that, the oth frame is the current frame Y. At the key frame Y o The reprojection error e between the current frame Y p It is expressed as the following formula:

[0014] e p =τ-π(I o ,K,L(R+n R ,t+n t )),

[0015] Among them, τ is the projection of feature point p to key frame Y o 2D pixel coordinates on the image, π is the projection function, I o is the feature point p in key frame Y o 3D coordinates on the camera coordinate system, L(R,t) is the relative transformation matrix from the world coordinate system to the camera coordinate system, R is the rotation matrix, t is the translation matrix, n R and n t They are rotation matrix noise and translation matrix noise respectively, K is the camera intrinsic parameter matrix;

[0016] Construct a cost function J, taking into account the reprojection error e of all feature points p , using the weighted residual w iTo reflect the importance of different feature points, an error covariance matrix Σp is introduced to represent the uncertainty of the feature point position; based on the distribution of the prior information P0 about the initial posture, the cost function J uses the maximum a posteriori probability estimation under the Bayesian framework to optimize the camera's posture parameters R and t; the corresponding cost function J is expressed as:

[0017]

[0018] Among them, A represents the set of feature points, λ is the regularization parameter, and the regularization term R(L) is introduced to constrain the change of camera posture L.

[0019] Preferably, in step S2, for each feature point p, the step calculates its average viewing angle deviation between a series of consecutive frames. Average viewing angle deviation The calculation formula is:

[0020]

[0021] in represents the average value of the viewing angle deviation of feature point p between multiple frames, ω is a function that measures the angle, and m j and are the positions of feature point p and its projection point in the jth frame, respectively, and the subscript j+o represents the j+oth frame, m j+o and Respectively represent the positions of feature point p and its projection point in the j+oth frame;

[0022] This step averages the viewing angle deviation Compare with the corresponding threshold θ; at the same time, introduce the velocity vector v p and the acceleration vector a p To characterize the motion state of the feature point, a feature point decision rule is obtained; the calculation formula of the decision rule is:

[0023]

[0024] Among them, α, β, γ are weight coefficients, corresponding to the average viewing angle deviation Velocity vector v p and the acceleration vector a p , s is the indicator function, v th and a th They are the velocity vector v p and the acceleration vector a p The threshold value, f th is the threshold of the final decision. If f p >f th , it is determined to be a dynamic point; otherwise it is a static point.

[0025] Preferably, the background restoration algorithm is used in step S3, including similarity search and pixel matching; the similarity search includes:

[0026] Step 1: Information estimation: Taking the target frame as the reference, move a time window containing several frames of images in both the forward and reverse directions as the repair window. Within the repair window, P = [x, y, z] T is a point in space, and its projections in the reference frame and target frame are p1 and p2 respectively. The pixel positions of the pixels p1 and p2 can be obtained from the pinhole camera model. The three-dimensional vector of the point P movement is calculated, and then the three-dimensional vector is projected onto the two-dimensional plane, and then the displacement change Δp and rotation change Δθ between any two frames are obtained. ORB feature extraction and counting operations are performed on each frame of the image, where the feature point count is ∑W i ;

[0027] Step 2, reference frame selection: extract features from the repair window and match them. When the displacement change Δp between an image frame and the target frame exceeds the threshold τ and the rotation change Δθ is greater than the threshold γ, this frame is used as a candidate reference frame to screen out candidate reference frames that meet specific conditions; next, these candidate reference frames are further screened, and feature points of each candidate reference frame are matched with the target frame, and the number of successfully matched feature points is recorded as ∑Z i ; If ∑W i and ∑Z i If the judgment condition is met, the candidate reference frame that meets the condition is added to the reference frame library; the judgment condition is expressed as:

[0028] |∑W i -∑Z i |<E1, where E1 represents the image information judgment threshold.

[0029] Preferably, the background repair algorithm used in pixel matching includes: an initialization stage, the target frame image to be repaired is set as image A, and the reference frame is set as image B. Subsequently, a 3×3 pixel block is randomly selected in image A as a matching block, and an offset is randomly assigned to it, and a matching block corresponding to the matching block in image A is found in image B; a propagation stage: the offset difference between the matching block in image A and the matching block in image B is calculated, and the minimum offset value is found; a search stage: for each pixel point in image B, a more matching offset is found within the concentric circle centered on the current matching block to replace the current offset, the initial search radius is set to the size of the image, and then the radius is gradually reduced at half the rate until the search ends; this process repeats the two steps of propagation and search until each pixel block finds the most appropriate and accurate offset; finally, the pixel values ​​corresponding to these offsets are assigned to the corresponding pixel blocks in image A.

[0030] The present invention also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of a robot algorithm for a dynamic occlusion scenario as described above are implemented.

[0031] The present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein when the processor executes the computer program, the steps of a robot algorithm for dynamic occlusion scenarios as described above are implemented.

[0032] The present invention has the following advantages:

[0033] 1. The present invention uses the Oneformer network for improvement and adds multi-directional feature enhancement convolution, which is designed to solve the situation where dynamic objects may present multiple posture states or irregular parts are blocked during movement. Through multi-directional convolution operations, the network can capture the extended features of objects in different directions. With the hole convolution, the receptive field of the convolution kernel is significantly expanded and the range of the receptive field can be flexibly controlled, thereby enhancing the understanding of the all-round features of dynamic objects. This design helps the network to infer the information of the occluded part through features in other directions when the object is blocked, thereby improving the accuracy of segmentation.

[0034] The dual-pooling feature enhancement module is designed to improve the network's ability to capture weak semantic features. By combining different pooling strategies such as maximum pooling and average pooling, the network can obtain richer feature expressions, thereby enhancing the understanding of occluded scenes. This design helps the network to accurately identify the boundaries of objects even when the boundary semantic information is weak, thereby improving the accuracy of segmentation.

[0035] The multi-attention module is designed to better focus on occluded areas. By introducing the attention mechanism, the network can adaptively adjust the attention to different areas, enhance the pixel weights of dynamic object areas, and thus more accurately identify occluded objects. This design helps the network capture the key features of occluded objects in the case of occlusion, thus improving the robustness of segmentation.

[0036] 2. In step S2, the present invention not only takes into account the reprojection errors of all feature points, but also introduces a regularization term to constrain the change of camera posture, thereby avoiding overfitting. At the same time, weighted residuals are used to reflect the importance of different feature points. In addition, the maximum a posteriori probability (MAP) estimation under the Bayesian framework is used to optimize the change parameters of the camera posture. This algorithm also characterizes the motion state of feature points from three aspects: average viewing angle deviation, velocity vector, and acceleration vector. By designing feature point decision rules, the motion consistency of feature points between multiple consecutive frames is determined, and the motion state of feature points is accurately identified, and then dynamic objects are eliminated using their semantic information.

[0037] 3. In order to extract a sufficient number of feature points and improve the accuracy of feature information and pose estimation, the present invention uses a background repair algorithm based on approximate nearest neighbor matching. The algorithm uses the static information in the adjacent frames to repair the static background of the current frame RGB image after the dynamic area is eliminated. The background repair algorithm of the optimal nearest neighbor pixel matching is specifically divided into two steps: similarity search and pixel matching. The similarity search is further divided into two sub-steps: information estimation and reference frame selection. Through these steps, the algorithm can select reference frames with sufficiently rich image information in the set repair window based on the richness of the image information, and can effectively restore the static background after the dynamic objects in the target frame are eliminated based on these reference frames, thereby improving the quality and quantity of feature points, thereby improving the performance of the SLAM algorithm in a dynamic environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 This is a basic flow chart of a robot algorithm for dynamic occlusion scenarios according to the present invention.

[0039] Figure 2 The flowchart of the SLAM system of the present invention is applied.

[0040] Figure 3 It is a schematic diagram of the structure of the improved Oneformer segmentation network in the present invention.

[0041] Figure 4 This is a flow chart of the multi-directional feature enhancement convolution module in the present invention.

[0042] Figure 5 Flow chart of the dual-pooling feature enhancement module in the present invention.

[0043] Figure 6 This is a flowchart of the improved multi-attention mechanism module in the present invention.

[0044] Figure 7 It is a schematic diagram of the principle of the background restoration algorithm in the present invention.

[0045] Figure 8This is a comparison chart of the object segmentation effects of the present invention and the existing technologies such as YOLOv8 and Oneformer.

[0046] Fig. 9 This is a comparison chart of the feature point judgment effects of the present invention and the existing algorithms of ORB-SLAM3 and DynaSLAM.

[0047] Fig.10 The figure is a background restoration effect diagram obtained by using the present invention.

[0048] Fig.11 This is a comparison chart of the feature point extraction effects of the present invention and the existing algorithms of ORB-SLAM3 and DynaSLAM.

[0049] Fig.12 Comparison diagram of the trajectories obtained by the present invention and ORB-SLAM3, DS-SLAM, and DynaSLAM algorithms in SLAM experiments.

[0050] Fig.13 In order to verify the feasibility of the present invention, a real scene graph with dynamic object occlusion is set.

[0051] Fig.14 This is a comparison chart of the trajectories obtained by the present invention and the ORB-SLAM3 algorithm in real scene experiments. DETAILED DESCRIPTION

[0052] The specific implementation modes of the present invention will be further explained in detail below by describing the embodiments with reference to the accompanying drawings, so as to help those skilled in the art to have a more complete, accurate and in-depth understanding of the inventive concept and technical solution of the present invention.

[0053] Embodiment 1.

[0054] like Figure 1-Figure 7 As shown, the present invention provides a robot algorithm for dynamic occlusion scenes, including the following steps.

[0055] Step S1: collect image information and use an improved segmentation algorithm to completely segment the blocked object.

[0056] In this step, the robot's built-in RGB-D camera is used to collect image information, and then the obtained image information is input into the segmentation network. Multiple modules are integrated into the backbone network to improve the accuracy of segmentation, including: 1) Integrating multi-directional feature enhancement convolution into the backbone network, which consists of four strip convolutions in different directions and dilated convolutions, significantly expanding the receptive field of the convolution kernel and being able to flexibly control the range of the receptive field; 2) Adding a dual-pooling feature enhancement module to the Neck part of the network to perform feature enhancement processing on the input feature layer, aiming to strengthen the weak semantic information in these feature maps; 3) Sending the high-dimensional feature map extracted by the backbone network to the multi-attention mechanism module, which can enhance the pixel weights of the dynamic object area, thereby improving the recognition accuracy of dynamically occluded objects.

[0057] Specifically, the segmentation network introduces a residual structure to enhance the performance of the backbone network based on the feature pyramid network (FPN) constructed by the backbone, forming a residual-feature pyramid network, and uses it as the backbone network; multiple modules are integrated into the backbone network to improve the accuracy of segmentation. A multi-directional feature enhancement convolution module is set between the C1 block and the C2 block, a multi-attention module is set on the lateral connection between the C4 block and the P4 block, and a double-pooling feature enhancement module is set on the lateral connection between the C3 block and the P3 block.

[0058] Among them, the method of multi-directional feature enhancement convolution includes: first, the input image is preliminarily processed using a 3×3 depth convolution. This step can accurately capture the local features of the input image and provide the necessary local information basis for the subsequent multi-directional strip convolution. Secondly, the strip convolutions in four directions (including horizontal, vertical, anti-diagonal and main diagonal) are applied in parallel to the results obtained from the preliminary processing, and each direction is matched with a corresponding hole convolution. This parallel processing method enables the network to simultaneously obtain rich feature maps from multiple scales and different levels, so as to more comprehensively understand the feature structure of the input data. Finally, the features of the strip convolutions and hole convolutions from these four different directions are fused. This fusion process not only retains the key feature information in each direction, but also generates a richer and more comprehensive multi-scale feature representation. This fusion strategy enables the network to pay attention to local details while taking into account global features, thereby improving the overall feature extraction capability.

[0059] The calculation method of multi-directional feature enhancement convolution is as follows:

[0060]

[0061] Among them, M in Represents the input of the MDFEC (Multi-Directional Feature Enhanced Convolution) module, represents the depth of convolution, 3×3 represents the size of the convolution kernel, Directioni Represents the convolution operation of multi-directional convolution in different directions, M dfec Represents the result of multi-directional feature enhancement convolution. After multi-directional feature enhancement convolution, the result is first batch normalized, then processed by the multi-layer perceptron classifier, and then element-by-element fused with the input image to finally generate the MDFEC module output.

[0062] In the dual pooling feature enhancement module, when the input feature map D is input, pooling kernels of various sizes are first used to capture the multi-scale features of the image. The pooling operation of each size pooling kernel has two parallel branches: one branch performs average pooling to fuse the information of the pooling area; the other branch performs maximum pooling to accurately capture the salient features in the pooling area. Next, the feature maps generated by the two branches are merged in the channel dimension, and the merged feature map is obtained by batch normalization (BN) layer and ReLU activation function in sequence, which can enhance the stability and nonlinear expression ability of the features. Subsequently, a 1×1 convolution, BN layer and ReLU activation function are applied to the merged feature map to achieve further feature extraction and dimensionality reduction. After that, these feature maps are upsampled to the same size as the input feature map using the bilinear interpolation method to form the pooling result of the corresponding scale features. The input feature map is then processed by a 1×1 convolution, BN layer and ReLU activation function, and its number of channels is compressed to one-fourth of the original input channel number. Then, this processed feature map is merged with the feature map obtained by multi-scale pooling in the channel dimension. Finally, a 1×1 convolution, BN layer and ReLU activation function are applied again to further fuse these multi-scale features to generate the output feature map of the DPEM (Dual Pooling Feature Enhancement Module) module. The pooling kernel uses three sizes of pooling operations: global pooling, pooling with a kernel size of 4×4 and a stride of 4, and pooling with a kernel size of 2×2 and a stride of 2. These pooling operations of different sizes can capture feature information of different scales.

[0063] The spatial attention mechanism aims to increase the weight of pixel values ​​in the occluded areas of the feature map. When a high-dimensional feature map is input into the spatial attention mechanism, average pooling and maximum pooling are performed respectively, and then the two sets of pooled features are concat fused to form a new feature map f. c Next, the feature map passes through a 3×3×1 convolution layer and a sigmoid function to obtain a weighted spatial attention map f u The formula of the spatial attention mechanism is as follows:

[0064] f u =sigmoid(c(f c)), where c is a 3×3×1 convolutional network and sigmoid() represents the sigmoid function. Spatial attention map f u It can guide the network to pay more attention to the area where the dynamic objects are located. u Connect with the original feature map Q to get the feature map Q weighted by spatial attention u .

[0065] The channel attention mechanism is used to assign corresponding weights to each layer of channels in the feature map. This algorithm inputs the H×W×C feature map F into the channel attention mechanism, performs global attention average pooling and maximum pooling operations on the feature map, and obtains the information of each channel of the feature map. The feature F obtained by average pooling and maximum pooling avg With F max The fully connected layer FC module strengthens the correlation between channels and redistributes the weights of each channel to obtain the weighted channel attention map f v , to better learn the occlusion features. The formula of the channel attention mechanism is as follows:

[0066] f v =sigmoid(Pη(F avg +F max ))

[0067] Where η represents the ReLU function and P is the parameter of the fully connected layer. Channel attention map f v Connect with the original feature map Q to get the feature map Q weighted by channel attention v .

[0068] The feature map Q weighted by channel attention v and the feature map Q weighted by spatial attention u After concatenation, we get the output of the multi-attention mechanism module.

[0069] Step S2: Analyze the motion state of the object by using the object motion estimation method, and remove the accurately identified dynamic area.

[0070] This step uses the camera pose estimation to preliminarily determine the camera position, and then analyzes the motion state of the object. During the analysis process, the feature points in space are dynamically judged. For a feature point p in space, the current frame Y, and the key frame Y o , calculate the feature point p in key frame Y o The reprojection error between the current frame Y needs to take into account the uncertainty of the camera's attitude parameters R, t, and the camera's intrinsic parameter matrix K, and the noise n of the camera attitude estimation. R , n t Among them, the key frame Y oThe oth frame is then the current frame Y. In this embodiment, o is 10. o The reprojection error e between the current frame Y p It is expressed as the following formula: p =τ-π(I o ,K,L(R+n R ,t+n t )),

[0071] Among them, τ is the projection of feature point p to key frame Y o 2D pixel coordinates on the image, π is the projection function, I o is the feature point p in key frame Y o 3D coordinates on the camera coordinate system, L(R,t) is the relative transformation matrix from the world coordinate system to the camera coordinate system, R is the rotation matrix, t is the translation matrix, n R and n t They are rotation matrix noise and translation matrix noise respectively, and K is the camera intrinsic parameter matrix.

[0072] In order to optimize the reprojection error, a cost function J is constructed, which not only takes into account the reprojection error e of all feature points p , and a regularization term R(L) is introduced to constrain the change of camera pose L to avoid overfitting. At the same time, the weighted residual w is used i To reflect the importance of different feature points, an error covariance matrix Σp is introduced to represent the uncertainty of the feature point position. In addition, based on the distribution of the prior information P0 about the initial posture, the cost function J uses the maximum a posteriori probability (MAP) estimation under the Bayesian framework to optimize the camera's posture parameters R and t. Therefore, the corresponding cost function J is expressed as:

[0073]

[0074] Among them, A represents the set of feature points, and λ is the regularization parameter, which is used to balance the influence of data term and regularization term. represents the transpose of the reprojection error, Σp -1 is the inverse matrix of the error covariance matrix Σp.

[0075] In order to more accurately determine whether a feature point belongs to a dynamic object, in addition to the perspective deviation In addition, the continuity in the time dimension can also be considered, that is, the consistency of the motion of the feature points between multiple consecutive frames. For each feature point p, this step calculates its average viewing angle deviation between a series of consecutive frames. Average viewing angle deviation The calculation formula is:

[0076]

[0077] in represents the average value of the viewing angle deviation of feature point p between multiple frames, ω is a function that measures the angle, and m j and are the positions of feature point p and its projection point in the jth frame, respectively, and the subscript j+o represents the j+oth frame, m j+o and Respectively represent the positions of the feature point p and its projection point in the j+oth frame.

[0078] This step averages the viewing angle deviation Compare with the corresponding threshold θ; at the same time, introduce the velocity vector v p and the acceleration vector a p To characterize the motion state of the feature point, a feature point decision rule is obtained. The calculation formula of the decision rule is:

[0079]

[0080] Among them, α, β, γ are weight coefficients, corresponding to the average viewing angle deviation Velocity vector v p and the acceleration vector a p , s is the indicator function, v th and a th They are the velocity vector v p and the acceleration vector a p The threshold value, f th is the threshold of the final decision. If f p >f th , it is determined to be a dynamic point; otherwise it is a static point.

[0081] After completing the dynamic judgment of the feature points, the accurately identified dynamic area is eliminated. So far, after the occluded object is completely segmented and accurately determined to be a dynamic object, the algorithm uses its semantic information to achieve accurate elimination of dynamic objects.

[0082] Step S3: Use the static information in several adjacent reference frames to complete the repair of the dynamic area, extract feature points from the repaired data set, and calculate the pose.

[0083] In this step, an optimal neighbor pixel matching background restoration algorithm based on approximate nearest neighbor matching is used. This algorithm uses static information from several adjacent frames to repair the static background in the RGB image of the current frame after removing the dynamic area. The optimal neighbor pixel matching background restoration algorithm is specifically divided into two steps: similarity search and pixel matching. Similarity search is further divided into two sub-steps: information estimation and reference frame selection.

[0084] Step 1: Information estimation:

[0085] Given that the richness of image information has a decisive influence on the effectiveness of image restoration, the method of simply mapping the previous frame to the target frame is difficult to effectively extract image information. This algorithm uses inter-frame feature matching technology to evaluate and obtain image information. The specific method is to use the target frame as a reference and move a time window containing several frames (15 frames in this embodiment) of images along the arrow direction (including both positive and negative directions) as the restoration window. In this restoration window, ORB feature extraction and counting operations are performed on each frame of the image, where the feature point count is ∑W i .

[0086] Pose is one of the important indicators for expressing the relationship between adjacent frames. Epipolar geometry is used to estimate the relationship between adjacent frames. T is a point in space, whose projections in the reference frame and target frame are p1 and p2 respectively. The pixel positions of pixel points p1 and p2 can be obtained from the pinhole camera model. The three-dimensional vector of the motion of point P is calculated, and then the three-dimensional vector is projected onto the two-dimensional plane, and then the displacement change Δp and rotation change Δθ between any two frames are obtained.

[0087] Step 2: Reference frame selection:

[0088] In order to optimize the reference frame selection process and ensure the effectiveness of approximate nearest neighbor matching, this step combines the image information and posture information calculated in step 1. This step takes a time window of 15 forward and reverse frames as a group (i.e., a repair window), extracts features from the window and matches them to screen out candidate reference frames that meet specific conditions. The specific conditions include Δp exceeding the threshold τ and Δθ greater than the threshold γ. If they are met, this frame is used as a candidate reference frame. Next, these candidate reference frames are further screened, and each candidate reference frame is matched with the target frame for feature points, and the number of successfully matched feature points is recorded as ∑Z i , then |∑W i -∑Z i | can reflect the richness of the image information of the candidate reference frame. i and ΣZ i If the judgment condition is met, the candidate reference frame is determined to contain efficient and rich image information, and the candidate reference frame that meets the condition is added to the reference frame library. The judgment condition is expressed as:

[0089] |ΣW i -ΣZ i |<E1, where E1 represents the image information judgment threshold, which is set to 100 in this embodiment.

[0090] The reference frame library is obtained by the similarity search process mentioned above, and the pixel matching is performed by improving the traditional approximate nearest neighbor matching algorithm to perform background repair. The background repair algorithm mainly includes three stages: initialization, propagation, and search. In the initialization stage, the target frame image to be repaired is set as image A, and the reference frame is set as image B. Subsequently, a 3×3 pixel block is randomly selected in image A as the matching block, and an offset is randomly assigned to it, and a matching block corresponding to the matching block in image A is found in image B. In the propagation stage, the offset difference between the matching block in image A and the matching block in image B is calculated, and the value with the smallest offset is found. In the search stage, for each pixel point in image B, a more matching offset is found within the concentric circle centered on the current matching block to replace the current offset. The initial radius of the search is set to the size of the image, and then the radius is gradually reduced at half the rate until the search ends. This process repeats the two steps of propagation and search until each pixel block finds the most appropriate and accurate offset. Finally, the pixel values ​​corresponding to these offsets are assigned to the corresponding pixel blocks in image A. Experiments show that the reference frame library usually consists of 2-4 frames of images, and the number of iterations is generally 3, which can achieve a better restoration effect.

[0091] Through these steps, this algorithm can effectively restore the static background after the dynamic area is removed, improve the quality and quantity of feature points, and thus improve the performance of the SLAM algorithm in a dynamic environment. After completing the background repair, feature points are extracted from the repaired data set and the pose is calculated.

[0092] The following is an explanation of the process of the above-mentioned robot algorithm for dynamic occlusion scenarios in combination with specific experiments.

[0093] like Figure 8 As shown in the figure, in order to verify the segmentation effect of the improved segmentation network algorithm, a comprehensive verification of occlusion in four different situations was carried out on the dynamic subset of the TUM dataset, in which the red box in the figure is used as an auxiliary mark. The segmentation ability of highly occluded objects can be verified from the first and second rows in the figure, and the low occlusion objects can be verified from the third and fourth rows. Through segmentation comparison, it is found that since YOLOv8 does not have a special object feature enhancement layer, it cannot complete the complete segmentation of the boundary of the occluded object, resulting in interruption in the segmentation process or low segmentation accuracy. The unimproved Oneformer has improved segmentation effect compared with YOLOv8, but the segmentation accuracy is still insufficient. Therefore, a multi-directional feature enhancement convolution and feature enhancement module are designed in this algorithm to enhance the underlying subject information. At the same time, a multi-attention module is integrated to focus on the occluded area, while weakening the influence of other non-essential information and enhancing the accuracy of segmentation. In general, in dynamic occlusion scenes, this algorithm has a good effect on the segmentation of potential dynamic objects with few pixel weights and weak boundary semantic information.

[0094] like Fig. 9 As shown in Figure 2, four sub-datasets of dynamic sequences in the TUM dataset are selected to judge and evaluate the feature points of objects, such as Figure 8 As shown. The first and second rows are feature point judgments of the ORB-SLAM3 algorithm and the Dyna-SLAM algorithm respectively, and the third row is the feature point judgment of this algorithm. It can be seen that the ORB-SLAM3 algorithm judges the feature points on both dynamic objects and static objects as static feature points (people are dynamic targets, green is the mark of static feature points, and red is the mark of dynamic feature points). The feature point judgment effect of Dyna-SLAM is obviously better than the former, but it will make misjudgments when the object moves at high speed. Compared with the first two algorithms, this algorithm has a more accurate ability to judge feature points.

[0095] like Fig.10 As shown, Fig. 9 (a) to (d) are the original RGB images containing dynamic characters. Fig. 9 (e)~(h) are the restored RGB images. It can be seen that the restored images only contain the original static background in the scene.

[0096] like Fig.11 As shown in the figure, ORB-SLAM3 and DynaSLAM algorithms are selected to compare with this algorithm for feature point extraction. ORB-SLAM3 cannot distinguish between dynamic and static objects, so feature points on dynamic objects cannot be removed. Dyna-SLAM uses instance segmentation technology to mark the area where dynamic objects are located and remove them from the environment to reduce the impact of dynamic objects on the global structure. Figure 1 However, after removing dynamic objects, this algorithm reduces the number of static feature points used for pose estimation and map construction. This algorithm removes dynamic areas and performs background repair on the removed areas to extract more abundant static feature points, thereby being able to construct a more accurate trajectory map.

[0097] like Fig.12 As shown in the figure, ORB-SLAM3 cannot distinguish between dynamic and static objects, so its robustness is very poor. Dyna-SLAM uses instance segmentation technology to mark the area where dynamic objects are located and remove them from the environment to reduce the impact of dynamic objects on the global structure. Figure 1However, in a high-occlusion environment or when the running speed is too fast, it is difficult to accurately identify potential dynamic objects and their motion state. Moreover, after removing dynamic objects, this algorithm does not perform background repair, which reduces the number of static feature points used for pose estimation and map construction, thus affecting the positioning accuracy of the algorithm. This algorithm uses an improved Oneformer segmentation network, a dynamic judgment method that combines camera pose estimation with object motion state judgment, and background repair processing to better solve the impact of dynamic feature points, repair dynamic areas, and increase the number of feature points used for pose estimation.

[0098] like Fig.13 As shown in the figure, to verify the feasibility of this method, a real scene with dynamic object occlusion is set for SLAM experiment. The Husky wheeled robot is used. The platform configuration is i7-10875H CPU, 8GB memory, GTX1080 GPU, Ubuntu18.04 operating system, and the data set is collected in the real scene. The experimental results are shown in the figure. Fig.14 As shown in the figure, it can be seen that the ORB-SLAM3 algorithm has poor trajectory accuracy due to its inability to determine feature points. This method combines an improved segmentation network, accurately determines feature points, and performs dynamic background repair, which improves its trajectory accuracy.

[0099] Embodiment 2.

[0100] Corresponding to the first embodiment of the present invention, the second embodiment of the present invention provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the following steps are implemented according to the method of the first embodiment.

[0101] Step S1: collect image information and use an improved segmentation algorithm to completely segment the blocked object.

[0102] Step S2, analyzing the motion state of the object by using the object motion estimation method, and removing the accurately identified dynamic area.

[0103] Step S3, using the static information in several adjacent reference frames to complete the repair of the dynamic area, extracting feature points from the repaired data set, and calculating the pose.

[0104] The above storage media include: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), optical disk and other media that can store program codes.

[0105] The specific limitations on the steps implemented after the program in the computer-readable storage medium is executed can be found in Example 1, and will not be described in detail here.

[0106] Embodiment three.

[0107] Corresponding to the first embodiment of the present invention, the third embodiment of the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the program, the following steps are implemented according to the method of the first embodiment.

[0108] Step S1: collect image information and use an improved segmentation algorithm to completely segment the blocked object.

[0109] Step S2, analyzing the motion state of the object by using the object motion estimation method, and removing the accurately identified dynamic area.

[0110] Step S3, using the static information in several adjacent reference frames to complete the repair of the dynamic area, extracting feature points from the repaired data set, and calculating the pose.

[0111] The specific limitations on the above steps of implementing the computer device can be found in Example 1, and will not be described in detail here.

[0112] It should be noted that each box in the block diagram and / or flow chart in the accompanying drawings of the specification of the present invention, as well as the combination of boxes in the block diagram and / or flow chart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and machine instructions.

[0113] The present invention is described above by way of example in conjunction with the accompanying drawings. It is obvious that the specific implementation of the present invention is not limited to the above-mentioned method. As long as various non-substantial improvements are made using the inventive concept and technical solution of the present invention, or the inventive concept and technical solution are directly applied to other occasions without improvement, they are all within the protection scope of the present invention.

Claims

1. A robot algorithm for dynamic occlusion scenes, characterized by: The following steps are involved: Step S1, collecting image information, and using an improved segmentation algorithm to completely segment the obscured object; Step S2, analyzing the motion state of the object by an object motion estimation method, and removing the accurately identified dynamic area; Step S3, using static information in several adjacent reference frames to complete the repair of the dynamic area, extracting feature points from the repaired data set, and calculating the pose; Among them, the step S1 includes: 1) integrating multi-directional feature enhancement convolution into the backbone network, which consists of four strip convolutions in different directions and a dilated convolution, which significantly expands the receptive field of the convolution kernel and can flexibly control the range of the receptive field; 2) adding a double-pooling feature enhancement module to the Neck part of the network to perform feature enhancement processing on the input feature layer, aiming to strengthen the weak semantic information in these feature maps; 3) sending the high-dimensional feature map extracted by the backbone network to the multi-attention mechanism module, which can enhance the pixel weights of the dynamic object area, thereby improving the recognition accuracy of dynamically occluded objects.

2. The robot algorithm for dynamic occlusion scenarios according to claim 1, characterized in that: The calculation method of multi-directional feature enhancement convolution is as follows: Among them, M in Indicates input, represents the depth of convolution, 3×3 represents the size of the convolution kernel, Direction i Represents the convolution operation of multi-directional convolution in different directions, M dfec Represents the result of multi-directional feature enhancement convolution; the result M dfec First, batch normalization is performed, then it is processed by a multi-layer perceptron classifier, and then it is fused element by element with the input image for output.

3. The robot algorithm for dynamic occlusion scenarios according to claim 1, characterized in that: In the dual pooling feature enhancement module, when the feature map D is input, the module first uses pooling kernels of various sizes to capture the multi-scale features of the image; For each pooling operation of each size, two parallel branches are designed: one branch performs average pooling to fuse the information of the pooling area; the other branch performs maximum pooling to accurately capture the salient features in the pooling area; then, the feature maps generated by the two branches are merged in the channel dimension and passed through the batch normalization (BN) layer and the ReLU activation function in turn to enhance the stability and nonlinear expression ability of the features; Subsequently, a 1×1 convolution, BN layer and ReLU activation function are applied to further extract features and reduce the dimension of the merged feature map; after that, the bilinear interpolation method is used to upsample these feature maps to the same size as the module input feature map; the module input feature map is then processed by a 1×1 convolution, BN layer and ReLU activation function, and its number of channels is compressed to one-fourth of the original input channel number; then, this processed feature map is merged with the feature map obtained by multi-scale pooling in the channel dimension; finally, a 1×1 convolution, BN layer and ReLU activation function are applied again to further fuse these multi-scale features, thereby generating the output feature map of the dual pooling feature enhancement module.

4. The robot algorithm for dynamic occlusion scenarios according to claim 1, characterized in that: In step S2, for a feature point p in space, the current frame Y, the key frame Y o , where key frame Y o After that, the oth frame is the current frame Y. At the key frame Y o The reprojection error e between the current frame Y p It is expressed as the following formula: e p =τ-π(I o ,K,L(R+n R ,t+n t )), Among them, τ is the projection of feature point p to key frame Y o 2D pixel coordinates on the image, π is the projection function, I o is the feature point p in key frame Y o 3D coordinates on the camera coordinate system, L(R,t) is the relative transformation matrix from the world coordinate system to the camera coordinate system, R is the rotation matrix, t is the translation matrix, n R and n t They are rotation matrix noise and translation matrix noise respectively, K is the camera intrinsic parameter matrix; Construct a cost function J, taking into account the reprojection error e of all feature points p , using the weighted residual w i To reflect the importance of different feature points, an error covariance matrix Σp is introduced to represent the uncertainty of the feature point position; based on the distribution of the prior information P0 about the initial posture, the cost function J uses the maximum a posteriori probability estimation under the Bayesian framework to optimize the camera's posture parameters R and t; the corresponding cost function J is expressed as: Among them, A represents the set of feature points, λ is the regularization parameter, and the regularization term R(L) is introduced to constrain the change of camera posture L.

5. The robot algorithm for dynamic occlusion scenarios according to claim 4, characterized in that: In step S2, for each feature point p, the average viewing angle deviation between a series of consecutive frames is calculated. Average viewing angle deviation The calculation formula is: in represents the average value of the viewing angle deviation of feature point p between multiple frames, ω is a function that measures the angle, and m j and are the positions of feature point p and its projection point in the jth frame, respectively, and the subscript j+o represents the j+oth frame, m j+o and Respectively represent the positions of feature point p and its projection point in the j+oth frame; This step averages the viewing angle deviation Compare with the corresponding threshold θ; at the same time, introduce the velocity vector v p and the acceleration vector a p To characterize the motion state of the feature point, a feature point decision rule is obtained; the calculation formula of the decision rule is: Among them, α, β, γ are weight coefficients, corresponding to the average viewing angle deviation Velocity vector v p and the acceleration vector a p , s is the indicator function, v th and a th They are the velocity vector v p and the acceleration vector a p The threshold value, f th is the threshold of the final decision. If f p >f th , it is determined to be a dynamic point; otherwise it is a static point.

6. The robot algorithm for dynamic occlusion scenarios according to claim 1, characterized in that: In step S3, a background restoration algorithm is used, including similarity search and pixel matching; Similar searches include: Step 1: Information estimation: Taking the target frame as the reference, move a time window containing several frames of images in both the forward and reverse directions as the repair window. Within the repair window, P = [x, y, z] T is a point in space, and its projections in the reference frame and target frame are p1 and p2 respectively. The pixel positions of the pixels p1 and p2 can be obtained from the pinhole camera model. The three-dimensional vector of the point P movement is calculated, and then the three-dimensional vector is projected onto the two-dimensional plane, and then the displacement change Δp and rotation change Δθ between any two frames are obtained. ORB feature extraction and counting operations are performed on each frame of the image, where the feature point count is ∑W i ; Step 2, reference frame selection: extract features from the repair window and match them. When the displacement change Δp between an image frame and the target frame exceeds the threshold τ and the rotation change Δθ is greater than the threshold γ, this frame is used as a candidate reference frame to screen out candidate reference frames that meet specific conditions; next, these candidate reference frames are further screened, and feature points of each candidate reference frame are matched with the target frame, and the number of successfully matched feature points is recorded as ∑Z i ; If ∑W i and ∑Z i If the judgment condition is met, the candidate reference frame that meets the condition is added to the reference frame library; the judgment condition is expressed as: ∑W i -∑Z i |<E1, Wherein, E1 represents the image information judgment threshold.

7. The robot algorithm for dynamic occlusion scenarios according to claim 6, characterized in that: The background repair algorithm used in pixel matching includes: initialization stage, the target frame image to be repaired is set as image A, and the reference frame is set as image B. Subsequently, a 3×3 pixel block is randomly selected in image A as the matching block, and an offset is randomly assigned to it, and a matching block corresponding to the matching block in image A is found in image B; propagation stage: the offset difference between the matching block in image A and the matching block in image B is calculated, and the minimum offset value is found; search stage: for each pixel point in image B, a more matching offset is found within the concentric circle centered on the current matching block to replace the current offset, the initial search radius is set to the size of the image, and then the radius is gradually reduced at half the rate until the search ends; this process repeats the two steps of propagation and search until each pixel block finds the most appropriate and accurate offset; finally, the pixel values ​​corresponding to these offsets are assigned to the corresponding pixel blocks in image A.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a robot algorithm for dynamic occlusion scenarios as described in any one of claims 1 to 7 are implemented.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: When the processor executes the computer program, the steps of a robot algorithm for a dynamic occlusion scene as described in any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • Robot target recognition and motion detection method based on deep learning, storage medium and equipment

    CN114782691A

  • Robot map construction method based on double constraints, storage medium and equipment

    CN116067360A

  • Deep learning based robot target recognition and motion detection method, storage medium and apparatus

    US11763485B1

Cited By

  • Robot algorithm facing depth constraint in dynamic scene, storage medium and equipment

    CN121582336A