A Robot Algorithm, Storage Medium and Device for Dynamic Occlusion Scenarios
Through improved segmentation algorithm and background repair technology, the problem of inaccurate feature point recognition in dynamic occlusion scenarios is solved, more efficient feature information recovery and pose estimation are achieved, and the dynamic environment adaptability of the SLAM algorithm is improved.
Patent Information
- Application Number
- CN202510090444.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-01-21
AI Technical Summary
The prior art cannot evaluate the real-time motion state of potential dynamic objects very accurately, resulting in the removal of dynamic feature points, resulting in the reduction of feature information, and the static background behind the dynamic area cannot be effectively repaired, affecting the accuracy of feature information and pose estimation.
An improved segmentation algorithm is used to segment objects, and the feature recognition accuracy is improved through multi-directional feature enhancement convolution and dual-pooling feature enhancement modules, and the recognition of dynamic object areas is enhanced with multi-attention mechanisms; dynamic area repair is used to repair dynamic areas, and static background is restored through background repair algorithms.
It improves the accuracy of segmentation and identification of feature points in dynamic occlusion scenarios, enhances the robustness and quantity of feature information, and improves the performance of SLAM algorithm in dynamic environments.
Smart Images

Figure CN120014597B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of Simultaneous Location And Mapping (SLAM), and particularly relates to a robot algorithm, a storage medium, and a device for a dynamic occlusion scenario. Background Art
[0002] Simultaneous Location And Mapping (SLAM) technology focuses on enabling a mobile robot to achieve autonomous navigation and exploration in an unknown environment. During the movement, this technology can accurately construct a model and a map of the environment, which is crucial for the robot to understand and adapt to various complex environments, and has far-reaching significance for improving the task execution efficiency and enhancing the environmental interactivity. When a camera is used as a sensor, this technology is called visual SLAM. If a laser sensor is carried, it is called laser SLAM. However, laser sensors have limitations in terms of information acquisition, volume, etc., which to a certain extent restricts the wide application of laser SLAM technology. In contrast, visual SLAM has gradually become a research hotspot in the SLAM field due to its small volume, low cost, and richer information acquisition ability. However, in actual scenarios, due to object occlusion, it is difficult to completely label or accurately segment the occluded potential dynamic objects. During the segmentation stage of potential dynamic objects, although possible dynamic elements can be initially identified, there is a lack of an effective verification mechanism to confirm whether these candidate objects have actually moved, and the motion consistency of feature points between multiple consecutive frames cannot be guaranteed, such as the existing Oneformer segmentation network. Therefore, the prior art cannot very accurately evaluate the real-time motion state of potential dynamic objects, which leads to the problem that removing dynamic feature points by traditional algorithms will cause a reduction in feature information. In addition, the prior art also has deficiencies in repairing the static background of the current frame RGB image after removing the dynamic area, and fails to screen the most suitable image frames, and the static information in adjacent frames may not be the best choice for repairing the static background, resulting in poor repair effect of the image static background and affecting the accuracy of feature information and pose estimation. Summary of the Invention
[0003] The object of the present invention is to provide a robot algorithm for a dynamic occlusion scenario, which is used to solve the technical problem in the prior art that the real-time motion state of potential dynamic objects cannot be very accurately evaluated, and removing dynamic feature points will cause a reduction in feature information.
[0004] The described robot algorithm for a dynamic occlusion scenario includes the following steps.
[0005] Step S1: Collect image information and use an improved segmentation algorithm to perform a complete segmentation of the occluded object.
[0006] Step S2: Analyze the motion state of the object by means of object motion estimation, and eliminate the accurately identified dynamic regions.
[0007] Step S3: Use the static information in several adjacent reference frames to complete the repair of the dynamic regions, extract feature points from the repaired dataset, and calculate the pose.
[0008] Among them, the step S1 includes: 1) Incorporate multi-directional feature enhancement convolution into the backbone network, which consists of four bar-shaped convolutions in different directions combined with dilated convolution, significantly expanding the receptive field of the convolution kernel and being able to flexibly control the range of the receptive field; 2) Add a double pooling feature enhancement module in the Neck part of the network to perform feature enhancement processing on the input feature layers, aiming to strengthen the weak semantic information in these feature maps; 3) Feed the high-dimensional feature maps extracted by the backbone network into a multi-attention mechanism module, which can enhance the pixel weights of the dynamic object regions, thereby improving the recognition accuracy of dynamic occluded objects.
[0009] Preferably, the calculation method of the multi-directional feature enhancement convolution is shown as follows:
[0010]
[0011] Where M in represents the input, represents the depth convolution, 3×3 represents the size of the convolution kernel, Direction i represents the convolution operation of the multi-directional convolution in different directions, M dfec represents the result of the multi-directional feature enhancement convolution; the result M dfec is first subjected to batch normalization processing, then processed by a multi-layer perceptron classifier, and then fused with the input image by element-wise addition for output.
[0012] Preferably, in the dual pooling feature enhancement module, when the feature map D is input, the module first uses pooling kernels of various sizes to capture the multi-scale features of the image; for each size of pooling operation, two parallel branches are designed: one branch performs average pooling to fuse the information in the pooling region; the other branch performs max pooling to accurately capture the significant features in the pooling region; next, the feature maps generated by these two branches are merged in the channel dimension and sequentially passed through a batch normalization (BN) layer and a ReLU activation function to enhance the stability and non-linear expression ability of the features; subsequently, a 1×1 convolution, a BN layer, and a ReLU activation function are applied to further extract features and reduce the dimension of the merged feature map; then, the bilinear interpolation method is used to upsample these feature maps to the same size as the input feature map of the module; the input feature map of the module is then processed through a 1×1 convolution, a BN layer, and a ReLU activation function, and its number of channels is compressed to one-fourth of the original input number of channels; then, this processed feature map is merged with the feature map obtained by multi-scale pooling in the channel dimension; finally, a 1×1 convolution, a BN layer, and a ReLU activation function are applied again to further fuse these multi-scale features, thereby generating the output feature map of the dual pooling feature enhancement module.
[0013] Preferably, in step S2, for a feature point p in space, the current frame Y, and the key frame Y o , where the o-th frame after the key frame Y o is the current frame Y, the reprojection error e o between the key frame Y p and the current frame Y is expressed as the following formula:
[0014] e p = τ - π(I o , K, L(R + n R , t + n t ))
[0015] where τ is the 2D pixel coordinate of the feature point p projected onto the key frame Y o , π is the projection function, I o is the 3D coordinate of the feature point p on the key frame Y o , L(R, t) is the relative transformation matrix from the world coordinate system to the camera coordinate system, R is the rotation matrix, t is the translation matrix, n R and n t are the rotation matrix noise and the translation matrix noise respectively, and K is the camera intrinsic matrix;
[0016] Construct a cost function J. Considering the reprojection error e p of all feature points, use the weighted residual w iTo reflect the importance of different feature points, an error covariance matrix Σp is introduced to represent the uncertainty of the feature point positions; based on the prior information P0 about the distribution of the initial pose, the cost function J uses the maximum a posteriori probability estimation in the Bayesian framework to optimize the camera pose parameters R and t; the corresponding cost function J is expressed as:
[0017]
[0018] where A represents the set of feature points, λ is the regularization parameter, and a regularization term R(L) is introduced to constrain the change of the camera pose L.
[0019] Preferably, in the step S2, for each feature point p, this step calculates its average viewing angle deviation among a series of consecutive frames Average viewing angle deviation The calculation formula is:
[0020]
[0021] where represents the average value of the viewing angle deviation of the feature point p among multiple frames, ω is a function for measuring the included angle, m j and are the positions of the feature point p and its projection point in the j-th frame respectively, and the subscript j + o represents the (j + o)-th frame, m j+o and represent the positions of the feature point p and its projection point in the (j + o)-th frame respectively;
[0022] This step compares the average viewing angle deviation with the corresponding threshold θ; meanwhile, a velocity vector v p and an acceleration vector a p are introduced to characterize the motion state of the feature point, and a feature point decision rule is obtained; the calculation formula of the decision rule is:
[0023]
[0024] where α, β, γ are weight coefficients, corresponding to the average viewing angle deviation velocity vector v p and acceleration vector a p respectively, s is the indicator function, v th and a th are the thresholds of the velocity vector v p and acceleration vector a p respectively, f th is the threshold of the final decision. If f p > f th , it is determined as a dynamic point; otherwise, it is a static point.
[0025] Preferably, the background repair algorithm used in step S3 includes similarity search and pixel matching; the similarity search includes:
[0026] Step 1, information estimation: Taking the target frame as a reference, move a time window containing several frames of images in both the forward and reverse directions as the repair window. Within this repair window, let P = [x, y, z] T be a point in space, and its projections in the reference frame and the target frame be p1 and p2 respectively. The pixel positions of pixel points p1 and p2 can be obtained from the pinhole camera model; from this, calculate the three-dimensional vector of the movement of point P, then project the three-dimensional vector onto the two-dimensional plane, and then obtain the displacement change amount Δp and the rotation change amount Δθ between any two frames; perform ORB feature extraction and counting operations on each frame of the image, where the number of feature points counted is ∑W i ;
[0027] Step 2, reference frame selection: Extract features from the repair window and perform matching. When the displacement change amount Δp exceeds the threshold τ and the rotation change amount Δθ is greater than the threshold γ between an image frame and the target frame, use this frame as a candidate reference frame to screen out candidate reference frames that meet specific conditions; next, further screen these candidate reference frames, perform feature point matching between each candidate reference frame and the target frame, and record the number of successfully matched feature points as ∑Z i ; if ∑W i and ∑Z i meet the judgment condition, then add the candidate reference frames that meet the condition to the reference frame library; the judgment condition is expressed as:
[0028] |∑W i -∑Z i | < E1, where E1 represents the image information judgment threshold.
[0029] Preferably, the background repair algorithm used for pixel matching includes: In the initialization stage, set the target frame image to be repaired as Figure A and the reference frame as Figure B. Subsequently, randomly select a 3×3 pixel block in Figure A as the matching block and randomly assign an offset to it, and find a matching block corresponding to the matching block in Figure A in Figure B; Propagation stage: Calculate the offset difference between the matching block in Figure A and the matching block in Figure B, and find the value with the smallest offset: Search stage: For each pixel point in Figure B, find a more matching offset within the concentric circle centered on the current matching block to replace the current offset. The initial radius of the search is set to the size of the picture, and then the radius is gradually reduced at a rate of half until the search ends; This process will repeat the two steps of propagation and search until the most suitable and accurate offset is found for each pixel block; Finally, assign the pixel values corresponding to these offsets to the corresponding pixel blocks in Figure A.
[0030] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of a robot algorithm for a dynamic occlusion scenario as described above are implemented.
[0031] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and capable of running on the processor. When the processor executes the computer program, the steps of a robot algorithm for a dynamic occlusion scenario as described above are implemented.
[0032] The present invention has the following advantages:
[0033] 1. The present invention uses the Oneformer network for improvement and adds multi-directional feature enhancement convolution. Its design aims to solve the situation where dynamic objects may present various pose states during movement or irregular parts are occluded. Through multi-directional convolution operations, the network can capture the extended features of objects in different directions. Combined with dilated convolution, the receptive field of the convolution kernel is significantly expanded and the range of the receptive field can be flexibly controlled, thereby enhancing the understanding of the omnidirectional features of dynamic objects. This design helps the network to still infer the information of the occluded part through features in other directions when the object is occluded, improving the accuracy of segmentation.
[0034] The design of the dual pooling feature enhancement module aims to improve the network's ability to capture weak semantic features. By combining different pooling strategies such as max pooling and average pooling, the network can obtain richer feature expressions, thereby enhancing the understanding of the occlusion scenario. This design helps the network to still accurately identify the boundaries of objects when the boundary semantic information is weak, improving the segmentation accuracy.
[0035] The design of the multi-attention module is to better focus on the occluded areas. By introducing the attention mechanism, the network can adaptively adjust the attention to different areas, enhance the pixel weights of the dynamic object areas, and thus more accurately identify the occluded objects. This design helps the network to still capture the key features of the occluded objects in the case of occlusion, improving the robustness of segmentation.
[0036] 2. In step S2 of the present invention, not only the reprojection error of all feature points is considered, but also a regularization term is introduced to constrain the change of the camera pose, thereby avoiding overfitting. At the same time, weighted residuals are used to reflect the importance of different feature points. In addition, the maximum a posteriori probability (MAP) estimation under the Bayesian framework is adopted to optimize the change parameters of the camera pose. This algorithm also characterizes the motion state of feature points from three aspects: average view deviation, velocity vector, and acceleration vector. By designing a feature point decision rule, the motion consistency of feature points among multiple consecutive frames is determined, the accurate recognition of the motion state of feature points is realized, and then the semantic information is used to eliminate dynamic objects.
[0037] 3. In order to ensure an adequate number of feature points are extracted, improve the feature information, and enhance the accuracy of pose estimation, the present invention uses a background repair algorithm improved based on approximate nearest neighbor matching. This algorithm repairs the static background of the current frame RGB image after removing the dynamic area through the static information in adjacent frames. The background repair algorithm based on optimal nearest neighbor pixel matching is specifically divided into two steps: similarity search and pixel matching. The similarity search is further divided into two sub-steps: information estimation and reference frame selection. Through these steps, the algorithm can select a reference frame with rich enough image information in the set repair window based on the richness of the image information, and can effectively restore the static background after removing the dynamic objects in the target frame based on these reference frames, improving the quality and quantity of feature points, thereby enhancing the performance of the SLAM algorithm in a dynamic environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is the basic flowchart of a robot algorithm for a dynamic occlusion scenario according to the present invention.
[0039] Figure 2 is the flowchart of the SLAM system applying the present invention.
[0040] Figure 3 is the structural schematic diagram of the improved Oneformer segmentation network in the present invention.
[0041] Figure 4 is the flowchart of the multi-directional feature enhancement convolution module in the present invention.
[0042] Figure 5 is the flowchart of the double pooling feature enhancement module in the present invention.
[0043] Figure 6 is the flowchart of the improved multi-attention mechanism module in the present invention.
[0044] Figure 7 is the schematic diagram of the principle of the background repair algorithm in the present invention.
[0045] Figure 8This is a comparison chart of the object segmentation effect between the present invention and existing technologies such as YOLOv8 and Oneformer.
[0046] Figure 9 This is a comparison chart of the feature point judgment effect between the present invention and existing algorithms such as ORB-SLAM3 and DynaSLAM.
[0047] Figure 10 This is the background restoration effect diagram obtained by using the present invention.
[0048] Figure 11 This is a comparison chart of the feature point extraction effect between the present invention and existing algorithms such as ORB-SLAM3 and DynaSLAM.
[0049] Figure 12 This is a comparison chart of the trajectories obtained by the present invention and algorithms such as ORB-SLAM3, DS-SLAM, and DynaSLAM in the SLAM experiment.
[0050] Figure 13 To verify the feasibility of the present invention, a real scene diagram with dynamic object occlusion is set up.
[0051] Figure 14 This is a comparison chart of the trajectories obtained by the present invention and the ORB-SLAM3 algorithm in the real scene experiment. Detailed implementation manners
[0052] The following is a more detailed description of the specific implementation manners of the present invention by referring to the accompanying drawings and describing the embodiments, so as to help those skilled in the art have a more complete, accurate and in-depth understanding of the inventive concept and technical solution of the present invention.
[0053] Embodiment 1.
[0054] As Figures 1-7 shown, the present invention provides a robot algorithm for dynamic occlusion scenarios, including the following steps.
[0055] Step S1: Collect image information and use an improved segmentation algorithm to perform complete segmentation on the occluded object.
[0056] In this step, first, the RGB-D camera carried by the robot is used to collect image information, and then the obtained image information is input into the segmentation network. Multiple modules are incorporated into the backbone network to improve the accuracy of segmentation, including: 1) Incorporating multi-directional feature enhancement convolution into the backbone network, which consists of four bar-shaped convolutions in different directions combined with dilated convolutions, significantly expanding the receptive field of the convolution kernel and being able to flexibly control the range of the receptive field; 2) Adding a dual pooling feature enhancement module in the Neck part of the network to perform feature enhancement processing on the input feature layers, aiming to strengthen the weak semantic information in these feature maps; 3) Feeding the high-dimensional feature maps extracted by the backbone network into the multi-attention mechanism module, which can enhance the pixel weights in the dynamic object areas, thereby improving the recognition accuracy of dynamically occluded objects.
[0057] Specifically, on the basis of constructing a Feature Pyramid Network (FPN) in the backbone of this segmentation network, a residual structure is introduced to enhance the performance of the backbone network, forming a Residual-Feature Pyramid Network, which is used as the backbone network; multiple modules are incorporated into the backbone network to improve the accuracy of segmentation. A multi-directional feature enhancement convolution module is set between the C1 block and the C2 block, a multi-attention module is set on the lateral connection from the C4 block to the P4 block, and a dual pooling feature enhancement module is set on the lateral connection from the C3 block to the P3 block.
[0058] Among them, the method of multi-directional feature enhancement convolution includes: First, a 3×3 depth convolution is used to preliminarily process the input image. This step can accurately capture the local features of the input image and provide the necessary local information basis for the subsequent multi-directional bar-shaped convolutions. Second, bar-shaped convolutions in four directions (including horizontal, vertical, anti-diagonal, and main diagonal) are applied in parallel to the results obtained from the preliminary processing, and each direction is combined with a corresponding dilated convolution. This parallel processing method enables the network to obtain rich feature maps from multiple scales and different levels simultaneously, thereby more comprehensively understanding the feature structure of the input data. Finally, the features from the bar-shaped convolutions and dilated convolutions in these four different directions are fused. This fusion process not only retains the key feature information in each direction but also generates a more rich and comprehensive multi-scale feature representation. This fusion strategy enables the network to take into account both local details and global features while paying attention to local details, thereby improving the overall feature extraction ability.
[0059] The calculation method of multi-directional feature enhancement convolution is shown as follows:
[0060]
[0061] Where M in represents the input of the MDFEC (multi-directional feature enhancement convolution) module, represents the depth convolution, 3×3 represents the size of the convolution kernel, Directioni represents the convolution operations of multi-directional convolution in different directions, M dfec represents the result of multi-directional feature enhancement convolution. After multi-directional feature enhancement convolution, the obtained result is first subjected to batch normalization processing, then processed by a multi-layer perceptron classifier, and then element-wise added and fused with the input image to finally generate the output of the MDFEC module.
[0062] In the dual pooling feature enhancement module, when the input feature map D is input, first, pooling kernels of multiple sizes are used to capture the multi-scale features of the image. The pooling operations of each size of pooling kernel have two parallel branches: one branch performs average pooling to fuse the information in the pooling area; the other branch performs max pooling to accurately capture the significant features within the pooling area. Next, the feature maps generated by these two branches are merged in the channel dimension and passed through a batch normalization (BN) layer and a ReLU activation function in sequence to obtain the merged feature map, which can enhance the stability and non-linear expression ability of the features. Subsequently, a 1×1 convolution, a BN layer, and a ReLU activation function are applied to the merged feature map to achieve further feature extraction and dimensionality reduction processing. Then, the bilinear interpolation method is used to upsample these feature maps to the same size as the input feature map to form the pooling results of the corresponding scale features. The input feature map is then processed through a 1×1 convolution, a BN layer, and a ReLU activation function, and its number of channels is compressed to one-fourth of the original input number of channels. Then, this processed feature map is merged with the feature map obtained by multi-scale pooling in the channel dimension. Finally, a 1×1 convolution, a BN layer, and a ReLU activation function are applied again to further fuse these multi-scale features, thereby generating the output feature map of the DPEM (dual pooling feature enhancement module) module. Three sizes of pooling operations are used for the pooling kernels: global pooling, pooling with a kernel size of 4×4 and a stride of 4, and pooling with a kernel size of 2×2 and a stride of 2. These different sizes of pooling operations can capture feature information of different scales.
[0063] The spatial attention mechanism aims to increase the weights of the pixel values in the occluded regions of the feature map. When the high-dimensional feature map is input into the spatial attention mechanism, average pooling and max pooling are respectively performed, and then the two sets of pooled features are concat-fused to form a new feature map f c . Then, this feature map passes through a 3×3×1 convolutional layer and a sigmoid function to obtain the spatial attention map f for weight allocation u . The formula of the spatial attention mechanism is as follows:
[0064] f u = sigmoid(c(f c)) where c is a 3×3×1 convolutional network and sigmoid() represents the sigmoid function. The spatial attention map f u can guide the network to pay more attention to the regions where dynamic objects are located. Connect f u to the original feature map Q to obtain the feature map Q weighted by spatial attention u .
[0065] The channel attention mechanism functions to assign corresponding weights to each layer of channels in the feature map. In this algorithm, the H×W×C feature map F is input into the channel attention mechanism, and global attention average pooling and max pooling operations are performed on the feature map to obtain the information of each channel of the feature map. The features F avg and F max obtained through average pooling and max pooling are passed through the fully connected layer FC module to strengthen the correlation between channels and re - assign the weights of each channel to obtain the channel attention map f v for better learning of occluded features. The formula of the channel attention mechanism is as follows:
[0066] f v = sigmoid(Pη(F avg +F max ))
[0067] where η represents the ReLU function and P are the parameters of the fully connected layer. The channel attention map f v is connected to the original feature map Q to obtain the feature map Q weighted by channel attention v .
[0068] Connect the feature map Q v weighted by channel attention and the feature map Q u weighted by spatial attention to obtain the output of the multi - attention mechanism module.
[0069] Step S2: Analyze the motion state of the object by the method of object motion estimation, and eliminate the accurately identified dynamic regions.
[0070] In this step, the camera pose estimation is used to initially determine the camera position, and then the motion state of the object is analyzed. During the analysis, dynamic judgment is performed on the feature points in space. For a feature point p in space, the current frame Y, and the key frame Y o , calculate the reprojection error of the feature point p between the key frame Y o and the current frame Y, considering the uncertainty of the camera pose parameters R, t, and the camera intrinsic matrix K, and also considering the noise n R of the camera pose estimation, n t . Among them, the key frame Y oThe o-th frame after that is the current frame Y, and o takes 10 in this embodiment. Among the key frame Y o and the reprojection error e p between the current frame Y is expressed as the following formula: e p = τ - π(I o , K, L(R + n R , t + n t ))
[0071] where τ is the 2D pixel coordinate of the feature point p projected onto the key frame Y o , π is the projection function, I o is the 3D coordinate of the feature point p on the key frame Y o , L(R, t) is the relative transformation matrix from the world coordinate system to the camera coordinate system, R is the rotation matrix, t is the translation matrix, n R and n t are the rotation matrix noise and the translation matrix noise respectively, and K is the camera internal parameter matrix
[0072] To optimize the reprojection error, a cost function J is constructed, which not only considers the reprojection error e p of all feature points, but also introduces a regularization term R(L) to constrain the change of the camera pose L, thus avoiding overfitting. At the same time, a weighted residual w i is used to reflect the importance of different feature points, and an error covariance matrix Σp is introduced to represent the uncertainty of the feature point position. In addition, based on the prior information P0 about the distribution of the initial pose, the cost function J adopts the maximum a posteriori probability (MAP) estimation under the Bayesian framework to optimize the camera pose parameters R, t. Thus, the corresponding cost function J is expressed as:
[0073]
[0074] where A represents the set of feature points, λ is the regularization parameter used to balance the influence of the data term and the regularization term represents the transpose of the reprojection error, and Σp -1 is the inverse matrix of the error covariance matrix Σp
[0075] To more accurately determine whether a feature point belongs to a dynamic object, in addition to the perspective deviation , the continuity in the time dimension can also be considered, that is, the motion consistency of the feature point among multiple consecutive frames. For each feature point p, this step calculates its average perspective deviation Average perspective deviation The calculation formula of is:
[0076]
[0077] where represents the average perspective deviation of feature point p among multiple frames. ω is a function for measuring the included angle, m j and are the positions of feature point p and its projection point in the j-th frame respectively, and the subscript j + o represents the (j + o)-th frame, m j+o and represent the positions of feature point p and its projection point in the (j + o)-th frame respectively.
[0078] This step compares the average perspective deviation with the corresponding threshold θ; meanwhile, the velocity vector v p and acceleration vector a p are introduced to characterize the motion state of the feature point, obtaining a decision rule for the feature point. The calculation formula of the decision rule is:
[0079]
[0080] where α, β, γ are weight coefficients, corresponding to the average perspective deviation velocity vector v p and acceleration vector a p respectively, s is an indicator function, v th and a th are the thresholds of the velocity vector v p and acceleration vector a p respectively, f th is the threshold for the final decision. If f p > f th , it is determined as a dynamic point; otherwise, it is a static point.
[0081] After completing the dynamic judgment of the feature points, the accurately identified dynamic regions are removed. Thus, after this algorithm completely segments the occluded object and accurately determines that the object is a dynamic object, it realizes the accurate removal of the dynamic object using its semantic information.
[0082] Step S3: Use the static information in several adjacent reference frames to complete the repair of the dynamic region, extract feature points from the repaired dataset, and calculate the pose.
[0083] In this step, an optimal nearest-neighbor pixel matching background repair algorithm based on approximate nearest-neighbor matching is used. This algorithm repairs the static background in the RGB image of the current frame after removing the dynamic region through the static information in several adjacent frames. The optimal nearest-neighbor pixel matching background repair algorithm is specifically divided into two steps: similarity search and pixel matching. The similarity search is further divided into two sub-steps: information estimation and reference frame selection.
[0084] Step 1: Information estimation:
[0085] Since the richness of image information has a decisive impact on the effectiveness of image restoration, it is difficult to effectively extract image information by simply mapping the previous frame to the target frame. This algorithm uses inter-frame feature matching technology to evaluate and obtain image information. Specifically, taking the target frame as a reference, a time window containing several frames (15 frames in this embodiment) of images is moved in the arrow direction (including both forward and reverse directions) as the restoration window. In this restoration window, ORB feature extraction and counting operations are performed on each frame of the image, where the feature point count is ∑W i .
[0086] Pose is one of the important expression indexes of the connection between adjacent frames. The epipolar geometry is used to estimate the connection between adjacent frames. Let P = [x, y, z] T be a point in space, and its projections in the reference frame and the target frame are p1 and p2 respectively. The pixel positions of the pixel points p1 and p2 can be obtained from the pinhole camera model; from this, the three-dimensional vector of the movement of point P is calculated, and then the three-dimensional vector is projected onto a two-dimensional plane, and then the displacement change amount Δp and the rotation change amount Δθ between any two frames are obtained.
[0087] Step 2: Selection of the reference frame:
[0088] This step is to optimize the selection process of the reference frame and ensure the effectiveness of approximate nearest neighbor matching. Therefore, the image information and pose information calculated in Step 1 are combined. In this step, a time window of 15 frames in both forward and reverse directions is taken as a group (i.e., the restoration window), features are extracted from this window and matched to screen out candidate reference frames that meet specific conditions. The specific conditions include that Δp exceeds the threshold τ and Δθ is greater than the threshold γ. If satisfied, this frame is used as a candidate reference frame. Next, these candidate reference frames are further screened. Each candidate reference frame is feature-point matched with the target frame, and the number of successfully matched feature points is recorded as ∑Z i , then |∑W i -∑Z i | can reflect the richness of the image information of the candidate reference frame. If ΣW i and ΣZ i meet the judgment conditions, it is determined that this candidate reference frame contains efficient and rich image information, and the satisfied candidate reference frames are added to the reference frame library. The judgment conditions are expressed as:
[0089] |ΣW i -ΣZ i | < E1, where E1 represents the image information judgment threshold, which is set to 100 in this embodiment.
[0090] The reference frame library is obtained through the above similar search process, and pixel matching is used to repair the background by improving the traditional approximate nearest neighbor matching algorithm. The background repair algorithm mainly includes three stages: initialization, propagation, and search. In the initialization stage, the target frame to be repaired is set as Figure A, and the reference frame is set as Figure B. Subsequently, a 3×3 pixel block is randomly selected in Figure A as the matching block, and a random offset is assigned to it. A matching block corresponding to the matching block in Figure A is found in Figure B. In the propagation stage, the offset difference between the matching block in Figure A and the matching block in Figure B is calculated, and the value with the smallest offset is found. In the search stage, for each pixel point in Figure B, a more matching offset is searched for within the concentric circle centered on the current matching block to replace the current offset. The initial radius of the search is set to the size of the image, and then the radius is gradually reduced at a rate of half until the search ends. This process will repeat the two steps of propagation and search until the most suitable and accurate offset is found for each pixel block. Finally, the pixel values corresponding to these offsets are assigned to the corresponding pixel blocks in Figure A. It is known from experiments that the reference frame library usually consists of 2-4 frames of images, and the general number of iterations is 3 times, which can achieve a good repair effect.
[0091] Through these steps, this algorithm can effectively restore the static background after removing the dynamic area, improve the quality and quantity of feature points, and thus improve the performance of the SLAM algorithm in a dynamic environment. After completing the background repair, feature points are extracted from the repaired dataset, and the pose is calculated.
[0092] The following will illustrate the process of the above robot algorithm for dynamic occlusion scenarios in combination with specific experiments.
[0093] As Figure 8 shown, in order to verify the segmentation effect of the improved segmentation network algorithm, a comprehensive verification of occlusions in four different situations was carried out on the dynamic subset of the TUM dataset, where the red boxes in the figure are used as auxiliary marks. The segmentation ability of high-occlusion objects can be verified from the first and second rows of the figure, and low-occlusion objects can be verified from the third and fourth rows. Through segmentation comparison, it is found that since YOLOv8 does not have a dedicated object feature enhancement layer, it is unable to complete the complete segmentation of the occluded object boundary, resulting in interruption or low segmentation accuracy during the segmentation process. The unimproved Oneformer has a better segmentation effect than YOLOv8, but the segmentation accuracy is still insufficient. Therefore, this algorithm designs a multi-directional feature enhancement convolution and feature enhancement module to enhance the underlying main body information. At the same time, a multi-attention module is incorporated to focus on the occluded area, while weakening the influence of other unnecessary information and enhancing the segmentation accuracy. Generally speaking, in a dynamic occlusion scenario, this algorithm has a good effect on segmenting potential dynamic objects with few pixel weights and weak boundary semantic information.
[0094] like Figure 9 As shown in Figure 2, four sub-datasets of dynamic sequences in the TUM dataset are selected to judge and evaluate the feature points of objects, such as Figure 8 As shown. The first and second rows are feature point judgments of the ORB-SLAM3 algorithm and the Dyna-SLAM algorithm respectively, and the third row is the feature point judgment of this algorithm. It can be seen that the ORB-SLAM3 algorithm judges the feature points on both dynamic objects and static objects as static feature points (people are dynamic targets, green is the mark of static feature points, and red is the mark of dynamic feature points). The feature point judgment effect of Dyna-SLAM is obviously better than the former, but it will make misjudgments when the object moves at high speed. Compared with the first two algorithms, this algorithm has a more accurate ability to judge feature points.
[0095] like Figure 10 As shown, Figure 9 (a) to (d) are the original RGB images containing dynamic characters. Figure 9 (e)~(h) are the restored RGB images. It can be seen that the restored images only contain the original static background in the scene.
[0096] like Figure 11 As shown in the figure, ORB-SLAM3 and DynaSLAM algorithms are selected to compare with this algorithm for feature point extraction. ORB-SLAM3 cannot distinguish between dynamic and static objects, so feature points on dynamic objects cannot be removed. Dyna-SLAM uses instance segmentation technology to mark the area where dynamic objects are located and remove them from the environment to reduce the impact of dynamic objects on the global structure. Figure 1 However, after removing dynamic objects, this algorithm reduces the number of static feature points used for pose estimation and map construction. This algorithm removes dynamic areas and performs background repair on the removed areas to extract more abundant static feature points, thereby being able to construct a more accurate trajectory map.
[0097] like Figure 12 As shown in the figure, ORB-SLAM3 cannot distinguish between dynamic and static objects, so its robustness is very poor. Dyna-SLAM uses instance segmentation technology to mark the area where dynamic objects are located and remove them from the environment to reduce the impact of dynamic objects on the global structure. Figure 1However, in a high-occlusion environment or when the running speed is too fast, it is difficult to accurately identify potential dynamic objects and their motion state. Moreover, after removing dynamic objects, this algorithm does not perform background repair, which reduces the number of static feature points used for pose estimation and map construction, thus affecting the positioning accuracy of the algorithm. This algorithm uses an improved Oneformer segmentation network, a dynamic judgment method that combines camera pose estimation with object motion state judgment, and background repair processing to better solve the impact of dynamic feature points, repair dynamic areas, and increase the number of feature points used for pose estimation.
[0098] like Figure 13 As shown in the figure, to verify the feasibility of this method, a real scene with dynamic object occlusion is set for SLAM experiment. The Husky wheeled robot is used. The platform configuration is i7-10875H CPU, 8GB memory, GTX1080 GPU, Ubuntu18.04 operating system, and the data set is collected in the real scene. The experimental results are shown in the figure. Figure 14 As shown in the figure, it can be seen that the ORB-SLAM3 algorithm has poor trajectory accuracy due to its inability to determine feature points. This method combines an improved segmentation network, accurately determines feature points, and performs dynamic background repair, which improves its trajectory accuracy.
[0099] Embodiment 2.
[0100] Corresponding to the first embodiment of the present invention, the second embodiment of the present invention provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the following steps are implemented according to the method of the first embodiment.
[0101] Step S1: collect image information and use an improved segmentation algorithm to completely segment the blocked object.
[0102] Step S2, analyzing the motion state of the object by using the object motion estimation method, and removing the accurately identified dynamic area.
[0103] Step S3, using the static information in several adjacent reference frames to complete the repair of the dynamic area, extracting feature points from the repaired data set, and calculating the pose.
[0104] The above storage media include: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), optical disk and other media that can store program codes.
[0105] For the specific limitations of the steps implemented after the program in the computer-readable storage medium, reference can be made to Embodiment 1, and details will not be described herein again.
[0106] Embodiment 3.
[0107] Corresponding to Embodiment 1 of the present invention, Embodiment 3 of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the following steps are implemented according to the method of Embodiment 1.
[0108] Step S1: Collect image information, and use an improved segmentation algorithm to perform complete segmentation on the occluded object.
[0109] Step S2: Analyze the motion state of the object by means of object motion estimation, and eliminate the accurately identified dynamic region.
[0110] Step S3: Use the static information in a number of adjacent reference frames to complete the repair of the dynamic region, extract feature points from the repaired data set, and calculate the pose.
[0111] For the specific limitations of the steps implemented by the computer device, reference can be made to Embodiment 1, and details will not be described herein again.
[0112] It should be noted that each block in the block diagram and / or flowchart in the accompanying drawings of the present invention, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that executes the specified functions or actions, or can be implemented by a combination of dedicated hardware and machine instructions obtained.
[0113] The present invention has been described exemplarily above with reference to the accompanying drawings. Obviously, the specific implementation of the present invention is not limited by the above-mentioned manner. As long as various non-substantive improvements are made by adopting the inventive concept and technical solution of the present invention, or the inventive concept and technical solution of the present invention are directly applied to other occasions without improvement, they are all within the protection scope of the present invention.
Claims
1. A robot algorithm for dynamic occlusion scenarios, characterized by: The following steps are involved: Step S1, collecting image information, and using an improved segmentation algorithm to completely segment the obscured object; Step S2, analyzing the motion state of the object by an object motion estimation method, and removing the accurately identified dynamic area; Step S3, using static information in several adjacent reference frames to complete the repair of the dynamic area, extracting feature points from the repaired data set, and calculating the pose; Among them, the step S1 includes: 1) incorporating multi-directional feature enhancement convolution into the backbone network, which is composed of strip convolutions in four different directions of horizontal, vertical, anti-diagonal and main diagonal, and hollow convolution; 2) adding a double pooling feature enhancement module to the Neck part of the network to perform feature enhancement processing on the input feature layer; 3) sending the high-dimensional feature map extracted by the backbone network to the multi-attention mechanism module, which can enhance the pixel weight of the dynamic object area; The improved segmentation algorithm introduces a residual structure to form a residual-feature pyramid network based on the backbone feature pyramid network. A multi-directional feature enhancement convolution module is set between the C1 block and the C2 block, a multi-attention module is set on the lateral connection between the C4 block and the P4 block, and a double-pooling feature enhancement module is set on the lateral connection between the C3 block and the P3 block. The calculation method of multi-directional feature enhancement convolution is as follows: Among which M in represents the input, represents depth convolution, 3×3 represents the size of the convolution kernel, Direction i represents the convolution operations of multi-directional convolution in different directions, M dfec represents the result of multi-directional feature enhancement convolution; the result M dfec first undergoes batch normalization processing, then passes through a multi-layer perceptron classifier, and then is element-wise added to the input image for fusion output; In the dual pooling feature enhancement module, when the feature map D is input, the module first uses pooling kernels of various sizes to capture the multi-scale features of the image; for each size of pooling operation, two parallel branches are designed: one branch performs average pooling to fuse the information of the pooling area; the other branch performs maximum pooling; next, the feature maps generated by the two branches are merged in the channel dimension and passed through the batch normalization BN layer and the ReLU activation function in turn; then, a 1×1 convolution, BN layer and ReLU activation function are applied to further extract features from the merged feature map and dimension reduction processing; then, the bilinear interpolation method is used to upsample these feature maps to the same size as the module input feature map; the module input feature map is then processed by a 1×1 convolution, a BN layer, and a ReLU activation function, and its number of channels is compressed to one-fourth of the original input channel number; then, this processed feature map is merged with the feature map obtained by multi-scale pooling in the channel dimension; finally, a 1×1 convolution, a BN layer, and a ReLU activation function are applied again to further fuse these multi-scale features, thereby generating the output feature map of the dual pooling feature enhancement module.
2. The robot algorithm for a dynamic occlusion scenario according to claim 1, wherein: In the step S2, for a feature point p in space, the current frame Y, and the key frame Y o , where the key frame Y o , the o-th frame after it is the current frame Y. The reprojection error e o between the key frame Y p and the current frame Y is expressed as the following formula: e p = τ - π(I o , K, L(R + n R , t + n t )) where τ is the 2D pixel coordinates of the feature point p projected onto the key frame Y o ; π is the projection function; I o is the 3D coordinates of the feature point p on the key frame Y o ; L(R, t) is the relative transformation matrix from the world coordinate system to the camera coordinate system; R is the rotation matrix; t is the translation matrix; n R and n t are the rotation matrix noise and the translation matrix noise respectively; K is the camera intrinsic parameter matrix Construct a cost function \(J\) that takes into account the reprojection error \(e\) of all feature points p , and use the weighted residual \(w\) i to reflect the importance of different feature points, introduce an error covariance matrix \(\sum_p\) to represent the uncertainty of the feature point positions; based on the prior information \(P_0\) about the distribution of the initial pose, the cost function \(J\) uses the maximum a posteriori probability estimation under the Bayesian framework to optimize the camera pose parameters \(R\), \(t\); the corresponding cost function \(J\) is expressed as: Among them, A represents the set of feature points, λ is the regularization parameter, and the regularization term R(L) is introduced to constrain the change of camera posture L.
3. The robot algorithm for a dynamic occlusion scenario according to claim 2, wherein: In the step S2, for each feature point p, the average view angle deviation of the feature point p among a series of consecutive frames is calculated in this step Average view angle deviation The calculation formula is as follows: Among them represents the average value of the perspective deviation of the feature point p among multiple frames. ω is a function for measuring the included angle, and m j and are respectively the positions of the feature point p and its projection point in the j-th frame, and the subscript j + o represents the (j + o)-th frame, and m j+o and respectively represent the positions of the feature point p and its projection point in the (j + o)-th frame; This step compares the average perspective deviation with the corresponding threshold θ; meanwhile, the velocity vector v p and the acceleration vector a p are introduced to characterize the motion state of the feature point, and a feature point decision rule is obtained; the calculation formula of the decision rule is as follows: Among them, α, β, γ are weight coefficients, corresponding to the average view deviation respectively velocity vector v p and acceleration vector a p , s is an indicator function, v th and a th are the thresholds of the velocity vector v p and acceleration vector a p respectively, f th is the threshold for the final decision. If f p > f th , it is determined as a dynamic point; otherwise, it is a static point.
4. The robot algorithm for a dynamic occlusion scenario according to claim 1, characterized in that: In step S3, a background restoration algorithm is used, including similarity search and pixel matching; Similar searches include: Step 1, Information Estimation: Taking the target frame as a reference, move a time window containing several frame images in both the forward and reverse directions as a repair window. Within this repair window, let \(P = [x, y, z]\) T be a point in space, and its projections in the reference frame and the target frame be \(p1\) and \(p2\) respectively. Obtain the pixel positions of the pixel points \(p1\) and \(p2\) from the pinhole camera model; thereby calculate the three-dimensional vector of the movement of point \(P\), then project the three-dimensional vector onto a two-dimensional plane, and then obtain the displacement change \(\Delta p\) and the rotation change \(\Delta\theta\) between any two frames; perform ORB feature extraction and counting operations on each frame image, where the number of feature points is \(\sum W\) i ; Step 2, Reference Frame Selection: Extract features from within the repair window and perform matching. When the displacement change amount Δp exceeds the threshold τ and the rotation change amount Δθ is greater than the threshold γ between an image frame and the target frame, this frame is used as a candidate reference frame to screen out candidate reference frames that meet specific conditions; Next, further screen these candidate reference frames, perform feature point matching between each candidate reference frame and the target frame, and record the number of successfully matched feature points as ∑Z i ; If ∑W i and ∑Z i meet the judgment conditions, then add the candidate reference frames that meet the conditions to the reference frame library; The judgment conditions are expressed as: |∑W i -∑Z i |<E1, Wherein, E1 represents the image information judgment threshold.
5. The robot algorithm for a dynamic occlusion scenario according to claim 4, characterized in that: The background repair algorithm adopted for pixel matching includes: In the initialization stage, set the target frame image to be repaired as Image A and the reference frame as Image B; Subsequently, randomly select a 3×3 pixel block in Image A as the matching block, and randomly assign an offset to it, then find a matching block corresponding to the matching block in Image A in Image B; Propagation stage: Calculate the offset difference between the matching block in Image A and the matching block in Image B, and find the value with the smallest offset; Search stage: For each pixel point in Image B, find a more suitable offset within the concentric circle centered on the current matching block to replace the current offset. The initial radius of the search is set to the size of the image, and then the radius is gradually reduced at a rate of half until the search ends; This process will repeat the two steps of propagation and search until the most suitable and accurate offset is found for each pixel block; Finally, assign the pixel values corresponding to these offsets to the corresponding pixel blocks in Image A.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of a robot algorithm for a dynamic occlusion scenario as described in any one of claims 1-5.
7. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the steps of a robot algorithm for a dynamic occlusion scenario as described in any one of claims 1-5.