Dynamic visual identification method based on lightweight instance segmentation network

By introducing a lightweight instance segmentation network into the visual SLAM system, separating dynamic and static areas, eliminating dynamic feature points and performing background repair, the problem of poor positioning accuracy and mapping effect of visual SLAM in dynamic scenes is solved, and more efficient feature point classification and mapping effect are achieved.

CN119942115APending Publication Date: 2025-05-06HECHI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510020786.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing visual SLAM technology is difficult to effectively process dynamic objects in dynamic scenarios, resulting in poor positioning accuracy and mapping effects.

Method used

The dynamic visual recognition method based on lightweight instance segmentation network is adopted, and the scene data is separated by the instance segmentation model, dynamic object images are extracted and dynamic feature points are eliminated. Only static feature points are used for pose estimation and environmental map construction, and background repair is carried out to improve the mapping effect.

Benefits of technology

It improves the accuracy and processing speed of the instance segmentation network, enhances the accuracy of feature point classification and graph construction effect, and ensures the real-time and positioning accuracy of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942115A_ABST
    Figure CN119942115A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of visual identification, in particular to a dynamic visual identification method based on a lightweight instance segmentation network, which comprises the following steps: S1, acquiring scene data according to sensor data and camera data of a mobile robot; s2, adding an instance segmentation model in a tracking thread of the ORB-SLAM3 to obtain a dynamic mask; s3, extracting the data in the step S1 in a tracking thread to obtain rough classification of feature points; s4, combining the dynamic mask in the step S2 with the feature point coarse classification in the step S3 to obtain final feature classification so as to construct an environment map; s5, carrying out background restoration on the blocked background filled in the key frame with the dynamic object removed in the final feature classification in the step S4; s6, acquiring data of a tracking thread through a local map construction thread of ORB-SLAM3; and S7, fusing the data of the local map building thread through a loop-back thread of the ORB-SLAM3. According to the method, the precision and processing speed of the instance segmentation network can be improved, the feature point classification accuracy is improved, and the mapping effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of visual recognition technology, and in particular to a dynamic visual recognition method based on a lightweight instance segmentation network. Background Art

[0002] SLAM refers to the cognitive process of a mobile carrier in an unknown environment. The process can be described as a mobile carrier moving from a certain position in the scene in the absence of any prior knowledge. During the movement, the carried sensor is used to obtain environmental information, and the motion state is estimated based on the environmental information data, so as to obtain the position and posture of the mobile carrier, realize accurate positioning function and build an environmental map at the same time. SLAM technology is mainly divided into two categories, namely visual SLAM and laser SLAM. Laser SLAM obtains data through laser radar, which is relatively convenient and has high positioning accuracy when calculating the position and posture. However, due to the expensiveness of laser radar, it is subject to certain limitations in practical applications. Visual SLAM uses visual sensors as the main data source, and uses the characteristics of rich feature differentiation of image color and texture information to estimate the position and posture. The camera used in visual SLAM has more cost advantages than laser radar, and can obtain richer environmental information, and has a wider range of applications.

[0003] For dynamic objects in the environment, a direct idea for visual SLAM is to extract dynamic components from the input data and clearly identify them as outliers, so as to exclude them in the pose estimation and mapping process. After years of development, traditional model-based visual SLAM has the advantages of a complete mathematical model, no dependence on data sets, and low computational complexity, but it also has the above-mentioned problems. In recent years, with the rapid development of deep learning technology, more and more researchers have begun to integrate deep learning technologies such as target detection, semantic segmentation, and instance segmentation into dynamic visual SLAM to solve problems such as dynamic object interference in dynamic scenes in visual SLAM. Deep learning-based methods use semantic labeling or target detection to preprocess potential dynamic objects, which can effectively remove dynamic feature points and improve positioning accuracy.

[0004] Instance segmentation is an advanced segmentation task that combines object detection with semantic segmentation. There are two methods, two-stage and one-stage: In the two-stage method, the first route is a top-down method based on object detection, which locates the box where each instance is located by using object detection, and then performs semantic segmentation within the box to determine the mask of each instance; the second is a bottom-up method based on semantic segmentation, which first identifies pixels through semantic segmentation, and then uses clustering or metric learning methods to distinguish different instances of the same kind. The two-stage method maintains better low-level features, but the subsequent steps (such as clustering or metric calculation) are more cumbersome, and the generalization ability of the model may be poor.

[0005] In the one-stage method, there are two types of methods: global image-based and local image-based. The global image-based method does not need to be cropped and aligned. It first forms a feature map of the entire instance, and then combines the features to obtain the final mask of each instance. The method based on local information directly outputs the segmentation result. The method based on contour prediction mask usually uses 20 to 40 coefficients to parameterize the mask contour. These methods have fast inference speed and are easy to optimize, but they cannot accurately depict the mask or objects with holes in the center. Summary of the invention

[0006] In order to solve the above problems, the present invention provides a dynamic visual recognition method based on a lightweight instance segmentation network, which can improve the accuracy and processing speed of the instance segmentation network, improve the accuracy of feature point classification and improve the mapping effect.

[0007] In order to achieve the above object, the technical solution adopted by the present invention is:

[0008] A dynamic visual recognition method based on a lightweight instance segmentation network comprises the following steps:

[0009] S1. Obtain scene data based on sensor data and camera data of the mobile robot;

[0010] S2. Add an instance segmentation model in the tracking thread of ORB-SLAM3, separate the dynamic area and the static area of ​​the scene data of step S1 through the instance segmentation model to extract the dynamic object image, process the dynamic object image and obtain a priori dynamic object mask, and remove the dynamic area feature points of the priori dynamic object mask to obtain a dynamic mask;

[0011] S3. The data of step S1 is extracted in the tracking thread to obtain the ORB feature of each frame image, and the dynamic points of the ORB feature are extracted by multi-view geometry to obtain a rough classification of feature points;

[0012] S4. Combining the dynamic mask of step S2 with the rough classification of feature points of step S3 to obtain a final classification of features, removing feature points regarded as dynamic points, and using feature points regarded as static points for camera pose estimation to construct an environment map;

[0013] S5. Fill the occluded background with the key frame of the dynamic object removed in the final classification of the features in step S4 to perform background repair;

[0014] S6. Obtain the data of the tracking thread through the local map building thread of ORB-SLAM3 to optimize the key frame poses and corresponding map points in the environment map, and perform local bundling adjustments to ensure consistency between the environment map and the camera trajectory;

[0015] S7. The data of the local map building thread is fused through the loop thread of ORB-SLAM3 to merge the environment map and perform global BA optimization.

[0016] Furthermore, in step S2, separating the dynamic area and the static area of ​​the scene data in step S1 by using the instance segmentation model includes the following steps:

[0017] S2.1 inserting a BoT block into the VoVNetV2 feature extraction network of the SOLOv2 model, and replacing the convolution in VoVNetV2 with a Ghost Conv module to form an improved feature extraction network, dividing the scene data into an S×S grid and inputting it into the improved feature extraction network, and obtaining a feature map by extracting image features;

[0018] S2.2 replaces the original network of the Neck part in the SOLOv2 model with a Neck part of the FPG type to form an improved Neck part, inputs the feature map into the improved Neck part to obtain feature maps of different levels of size, and fuses the feature maps of each level to obtain a high-resolution feature map with semantic information;

[0019] S2.3 Add a CBAM module to the Head part of the original network in the SOLOv2 model to obtain an improved Head part, input the high-resolution feature map into the improved Head part, predict the instance category through the CategoryBranch of the Head part, predict the instance mask through the MaskBranch of the Head part, and obtain the instance segmentation result by matching the instance category with the instance mask one by one to obtain a dynamic mask.

[0020] Further, in step S2.1, the BoT block uses the MHSA module to adjust the spatial dimension of the two-dimensional image x to a one-dimensional sequence x through the PatchEmebd module before inputting it into the improved feature extraction network. p , where x∈R (H ×W×C) , x p ∈R N×(P·C) , (H, W) is the resolution of the input image; C is the number of channels; (P, P) is the resolution of each image block; N = HW / P 2 is the number of image blocks, N is the effective input sequence length of the MHSA module;

[0021] The position information of the sequence is transferred to the feature sequence, and three identical feature matrices Q, K, and V are generated. The feature matrices are projected h times to C by linear projection. q, C k , C v Dimension, calculate parallel dot product attention so that the MHSA module can utilize the order information of the sequence. The parallel dot product attention calculation method is:

[0022]

[0023] Furthermore, in step S2.1, the Ghost Conv module generates a small channel feature map with a smaller channel by traditional convolution, the small channel feature map generates a new feature map through a cheap linear transformation operation, and then the small channel feature map is spliced ​​with the new feature map to obtain a feature map.

[0024] Furthermore, in step 2.3, the CBAM module performs global maximum pooling and global average pooling on the high-resolution feature map to obtain the channel dimension weight and spatial dimension weight of the high-resolution feature map, and inputs the weight vectors of the channel dimension weight and the spatial dimension weight into the same multi-layer perceptron, adds the mapping weights output by the MLP at the pixel level, and then outputs the channel attention feature M after activation by the sigmoid function. C (F), then:

[0025]

[0026]

[0027] Among them, σ is the sigmoid function; W0 and W1 are the weights shared by MLP; W0 is the ReLU activation function, which converts M C Perform channel multiplication with the high-resolution feature map to obtain F', and use F' as the input of the spatial attention module;

[0028] In the spatial attention module, global maximum pooling and global average pooling are performed based on the channel to obtain two feature maps with a channel number of 1. After channel splicing, a 7×7 convolution operation is performed to reduce the number of channels to 1, and then the spatial attention feature M is obtained through the Sigmoid function. S (F'), then:

[0029]

[0030] Where σ is the sigmoid function; f 7×7 Represents a 7×7 convolution; the spatial attention feature M S (F') is multiplied by F' to obtain a dynamic mask.

[0031] Further, in step S3, x is a feature point on the previous key frame, and the projection of point x to the position x' on the current key frame CF and the projection depth Z can be calculated based on the camera motion. proj , the calculation method is:

[0032]

[0033] Among them, u, v, w are the coordinates of the x image coordinate system; X, Y, Z are the coordinates of the x' image coordinate system; K is the camera intrinsic parameter matrix;

[0034] By calculating the back-projection angle α between x and x', the back-projection angle α and the depth of the feature point Z' and Z proj The difference between and determines whether x is a dynamic point.

[0035] Furthermore, when the back-projection angle α is greater than 30° and Z proj - When z' is greater than the threshold, x is a dynamic point.

[0036] Furthermore, in step S4, according to the positions of the previous frame and the current frame, the RGB and depth channels of the previous key frame are projected to the dynamic part of the current frame, and a realistic image without dynamic objects is synthesized by using the static information of the previous view to repair the occluded background.

[0037] The beneficial effects of the present invention are:

[0038] In instance segmentation, the BoT block is inserted into the VoVNetV2 feature extraction network, and Ghost Conv is used to replace the traditional convolution in VoVNetV2. The BoT module obtains the global dependency of the image by aggregating the local interaction information of the image, enabling the network to obtain the associated features of a long sequence. Ghost Conv significantly reduces the computational cost and memory usage of the network by replacing some traditional convolutions with cheap linear operations. The improved backbone network not only enhances the feature extraction capability of the model, but also reduces the computational amount and parameter amount of the model; FPG is used to replace the Neck part of the original network. FPG is a deep multi-channel feature pyramid that fuses features in multiple directions in multi-scale space to obtain high-resolution features with semantic information; the convolutional block attention module CBAM module is added to the Head part of the original network. The attention mechanism of the CBAM module first learns feature information from the two main dimensions of space and channel, and then fuses this information to obtain refined feature information.

[0039] In the tracking thread, instance segmentation model and multi-view geometry model are added in parallel. In order to solve the problem that ORB-SLAM3 does not process dynamic objects in the scene, parallel instance segmentation model and multi-view geometry model are added to segment the static area and dynamic area in the scene, obtain feature point classification, and eliminate dynamic feature points and only use static feature points to estimate the camera pose to improve the positioning accuracy and robustness of the system. At the same time, parallel design and lightweight instance segmentation network will not significantly increase the time spent on this part of the system to ensure the real-time performance of the system. The background of the key frames that eliminate dynamic objects after instance segmentation is repaired. Currently, many similar methods do not perform background repair on the scene after segmenting dynamic objects, which leads to distortion of the mapping results and cannot truly reflect the current situation of the scene. After eliminating dynamic points, a background repair thread is added to improve the mapping effect and the repositioning effect after map creation. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a flow chart of a preferred embodiment of the present invention.

[0041] Figure 2 It is a structural diagram of MHSA of a preferred embodiment of the present invention.

[0042] Figure 3 It is a schematic diagram of the Ghost Conv module structure of a preferred embodiment of the present invention.

[0043] Figure 4 It is a schematic diagram of the FPG module structure of a preferred embodiment of the present invention.

[0044] Figure 5 It is a schematic diagram of the CBAM module structure of a preferred embodiment of the present invention.

[0045] Figure 6 This is a flow chart of example segmentation and extraction of dynamic objects according to a preferred embodiment of the present invention.

[0046] Figure 7 It is a schematic diagram of determining dynamic points using a multi-view geometry method according to a preferred embodiment of the present invention.

[0047] Figure 8 This is a background restoration effect diagram of a preferred embodiment of the present invention.

[0048] Fig. 9 It is a part of the Cityscapes dataset of a preferred embodiment of the present invention.

[0049] Fig.10 It is a comparison diagram of segmentation accuracy of different categories according to a preferred embodiment of the present invention.

[0050] Fig.11It is a schematic diagram of the visualization results of the improved algorithm and the original algorithm on the Cityscapes validation set according to a preferred embodiment of the present invention.

[0051] Fig.12 It is a comparison diagram of trajectories in different sequences according to a preferred embodiment of the present invention.

[0052] Fig.13 This is a real scene operation effect diagram of a preferred embodiment of the present invention.

[0053] Fig.14 It is a trajectory comparison diagram of a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0054] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0055] Unless otherwise defined, all technical and scientific terms used in this embodiment have the same meaning as those generally understood by those skilled in the art of the present invention. The terms used in this embodiment in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used in this embodiment includes any and all combinations of one or more related listed items.

[0056] See also Figures 1 to 14 A dynamic visual recognition method based on a lightweight instance segmentation network according to a preferred embodiment of the present invention comprises the following steps:

[0057] S1. Obtain scene data based on the sensor data and camera data of the mobile robot. In this embodiment, the dynamic object segmentation SLAM algorithm based on ORB-SLAM3 is implemented around ORB feature points. The system supports multiple mode inputs such as monocular, binocular, RGB-D camera, and camera + MIU. The main threads include tracking thread, local map construction thread, loop thread, and Pangolin visualization thread. The data captured by the sensor enters the tracking thread after initialization. The tracking thread first extracts ORB feature points for each frame image and calculates the corresponding descriptors.

[0058] S2. Add an instance segmentation model in the tracking thread of ORB-SLAM3, separate the dynamic area and the static area of ​​the scene data in step S1 through the instance segmentation model to extract the dynamic object image, process the dynamic object image and obtain the prior dynamic object mask, and remove the dynamic area feature points of the prior dynamic object mask to obtain the dynamic mask.

[0059] In step S2, the scene data of step S1 is separated into dynamic areas and static areas by using an instance segmentation model, including the following steps:

[0060] S2.1 inserts the BoT block into the VoVNetV2 feature extraction network of the SOLOv2 model, and replaces the convolution in VoVNetV2 with the Ghost Conv module to form an improved feature extraction network. The scene data is divided into an S×S grid and input into the improved feature extraction network to obtain a feature map by extracting image features.

[0061] In step S2.1, the BoT block uses the MHSA module, and before the grid is input to the improved feature extraction network, in order to process the two-dimensional image, the spatial dimension of the two-dimensional image x is adjusted to a one-dimensional sequence x through the PatchEmebd module. p , where x∈R (H×W×C) , x p ∈R N×(P·C) , (H, W) is the resolution of the input image; C is the number of channels; (P, P) is the resolution of each image block; N = HW / P 2 is the number of image blocks, and N is the effective input sequence length of the MHSA module.

[0062] In step S2.1, the GhostConv module generates a small channel feature map with a smaller channel through traditional convolution. The small channel feature map generates a new feature map through a cheap linear transformation operation, and then the small channel feature map is concatenated with the new feature map to obtain a feature map.

[0063] The position information of the sequence is transferred to the feature sequence, and three identical feature matrices Q, K, and V are generated. The feature matrices are projected h times to C by linear projection. q , C k , C v Dimension, calculate parallel dot product attention so that the MHSA module can utilize the order information of the sequence. The parallel dot product attention calculation method is:

[0064]

[0065] S2.2 replaces the original network of the Neck part in the SOLOv2 model with an FPG type Neck part to form an improved Neck part, inputs the feature map into the improved Neck part to obtain feature maps of different sizes at each level, and fuses the feature maps at each level to obtain a high-resolution feature map with semantic information.

[0066] S2.3 adds a CBAM module to the Head part of the original network in the SOLOv2 model to obtain an improved Head part, inputs the high-resolution feature map into the improved Head part, predicts the instance category through the Category Branch of the Head part, predicts the instance mask through the Mask Branch of the Head part, and obtains the instance segmentation result by matching the instance category with the instance mask one by one to obtain a dynamic mask.

[0067] In step 2.3, the CBAM module performs global maximum pooling and global average pooling on the high-resolution feature map to obtain the channel dimension weight and spatial dimension weight of the high-resolution feature map, and inputs the weight vectors of the channel dimension weight and spatial dimension weight into the same multi-layer perceptron, adds the mapping weights output by the MLP at the pixel level, and then outputs the channel attention feature M after activation by the sigmoid function. C (F), then:

[0068]

[0069] Among them, σ is the sigmoid function; W0 and W1 are the weights shared by MLP; W0 is the ReLU activation function, which converts M C Multiply the channel with the high-resolution feature map to get F', and use F' as the input of the spatial attention module;

[0070] In the spatial attention module, global maximum pooling and global average pooling are performed based on the channel to obtain two feature maps with a channel number of 1. After channel splicing, a 7×7 convolution operation is performed to reduce the number of channels to 1, and then the spatial attention feature M is obtained through the Sigmoid function. S (F'), then:

[0071]

[0072] Where σ is the sigmoid function; f 7×7 Represents a 7×7 convolution; the spatial attention feature M S (F') is multiplied by F' to obtain a dynamic mask.

[0073] SOLOv2 is one of the instance segmentation algorithms with relatively good comprehensive performance at this stage, but its detection and segmentation effect on small objects is poor. This embodiment improves the SOLOv2 network in terms of small target object detection and algorithm processing speed:

[0074] 1. Insert the BoT block into the VoVNetV2 feature extraction network and use Ghost Con to replace the traditional convolution in VoVNetV2. The BoT module obtains the global dependency of the image by aggregating the local interaction information of the image, enabling the network to obtain the associated features of long sequences. Ghost Conv significantly reduces the computational cost and memory usage of the network by replacing some traditional convolutions with cheap linear operations. The improved backbone network not only enhances the feature extraction capability of the model, but also reduces the computational complexity and parameter amount of the model.

[0075] VoVNetV2 is an efficient and lightweight backbone network. First, a stemblock is formed by three 3×3 convolutional layers to complete the downsampling operation. Then, after a One-Shot Aggregation (OSA) module containing 4 stages, the OSA module consists of 5 consecutive convolutional layers, aggregates the feature map, and is fused with the Effective Squeeze-Excitation (eSE) module before adding residual connections to obtain the final output. The eSE module addresses the problem of channel information loss due to size reduction in the Squeeze-and-Excitation (SE) attention mechanism. Two fully connected layers (Full Connection, FC) that do not reduce the channel size are used to retain channel information, and the interdependence between feature map channels is explicitly modeled to enhance the representation ability of the feature map. The addition of the eSE module improves the network's ability to interact with information between image channels, but lacks the impact of global information on features. Therefore, this embodiment adds BoTblock to the OSA module to obtain global dependencies by aggregating local interactions.

[0076] The BoT block moves self-attention into computer vision tasks, adding 1×1 convolutional layers before and after the Multi-Head Self-Attention (MHSA) structure. The BoT block greatly improves small object detection. The OSA module with the BoT block structure not only focuses on the aggregation of local information, but also improves the network's attention to global information, effectively combining local information with global information, making the features of small objects in the image more prominent. The MHSA structure used in the BoT block is shown in the figure below. Figure 2 shown.

[0077] Since the addition of BoT blocks will inevitably lead to an increase in model parameters and computational complexity, in order to solve this problem, this embodiment uses Ghost Conv to replace traditional convolution to reduce model parameters and computational complexity. Ghost Conv is a method of compressing models that can generate more feature maps with fewer parameters, while ensuring network accuracy while reducing network parameters and computational complexity. In order to reduce network computational complexity, Ghost Conv divides traditional convolution into two steps, such as Figure 3 As shown in the figure, first a feature map with a smaller channel is generated through traditional convolution, then a new feature map is generated based on the obtained feature map through a cheap linear transformation operation (depthwise convolution), and finally the two sets of feature maps are concatenated together to obtain the final output feature map.

[0078] like Figure 3 As shown in the figure, c, h, w represent the number of channels, height and width of the input image, respectively; m, h', w' represent the number of channels, height and width of the intrinsic feature map obtained after traditional convolution, respectively; n represents the number of channels of the final output feature map, the size of the traditional convolution kernel is k, Φ represents the depth convolution, the size of the depth convolution kernel is d, and after s times of transformation, the speedup ratio between the traditional convolution calculation amount and the Ghost Conv calculation amount is r s (speed up ratio) is shown in formula (2-4):

[0079]

[0080] After s transformations, the compression ratio between the traditional convolution kernel parameters and the Ghost Conv convolution kernel parameters is r c (compression ratio) is shown in formula (2-5).

[0081]

[0082] Among them, since n is the number of channels of the final output feature map, n = m*s, so m = n / s, where s-1 is because the identity part does not need to be calculated. From the above calculation formula, it can be seen that compared with traditional convolution, GhostConv can reduce the amount of calculation and parameters of the model and achieve faster calculation speed.

[0083] 2. In step S2.2, each level of feature map enters the prediction head to predict the semantic category and instance mask. In the feature map outputs of each size of the feature pyramid network, the small-size feature map has a large receptive field and rich semantic information, but the image resolution is low, the target position is relatively rough, and the small target information is seriously missing; while the large-size feature map has a higher resolution and accurate target position, but the receptive field is small and lacks semantic information. In order to solve this problem, this embodiment constructs fine resolution features by using FPG instead of FPN.

[0084] FPG is a deep multi-path feature pyramid with a structure like Figure 4 As shown in Figure 2. The feature scale space is represented as a regular grid of parallel bottom-up paths and fused through multi-directional lateral connections. Unlike FPN, all independent pathways of FPG are built bottom-up, similar to the backbone path from input image to predicted output. In order to achieve information exchange at all levels of the image, FPG interweaves pyramid paths between and within scales through various lateral connections, forming a deep feature pyramid network. Figure 4 In the figure, the green arrow realizes the feature fusion between adjacent channels, the blue arrow shortens the path from low-level features to high-level features, the purple arrow fuses the up-sampled features and the down-sampled features together, and the red arrow directly connects and shortens the training time of the network.

[0085] 3. CBAM is an efficient and lightweight attention module with the following structure: Figure 5 As shown in Figure 2, the image feature weights are obtained from the two dimensions of channel and space, and then the feature map is refined to enhance the representation ability of the feature map.

[0086] The loss function plays a vital role in machine learning and deep learning. It is a key part of model optimization. The loss function is the objective function in the model training and optimization process. By defining, measuring and minimizing the loss, the model can better adapt to the training data and improve its generalization ability on new data. The loss function is used to measure the difference or error between the output of the model and the true label. A lower loss value indicates that the model performs better on the training data, while a higher loss value indicates that the model needs to be adjusted to improve performance. In this embodiment, the commonly used loss functions for instance segmentation models are cross entropy loss, boundary loss, and Dice loss.

[0087] In a dynamic SLAM system, the pose estimation result obtained by using feature points on objects with relatively low dynamic probability is more accurate than that obtained by using feature points on objects with high dynamic probability. Therefore, in the system proposed in this paper, an instance segmentation network is used to extract objects with relatively high dynamic probability in dynamic scenes, and feature points on moving objects are removed. Static feature points are used for pose estimation. The specific process of instance segmentation to extract dynamic objects is as follows: Figure 6 shown.

[0088] S3. The data of step S1 is extracted in the tracking thread to obtain the ORB features for each frame image, and the dynamic points of the ORB features are extracted through multi-view geometry to obtain a rough classification of feature points.

[0089] In step S3, x is a feature point on the previous key frame. Based on the camera motion, the projection of point x to the position x' on the current key frame CF and the projection depth Z can be calculated. proj , the calculation method is:

[0090]

[0091] Among them, u, v, w are the coordinates of the x image coordinate system; X, Y, Z are the coordinates of the x' image coordinate system; K is the camera intrinsic parameter matrix;

[0092] By calculating the back-projection angle α between x and x', the back-projection angle α and the depth of the feature point Z' and Z proj The difference between and determines whether x is a dynamic point.

[0093] When the back-projection angle α is greater than 30° and Z proj - When Z' is greater than the threshold, x is a dynamic point.

[0094] By using an instance segmentation network, we can segment out potential dynamic objects, but there are still many non-prior dynamic objects in most dynamic scenes. It is impossible to obtain dynamic masks through the instance segmentation network. Multi-view geometry methods are used to extract non-prior dynamic objects.

[0095] like Figure 7 As shown, x is the feature point on the previous key frame. Based on the camera motion, the projection of point x to the position x' on the current key frame CF and the projection depth Z can be calculated. proj For each feature point on the previous frame, calculate the back-projection angle α between x and x'. If α>30°, it is considered that the point is very likely to be occluded in the current frame, resulting in an incorrect match in the feature point matching, so it is considered as a dynamic point to be removed; when α>30°, compare the feature point depth Z' and Z obtained by the depth camera on the current frame proj The difference in Figure 6 As can be seen from the figure, when x is a static point, Z'=Zproj ; When x is a dynamic point, Z' is much smaller than Z proj Therefore, all points whose Zproj-Z' is greater than the threshold of 0.4m are considered dynamic points.

[0096] S4. Combine the dynamic mask of step S2 with the coarse classification of feature points in step S3 to obtain the final classification of features, remove feature points that are considered dynamic points, and use feature points that are considered static points for camera pose estimation to construct an environment map.

[0097] In step S4, the RGB and depth channels of the previous key frame are projected to the dynamic part of the current frame according to the positions of the previous frame and the current frame, and a realistic image without dynamic objects is synthesized by using the static information of the previous view to repair the occluded background.

[0098] S5. Perform background restoration by filling the blocked background in the key frame where the dynamic object is eliminated in the final classification of the features in step S4.

[0099] Background inpainting is performed by filling the occluded background in keyframes that remove dynamic objects after instance segmentation. By using static information from previous views, a realistic image without dynamic objects can be synthesized. Such synthesized frames, containing only static objects in the environment, are extremely helpful for virtual and augmented reality applications, as well as relocalization and camera tracking after map creation.

[0100] Since we know the positions of the previous and current frames, we project the RGB and depth channels of the first 20 keyframes to the dynamic part of the current frame. For some gaps there is no correspondence and they are left empty. Some areas cannot be filled because their corresponding part of the scene has not appeared in the keyframes so far, or if it has appeared, it has no valid depth information. Figure 4-5 shows a synthetic image of the input frame from the TUM dataset, Figure 8 (a) is the original image, Figure 8 (b) is the image after background restoration. The image after background restoration removes the dynamic content and effectively fills the segmented part with the information of the static background.

[0101] S6. Obtain the data of the tracking thread through the local map building thread of ORB-SLAM3 to optimize the key frame poses and corresponding map points in the environment map, and perform local bundling adjustments to ensure the consistency of the environment map and the camera trajectory.

[0102] S7. The data of the local map building thread is fused through the loop closure thread of ORB-SLAM3 to merge the environment map and perform global BA optimization.

[0103] The loop thread includes three parts: closed-loop detection, closed-loop correction and map fusion. The bag of words is used to obtain candidate closed-loop frames, which are matched with the current key frames to correct the accumulated errors of the system. Finally, the maps are merged for global BA optimization. The visualization thread mainly visualizes the system's key frames, map points, graph structures and some system operation buttons by calling Pangolin based on OpenGL.

[0104] In order to verify the effectiveness of the instance segmentation model, this embodiment conducts the following experiments:

[0105] Experimental environment configuration: The hyperparameters are set as follows: the optimizer is AdamW; the weight decay coefficient is 0.0001; the number of training iterations is 50; the batch size is 4; the learning rate (Lr) is 0.0001; the learning rate is adjusted at the 20th, 30th and 40th cycles; the image scale is randomly sampled from 1024 to 2048.

[0106] Experimental Dataset: The Cityscapes dataset was used for experiments. The Cityscapes dataset is a road scene target segmentation dataset that focuses on providing training and performance testing for autonomous driving environment perception models. The images in this dataset come from video sequences of different road scenes in 50 cities in Germany and neighboring countries, covering different street scenes, road scenes, and seasons. The dataset contains a total of 5,000 finely annotated images and 2,000 coarsely annotated images. Since instance segmentation requires a high level of data annotation, this experiment only used finely annotated images, including 2,975 images for training, 1,525 for testing, and 500 for verification. This embodiment mainly tests the instance segmentation algorithm in road scenes, so we selected the eight most common road scene categories that contain instance segmentation labels in the Cityscapes dataset, namely pedestrians, riders, cars, trucks, buses, trains, motorcycles, and bicycles. The number of images and instances of each category is shown in Table 1-1. The content of part of the Cityscapes dataset is shown in the following table. Fig. 9 As shown:

[0107] Table 1-1 Number of images and instances for each category in the dataset

[0108]

[0109] Experimental evaluation indicators: Since this experiment is based on the improvement of SOLOv2, we will continue to use the original evaluation indicators of SOLOv2, that is, use the mean average precision (mAP) and mean average recall (mAR) to comprehensively evaluate the algorithm, and the calculations are shown in formulas (7) and (8) respectively. In addition, the number of frames per second (Frames Per Second, FPS) is used to evaluate the segmentation speed of the algorithm.

[0110]

[0111] Among them, C represents the total number of categories, c represents the current category, T represents the threshold, t represents the current threshold, true positive (TP) represents the positive samples correctly predicted by the model as positive samples; false positive (FP) represents the positive samples and negative samples incorrectly predicted by the model as positive samples; false negative (FN) represents the positive samples incorrectly predicted by the model as negative samples.

[0112] Comparison of different algorithms: To verify the effectiveness of the algorithm proposed in this embodiment, we compared it with the current mainstream instance segmentation algorithms Mask R-CNN, YOLACT, PolarMask, SOLO and SOLOv2 on the same data set. For the comparison algorithm, we selected ResNet_50 as the backbone network. The experimental results are shown in Table 1-2. It can be seen that compared with the two-stage excellent instance segmentation algorithm Mask R-CNN, although the mAP and AP of the algorithm proposed in this embodiment are S 、AP M and AP L They decreased by 3.2%, 8.4%, 3.1% and 0.5% respectively, but the FPS increased by 5.1. The algorithm proposed in this embodiment greatly improved the algorithm segmentation speed with a small loss of accuracy. Compared with YOLACT, the mAP and AP of the algorithm proposed in this embodiment are S 、AP M and AP L The segmentation accuracy and speed have been improved to varying degrees. Compared with PolarMask, the mAP and AP of the algorithm proposed in this embodiment are S 、AP M and AP L The segmentation accuracy of the proposed algorithm is significantly improved by 8.9%, 2.6%, 9.7% and 11.4% respectively, and the FPS is reduced by 0.6. The algorithm proposed in this embodiment significantly improves the segmentation accuracy of the algorithm with a small loss in segmentation speed. Compared with SOLO, the mAP and AP S 、AP M and AP LThe segmentation accuracy and speed have been improved to varying degrees. Compared with SOLOv2, the mAP and AP of the algorithm proposed in this embodiment are improved by 6.8%, 2.7%, 9.6% and 5.2% respectively, and the FPS is improved by 3. S 、AP M and AP L The segmentation accuracy and speed have been improved by 4.2%, 2.6%, 6.9% and 2.5% respectively, and the FPS has been improved by 1.7. Both the segmentation accuracy and speed have been improved to varying degrees. From the experimental results, it can be seen that except for the two-stage instance segmentation algorithm Mask R-CNN, which has higher segmentation accuracy than the algorithm proposed in this embodiment, the segmentation accuracy of other single-stage instance segmentation algorithms is lower than that of the algorithm proposed in this embodiment, especially in the segmentation of small objects. The algorithm proposed in this embodiment has obvious advantages. In terms of segmentation speed, the FPS of the algorithm proposed in this embodiment is higher than that of other algorithms except PolarMask, which reflects a strong processing speed. The comprehensive algorithm segmentation accuracy and speed verify the effectiveness of the algorithm proposed in this embodiment.

[0113] Table 1-2 Comparison of performance of different algorithms

[0114]

[0115] This example also compares the segmentation accuracy of the dataset category between the proposed algorithm and the original algorithm. The results are as follows: Fig.10 As shown in the figure, among the eight categories, compared with the original algorithm, the algorithm proposed in this embodiment improves by 9.9%, 9.9%, 2.3%, 8.1% and 7.5% in the five categories of people, riders, trucks, motorcycles and bicycles respectively. In the three categories of cars, buses and trains, it is slightly lower than the original algorithm.

[0116] Ablation experiment: In order to further verify the effectiveness of the proposed structure, the detection effect of the proposed structure is compared with that of the original module, and the results are shown in Tables 1-3. After replacing the backbone network with VoVNetV2, the mAP decreased by 2.9% compared with the original model, but the FPS increased by 2.9. Although the addition of VoVNetV2 reduced the segmentation accuracy of the model, it accelerated the processing speed of the model. When FPN was replaced by FPG, the mAP of the model increased by 1.9%, while the FPS decreased by 0.3. After adding the BoT module, the mAP of the model increased by 2.2%, but the FPS decreased by 1.4. When the convolution was replaced by Ghost Conv, although the mAP value of the model increased by only 0.2%, the FPS increased by 1.1. Finally, after adding the CBAM module, the mAP increased by 4.2% and the FPS increased by 1.7 compared with the original model, verifying the effectiveness of the modules added in this embodiment.

[0117] Table 1-3 Comparison of ablation test performance

[0118] Table 3-4 Comparison of ablation experiment performance

[0119]

[0120] In order to more intuitively demonstrate the effectiveness of the improved algorithm, we made a visual comparison of the segmentation results of the improved algorithm and the original SOLOv2 algorithm, as shown in Fig.11 As shown, from top to bottom: 1- input RGB image; 2- SOLOv2; 3- our algorithm. The red boxes indicate areas that SOLOv2 failed to detect or achieved incomplete segmentation compared to the improved algorithm. In the figure, from top to bottom, one can observe the original image, the SOLOv2 segmentation result image, and the improved segmentation result image. The red boxes indicate areas where SOLOv2 has problems in detection or incomplete segmentation compared to the improved algorithm. Obviously, the improved algorithm performs well in segmentation accuracy and mask quality compared to the original algorithm. Especially for smaller objects, the original algorithm often has missed detections and incorrect segmentations in these cases.

[0121] In order to verify the effectiveness of the improved algorithm based on ORB-SLAM3 in this embodiment, the following experiments are conducted:

[0122] Experimental data set: Experiments were conducted using the TUM data set. The TUM data set is a series of data sets for visual SLAM research provided by the Technical University of Munich. These data sets usually contain image sequences collected in different environments, as well as the corresponding camera motion and ground truth real position data. The dynamic object sequence in the TUM data set is considered to be a benchmark for evaluating the performance of SLAM systems in dynamic environments. The sequence can be divided into two categories, namely high-dynamic scenes (marked as walk) and low-dynamic scenes (marked as static). In the high-dynamic scene sequence, the main dynamic object is a person. In this scene, two people move back and forth around a table. During this period, they may move a stool to sit down or stand up, and there is a situation where the table blocks the moving person. In the low-dynamic scene sequence, two people sit on a chair indoors and talk. They communicate only through verbal communication and gestures.

[0123] The dynamic object sequence shows a total of four types of camera motion: (1) xyz, the camera moves along the x, y, and z directions while maintaining its initial orientation; (2) rpy, which means that the camera rotates around the roll-pitch-yaw axes while remaining in the same position; (3) static, which means that the camera is manually kept stationary; (4) halfsphere, which means that the camera moves on a small hemisphere with an approximate diameter of one meter. The size of all images in the TUM sequence is 640x 480. For ease of expression, we use fr3 / sit / half, fr3 / sit / xyz, fr3 / walk / hs, fr3 / walk / static, fr3 / walk / rpy and fr3 / walk / xyz to represent six groups of image sequences, where fr3 represents the scene category, sit / wlak represents low-dynamic or high-dynamic scenes, and the last parameter represents the four types of camera motion.

[0124] Experimental evaluation indicators: In the experiments in this chapter, absolute trajectory error (ATE) and relative pose error (RPE) are mainly used to evaluate the accuracy of our system.

[0125] ATE is the difference between the estimated global robot path and the true path, calculated as shown in formula (9).

[0126]

[0127] Where N is the total number of time steps; P est (t) represents the t time step, calculating the camera pose on the estimated trajectory; P gt(t) represents the camera pose on the true trajectory. The smaller the ATE is, the closer the estimated trajectory is to the true trajectory. Otherwise, it means that there is a large error in the estimation.

[0128] RPE is different from ATE. RPE focuses on the error between camera poses between adjacent frames without considering the alignment of the overall trajectory. The calculation is shown in formula (10). RPE facilitates a more detailed understanding of the performance of the SLAM algorithm, especially in local areas. In this experiment, translation RPE (TranslationRelative Pose Error) and rotation RPE (RotationRelativePose Error) are selected for evaluation.

[0129]

[0130] Among them, Q i represents the true pose of the i-th frame image, Q i ∈SE(3);P i represents the estimated pose of the i-th frame image, P i ∈SE(3).

[0131] We conducted experiments on four high-dynamic sequences and two low-dynamic sequences in the TUM dataset, and selected the median error (Mean Error), root mean square error (RMSE) and standard deviation (SD) to quantitatively evaluate the accuracy of the algorithm.

[0132] Comparison of different algorithms: In order to verify the effectiveness of the algorithm proposed in this embodiment, we experimentally compared it with the basic algorithm ORB-SLAM3 and the current excellent semantic SLAM algorithm DS-SLAM and DynaSLAM in the dynamic sequence of the TUM dataset.

[0133] The experimental results compared with the ORB-SLAM3 algorithm are shown in Table 2-1, Table 2-2, and Table 2-3. Compared with ORB-SLAM3, the algorithm proposed by us shows improvements in ATE and RPE of each sequence. Especially in highly dynamic sequences, our algorithm has achieved significant improvements in both ATE and RPE relative to ORB-SLAM3. Specifically, in terms of ATE, the highest improvement sequence in Mean and RMSE values ​​reached 97.84% and 97.76% respectively; in terms of RPE, the highest improvement sequence in Mean and RMSE values ​​reached 96.10% and 95.25% respectively. This is mainly attributed to the fact that the algorithm proposed by us introduces an improved instance segmentation network in the ORB-SLAM3 framework, which can effectively remove dynamic objects and reduce the mismatch of feature points, thereby significantly improving the positioning accuracy and robustness of the system in highly dynamic environments. In low-dynamic sequences, the performance improvement of our algorithm is smaller, mainly because in this scenario, ORB-SLAM3 uses the RANSAC algorithm to identify slightly moving objects as outliers and remove them, so it can handle low-dynamic scenes well and achieve good results.

[0134] Table 2-1 Comparison of results between ORB-SLAM3 and this article's algorithm ATE(m)

[0135]

[0136] Table 4-3Comparison of translation RPE(m) results between ORB-SLAM3and this algorithm

[0137]

[0138] Table 2-3 Comparison of rotation RPE(m) results between ORB-SLAM3 and this algorithm

[0139]

[0140] The experimental results compared with the DS-SLAM algorithm are shown in Tables 2-4, 2-5, and 2-6. Compared with DS-SLAM, the algorithm we proposed has certain improvements in ATE and RPE of each sequence. Specifically, in terms of ATE, the highest improvement sequence in Mean and RMSE values ​​reached 93.66% and 92.32%, respectively, and the lowest improvement sequence was 2.11% and 4.46%, respectively; in terms of RPE, the highest improvement sequence in Mean and RMSE values ​​reached 84.18% and 88.69%, respectively, and the lowest improvement sequence was 2.85% and 2.27%, respectively. This is mainly because the SegNet network used by DS-SLAM has a poor segmentation effect, resulting in incorrect matching points. At the same time, because DS-SLAM does not perform background repair on the image after segmentation and elimination, the system has insufficient number of matching feature points for pose estimation, which affects the accuracy of the algorithm.

[0141] Table 2-4 Comparison of results between DS-SLAM and this algorithm ATE(m)

[0142]

[0143] Table 2-5 Comparison of translation RPE(m) results between DS-SLAM and this algorithm

[0144]

[0145] Table 2-6 Comparison of rotation RPE (m) results between DS-SLAM and this algorithm

[0146]

[0147] The experimental results compared with the DynaSLAM algorithm are shown in Tables 2-7, 2-8, and 2-9. Compared with DynaSLAM, the algorithm we proposed has certain improvements in ATE and RPE of some sequences. Specifically, in terms of ATE, the highest improvement sequence in Mean and RMSE values ​​reached 44.62% and 18.89% respectively, and the lowest improvement sequence was -40.91% and -23.33% respectively; in terms of RPE, the highest improvement sequence in Mean and RMSE values ​​reached 29.49% and 28.72% respectively, and the lowest improvement sequence was -9.89% and -14.55% respectively. Although the algorithm proposed in this embodiment does not perform as well as DynaSLAM in some sequences, it performs better in most sequences. This is mainly because the Mask R-CNN network segmentation effect used by DynaSLAM is relatively outstanding, but Mask R-CNN takes a long time, resulting in slow system operation. The algorithm proposed in this embodiment has certain advantages in comprehensive accuracy and speed.

[0148] Table 2-7 Comparison of results between DynaSLAM and this article's algorithm ATE(m)

[0149]

[0150] Table 2-8 Comparison of translation RPE (m) results between DynaSLAM and this algorithm

[0151]

[0152] Table 2-9 Comparison of rotation RPE (m) results between DynaSLAM and this algorithm

[0153]

[0154]

[0155] In order to more intuitively demonstrate the effectiveness of the algorithm proposed in this embodiment, we also compared the trajectory diagrams of the proposed algorithm with those of ORB-SLAM3, DS-SLAM, and DynaSLAM in six dynamic sequences of the TUM dataset. The results are as follows: Fig.12 shown.

[0156] exist Fig.12 In the figure, the blue line represents the trajectory estimated by the algorithm, the black line is the ground truth trajectory, and the red line segment represents the error between the estimated trajectory and the true trajectory. The longer the red line, the greater the error. It can be clearly seen that the semantic SLAM algorithms DS-SLAM, DynaSLAM and our algorithm based on deep learning perform better than ORB-SLAM3 in high dynamic sequences. The trajectories of the three semantic SLAM algorithms are roughly similar, but in the fr3 / walk / rpy sequence, there are obvious differences. In this sequence, DS-SLAM performs abnormally in this sequence, and although the estimated trajectory of DynaSLAM is close to the ground truth trajectory, the convergence degree of the trajectory in the lateral movement direction of the camera is not as obvious as the algorithm proposed in this embodiment. This is because in this sequence, the camera motion is large and the frame rate is low. In addition, the semantic segmentation network used by DynaSLAM and DS-SLAM lacks sufficient robustness to motion blur, resulting in erroneous data association that reduces tracking accuracy. The instance segmentation model used by the algorithm proposed in this embodiment performs well in processing motion blur, and its secondary filtering of the geometric method can effectively suppress dynamic points, so it has the best tracking accuracy among the four algorithms. In the two low-dynamic sequences, the trajectory convergence of ORB-SLAM3 is very close to that of the other three dynamic SLAM algorithms, and even better than DynaSLAM and DS-SLAM. This can be attributed to the fact that the characters only made slight gestures in these two low-dynamic sequences.

[0157] In order to verify the effectiveness of the algorithm proposed in this embodiment in real scenarios, we carried out experiments on the algorithm on the TARKBOT R20-MEC series ROS robot platform of Tucker Innovation in laboratory scenarios. The robot is equipped with an Orbbec AstraPro Plus depth camera, a Blue Ocean E300 laser radar, a Jeston Xavier NX host, a 12V lithium battery, 4 DC reduction motors, 4 Mecanum wheels, a ROS robot expansion board, and a touch screen.

[0158] The results are as follows Fig.13 shown. Fig.13 (a) is the result of feature point extraction by the ORB-SLAM3 algorithm. The algorithm does not distinguish between dynamic and static feature points. Using feature points on a walking person for pose estimation will result in pose estimation errors. Fig.13(b) is the dynamic object mask extracted by the algorithm proposed in this embodiment, and all dynamic objects in the scene are segmented. Fig.13 (c) is the result of feature point extraction by the algorithm proposed in this embodiment. The blue dots are static feature points and the red dots are dynamic feature points. It can be seen that the static feature points and dynamic feature points in the scene have been distinguished. The dynamic feature points will be removed subsequently and only the static points will be used for pose estimation to improve the positioning accuracy of the algorithm.

[0159] This example also compares the estimated trajectory and the actual trajectory of the ROS robot equipped with different algorithms in real scenes. The experimental results are as follows: Fig.14 As shown, Fig.14 (a) is a comparison of the estimated trajectory and the actual trajectory of the ROS robot equipped with the ORB-SLAM3 algorithm. Fig.14 (b) is a comparison diagram of the estimated trajectory and the actual trajectory of the ROS robot equipped with the SLAM algorithm proposed in this embodiment. The blue line represents the estimated trajectory, the black line represents the actual trajectory, and the red line represents the difference between the estimated trajectory and the actual trajectory. It can be seen that the difference between the estimated trajectory and the actual trajectory in the left figure is significantly greater than that in the right figure, which confirms the effectiveness of the algorithm proposed in this embodiment in real dynamic scenes.

Claims

1. A dynamic visual recognition method based on a lightweight instance segmentation network, characterized in that: The steps include: S1. Obtain scene data based on sensor data and camera data of the mobile robot; S2. Add an instance segmentation model in the tracking thread of ORB-SLAM3, separate the dynamic area and the static area of ​​the scene data of step S1 through the instance segmentation model to extract the dynamic object image, process the dynamic object image and obtain a priori dynamic object mask, and remove the dynamic area feature points of the priori dynamic object mask to obtain a dynamic mask; S3. The data of step S1 is extracted in the tracking thread to obtain the ORB feature of each frame image, and the dynamic points of the ORB feature are extracted by multi-view geometry to obtain a rough classification of feature points; S4. Combining the dynamic mask of step S2 with the rough classification of feature points of step S3 to obtain a final classification of features, removing feature points regarded as dynamic points, and using feature points regarded as static points for camera pose estimation to construct an environment map; S5. Fill the occluded background with the key frame of the dynamic object removed in the final classification of the features in step S4 to perform background repair; S6. Obtain the data of the tracking thread through the local map building thread of ORB-SLAM3 to optimize the key frame poses and corresponding map points in the environment map, and perform local bundling adjustments to ensure consistency between the environment map and the camera trajectory; S7. The data of the local map building thread is fused through the loop thread of ORB-SLAM3 to merge the environment map and perform global BA optimization.

2. The dynamic visual recognition method based on a lightweight instance segmentation network according to claim 1, characterized in that: In step S2, the separation of the dynamic area and the static area of ​​the scene data in step S1 by the instance segmentation model includes the following steps: S2.1 inserting a BoT block into the VoVNetV2 feature extraction network of the SOLOv2 model, and replacing the convolution in VoVNetV2 with a Ghost Conv module to form an improved feature extraction network, dividing the scene data into an S×S grid and inputting it into the improved feature extraction network, and obtaining a feature map by extracting image features; S2.2 replaces the original network of the Neck part in the SOLOv2 model with a Neck part of the FPG type to form an improved Neck part, inputs the feature map into the improved Neck part to obtain feature maps of different levels of size, and fuses the feature maps of each level to obtain a high-resolution feature map with semantic information; S2.3 Add a CBAM module to the Head part of the original network in the SOLOv2 model to obtain an improved Head part, input the high-resolution feature map into the improved Head part, predict the instance category through the Category Branch of the Head part, predict the instance mask through the MaskBranch of the Head part, and obtain the instance segmentation result by matching the instance category with the instance mask one by one to obtain a dynamic mask.

3. The dynamic visual recognition method based on a lightweight instance segmentation network according to claim 2, characterized in that: In step S2.1, the BoT block uses the MHSA module to adjust the spatial dimension of the two-dimensional image x to a one-dimensional sequence x through the PatchEmebd module before inputting it into the improved feature extraction network. p , where x∈R (H×W×C) , x p ∈R N ×(P·C) , (H, W) is the resolution of the input image; C is the number of channels; (P,P) is the resolution of each image block; N = HW / P 2 is the number of image blocks, N is the effective input sequence length of the MHSA module; The position information of the sequence is added to the feature sequence, and three identical feature matrices Q, K, and V are generated. The feature matrices are projected h times to C by linear projection. q , C k , C v Dimension, calculate parallel dot product attention so that the MHSA module can utilize the order information of the sequence. The parallel dot product attention calculation method is:

4. The dynamic visual recognition method based on a lightweight instance segmentation network according to claim 2, characterized in that: In step S2.1, the Ghost Conv module generates a small channel feature map with a smaller channel by traditional convolution, and the small channel feature map generates a new feature map through a cheap linear transformation operation, and then the small channel feature map is spliced ​​with the new feature map to obtain a feature map.

5. The dynamic visual recognition method based on a lightweight instance segmentation network according to claim 2, characterized in that: In step 2.3, the CBAM module performs global maximum pooling and global average pooling on the high-resolution feature map to obtain the channel dimension weight and spatial dimension weight of the high-resolution feature map, and inputs the weight vectors of the channel dimension weight and the spatial dimension weight into the same multi-layer perceptron, adds the mapping weights output by the MLP at the pixel level, and then outputs the channel attention feature M after activation by the sigmoid function. C (F), then: Among them, σ is the sigmoid function; W0 and W1 are the weights shared by MLP; W0 is the ReLU activation function, which converts M C Perform channel multiplication with the high-resolution feature map to obtain F', and use F' as the input of the spatial attention module; In the spatial attention module, global maximum pooling and global average pooling are performed based on the channel to obtain two feature maps with a channel number of 1. After channel splicing, a 7×7 convolution operation is performed to reduce the number of channels to 1, and then the spatial attention feature M is obtained through the Sigmoid function. S (F'), then: Where σ is the sigmoid function; f 7×7 Represents a 7×7 convolution; the spatial attention feature M S (F') is multiplied by F' to obtain a dynamic mask.

6. The dynamic visual recognition method based on a lightweight instance segmentation network according to claim 1, characterized in that: In step S3, x is a feature point on the previous key frame. Based on the camera motion, the projection of point x to the position x' on the current key frame CF and the projection depth Z can be calculated. proj , the calculation method is: Among them, u, v, w are the coordinates of the x image coordinate system; X, Y, Z are the coordinates of the x' image coordinate system; K is the camera intrinsic parameter matrix; By calculating the back-projection angle α between x and x', the back-projection angle α and the depth of the feature point Z' and Z proj The difference between and determines whether x is a dynamic point.

7. The method for dynamic visual recognition based on a lightweight instance segmentation network according to claim 6, characterized in that: When the back-projection angle α is greater than 30° and Z proj - When Z' is greater than the threshold, x is a dynamic point.

8. The dynamic visual recognition method based on a lightweight instance segmentation network according to claim 1, characterized in that: In step S4, the RGB and depth channels of the previous key frame are projected to the dynamic part of the current frame according to the positions of the previous frame and the current frame, and a realistic image without dynamic objects is synthesized by using the static information of the previous view to repair the occluded background.

Citation Information

Cited By

  • SLAM method based on double-flow feature fusion

    CN120740570A