A method for improving the robustness of visual SLAM applied to dynamic indoor scenes
Through semantic segmentation neural network and lightweight frame tracking combined with pole geometric constraints, the stability and accuracy of visual SLAM in dynamic indoor environments is solved, and the effective elimination of moving objects and the accuracy of pose estimation is achieved.
Patent Information
- Application Number
- CN202211454580.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-21
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-11-21
AI Technical Summary
The stability and accuracy of existing visual SLAM technology in dynamic indoor environments are disturbed by moving objects, especially the inability to effectively distinguish between active and passive moving objects, resulting in a significant reduction in system stability.
Semantic segmentation neural network is used to extract the region of a priori moving object in the scene, generate and finely process the mask outer contour, combine lightweight frame tracking and geometric constraints, segment feature points through Euro-style distance and depth map information, and eliminate interference from moving objects using optoelectrode geometry and scene flow constraints to perform accurate pose estimation.
The stability and accuracy of the visual SLAM system in a dynamic indoor environment is improved, and the accuracy and robustness of the system in the presence of moving objects is ensured through the combination of semantic segmentation and geometric constraints.
Smart Images

Figure CN115830066B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of SLAM, and in particular relates to a method for improving the robustness of visual SLAM applied to dynamic indoor scenes. Background Art
[0002] Simultaneous Localization and Mapping (SLAM) is an important algorithm in the field of robotics. It is an important basic technology for robots to perceive the outside world and clearly understand their own position and the appearance of the environment. Currently, SLAM technology plays an important role in aerospace, smart transportation, smart home, and autonomous driving. Visual SLAM technology refers to the simultaneous positioning and mapping technology using cameras as the main sensor. The current visual SLAM can maintain high accuracy in static environments, but in dynamic environments, due to the presence of autonomously moving objects in the scene, such as people, animals, chairs, tables, etc., the autonomous movement of these objects will interfere with the work of the visual SLAM system, resulting in a significant reduction in the stability of the system. Therefore, it is very meaningful to design a moving area removal module.
[0003] There are currently three main methods for improving the robustness of visual SLAM for dynamic indoor scenes: (1) Only the areas of objects that actively move, such as people and animals, in the scene are extracted through a semantic segmentation neural network, and then the feature points distributed in these areas are eliminated and do not participate in the calculation of the frame tracking thread; (2) Without using a neural network to extract prior information, the set of current moving feature points is obtained through the native geometric information in the image, through graph clustering, foreground culling and other algorithms, and the remaining set of static feature points after eliminating the moving points participates in the calculation of the frame tracking thread; (3) With the help of a neural network, the prior information of the moving area in the scene is extracted, and then combined with the geometric constraints between image frames, a more accurate set of static feature points is obtained and participates in the calculation of the frame tracking thread.
[0004] Technology (1) is a relatively simple processing method, and its main disadvantage is that the semantic information extracted by the semantic segmentation network is usually inaccurate and cannot identify objects that move passively, such as quilts, chairs, monitors, etc.
[0005] Technology (2) is a lightweight processing method. Its main disadvantage is that the native information in the image is difficult to express and extract well by mathematical formulas. When the camera and objects in the scene move at the same time, it is often impossible to determine whether it is the camera that is moving or there are really objects in the scene that are moving simply by using the native information in the image.
[0006] The technique (3) is a method that combines prior information and inter-frame geometric constraints. In theory, this method can more accurately eliminate the interference of moving objects on the system. However, there are still many problems in the current technical achievements: (1) The semantic information extraction method is too simple and does not consider the positional relationship between passive moving objects and active moving objects. In fact, the movement of passive objects is often caused by physical contact with active moving objects; (2) The computing power consumption of the calculation method for the inter-frame initial pose used to calculate geometric constraint data is relatively large, and many technical achievements do not eliminate the influence of moving objects in this step, resulting in a large error in the obtained initial inter-frame transformation matrix, and finally leading to inaccurate geometric constraint data; (3) Since there is significant noise interference in single geometric constraint data, but the vast majority of current technical achievements only use single geometric constraint data as the standard for finally separating static feature points and moving feature points. Summary of the Invention
[0007] To solve the above technical problems, the present invention proposes a method for improving the robustness of visual SLAM applied to dynamic indoor scenes, which enhances the stability of visual SLAM in indoor scenes with moving objects through prior information extracted by a neural network, lightweight frame tracking, and geometric constraints.
[0008] The technical solution adopted by the present invention is as follows: A method for improving the robustness of visual SLAM applied to dynamic indoor scenes, and the specific steps are as follows:
[0009] S1. Use a semantic segmentation neural network to extract the prior moving object region in the scene, generate a corresponding mask according to the semantic segmentation result, and then use the morphological operations of the image: erosion and dilation, to refine the mask, eliminate incorrect segmentation results, and extract the outer contour of the refined semantic mask;
[0010] S2. Based on the outer contour of the semantic mask extracted in step S1, traverse all feature points in the current frame to calculate the Euclidean distance from each feature point to the outer contour of the semantic mask, and divide the feature points in the current frame into three sets according to this distance combined with the depth map information corresponding to the current frame: prior dynamic feature point set, prior unknown feature point set, and prior static feature point set;
[0011] S3. Input the prior static feature point sets of the current frame and the previous frame into the lightweight frame tracking process in the motion consistency detection process to obtain an inter-frame transformation matrix, and use the obtained inter-frame transformation matrix and the geometric constraints in the motion consistency detection process to act on the prior unknown feature point set with good data association in the current frame and the previous frame to obtain epipolar geometric constraint and scene flow constraint data;
[0012] S4. Combine the data information of the epipolar geometry constraint and the scene flow constraint into a two-dimensional vector, and then use the two-norm of this vector to represent the final error value. If the error value is greater than the set threshold, then this feature point belongs to the set of moving feature points; otherwise, it belongs to the set of static feature points. Then, use the obtained set of static feature points for accurate frame tracking, and finally obtain the pose estimation data that eliminates the interference of moving objects in the scene.
[0013] Further, in step S1, it is specifically as follows:
[0014] When the visual SLAM system starts running, the system obtains the original data through the RGB-D camera, including the RGB image and the corresponding depth map information.
[0015] First, input the RGB image into the semantic segmentation neural network. After being processed by the neural network, a semantic information image is obtained. Then, input the semantic information image into the mask processing process for grayscale and binary processing, and finally extract the outer contour information of the semantic mask.
[0016] Further, in step S1, use image morphological operations to eliminate the wrong areas, specifically as follows:
[0017] Process the semantic mask in three stages:
[0018] (1) The first erosion operation. Select a template with an elliptical structure to eliminate the prior motion area generated by errors. The result of the first erosion is AΘB1;
[0019] (2) The dilation operation. Also select a template with an elliptical structure to fully expand the prior motion area on the basis of the first operation, and obtain an over-expanded mask. The result of the first dilation is
[0020] (3) The second erosion operation. Select the template of the dilation operation and appropriately erode the mask on the basis of the second operation to reduce the over-expanded mask and obtain a refined semantic mask with a moderate mask size. The result of the second erosion is
[0021] Among them, A represents the mask to be filtered, Θ represents the erosion operator, represents the dilation operator, B1 represents the structural template of the first erosion operation, and B2 represents the structural template of the dilation operation, which is also an elliptical structure.
[0022] Further, in step S3, the motion consistency detection process includes a lightweight frame tracking process based on the Iterative Closest Point (ICP) algorithm and an epipolar geometry constraint and scene flow constraint data extraction process, specifically as follows:
[0023] Using the data association information of the prior stationary feature point sets in the current frame and the previous frame, combining the Iterative Closest Point (ICP) algorithm and the rigid body motion transformation equation, constructing the relationship between the inter-frame transformation matrix from the previous frame to the current frame and the data association information, establishing a least squares problem, and then using the Singular Value Decomposition (SVD) algorithm of the matrix to solve the solution of the least squares problem, thereby obtaining the inter-frame transformation matrix from the previous frame to the current frame. Then, using the obtained inter-frame transformation matrix to act on the prior unknown feature point sets with good data association in the current frame and the previous frame to obtain the epipolar geometry constraint and the scene flow constraint data.
[0024] Further, in the step S3, the process of extracting the epipolar geometry constraint data is specifically as follows:
[0025] P is a moving map point in the scene; I1 represents the previous frame image, p1 represents the pixel point of point P on the previous frame, l1 represents the epipolar line passing through point p1, and e1 represents the epipole; I2 represents the current frame image, p2 represents the pixel point of point P on the current frame, l2 represents the epipolar line passing through point p2, and e2 represents the epipole; P′ represents the point after point P moves, p2′ represents the projection point of P′ on the current frame, and d e represents the distance from p2′ to the epipolar line l2, represents the transformation matrix from the previous frame to the current frame, and O1O2 represents the baseline. The epipolar geometry constraint expression is as follows:
[0026] E = t^R (1)
[0027] F = K -T EK -1 (2)
[0028]
[0029] where, T represents the transpose of the matrix, K represents the internal parameter matrix of the camera, R represents the rotation matrix from the previous frame to the current frame, t represents the translation vector from the previous frame to the current frame, E represents the essential matrix, and F represents the fundamental matrix. According to the equation of the distance from a point to a line, the calculation formula of d e is as follows:
[0030]
[0031] Further, in the step S3, the process of extracting the scene flow constraint information is specifically as follows:
[0032] According to the pinhole imaging model of the camera, the rigid body motion transformation equation and the transformation equation between the world coordinate system and the camera coordinate system, combined with the depth map information of the current frame, the coordinates of a feature point corresponding to a map point in the world coordinate system are obtained.
[0033] Assume that the prior static feature point sets of the current frame and the previous frame have completed data association. Using this association information and the mapping equation from pixel points to map points, the displacement vector of the map points in the world coordinate system can be obtained, and this displacement vector is the scene flow.
[0034] Furthermore, in the step S3, the lightweight frame tracking process based on the ICP algorithm reduces the initialization time of the inter-frame transformation matrix, specifically as follows:
[0035] Assume there is a 3D point set in the camera coordinate system of the current frame There is a 3D point set in the camera coordinate system of the previous frame Where And Represents a 3D point in the camera coordinate system. Assume the rotation matrix from the previous frame to the current frame is R cl , and the translation vector is t cl . Assume that the above-mentioned 3D point sets have all undergone data association, and it is preset that Let Then the rotation matrix and the translation vector will be given by the following formula:
[0036]
[0037] t cl = p c - R cl p l (6)
[0038] Where Represents the i-th point of the 3D point set P c , Represents the i-th point of the 3D point set P l , and n represents the number of points in the 3D point set.
[0039] Solve Equation (5) using the singular value decomposition algorithm of the matrix. Define a 3×3 matrix W, and its calculation formula is as follows:
[0040]
[0041] Where Q, M, and ∑ represent diagonal matrices, and the matrix W is full rank.
[0042] Then the solution of Equation (5) is as follows:
[0043] R cl = QM T (8)
[0044] Substitute Equation (8) into Equation (6) to obtain the calculation formula for the translation vector:
[0045]
[0046] Advantages of the present invention: The method of the present invention uses a semantic segmentation neural network to extract the prior moving object regions in the scene, generates corresponding masks and performs refined processing, extracts the outer contour of the refined semantic mask, traverses all feature points in the current frame, calculates their Euclidean distances to the outer contour and the corresponding depth data, divides the current frame feature point set into three categories, then inputs the prior static feature point sets of the current frame and the previous frame into the lightweight frame tracking process in the motion consistency detection process to obtain the inter-frame transformation matrix, and then uses this transformation matrix and geometric constraints to act on the prior unknown feature point sets of the current frame and the previous frame to obtain the static feature point set in the current frame, and inputs it into the tracking thread for calculation to obtain a more accurate pose estimate. The method of the present invention extracts the prior motion region information in the scene through a semantic neural network to ensure that the system can obtain sufficient prior information; through morphological operations on the image: erosion and dilation operations, effectively filters out the incorrect parts in the semantic information and ensures the accuracy of the prior information; by extracting the outer contour information of the refined semantic mask, more accurately extracts the prior information according to the Euclidean distance from each feature point in the current frame to the contour and the depth information at the corresponding position in the depth map; uses the epipolar geometry constraint to measure the possibility of feature points coming from moving objects in real time, greatly improving the robustness of the geometric constraint; improves the stability and accuracy of the visual SLAM system when working in an indoor environment with moving objects. Description of the Drawings
[0047] Figure 1 It is a flowchart of a method for improving the robustness of visual SLAM applied to a dynamic indoor scene according to the present invention.
[0048] Figure 2 It is a schematic diagram of the epipolar geometry constraint in the embodiment of the method of the present invention.
[0049] Figure 3 It is a diagram of the prior information extraction result in the embodiment of the present invention.
[0050] Figure 4 It is a comparison diagram of the test results of the visual SLAM system with and without applying the present invention in a dynamic indoor environment in the embodiment of the present invention. Detailed Embodiments
[0051] The following further describes the content of the present invention in conjunction with the drawings and embodiments.
[0052] As Figure 1 shown, a flowchart of a method for improving the robustness of visual SLAM applied to a dynamic indoor scene according to the present invention is as follows:
[0053] S1. Use a semantic segmentation neural network to extract the prior moving object region in the scene, generate a corresponding mask according to the semantic segmentation result, and then use the morphological operations of the image: erosion and dilation, to refine the mask, eliminate incorrect segmentation results, and extract the outer contour of the refined semantic mask;
[0054] S2. Based on the outer contour of the semantic mask extracted in step S1, traverse all feature points in the current frame to calculate the Euclidean distance from each feature point to the outer contour of the semantic mask, and divide the feature points in the current frame into three sets according to this distance combined with the depth map information corresponding to the current frame: prior dynamic feature points, prior unknown feature points, and prior static feature points;
[0055] S3. Input the prior static feature point sets of the current frame and the previous frame into the lightweight frame tracking process in the motion consistency detection process to obtain the inter-frame transformation matrix, and use the obtained inter-frame transformation matrix and the geometric constraints in the motion consistency detection process to act on the set of prior unknown feature points with good data association in the current frame and the previous frame to obtain the epipolar geometric constraint and scene flow constraint data;
[0056] S4. Combine the data information of the epipolar geometric constraint and the scene flow constraint into a two-dimensional vector, and then use the two-norm of this vector to represent the final error value. If the error value is greater than the set threshold, then the feature point belongs to the moving feature point set, otherwise it belongs to the static feature point set. Then use the obtained static feature point set for accurate frame tracking, and finally obtain the pose estimation data that eliminates the interference of moving objects in the scene.
[0057] In this embodiment, in step S1, specifically as follows:
[0058] As Figure 1 shown, when the visual SLAM system starts to run, the system obtains the original data through an RGB-D camera, including the RGB image and the corresponding depth map information.
[0059] First, input the RGB image into the semantic segmentation neural network. After being processed by the neural network, a semantic information image is obtained. Then, input the semantic information image into the mask processing process for grayscale and binary processing, and finally extract the outer contour information of the semantic mask.
[0060] In this embodiment, in step S1, since there are small regions with incorrect segmentation in the semantic information image extracted by the semantic segmentation neural network, the incorrect regions are eliminated by using morphological operations of the image, specifically as follows:
[0061] The semantic mask is processed in three stages:
[0062] (1) The first erosion operation: In order to better erode the complex curve area, an elliptical structure template is selected to eliminate the prior motion area caused by errors. The first erosion result is AΘB1;
[0063] (2) Dilation operation, also using an elliptical structure template, fully expands the prior motion area based on the first operation to obtain an over-expanded mask. The result of the first dilation is
[0064] (3) The second erosion operation uses the template of the dilation operation. Based on the second operation, the mask is appropriately eroded to reduce the over-expanded mask and obtain a refined semantic mask with a moderate mask size. The result of the second erosion is: The purpose is to keep the area size of the original mask as much as possible.
[0065] Where A represents the mask to be filtered, Θ represents the corrosion operator, represents the expansion operator, B1 represents the structural template of the first corrosion operation, and B2 represents the structural template of the expansion operation, which is also an elliptical structure.
[0066] In this embodiment, in step S3, the motion consistency detection process includes a lightweight frame tracking process based on an iterative closest point (ICP) algorithm and an epipolar geometry constraint and scene flow constraint data extraction process, which are specifically as follows:
[0067] Using the data association information of the prior stationary feature point set in the current frame and the previous frame, combined with the iterative closest point (ICP) algorithm and the rigid body motion transformation equation, the relationship between the inter-frame transformation matrix from the previous frame to the current frame and the data association information is constructed, and a least squares problem is established. Then, the singular value decomposition (SVD) algorithm of the matrix is used to solve the solution of the least squares problem, thereby obtaining the inter-frame transformation matrix from the previous frame to the current frame. Then, the obtained inter-frame transformation matrix is used to act on the prior unknown feature point set with good data association in the current frame and the previous frame to obtain the epipolar geometry constraints and scene flow constraint data.
[0068] In this embodiment, in step S3, epipolar geometric constraint data is extracted, and epipolar geometric constraints are used to measure the possibility that feature points come from moving objects in real time, thereby improving the robustness of geometric constraints, as follows:
[0069] like Figure 2As shown in the figure, P is a moving map point in the scene; I1 represents the previous frame image, p1 represents the pixel point of point P in the previous frame, l1 represents the epipolar line passing through point p1, and e1 represents the epipole; I2 represents the current frame image, p2 represents the pixel point of point P in the current frame, l2 represents the epipolar line passing through point p2, and e2 represents the epipole; P′ represents the point after the movement of point P, p2′ represents the projection point of P′ in the current frame, and d e represents the distance from p2′ to the epipolar line l2, represents the transformation matrix from the previous frame to the current frame, and O1O2 represents the baseline. The epipolar geometry constraint expressions are as follows:
[0070] E = t^R (10)
[0071] F = K-TEK -1 (11)
[0072]
[0073] where T represents the transpose of the matrix, K represents the internal parameter matrix of the camera, R represents the rotation matrix from the previous frame to the current frame, t represents the translation vector from the previous frame to the current frame, E represents the essential matrix, and F represents the fundamental matrix. According to the equation for the distance from a point to a line, d e can be calculated as follows:
[0074]
[0075] In this embodiment, in the step S3, the process of extracting the scene flow constraint information is specifically as follows:
[0076] According to the pinhole imaging model of the camera, the rigid body motion transformation equation and the transformation equation between the world coordinate system and the camera coordinate system, combined with the depth map information of the current frame, the coordinates of a feature point corresponding to the map point in the world coordinate system are obtained.
[0077] Assume that the prior static feature point sets of the current frame and the previous frame have completed data association. Using this association information and the mapping equation from the pixel point to the map point, the displacement vector of the map point in the world coordinate system can be obtained, and this displacement vector is the scene flow.
[0078] In this embodiment, in the step S3, the lightweight frame tracking process based on the ICP algorithm reduces the initialization time of the inter-frame transformation matrix, specifically as follows:
[0079] Assume that there is a set of 3D points in the camera coordinate system of the current frame There is a set of 3D points in the camera coordinate system of the previous frame where and Represents a 3D point in the camera coordinate system. Assume that the rotation matrix from the previous frame to the current frame is R cl , and the translation vector is t cl . Assume that the above-mentioned 3D point sets have all undergone data association and are preset Let Then the rotation matrix and the translation vector will be given by the following formula:
[0080]
[0081] t cl = p c - R cl p l (15)
[0082] Where represents the i-th point of the 3D point set P c , represents the i-th point of the 3D point set P l , and n represents the number of points in the 3D point set.
[0083] Use the singular value decomposition algorithm of the matrix to solve Equation (5). Define a 3×3 matrix W, and its calculation formula is as follows:
[0084]
[0085] Where Q, M, and ∑ represent diagonal matrices, and the matrix W is full rank.
[0086] Then the solution of Equation (5) is as follows:
[0087] R cl = QM T (17)
[0088] Substitute Equation (8) into Equation (6) to obtain the calculation formula for the translation vector:
[0089]
[0090] As Figure 3 shown, it is the result of prior information extraction. Among them, the result of the original distribution of feature points is the result of not dividing the current frame feature point set. In the image where the prior information is obtained, it can be seen that the remaining feature point set is the union of the feature point set on the person and the feature point set within a certain range near the person. The feature point set from the person is the prior dynamic feature point set, and the feature point set near the person is the prior unknown feature point set.
[0091] As Figure 4As shown, the test results of the method of the present invention in a dynamic indoor scene. The dotted trajectory is the actual movement trajectory of the camera collected by the motion capture device, and the solid trajectory is the trajectory estimated by the visual SLAM system. By comparing the test results in the dynamic scene with and without using the solution proposed by the present invention, the experiment shows that the method of the present invention can significantly improve the stability of visual SLAM in a dynamic indoor scene.
[0092] Those of ordinary skill in the art will realize that the embodiments described herein are for helping the reader understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. For those skilled in the art, various changes and modifications can be made to the present invention. Any modifications, equivalent replacements, improvements, etc., made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.
Claims
1. A method for enhancing the robustness of visual SLAM applied to dynamic indoor scenes, the specific steps are as follows: S1. Use a semantic segmentation neural network to extract the prior moving object regions in the scene, generate corresponding masks according to the semantic segmentation results, and then use morphological operations on the images: erosion and dilation, to refine the masks, eliminate incorrect segmentation results, and extract the outer contours of the refined semantic masks; S2. Based on the outer contours of the semantic masks extracted in step S1, traverse all the feature points in the current frame to calculate the Euclidean distance from each feature point to the outer contours of the semantic masks, and divide the feature points in the current frame into three sets according to this distance combined with the depth map information corresponding to the current frame: prior dynamic feature point set, prior unknown feature point set, and prior static feature point set; S3. Input the prior static feature point sets of the current frame and the previous frame into the lightweight frame tracking process in the motion consistency detection process to obtain the inter-frame transformation matrix, and use the obtained inter-frame transformation matrix and the geometric constraints in the motion consistency detection process on the prior unknown feature point sets with good data association in the current frame and the previous frame to obtain the epipolar geometric constraint and scene flow constraint data; In step S3, the motion consistency detection process includes a lightweight frame tracking process based on the iterative closest point algorithm and an epipolar geometric constraint and scene flow constraint data extraction process; In step S3, the lightweight frame tracking process based on the ICP algorithm reduces the initialization time of the inter-frame transformation matrix, specifically as follows: Assume there is a 3D point set in the camera coordinate system in the current frame There is a 3D point set in the camera coordinate system in the previous frame where and represent a 3D point in the camera coordinate system; assume the rotation matrix from the previous frame to the current frame is R cl , and the translation vector is t cl ; assume that the above-mentioned 3D point sets have all undergone data association and are preset Let Then the rotation matrix and the translation vector will be given by the following formula: t cl = p c - R cl p l (2) Among them, represents the i-th point of the 3D point set P c and l also represents the i-th point of the 3D point set P, where n represents the number of points in the 3D point set; Use the singular value decomposition algorithm of the matrix to solve equation (1), define a 3×3 matrix W, and its calculation formula is as follows: Where Q, M, and ∑ represent diagonal matrices, the matrix W is full rank, and T represents the transpose operation of the matrix; Then the solution of equation (1) is as follows: R cl = QM T (4) Substitute equation (4) into equation (2) to obtain the calculation formula for the translation vector: S4. Combine the data information of the epipolar geometric constraint and the scene flow constraint into a two-dimensional vector, and then use the two-norm of the vector to represent the final error value. If the error value is greater than the set threshold, then the feature point belongs to the moving feature point set, otherwise it belongs to the static feature point set. Then use the obtained static feature point set for accurate frame tracking, and finally obtain the pose estimation data that eliminates the interference of moving objects in the scene.
2. A method for improving the robustness of visual SLAM applied to dynamic indoor scenes according to claim 1, characterized in that, In step S1, specifically as follows: When the visual SLAM system starts running, the system obtains the original data through an RGB-D camera, including RGB images and corresponding depth map information; First, input the RGB image into the semantic segmentation neural network, and after neural network processing, obtain the semantic information image. Then input the semantic information image into the mask processing process for grayscale and binary processing, and finally extract the outer contour information of the semantic mask.
3. A method for improving the robustness of visual SLAM applied to dynamic indoor scenes according to claim 2, characterized in that, In step S1, use morphological operations on the images to eliminate the incorrect regions, specifically as follows: Process the semantic mask in three stages: (1) The first erosion operation, select an elliptical structure template to eliminate the prior moving regions generated by errors, and the result of the first erosion is AΘB1; (2)Dilation operation, also using an elliptical template, fully expands the prior motion region based on the first operation to obtain an over-expanded mask, and the result of the first dilation is (3) Second corrosion operation: Select the template for dilation operation, and appropriately corrode the mask based on the second operation to reduce the overly expanded mask, obtaining a refined semantic mask with a moderate mask size. The result of the second corrosion is Where A represents the mask to be filtered, Θ represents the erosion operator, represents the dilation operator, B1 represents the structural template for the first erosion operation, and B2 represents the structural template for the dilation operation, which is also an elliptical structure.
4. A method for improving the robustness of visual SLAM applied to dynamic indoor scenes according to claim 1, characterized in that, In the step S3, the motion consistency detection process includes a lightweight frame tracking process based on the iterative closest point algorithm and a process of extracting epipolar geometry constraint and scene flow constraint data, which are specifically as follows: Using the data association information of the prior static feature point sets in the current frame and the previous frame, combining the iterative closest point algorithm and the rigid body motion transformation equation, constructing the relationship between the inter-frame transformation matrix from the previous frame to the current frame and the data association information, establishing a least squares problem, and then using the singular value decomposition algorithm of the matrix to solve the solution of the least squares problem, so as to obtain the inter-frame transformation matrix from the previous frame to the current frame. Then, using the obtained inter-frame transformation matrix to act on the prior unknown feature point sets with good data association in the current frame and the previous frame, the epipolar geometry constraint and scene flow constraint data are obtained.
5. A method for improving the robustness of visual SLAM applied to dynamic indoor scenes according to claim 4, characterized in that, In the step S3, the process of extracting epipolar geometry constraint data is specifically as follows: P is a moving map point in the scene; I1 represents the previous frame image, p1 represents the pixel point of point P on the previous frame, l1 represents the epipolar line passing through point p1, and e1 represents the epipole; I2 represents the current frame image, p2 represents the pixel point of point P on the current frame, l2 represents the epipolar line passing through point p2, and e2 represents the epipole; P′ represents the point after point P moves, p2′ represents the projection point of P′ on the current frame, and d e represents the distance from p2′ to the epipolar line l2, represents the transformation matrix from the previous frame to the current frame, and O1O2 represents the baseline; the epipolar geometry constraint expression is as follows: where T represents the transpose operation of a matrix, K represents the intrinsic matrix of the camera, R represents the rotation matrix from the previous frame to the current frame, t represents the translation vector from the previous frame to the current frame, E represents the essential matrix, and F represents the fundamental matrix; the formula for d can be obtained according to the equation of the distance from a point to a line e The calculation formula is as follows:
6. The method for enhancing the robustness of visual SLAM applied to dynamic indoor scenes according to claim 4, wherein In the step S3, the process of extracting scene flow constraint information is specifically as follows: According to the pinhole imaging model of the camera, the rigid body motion transformation equation, and the transformation equation between the world coordinate system and the camera coordinate system, combined with the depth map information of the current frame, the coordinates of a feature point corresponding to the map point in the world coordinate system are obtained; Assuming that the prior static feature point sets of the current frame and the previous frame have completed data association, using this association information and the mapping equation from the pixel point to the map point, the displacement vector of the map point in the world coordinate system can be obtained, and this displacement vector is the scene flow.