Dynamic scene SLAM (Simultaneous Localization and Mapping) method based on look-around camera

Through the dynamic scene SLAM method of surround-looking cameras, dynamic object segmentation and tracking is used using neural networks and BEV detectors, combined with ellipsoid modeling, the robustness problem of visual SLAM system in dynamic environment is solved, and efficient positioning and navigation planning is achieved.

CN120259668APending Publication Date: 2025-07-04CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510415948.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing visual SLAM system is not robust enough in dynamic environments, dynamic object segmentation is not accurate, and static constraint information is not fully utilized, especially semi-static objects are not processed sufficiently.

Method used

The dynamic scene SLAM method based on the surround view camera is adopted to perform semantic segmentation and detection through neural networks, and target tracking is performed using BEV detector and IMM algorithm. Combined with ellipsoid modeling and visual odometer optimization, a global map of static map points and semantic ellipsoids are output.

Benefits of technology

It improves the positioning robustness of the system in a dynamic environment, builds a semantic map with advanced navigation planning, and enhances the interactive ability of the agent in the environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259668A_ABST
    Figure CN120259668A_ABST
Patent Text Reader

Abstract

The invention relates to a dynamic scene SLAM method based on a look-around camera, and belongs to the field of automatic driving positioning mapping, and the method comprises the steps: carrying out the semantic segmentation and detection of six images inputted by the look-around camera through a neural network; calculating an initial pose of the current frame by using the front view image; detecting all potential moving objects by using a BEV detector to obtain a three-dimensional detection frame; performing target tracking on the detected three-dimensional detection frame by using 2D detection enhanced data association based on an IMM algorithm, and performing dynamic and static attribute judgment by using a tracking result; obtaining dynamic and semi-static object masks on all the images according to a dynamic and static attribute judgment result, removing dynamic objects according to the dynamic masks, and executing SLAM (Simultaneous Localization and Mapping) based on feature points; for a semi-static object, an ellipsoid is used for modeling; a visual odometer is optimized by observing an ellipsoid, and a global map with static map points and a semantic ellipsoid is output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous driving positioning and mapping, and relates to a dynamic scene SLAM method based on a surround-view camera. Background Art

[0002] In recent years, the technologies of mobile robots and intelligent vehicles have developed rapidly. Simultaneous Localization and Mapping (SLAM) is a key technology for intelligent vehicles and autonomous mobile robots. SLAM is a technology for real-time map construction and simultaneous positioning in an unknown environment, which was proposed in the early 1990s and is mainly applied to the military field and UAV navigation. Thanks to the rapid development of computer vision and deep learning, SLAM systems based on visual sensors have gradually become standard equipment for various mobile devices, UAVs, and robots.

[0003] However, the existing visual SLAM positioning methods have the following problems:

[0004] (1) The vast majority of traditional SLAM technical solutions are based on the assumption of a static environment. However, the presence of dynamic objects in actual application scenarios is inevitable, which makes traditional SLAM systems unable to work robustly in dynamic environments.

[0005] (2) Existing methods for removing dynamic objects such as geometric constraints and semantic segmentation are not accurate enough in segmenting dynamic objects.

[0006] (3) The static constraint information in the environment is not fully utilized, especially many SLAM systems do not handle semi-static objects sufficiently. Summary of the Invention

[0007] In view of this, the purpose of the present invention is to provide a dynamic scene SLAM method based on a surround-view camera, aiming to improve the robustness of the SLAM system in a dynamic environment and output a semantic map for high-order navigation planning.

[0008] To achieve the above object, the present invention provides the following technical solutions:

[0009] A dynamic scene SLAM method based on a surround-view camera, comprising the following steps:

[0010] S1: Performing semantic segmentation and detection on six images input by the surround-view camera through a neural network;

[0011] S2: Calculating the initial pose of the current frame using the front-view image;

[0012] S3: Detecting all potential moving objects using a BEV detector to obtain three-dimensional detection boxes;

[0013] S4: For the detected 3D detection boxes, use data association with enhanced 2D detection, perform target tracking based on the IMM algorithm, and use the tracking results to judge the static / dynamic attributes;

[0014] S5: According to the results of the static / dynamic attribute judgment, obtain the masks of dynamic and semi-static objects on all images, remove the dynamic objects according to the dynamic mask, and execute feature point-based SLAM;

[0015] S6: For semi-static objects, use an ellipsoid for modeling;

[0016] S7: Use the observations of the ellipsoid to optimize the visual odometer and output a global map with static map points and semantic ellipsoids.

[0017] Furthermore, the step S1 includes:

[0018] Perform semantic segmentation and detection on the six images input by the surround-view camera through a neural network to obtain the mask of potential moving objects and 2D detection boxes, and record the center coordinates (u 2c , v 2c ) of the 2D detection boxes; the potential moving objects refer to the objects with pre-set motion attributes.

[0019] Furthermore, the step S2 includes:

[0020] Use the front-view image among the six surround-view images to extract feature points. According to the guidance of the mask of potential moving objects, only retain the feature points of the static background part, and use the retained feature points to calculate the initial pose.

[0021] Furthermore, the step S3 includes:

[0022] Input the six surround-view images to obtain the detection results of potential moving targets. The result of a single detection is expressed as (P, D, θ c ), where P represents the coordinates (x c , y c , z c ) of the center point of the 3D box in the camera coordinate system, D represents the length, width, and height (l, w, h) of the 3D box, and θ c represents its rotation around the camera y-axis in the camera coordinate system;

[0023] Use the camera internal parameter K to project P onto the image to obtain the coordinates (u 3c , v 3c ) of the center point of the 3D box on the image.

[0024] Furthermore, the step S4 includes:

[0025] S41: For each 3D box of target detection, calculate the following two parameters of it on the image:

[0026] (1) Coordinates of the center point obtained by projection (u 3c , v 3c );

[0027] (2) The maximum bounding rectangle of the 3D detection box on the image (u max , u min , v max , v min );

[0028] S42: Calculate the distance between the center (u 2c , v 2c ) of the 2D detection box and the center (u 3c , v 3c ) of the 3D box projection, and normalize it to a score f d ; Calculate the IOU between the 2D detection box and the bounding box of the 3D detection box on the image, denoted as f iou ;

[0029] S43: Through semantic consistency and the maximum f d , f iou , obtain the one-to-one correspondence between the 2D detection result and the 3D detection result;

[0030] S44: On two adjacent frames of images, perform data association for each detection target in the current frame. The association basis is semantic consistency, IOU, and optical flow tracking, and parameterize the results of IOU and optical flow tracking, normalize them to a score f, and obtain the 3D box corresponding to each f through the method in step S43;

[0031] S45: Use the Hungarian algorithm for data association of 3D boxes, where the weight of the current 3D detection result is determined by f and its shape and IOU in 3D space;

[0032] S46: According to the result of data association, use the extended Kalman filter and the interactive multiple model algorithm to track the 3D detection result. Each tracked object may have three motion states: stationary, uniform velocity, and uniform rotational speed. The three motion models are as follows:

[0033]

[0034] S47: According to the tracking result, compare the object states at times t - 1 and t, including the difference in the center position and velocity, and judge the static and dynamic attributes of the object according to the magnitude of the difference.

[0035] Furthermore, the step S5 includes:

[0036] Based on the static / dynamic determination result of step S47 and the association method of step S43, obtain the motion attribute of the current image semantic mask, extract feature points based on static and semi-static objects to obtain a visual odometer; remove dynamic objects.

[0037] Further, the step S6 includes:

[0038] Use the result of 3D detection to initialize the parameters of the ellipsoid, and use the pose T of the camera wc Transfer the 3D detection result to the global coordinate system to obtain the parameters (x w , y w , z w , 0, 0, θ z , l1, l2, l3), and then use these parameters to initialize the pose of the ellipsoid and the lengths of the three semi-axes, where the rotation of the x and y axes is initialized to 0.

[0039] Further, the step S7 includes:

[0040] Simultaneously optimize the ellipsoid observed in six cameras and the feature point map obtained in the front view camera; the feature point residual is:

[0041] Δ x = u k - π k (Tc k c k-1 Tc k-1 x)

[0042] where π is the observation model of the camera, u is the observation of the camera on the map point, and Tc k c k-1 is the pose transformation from k - 1 to k;

[0043] For the ellipsoid, through the extrinsic parameter transformation, use the ellipsoids observed in all six camera images to construct a joint residual, and the residual of a single ellipsoid is:

[0044]

[0045] where β is the observation model of the ellipsoid and b is the observation of the camera on the ellipsoid.

[0046] The beneficial effects of the present invention are as follows:

[0047] (1) Improve positioning robustness: Based on the BEV detector and the 3D - 2D data association method, the present invention realizes the accurate segmentation of dynamic objects, ensuring the static environment assumption; and based on the surround-view camera, compared with traditional monocular and binocular visual SLAM systems, it increases the observations in a dynamic environment, greatly enhancing the positioning effect of the system in a dynamic environment.

[0048] (2) Construct a semantic map for high-order navigation and planning: Compared with constructing a sparse point map, the present invention can perform static semantic reconstruction on the environment observed by the surround-view camera, enabling the map to not only have a relocalization function but also play a role in high-order navigation and planning, enhancing the interaction ability of the intelligent agent in the environment.

[0049] Other advantages, objectives, and features of the present invention will to some extent be described in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. Brief Description of the Drawings

[0050] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:

[0051] Figure 1 is a flowchart of the dynamic scene SLAM method based on the surround-view camera according to the present invention;

[0052] Figure 2 is a schematic diagram of the association between 3D detection and 2D detection;

[0053] Figure 3 is a schematic diagram of the construction of the pose solution residual. Detailed Embodiments

[0054] The following illustrates the embodiments of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention schematically. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0055] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention schematically. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0056] In the following description, numerous specific details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the embodiments of the present invention.

[0057] Please refer to Figures 1-3 , a dynamic scene SLAM method based on a surround-view camera, the method comprising the following steps:

[0058] Embodiment 1:

[0059] A dynamic scene SLAM method based on a surround-view camera, comprising the following steps:

[0060] S1. Perform semantic segmentation and detection on six images input by the surround-view camera through a neural network.

[0061] S2. Calculate the initial pose of the current frame using the front-view image.

[0062] S3. Detect all potential moving objects using a BEV detector to obtain 3D detection boxes.

[0063] S4. For the detected 3D detection boxes, use data association enhanced by 2D detection, perform target tracking based on the IMM algorithm, and use the tracking results to judge the static and dynamic attributes.

[0064] S5. According to the results of the static and dynamic attribute judgment, obtain the masks of dynamic and semi-static objects on all images, remove the dynamic objects according to the dynamic mask, and perform feature-point-based SLAM.

[0065] S6. For semi-static objects, use an ellipsoid for modeling.

[0066] S7. Use the observations of the ellipsoid to optimize the visual odometer and output a global map with static map points and semantic ellipsoids.

[0067] Embodiment 2:

[0068] In this embodiment, step S1 includes: performing semantic segmentation and detection on six images input by the surround-view camera through a neural network. Obtain the mask and 2D detection box of potential moving objects, and record the center coordinates (u 2c , v 2c ) of the 2D detection box. Potential moving objects refer to objects with pre-set motion attributes, such as people and vehicles.

[0069] Step S2 includes: using the front view image among the six panoramic images to extract feature points, and according to the guidance of the mask of potential moving objects, only retaining the feature points of the static background part, and calculating the initial pose using the retained feature points.

[0070] Step S3 includes: inputting the six panoramic images to obtain the detection results of potential moving targets, and the result of a single detection is expressed as (P, D, θ c ) where P represents the coordinates (x c , y c , z c ) of the center point of the 3D box in the camera coordinates, D represents the length, width, and height (l, w, h) of the 3D box, and θ c represents its rotation around the camera y-axis in the camera coordinates; further, using the camera internal parameter K to project P onto the image to obtain the coordinates (u 3c , v 3c ) of the center point of the 3D box on the image.

[0071] Step S4 includes:

[0072] S41. For each 3D box of the target detection, calculate two parameters on the image, one is the projected center point coordinates (u 3c , v 3c ); the other is the maximum bounding rectangle box (u max , u min , v max , v min ) of its 3D detection box on the image.

[0073] S42. Calculate the distance between the center (u 2c , v 2c ) of the 2D detection box and the projected center (u 3c , v 3c ) of the 3D box, and normalize it to a score f d ; calculate the IOU of the bounding boxes of the 2D detection box and the 3D detection box on the image, denoted as f iou .

[0074] S43. Through semantic consistency and the maximum f d , f iou , obtain the one-to-one correspondence between the 2D detection result and the 3D detection result.

[0075] S44. On adjacent two frames of images, perform data association for each detection target in the current frame. The association basis is semantic consistency, IOU, and optical flow tracking, and parameterize the results of IOU and optical flow tracking, normalize them to a score f, and through the method of S43, obtain the 3D box corresponding to each f.

[0076] S45. Use the Hungarian algorithm for data association of 3D bounding boxes, where the weight of the current 3D detection result is determined by f, its shape in 3D space, and IOU weighting.

[0077] S46. According to the result of data association, use the extended Kalman filter and the interactive multiple model algorithm to track the 3D detection result. Each tracked object may have three motion states: stationary, uniform velocity, and uniform rotational speed. The three motion models are represented as follows:

[0078]

[0079] S47. According to the tracking result, compare the object states at times t - 1 and t, including the differences in the center position and velocity, and judge the static and dynamic attributes of the object based on the magnitude of the differences.

[0080] Step S5 includes:

[0081] According to the static and dynamic judgment result of S47 and the association method of S43, obtain the motion attribute of the current image semantic mask, extract feature points based on static and semi-static objects to obtain the visual odometer; directly remove dynamic objects.

[0082] Step S6 includes:

[0083] Use the result of 3D detection to initialize the parameters of the ellipsoid, and use the pose T of the camera wc Transfer the 3D detection result to the global coordinate system to obtain (x w , y w , z w , 0, 0, θ z , l1, l2, l3), and this parameter directly initializes the pose of the ellipsoid and the lengths of the three semi-axes, where the rotation of the x and y axes is initialized to 0.

[0084] Step S7 includes:

[0085] Simultaneously optimize the ellipsoid observed in the six cameras and the feature point map obtained in the front view camera. The feature point residual is Δ x = u k - π k (Tc k c k-1 Tc k-1 x), where π is the camera observation model, u is the camera observation of the map point, and Tc k c k-1 is the pose transformation from k - 1 to k; for the ellipsoid, through the extrinsic parameter transformation, use the ellipsoids observed in all six camera images to construct the joint residual, and the residual of a single ellipsoid is where β is the ellipsoid observation model and b is the camera observation of the ellipsoid.

[0086] In the above embodiments, the reference in the specification to "this embodiment" means that the specific features, structures, or characteristics described in connection with the embodiment are included in at least some embodiments, but not necessarily all embodiments. Multiple occurrences of "this embodiment" do not necessarily all refer to the same embodiment.

[0087] In the above embodiments, although the present invention has been described in connection with specific embodiments of the present invention, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art in light of the foregoing description. For example, other storage structures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed. Embodiments of the present invention are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims.

[0088] This embodiment also provides a computer-readable storage medium having stored thereon a computer program, which when executed by a processor implements any one of the methods in this embodiment.

[0089] This embodiment also provides an electronic terminal, including: a processor and a memory;

[0090] The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory so that the terminal executes any one of the methods in this embodiment.

[0091] For the computer-readable storage medium in this embodiment, those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to a computer program. The foregoing computer program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disk, or optical disc that can store program code.

[0092] The electronic terminal provided in this embodiment includes a processor, a memory, a transceiver, and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication therebetween. The memory is used to store a computer program, and the communication interface is used for communication. The processor and the transceiver are used to run the computer program so that the electronic terminal executes each step of the above method.

[0093] In this embodiment, the memory may include a random access memory (Random Access Memory, abbreviated as RAM), and may also include a non-volatile memory, such as at least one disk memory.

[0094] The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0095] The present invention can be used in numerous general-purpose or special-purpose computing system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on.

[0096] The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A dynamic scene SLAM method based on surround cameras, characterized in that: Including the following steps: S1: Perform semantic segmentation and detection on six images input by the surround-view camera through a neural network; S2: Calculate the initial pose of the current frame using the front-view image; S3: Use a BEV detector to detect all potential moving objects and obtain 3D detection boxes; S4: For the detected 3D detection boxes, use data association enhanced by 2D detection, perform target tracking based on the IMM algorithm, and use the tracking results to judge the static and dynamic attributes; S5: According to the results of the static and dynamic attribute judgment, obtain the masks of dynamic and semi-static objects on all images, remove dynamic objects according to the dynamic mask, and perform feature-point-based SLAM; S6: For semi-static objects, use an ellipsoid for modeling; S7: Utilize the observations of the ellipsoid to optimize the visual odometer and output a global map with static map points and semantic ellipsoids.

2. The dynamic scene SLAM method based on an omnidirectional camera according to claim 1, wherein: The step S1 includes: Six images input by the surround-view camera are subjected to semantic segmentation and detection through a neural network to obtain a mask of potential moving objects and 2D detection boxes, and the central coordinates (u 2c , v 2c ) of the 2D detection boxes are recorded; the potential moving objects refer to objects with pre-set motion attributes.

3. The method for dynamic scene SLAM based on an omnidirectional camera according to claim 1, wherein: The step S2 includes: Use the front-view image among the six surround-view images to extract feature points. According to the guidance of the mask of potential moving objects, only retain the feature points of the static background part, and use the retained feature points to calculate the initial pose.

4. The method for dynamic scene SLAM based on surround cameras according to claim 1, wherein: The step S3 includes: Six surround-view images are input, and the detection results of potential moving targets are obtained. The result of a single detection is represented as (P, D, θ c ), where P represents the coordinates (x c , y c , z c ) of the center point of the 3D bounding box in the camera coordinates, D represents the length, width, and height (l, w, h) of the 3D bounding box, and θ c represents its rotation around the camera y-axis in the camera coordinates; Project P onto the image using the camera intrinsic parameter K to obtain the coordinates (u 3c , v 3c ) of the center point of the 3D box on the image.

5. The method for dynamic scene SLAM based on surround cameras according to claim 1, wherein: The step S4 includes: S41: For each 3D box of object detection, calculate the following two parameters of it on the image: (1) The center point coordinates (u 3c , v 3c ) obtained by projection; (2) The maximum bounding rectangle of the 3D detection box on the image (u max , u min , v max , v min ); S42: Calculate the distance between the center (u 2c , v 2c ) of the 2D detection box and the projection center (u 3c , v 3c ) of the 3D box, and normalize it to a score f d ; Calculate the IOU of the bounding box of the 2D detection box and the 3D detection box on the image, denoted as f iou ; S43: Obtain the one-to-one correspondence between the 2D detection result and the 3D detection result through semantic consistency and the largest f d and f iou ; S44: On two adjacent frames of images, perform data association for each detection target in the current frame. The association basis is semantic consistency, IOU, and optical flow tracking. Parameterize the results of IOU and optical flow tracking, normalize them to a score f, and obtain the 3D box corresponding to each f through the method of step S43; S45: Use the Hungarian algorithm for data association of 3D boxes, where the weight of the current 3D detection result is determined by weighting f and its shape and IOU in 3D space; S46: According to the results of data association, use the extended Kalman filter and the interactive multiple model algorithm to track the 3D detection results. Each tracked object may have three motion states: stationary, uniform velocity, and uniform rotational speed. The three motion models are as follows: S47: According to the tracking results, compare the object states at times t-1 and t, including the differences in the center position and velocity, and judge the static and dynamic attributes of the object according to the magnitude of the differences.

6. The method for dynamic scene SLAM based on surround cameras according to claim 5, wherein: The step S5 includes: According to the static and dynamic judgment results of step S47 and the association method of step S43, obtain the motion attributes of the semantic mask of the current image, extract feature points based on static and semi-static objects to obtain the visual odometer; remove dynamic objects.

7. The method for dynamic scene SLAM based on an omnidirectional camera according to claim 1, wherein: The step S6 includes: Use the results of 3D detection to initialize the parameters of the ellipsoid, with the pose T of the camera wc Transfer the 3D detection results to the global coordinate system to obtain the parameters (x w , y w , z w , 0, 0, θ z , l1, l2, l3), and then use these parameters to initialize the pose and the lengths of the three semi-axes of the ellipsoid, where the rotations around the x and y axes are initialized to 0.

8. The dynamic scene SLAM method based on an omnidirectional camera according to claim 1, wherein: The step S7 includes: Optimize the ellipsoids observed in the six cameras and the feature point map obtained in the front-view camera simultaneously; the feature point residual is: Δ x = u k - π k (Tc k c k-1 Tc k-1 x) where π is the observation model of the camera, u is the camera's observation of the map point, and Tc k c k-1 is the pose transformation from k - 1 to k; For the ellipsoid, through extrinsic parameter transformation, construct a joint residual using the ellipsoids observed in all six camera images. The residual of a single ellipsoid is: where β is the observation model of the ellipsoid and b is the observation of the ellipsoid by the camera.