Visual inertia SLAM method and system suitable for complex dynamic environment

Through the time communication mechanism and semantic-assisted RANSAC method combined with Sampson geometric information, dynamic feature points are identified and eliminated, visual reprojection residuals are constructed, and positioning is optimized, which solves the positioning accuracy and robustness of the visual inertial SLAM system in complex dynamic environments, and improves the system operation efficiency.

CN120259359AActive Publication Date: 2025-07-04WUHAN UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510391883.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-04
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The existing visual inertial SLAM system is difficult to achieve high-precision pose estimation in complex dynamic environments, and is severely affected by the noise of characteristic points of dynamic objects, and the system operation efficiency is low after the introduction of deep learning technology.

Method used

The time communication mechanism is used to synchronize instance segmentation and feature extraction, combine semantic-assisted RANSAC method and Sampson geometric information to identify dynamic feature points, build IMU pre-integration and marginalized residuals, and optimize the pose through the solver.

Benefits of technology

It improves the accuracy and stability of dynamic feature recognition, improves the operating efficiency and positioning accuracy of the system, and enhances robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259359A_ABST
    Figure CN120259359A_ABST
Patent Text Reader

Abstract

The invention discloses a visual inertia SLAM method and system suitable for a complex dynamic environment, a storage medium and electronic equipment. The method comprises the following steps: acquiring a visual image sequence; the visual image sequence comprises a plurality of visual images; synchronously performing instance segmentation on the visual image by using a time communication mechanism to extract key frames and performing feature extraction and matching on the key frames to obtain feature matching pairs; static exterior point elimination is carried out on the feature matching pairs by using a semantic-assisted RANSAC method, and dynamic feature points are identified in combination with instance information and Sampson geometric information; and constructing an IMU pre-integration residual error and a marginalization residual error, rejecting the identified dynamic feature points, constructing a visual reprojection residual error by using the remaining static feature points, and calculating an optimized pose through a solver. According to the invention, the accuracy and stability of dynamic feature recognition in a complex dynamic scene can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision, and in particular, to a visual-inertial SLAM method, system, storage medium, and electronic device applicable to complex dynamic environments. Background Art

[0002] With the continuous progress in the fields of computer technology, sensor technology, artificial intelligence, Internet of Things, etc., mobile robots can perform autonomous navigation, decision-making, and operations in complex environments, thereby reducing the human workload, improving operation efficiency and safety, and being widely used in multiple fields such as manufacturing, medical care, service, security, and transportation. The market scale continues to grow and has become an important force driving social progress and development.

[0003] Simultaneous Localization and Mapping (SLAM) technology can help robots achieve autonomous localization and environmental reconstruction using multiple sensors in unknown environments and is a key technology for mobile robots to perform tasks such as autonomous navigation and environmental exploration. According to the types of sensors used, SLAM is mainly divided into laser SLAM, visual SLAM, and multi-source SLAM. Laser SLAM uses lidar to obtain point cloud information in the environment to achieve high-precision localization and mapping, but the cost is high. Visual SLAM uses camera sensors to obtain image features and uses computer vision technology to complete pose determination. It has a low cost and is widely used in robot systems. However, monocular visual SLAM has scale uncertainty and it is difficult to accurately recover the pose. An Inertial Measurement Unit (IMU) can obtain high-frequency inertial information of the carrier without being affected by the external environment and can be used to assist visual SLAM. Therefore, a visual-inertial SLAM system based on a visual camera and an IMU has become one of the core technologies for mobile robots to achieve autonomous navigation in complex environments. However, most of the existing visual-inertial SLAM systems are only applicable to static environments.

[0004] In a static scene, the feature points in the environment all remain stationary, satisfying the epipolar geometric constraint relationship in computer vision. The visual camera mounted on the mobile robot uses these static features to construct visual reprojection constraint information and achieves accurate pose estimation through non-linear optimization or filtering. However, in common urban complex scenes, there are a large number of moving objects, such as moving cars, walking pedestrians, etc. The feature points extracted from these dynamic objects no longer satisfy the epipolar geometric constraint relationship. These dynamic features will introduce additional noise and uncertainty, seriously affecting the accuracy and stability of back-end optimization, resulting in deviations in the pose estimation of the mobile robot and making it difficult to maintain the consistency of the environmental map. Therefore, it is difficult for a visual-inertial SLAM system for complex dynamic environments to achieve high-precision pose estimation.

[0005] For dynamic scenes, current SLAM technologies such as VINS-Mono use robust kernel functions to reduce the negative impact of dynamic feature outliers on optimization. However, this method can only reject slow-motion features and is difficult to work for fast-motion features. Some SLAM systems combine artificial intelligence technology and use deep learning methods to obtain high-level instance information in images, and combine traditional geometric information to identify dynamic features. They have strong adaptability in dynamic environments, but these methods do not consider system operation efficiency when introducing deep learning technology, which seriously affects the real-time performance of the system. At the same time, the combination of geometric information and instance information is relatively simple, and it is difficult to ensure the accuracy and stability of dynamic feature recognition.

[0006] In summary, the presence of moving objects in complex dynamic scenes seriously affects the positioning accuracy and robustness of the visual inertial SLAM system. Among the existing dynamic SLAM methods, the traditional geometric method is easily affected by observation noise, while the method that introduces deep learning technology has the problem of simple information utilization and low system operation efficiency. Summary of the invention

[0007] The embodiments of the present application provide a visual-inertial SLAM method, system, storage medium and electronic device suitable for complex dynamic environments, which can improve the accuracy and stability of dynamic feature recognition in complex dynamic scenes.

[0008] The present application embodiment provides a visual inertial SLAM method suitable for complex dynamic environments, including: Acquire a visual image sequence; the visual image sequence includes a plurality of visual images; Using a temporal communication mechanism, synchronously performing instance segmentation on the visual image to extract key frames and performing feature extraction and matching on the key frames to obtain feature matching pairs; Using the semantically assisted RANSAC method to remove static outliers from the feature matching pairs, and combining instance information and Sampson geometry information to identify dynamic feature points; Construct IMU pre-integration residuals and marginalization residuals, remove the identified dynamic feature points, use the remaining static feature points to construct visual reprojection residuals, and calculate the optimized pose through the solver.

[0009] Furthermore, the visual-inertial SLAM method applicable to complex dynamic environments, wherein the use of a time communication mechanism to synchronously perform instance segmentation on the visual image to extract key frames and perform feature extraction and matching on the key frames to obtain feature matching pairs, comprises: In the neural network thread, the time information corresponding to the received visual image is transmitted to the VIO thread, and instance segmentation is performed on the visual image to obtain instance pixels and instance IDs. An instance mask is constructed based on the instance pixels and the instance IDs, and the instance mask is transmitted to the VIO thread; In the VIO thread, according to the time information, the corresponding visual image is extracted from the image sequence. Feature tracking is performed on the feature points at the previous moment by means of optical flow to obtain feature matching pairs. The tracking points that fail to track and those that exceed the boundary are removed, and new corner points are extracted as supplementary feature points; The gray pixels at the corresponding positions in the instance mask are used as instance ID information, and the instance ID information is associated and matched with the corresponding feature points.

[0010] Further, in the above visual-inertial SLAM method applicable to a complex dynamic environment, wherein the constructing an instance mask based on the instance pixels and the instance IDs includes: Construct a grayscale image based on the instance pixels and the corresponding instance IDs; For the overlapping regions of different instance pixels, the instance IDs are overwritten on the grayscale image according to the segmentation order; The grayscale image after overwriting is used as the instance mask.

[0011] Further, in the above visual-inertial SLAM method applicable to a complex dynamic environment, wherein the using the semantic-assisted RANSAC method to remove static outliers from the feature matching pairs includes: Taking the feature matching pairs as the prior point set, and removing the feature points of the feature matching pairs that belong to potential dynamic semantics; According to the instance ID information to which each feature belongs, the feature matching pairs after removing the feature points of potential dynamic semantics are divided into different clusters; Calculate the fundamental matrix of adjacent key frames, remove the outliers in the feature matching pairs based on the fundamental matrix, and then add the previously removed feature points that belong to potential dynamic semantics.

[0012] Further, in the above visual-inertial SLAM method applicable to a complex dynamic environment, wherein identifying dynamic feature points by combining instance information and Sampson geometric information includes: Calculate the epipolar error information of each instance object based on the fundamental matrix, and add the epipolar error information to the geometric error sequence of the instance object; For each instance object, count the proportion of points in the geometric error sequence that are greater than a preset error threshold. If the proportion is greater than 80%, it is determined that the instance object is in a moving state, and all the feature points belonging to the instance object are regarded as dynamic features.

[0013] Furthermore, for the above visual-inertial SLAM method applicable to complex dynamic environments, the calculation formula for the epipolar error information is as follows:

[0014] Where and are the homogeneous coordinates of the pixel coordinates in adjacent key frames and respectively, is the fundamental matrix, and T represents the transpose.

[0015] Furthermore, for the above visual-inertial SLAM method applicable to complex dynamic environments, the construction of the IMU pre-integration residual and the marginalization residual, the elimination of the identified dynamic feature points, the construction of the visual reprojection residual using the remaining static feature points, and the calculation of the optimized pose through the solver include: Calculating the IMU pre-integration residual of adjacent key frames; Calculating the marginalization residual according to the marginalization information of the sliding window; Traversing all feature points, eliminating the identified dynamic features, and constructing the visual reprojection residuals of the starting observation frame and other observation frames using the remaining static features; Using the DogLeg method to solve the residual function in the solver to obtain the optimized pose.

[0016] The embodiments of the present application also provide a visual-inertial SLAM system applicable to complex dynamic environments, including: An acquisition module, configured to acquire a visual image sequence; the visual image sequence includes a plurality of visual images; A key frame elastic regulation module, configured to use the time communication mechanism to synchronously perform instance segmentation on the visual images to extract key frames and perform feature extraction and matching on the key frames to obtain feature matching pairs; A cascaded visual outlier elimination module, configured to use the semantic-assisted RANSAC method to eliminate static outliers from the feature matching pairs, and identify dynamic feature points in combination with instance information and Sampson geometric information; A backend optimization module, configured to construct an IMU pre-integration residual and a marginalization residual, eliminate the identified dynamic feature points, construct a visual reprojection residual using the remaining static feature points, and calculate the optimized pose through the solver.

[0017] The embodiments of the present application also provide a computer-readable storage medium, in which a plurality of instructions are stored, and the instructions are suitable for being loaded by a processor to execute any one of the above visual-inertial SLAM methods applicable to complex dynamic environments.

[0018] An embodiment of the present application also provides an electronic device, including a processor and a memory, the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used for the steps in the visual inertial SLAM method applicable to a complex dynamic environment described in any one of the above.

[0019] The visual inertial SLAM method, system, storage medium and electronic device applicable to a complex dynamic environment provided by the present application. In view of the problem that the operation efficiency decreases after introducing deep learning technology into dynamic SLAM, the present application proposes a key-frame elastic regulation architecture to realize the integrated operation of the instance segmentation network and VIO in a dual-thread manner, which can flexibly adapt to different neural networks, elastically regulate key frames, process multi-level visual observation information, and improve the system operation efficiency and stability. In addition, the present application also proposes a semantic-assisted RANSAC algorithm based on the RANSAC algorithm, which effectively reduces the outlier ratio of the initial point set and improves the static outlier rejection accuracy. At the same time, dynamic features are identified based on the characteristics of instance feature motion consistency and Sampson geometric error information, and the dynamic outlier rejection accuracy is improved. Significantly improve the positioning accuracy and robustness of visual inertial SLAM. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The following will clearly show the technical solutions and other beneficial effects of the present application by describing the specific embodiments of the present application in detail with reference to the drawings.

[0021] Figure 1 It is a flowchart of the visual inertial SLAM method applicable to a complex dynamic environment provided by an embodiment of the present application.

[0022] Figure 2 It is another flowchart of the visual inertial SLAM method applicable to a complex dynamic environment provided by an embodiment of the present application.

[0023] Figure 3 It is a flowchart of synchronously executing instance segmentation and feature extraction provided by an embodiment of the present application.

[0024] Figure 4 It is a flowchart of calculating the optimized pose provided by an embodiment of the present application.

[0025] Figure 5 It is a schematic structural diagram of the visual inertial SLAM system applicable to a complex dynamic environment provided by an embodiment of the present application.

[0026] Figure 6 It is a schematic structural diagram of the electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0028] The embodiments of the present application provide a visual inertial SLAM method, system, storage medium, and electronic device applicable to complex dynamic environments. A visual inertial SLAM system applicable to complex dynamic environments provided by the embodiments of the present application can be integrated into an electronic device, which can be a device such as a terminal, a server, etc. Among them, the terminal can include a tablet computer, a notebook computer, a personal computer (PC), a microprocessing box, or other devices, etc.

[0029] Please refer to Figure 1 And Figure 2 , Figure 1 is a flowchart of the visual inertial SLAM method applicable to complex dynamic environments provided by the embodiments of the present application. Figure 2 is another flowchart of the visual inertial SLAM method applicable to complex dynamic environments provided by the embodiments of the present application. It is applied to an electronic device. The visual inertial SLAM method applicable to complex dynamic environments includes the following steps: S1, obtain a visual image sequence; the visual image sequence includes a number of visual images.

[0030] Obtain a visual image sequence through a camera. The visual image sequence includes a number of visual images and the time information corresponding to each image.

[0031] S2, use a time communication mechanism to synchronously perform instance segmentation on the visual images to extract key frames and perform feature extraction and matching on the key frames to obtain feature matching pairs.

[0032] Specifically, use a time communication mechanism to synchronize instance segmentation key frames and key frame geometric feature extraction, and elastically adjust the key frame interval. Figure 3 is a flowchart of synchronously performing instance segmentation and feature extraction provided by the embodiments of the present application. As Figure 3 shown, the time communication mechanism synchronization is implemented through a Visual Inertial Odometry (VIO) thread.

[0033] In one embodiment, step S2 includes the following steps: S21. In the neural network thread, transmit the time information corresponding to the received visual image to the VIO thread, perform instance segmentation on the visual image to obtain instance pixels and instance IDs, construct an instance mask based on the instance pixels and instance IDs, and transmit the instance mask to the VIO thread.

[0034] Specifically, the MaskRCNN network can be used to perform instance segmentation on the visual image.

[0035] In one embodiment, the performing instance segmentation on the visual image in step S21 to obtain instance pixels and instance IDs specifically includes: S211. Construct a grayscale image based on the instance pixels and the corresponding instance IDs; S212. For the overlapping regions of different instance pixels, perform instance ID coverage on the grayscale image according to the segmentation order; S213. Use the covered grayscale image as the instance mask.

[0036] S22. In the VIO thread, according to the time information, extract the corresponding visual image from the image sequence, perform feature tracking on the feature points at the previous moment by the optical flow method to obtain feature matching pairs, eliminate the tracking points that fail to track and those that exceed the boundary, and extract new Shi-Tomasi corner points as supplementary feature points to meet the requirement of the number of feature points.

[0037] S23. Use the grayscale pixels at the corresponding positions in the instance mask as instance ID information, and perform associated matching between the instance ID information and the corresponding feature points.

[0038] S3. Use the semantic-assisted RANSAC method to eliminate static outliers from the feature matching pairs and identify dynamic feature points.

[0039] In one embodiment, the using the semantic-assisted RANSAC method to eliminate static outliers from the feature matching pairs in step S3 includes the following steps: S31. Use the feature matching pairs as the prior point set, and eliminate the feature points in the feature matching pairs that belong to potential dynamic semantics.

[0040] Specifically, use the feature matching pairs as the prior point set, and eliminate the feature points that belong to potential dynamic semantics therein, such as vehicles, pedestrians, etc., to reduce the outlier ratio of the prior point set.

[0041] S32. According to the instance ID information to which each feature belongs, divide the feature matching pairs after eliminating the feature points of potential dynamic semantics into different clusters.

[0042] The motion states of the feature points on an object should be consistent. Therefore, according to the instance ID information to which each feature in the image belongs, divide the features belonging to different instances into different clusters.

[0043] In S33, calculate the fundamental matrix of adjacent key frames, remove the outliers in the feature matching pairs based on the fundamental matrix, and then add the feature points that belonged to potential dynamic semantics and were removed previously.

[0044] Specifically, calculate the camera coordinate system of adjacent key frames and the fundamental matrix of the relative pose . Then, the calculation method of the fundamental matrix is as follows:

[0045] where is the external parameter rotation matrix of the camera and the IMU, is the camera internal parameter matrix, and the fundamental matrix describes the relative pose transformation relationship of visual pixels. and are respectively and the corresponding carrier coordinate systems, is the world coordinate system. The carrier pose corresponding to the previous frame time is , , where is the rotation matrix relative to , is the translation vector relative to . Similarly, the carrier pose corresponding to the current time is , .

[0046] In one embodiment, the identification of dynamic feature points by combining instance information and Sampson geometric information in step S3 includes the following steps: In S34, calculate the epipolar error information of each instance object based on the fundamental matrix, and add the epipolar error information to the geometric error sequence of the instance object.

[0047] where the geometric error is the point error between feature points, and the point error is obtained by calculating the Euclidean distance between two feature points.

[0048] For the feature points belonging to a certain instance object, and the homogeneous coordinates of the pixel coordinates in the adjacent key frames are respectively and . The calculation formula of the Sampson epipolar error information is:

[0049] Among them, is the fundamental matrix, and T represents the transpose.

[0050] S35. For each instance object, count the proportion of points in the geometric error sequence that are greater than a preset error threshold. If the proportion is greater than 80%, determine that the instance object is in a motion state, and regard all feature points belonging to the instance object as dynamic features.

[0051] In one embodiment, the Sampson epipolar error threshold can be 1.0.

[0052] S4. Construct the IMU pre-integration residual and the marginalization residual, remove the identified dynamic feature points, construct the visual reprojection residual using the remaining static feature points, and calculate the optimized pose through a solver.

[0053] In one embodiment, Figure 4 is the flowchart for calculating the optimized pose provided by the embodiment of the present application. As Figure 4 shown, step S4 includes the following steps: S41. Calculate the IMU pre-integration residual between adjacent key frames.

[0054] Specifically, first obtain the angular velocity and acceleration from the IMU, then perform preprocessing. At the first image key frame, initialize the velocity, position, and rotation. Then calculate the median integration for each pair of adjacent IMU data points. The median integration includes the median acceleration and the median angular velocity. Update the velocity, position, and rotation through the median acceleration and the median angular velocity. For two key frames, perform median integration on all IMU data points between them to obtain the change in relative motion (including relative velocity, relative position, and relative rotation). The change in relative motion is the IMU pre-integration between adjacent key frames. Calculate the IMU pre-integration residual by calculating the differences between the pre-integrated relative rotation and the visually estimated relative rotation, the pre-integrated relative velocity and the visually estimated relative velocity, and the pre-integrated relative position and the visually estimated relative position.

[0055] S42. Calculate the marginalization residual according to the marginalization information of the sliding window.

[0056] Specifically, first construct the objective function of the pose optimization problem, divide the state variables into the marginalized states and the remaining states, remove the marginalized states from the pose optimization problem through the Schur complement, and calculate the prior error term after removal.

[0057] S43. Traverse all feature points, remove the identified dynamic features, and construct the visual reprojection residuals of the starting observation frame and other observation frames using the remaining static features.

[0058] Specifically, first, calculate the coordinates of the map points projected into the camera coordinate system, then project these coordinates onto the image plane, and find the difference between the coordinates projected onto the image plane and the coordinates of the feature points observed in each frame of the image to obtain the reprojection error. For each map point, record its reprojection error in all observed frames to form a visual reprojection residual sequence.

[0059] S44. Use the DogLeg method to solve the residual function in the solver to obtain the optimized pose.

[0060] Take the IMU pre-integration residual, marginalization residual, and visual reprojection residual as the residual function. With the goal of minimizing the residual function, solve the residual function through the DogLeg method to obtain the optimized pose.

[0061] The present invention has the following advantages: 1) Aiming at the problem of the decreased operating efficiency after introducing deep learning technology into dynamic SLAM, the present invention proposes a key-frame elastic regulation architecture to realize the integrated operation of the instance segmentation network and VIO in a dual-thread manner. It can flexibly adapt to different neural networks, elastically regulate key frames, process multi-level visual observation information, and improve the system operating efficiency and stability.

[0062] 2) Based on the RANSAC algorithm, the present invention proposes a semantic-assisted RANSAC algorithm, which effectively reduces the proportion of outliers in the initial point set and improves the accuracy of static outlier rejection. At the same time, based on the characteristics of the instance feature motion consistency and Sampson geometric error information, dynamic features are identified to improve the accuracy of dynamic outlier rejection. It significantly improves the positioning accuracy and robustness of visual inertial SLAM.

[0063] According to the method described in the above embodiments, this embodiment will further describe from the perspective of a visual inertial SLAM system applicable to complex dynamic environments. The visual inertial SLAM system applicable to complex dynamic environments can be specifically implemented as an independent entity or integrated in an electronic device. The electronic device can be a terminal, a server, or other devices. Among them, the terminal can include a tablet computer, a laptop computer, a personal computer (PC), a micro-processing box, or other devices, etc.

[0064] Please refer to Figure 5 , Figure 5 Specifically describe the visual inertial SLAM system provided in the embodiments of the present application, which is applied to an electronic device. The visual inertial SLAM system applicable to complex dynamic environments may include: An acquisition module, configured to acquire a visual image sequence; the visual image sequence includes a plurality of visual images; The key-frame elastic regulation module is used to synchronously extract key frames from the visual image through the time communication mechanism for instance segmentation and perform feature extraction and matching on the key frames to obtain feature matching pairs; The cascaded visual outlier rejection module is used to perform static outlier rejection on the feature matching pairs using the semantics-assisted RANSAC method and identify dynamic feature points by combining instance information and Sampson geometric information; The backend optimization module is used to construct the IMU pre-integration residual and the marginalization residual, reject the identified dynamic feature points, construct the visual reprojection residual using the remaining static feature points, and calculate the optimized pose through a solver.

[0065] In specific implementation, each of the above modules and / or units can be implemented as an independent entity, or can be combined arbitrarily and implemented as the same or several entities. For the specific implementation of each of the above modules and / or units, reference can be made to the foregoing method embodiments. For the specific beneficial effects that can be achieved, reference can also be made to the beneficial effects in the foregoing method embodiments, which will not be elaborated herein.

[0066] In addition, an embodiment of the present application further provides an electronic device, which can be a device such as a computer or a tablet computer. The electronic device can implement the steps in any of the embodiments of the visual-inertial SLAM method applicable to a complex dynamic environment provided by the embodiments of the present application. Therefore, it can achieve the beneficial effects that any of the visual-inertial SLAM methods applicable to a complex dynamic environment provided by the embodiments of the present invention can achieve. For details, see the foregoing embodiments, which will not be elaborated herein.

[0067] Figure 6 The specific structural block diagram of the electronic device provided by the embodiment of the present invention is shown. The electronic device can be used to implement the visual-inertial SLAM method applicable to a complex dynamic environment provided in the foregoing embodiment. The electronic device 500 can be a device such as a terminal or a server. Among them, the terminal can include a tablet computer, a notebook computer, a personal computer (PC), a micro-processing box, or other devices, etc.

[0068] The RF circuit 510 is used to receive and transmit electromagnetic waves, enabling the mutual conversion between electromagnetic waves and electrical signals, so as to communicate with a communication network or other devices. The RF circuit 510 may include various existing circuit components for performing these functions. For example, antennas, radio frequency transceivers, digital signal processors, encryption / decryption chips, subscriber identity module (SIM) cards, memories, and so on. The RF circuit 510 can communicate with various networks such as the Internet, intranets, wireless networks or communicate with other devices through a wireless network. The above-mentioned wireless networks may include cellular phone networks, wireless local area networks or metropolitan area networks. The above-mentioned wireless networks can use various communication standards, protocols and technologies, including but not limited to Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as Institute of Electrical and Electronics Engineers standards IEEE 802.11a, IEEE 802.11b, IEEE 802.11g and / or IEEE 802.11n), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols for email, instant messaging and short messages, and any other suitable communication protocols, and may even include those protocols that have not yet been developed currently.

[0069] The memory 520 can be used to store software programs and modules, such as the corresponding program instructions / modules in the above embodiments. The processor 580 executes various functional applications and data processing by running the software programs and modules stored in the memory 520, that is, to implement functions such as taking pictures with the front camera, processing the captured images, and switching the display colors of the display content on the display screen. The memory 520 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 520 may further include a memory remotely disposed relative to the processor 580, and these remote memories can be connected to the electronic device 500 through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.

[0070] The input unit 530 can be used to receive input digital or character information, as well as generate a keyboard and a mouse related to user settings and function controls. The display unit 540 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces, and these graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. The display unit 540 may include a display panel 541. Optionally, the display panel 541 can be configured in the form of an LCD (Liquid Crystal Display) or an OLED (Organic Light-Emitting Diode).

[0071] The audio circuit 560, the speaker 561, and the microphone 562 can provide an audio interface between the user and the electronic device 500. The audio circuit 560 can transmit the converted electrical signal of the received audio data to the speaker 561, and the speaker 561 converts it into a sound signal for output; on the other hand, the microphone 562 converts the collected sound signal into an electrical signal, which is received by the audio circuit 560 and then converted into audio data. After the audio data is output to the processor 580 for processing, it is sent to another terminal, for example, through the RF circuit 510, or the audio data is output to the memory 520 for further processing. The audio circuit 560 may further include an earphone jack to provide communication between the peripheral earphone and the electronic device 500.

[0072] The electronic device 500 can help the user receive requests, send information, etc. through the transmission module 570 (such as a Wi-Fi module), and it provides the user with wireless broadband Internet access. Although the transmission module 570 is shown in the figure, it can be understood that it does not belong to the essential components of the electronic device 500 and can be omitted completely as needed without changing the essence of the invention.

[0073] The processor 580 is the control center of the electronic device 500, connecting various parts of the entire mobile phone through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 520, and by calling the data stored in the memory 520, it executes various functions of the electronic device 500 and processes data, thereby monitoring the electronic device as a whole. Optionally, the processor 580 may include one or more processing cores; in some embodiments, the processor 580 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 580 either.

[0074] The electronic device 500 also includes a power source 590 (such as a battery) for supplying power to each component. In some embodiments, the power source can be logically connected to the processor 580 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power source 590 may also include any components such as one or more DC or AC power sources, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0075] Although not shown, the electronic device 500 also includes a camera (such as a front camera, a rear camera), a Bluetooth module, etc., which will not be elaborated here. Specifically, in this embodiment, the display unit of the electronic device is a touch screen display, and the mobile terminal also includes a memory, and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors. One or more programs include instructions for performing the following operations: Obtain a visual image sequence; the visual image sequence includes a number of visual images; Use a time communication mechanism to synchronize the instance segmentation of the visual images to extract key frames and perform feature extraction and matching on the key frames to obtain feature matching pairs; Use the semantic-assisted RANSAC method to perform static outlier rejection on the feature matching pairs, and combine instance information and Sampson geometric information to identify dynamic feature points; Construct IMU pre-integration residuals and marginalization residuals, reject the identified dynamic feature points, construct visual reprojection residuals with the remaining static feature points, and calculate the optimized pose through a solver.

[0076] In specific implementation, the above-mentioned each module can be implemented as an independent entity, or can be combined arbitrarily to be implemented as the same or several entities. For the specific implementation of the above-mentioned each module, reference can be made to the previous method embodiments, which will not be elaborated here.

[0077] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling relevant hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. For this purpose, an embodiment of the present invention provides a storage medium storing multiple instructions that can be loaded by a processor to execute the steps of any of the embodiments of the visual-inertial SLAM method applicable to a complex dynamic environment provided by the embodiments of the present invention.

[0078] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.

[0079] Since the instructions stored in the storage medium can execute the steps of any of the embodiments of the visual-inertial SLAM method applicable to a complex dynamic environment provided by the embodiments of the present invention, the beneficial effects achievable by any of the visual-inertial SLAM methods applicable to a complex dynamic environment provided by the embodiments of the present invention can be achieved. For details, see the previous embodiments and will not be elaborated here.

[0080] The above has introduced in detail a visual-inertial SLAM method, system, storage medium, and electronic device applicable to a complex dynamic environment provided by the embodiments of the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A visual-inertial SLAM method applicable to complex dynamic environments, characterized in that, The method includes: Obtaining a visual image sequence; the visual image sequence includes a plurality of visual images; Using a time communication mechanism to synchronously perform instance segmentation on the visual images to extract key frames, and perform feature extraction and matching on the key frames to obtain feature matching pairs; Using a semantic-assisted RANSAC method to perform static outlier rejection on the feature matching pairs, and combining instance information and Sampson geometric information to identify dynamic feature points; Constructing an IMU pre-integration residual and a marginalization residual, rejecting the identified dynamic feature points, constructing a visual reprojection residual with the remaining static feature points, and calculating an optimized pose through a solver.

2. The visual inertial SLAM method applicable to complex dynamic environments according to claim 1, wherein The using a time communication mechanism to synchronously perform instance segmentation on the visual images to extract key frames, and perform feature extraction and matching on the key frames to obtain feature matching pairs includes: In a neural network thread, transmitting the time information corresponding to the received visual image to a VIO thread, performing instance segmentation on the visual image to obtain instance pixels and instance IDs, constructing an instance mask based on the instance pixels and the instance IDs, and transmitting the instance mask to the VIO thread; In the VIO thread, according to the time information, extracting the corresponding visual image from the image sequence, performing feature tracking on the feature points at the previous moment through an optical flow method to obtain feature matching pairs, rejecting the tracking points with failed tracking and those beyond the boundary, and extracting new corner points as supplementary feature points; Using the gray pixels at the corresponding positions in the instance mask as instance ID information, and associating and matching the instance ID information with the corresponding feature points.

3. The visual inertial SLAM method applicable to a complex dynamic environment according to claim 2, wherein The constructing an instance mask based on the instance pixels and the instance IDs includes: Constructing a gray image based on the instance pixels and the corresponding instance IDs; For the overlapping regions of different instance pixels, performing instance ID overlay on the gray image according to the segmentation order; Using the overlaid gray image as the instance mask.

4. The visual-inertial SLAM method applicable to a complex dynamic environment according to claim 2, wherein, The using a semantic-assisted RANSAC method to perform static outlier rejection on the feature matching pairs includes: Taking the feature matching pairs as a prior point set, and rejecting the feature points in the feature matching pairs that belong to potential dynamic semantics; According to the instance ID information to which each feature belongs, dividing the feature matching pairs after rejecting the feature points of potential dynamic semantics into different clusters; Calculating the fundamental matrix of adjacent key frames, rejecting the outliers in the feature matching pairs based on the fundamental matrix, and then adding the previously rejected feature points that belong to potential dynamic semantics.

5. The visual-inertial SLAM method applicable to complex dynamic environments according to claim 4, wherein, Combining instance information and Sampson geometric information to identify dynamic feature points includes: Calculating the epipolar error information of each instance object based on the fundamental matrix, and adding the epipolar error information to the geometric error sequence of the instance object; For each instance object, counting the proportion of points in the geometric error sequence that are greater than a preset error threshold. If the proportion is greater than 80%, it is determined that the instance object is in a motion state, and all feature points belonging to the instance object are regarded as dynamic features.

6. The visual-inertial SLAM method applicable to a complex dynamic environment according to claim 5, wherein, The calculation formula for the epipolar error information is: Among them, , are the homogeneous coordinates of the pixel coordinates in adjacent key frames and respectively, is the fundamental matrix, and T represents the transpose.

7. The visual-inertial SLAM method applicable to complex dynamic environments according to claim 1, characterized in that, Construct the IMU pre-integration residual and marginalization residual, remove the identified dynamic feature points, construct the visual reprojection residual using the remaining static feature points, and calculate the optimized pose through a solver, including: Calculate the IMU pre-integration residual between adjacent key frames; Calculate the marginalization residual according to the marginalization information of the sliding window; Traverse all feature points, remove the identified dynamic features, and construct the visual reprojection residuals of the starting observation frame and other observation frames using the remaining static features; Use the DogLeg method to solve the residual function in the solver to obtain the optimized pose.

8. A visual inertial SLAM system applicable to complex dynamic environments, characterized in that, Including: An acquisition module for acquiring a visual image sequence; the visual image sequence includes a plurality of visual images; A key frame elastic regulation module for synchronously performing instance segmentation on the visual images using a time communication mechanism to extract key frames and performing feature extraction and matching on the key frames to obtain feature matching pairs; A cascaded visual outlier removal module for removing static outliers from the feature matching pairs using a semantic-assisted RANSAC method and identifying dynamic feature points in combination with instance information and Sampson geometric information; A backend optimization module for constructing the IMU pre-integration residual and marginalization residual, removing the identified dynamic feature points, constructing the visual reprojection residual using the remaining static feature points, and calculating the optimized pose through a solver.

9. A computer-readable storage medium, characterized in that, Multiple instructions are stored in the computer-readable storage medium, and the instructions are suitable for being loaded by a processor to execute the visual inertial SLAM method applicable to a complex dynamic environment according to any one of claims 1 to 7.

10. An electronic device, characterized in that, Including a processor and a memory, the processor is electrically connected to the memory, the memory is used for storing instructions and data, and the processor is used for executing the steps in the visual inertial SLAM method applicable to a complex dynamic environment according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Monocular vision inertia SLAM method for dynamic scene

    CN111156984A

  • Vision and IMU sensor fusion positioning system based on dynamic object semantic segmentation

    CN113223045A

  • Multi-sensor fusion SLAM method and system applied to dynamic environment

    CN118209101A

  • Dynamic environment dense point cloud SLAM method and system based on YOLOv11 and ORB-SLAM3

    CN119540942A

  • Visual-inertial slam system

    WO2022183137A1