Visual inertia SLAM positioning method and positioning device

Through the visual-inertial SLAM positioning method, the point features and line features of the visual image are combined with the semantic segmentation model to solve the positioning accuracy and robustness problems caused by changes in lighting and dynamic objects in autonomous driving scenarios, achieving higher precision and stable positioning effects.

CN120800359APending Publication Date: 2025-10-17JIANG SU YI RUI QING LIAN QI CHE DIAN ZI YAN JIU YUAN YOU XIAN GONG SI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510908877.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing visual SLAM systems have insufficient positioning accuracy and robustness in autonomous driving scenarios due to changes in lighting, texture, and dynamic objects, making it difficult to meet the requirements of high-level tasks.

Method used

The visual-inertial SLAM positioning method is adopted to extract point features and line features of visual images, and combine them with the offline trained semantic segmentation model for semantic segmentation. The visual features of the salient areas are screened and tightly coupled with the IMU data for optimization, and the detection loop is used for global optimization.

Benefits of technology

It improves the performance of visual-inertial SLAM positioning, enhances the real-time and robustness of positioning, reduces cumulative errors, and improves positioning accuracy in autonomous driving scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120800359A_ABST
    Figure CN120800359A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of machine vision, provides a visual inertia SLAM (Simultaneous Localization and Mapping) positioning method and a positioning device, aims at an indoor parking lot scene, takes a line feature of a visual image as a supplementary visual feature, and solves the problems of poor positioning precision and robustness caused by changes of illumination, texture and dynamic objects in an automatic driving scene. A saliency area of visual features is obtained through semantic segmentation, the pose is calculated and optimized through features extracted by excluding unstable areas, the vehicle pose is optimized by screening the visual features of the saliency area and tightly coupling IMU measurement data, the real-time performance of positioning can be guaranteed, the global pose is optimized by detecting loops, the influence of accumulated errors is reduced, and the positioning accuracy is improved. Therefore, the visual inertia SLAM positioning performance is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of machine vision, in particular to a visual-inertial SLAM (Simultaneous Localization and Mapping) positioning method and a visual-inertial SLAM positioning device. BACKGROUND

[0002] Positioning is an indispensable technical link for automatic driving, and the SLAM (Simultaneous Localization and Mapping) technology enables an automatic driving vehicle to construct a map of the surrounding environment in real time and accurately position the vehicle in an environment, such as an indoor parking lot, that lacks effective GNSS (Global Navigation Satellite System) signals. The SLAM can be divided into laser SLAM and visual SLAM according to the used sensor, and the visual SLAM has advantages such as low sensor price and richer semantic information of visual images compared with the laser SLAM.

[0003] At present, the visual SLAM uses point features collected by vision to perform motion tracking and pose estimation. With the increasing demand for automatic driving vehicles, the visual SLAM system is not only required to have higher performance, but also expected to complete higher-level tasks, such as recognizing the information of objects in a scene, finding their positions and constructing a map containing semantic information. If only relying on feature points, the performance may be limited by the number and quality of feature points in some scenes. In addition, changes in light, texture and dynamic objects are common in the automatic driving scene, and the above method has obvious shortcomings in positioning accuracy and robustness. SUMMARY

[0004] In order to improve the positioning performance of the visual-inertial SLAM, a first object of the present application is to provide a visual-inertial SLAM positioning method.

[0005] A second object of the present application is to provide a visual-inertial SLAM positioning device.

[0006] The technical scheme adopted by the present application is as follows:

[0007] The first aspect embodiment of the present application provides a visual-inertial SLAM positioning method, comprising the following steps: extracting visual features of a visual image collected by a visual sensor, the visual features comprising point features and line features; performing semantic segmentation on the visual image by using an offline trained semantic segmentation model to obtain semantic regions; pre-integrating IMU (Inertial Measurement Unit) data, and initializing a visual-inertial navigation system in a visual-inertial loose coupling manner; screening visual features of the semantic regions, and performing tight coupling optimization on the screened visual features and the IMU data integration through a sliding window to update a visual sensor pose and a map landmark; and detecting a loop, and performing loop optimization to obtain a globally consistent map and pose.

[0008] The visual-inertial SLAM positioning method provided in the present application can further have the following additional technical features:

[0009] According to an embodiment of the present application, the visual features of the visual image collected by the visual camera are extracted, specifically comprising: using a DGTT (Dynamic Graph Traversal Technique) algorithm from OpenCV (a cross-platform computer vision and machine learning software library) to extract point features of the visual image.

[0010] According to an embodiment of the present application, the visual features of the visual image collected by the visual camera are extracted, specifically comprising: using a LSD (Linear Segment Detector) algorithm from OpenCV to extract line features of the visual image, and adjusting an image scale parameter s and a minimum density threshold d according to visual sensor parameters when extracting the line features; setting a minimum length threshold of the line features to filter out line features with a length less than the minimum length threshold; respectively calculating square differences of pixel coordinates of end points of line features to be matched in two frames of images, and if the square differences are less than or equal to a set threshold, further calculating an included angle of the line features to be matched, and if the included angle is less than or equal to a set threshold, establishing a matching relationship of the line features in the two frames of images.

[0011] According to an embodiment of the present application, the visual image is subjected to semantic segmentation by using an offline trained semantic segmentation model to obtain semantic regions, specifically comprising: using an encoder to extract hierarchical features of the visual image, wherein the encoder gradually reduces the number of channels of feature maps while gradually increasing the spatial size of the feature maps during operation; using a pyramid pooling module to perform global average pooling, convolution and up-sampling operations, and then adding and convolving the features to fuse the hierarchical features to generate fine features; and using a decoder to fuse, down-sample and up-sample the hierarchical features and the fine features to output the semantic regions.

[0012] According to one embodiment of the present application, the decoder is used to fuse the hierarchical features and the fine features, specifically including: calculating the weight alpha through the spatial attention mechanism and the channel attention mechanism, multiplying the hierarchical features and the fine features according to the weight alpha, and performing the addition operation to generate the fine features.

[0013] The embodiment of the second aspect of the present application proposes a visual-inertial SLAM positioning device, comprising: an extraction module, configured to extract visual features of a visual image collected by a visual sensor, the visual features comprising point features and line features; a segmentation module, configured to perform semantic segmentation on the visual image through an offline trained semantic segmentation model to obtain semantic regions; an initialization module, configured to pre-integrate IMU data, and initialize a visual-inertial navigation system in a visual-inertial loose coupling manner; an updating module, configured to filter the visual features of the semantic regions, and perform tight coupling optimization on the filtered visual features and the integrated IMU data through a sliding window to update the pose of the visual sensor and a map landmark; and an optimization module, configured to detect a loop, and perform loop optimization to obtain a globally consistent map and pose when the loop is detected.

[0014] The visual-inertial SLAM positioning device proposed in the present application can have the following additional technical features:

[0015] According to one embodiment of the present application, the extraction module is specifically configured to extract the point features of the visual image using a DGTT algorithm from OpenCV.

[0016] According to one embodiment of the present application, the extraction module is specifically configured to extract the line features of the visual image using an LSD algorithm from OpenCV, adjust the image scale parameter s and the minimum density threshold d according to the visual sensor parameters when extracting the line features, set a minimum length threshold for the line features, filter out the line features with a length less than the minimum length threshold, respectively calculate the square of the pixel coordinate difference of the end points of the line features to be matched in two frames of images, if the square is less than or equal to a set threshold, further calculate the included angle of the line features to be matched, and if the included angle is less than or equal to a set threshold, establish a matching relationship between the line features in the two frames of images.

[0017] According to one embodiment of the present application, the segmentation module is specifically configured to extract hierarchical features of the visual image using an encoder, the encoder gradually reduces the number of channels of the feature map while gradually increasing the spatial size of the feature map during operation; perform global average pooling, convolution, and up-sampling operations using a pyramid pooling module, and then perform addition and convolution operations on the features to fuse the hierarchical features to generate fine features; and fuse the hierarchical features and the fine features using a decoder, perform down-sampling and up-sampling operations, and output semantic regions.

[0018] According to one embodiment of the present application, the segmentation module is further configured to: calculate a weight alpha through a spatial attention mechanism and a channel attention mechanism, multiply the hierarchical features and the fine features according to the weight alpha, and perform an addition operation to generate fine features.

[0019] The present application has the following beneficial effects:

[0020] The present application uses line features of visual images as supplementary visual features in an indoor parking lot scene, solves the problem of poor positioning accuracy and robustness caused by changes in light, texture and dynamic objects in an autonomous driving scene, and further obtains significant regions of visual features through semantic segmentation, excludes features extracted from unstable regions from pose calculation and optimization, and optimizes the vehicle pose by tightly coupling the visual features of the selected significant regions and the IMU measurement data, which can ensure the real-time positioning of the vehicle, optimize the global pose by detecting the loop, and reduce the influence of cumulative errors, thereby significantly improving the positioning performance of visual-inertial SLAM. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 is a flowchart of a visual-inertial SLAM positioning method according to one embodiment of the present application;

[0022] Figure 2 is an architecture diagram of a semantic segmentation model according to one embodiment of the present application;

[0023] Figure 3 is a block schematic diagram of a visual-inertial SLAM positioning device according to one embodiment of the present application. DETAILED DESCRIPTION

[0024] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0025] Figure 1 is a flowchart of a visual-inertial SLAM positioning method according to one embodiment of the present application, as shown in Figure 1 the method comprises the following steps:

[0026] S1, extracting visual features of a visual image collected by a visual sensor, the visual features including point features and line features.

[0027] In one embodiment of the present application, the visual features of the visual image collected by the visual camera are extracted, specifically including: using the DGTT algorithm from OpenCV to extract the point features of the visual image.

[0028] In another embodiment of the present application, the visual features of the visual image captured by the visual camera are extracted, specifically including S11-S13:

[0029] S11, using the LSD algorithm from OpenCV to extract the line features of the visual image, and adjusting the image scale parameter s and the minimum density threshold d according to the visual sensor parameters when extracting the line features.

[0030] Specifically, the N-layer Gaussian pyramid generated in OpenCV is used to represent the original image, in which the image is down-sampled N-1 times, blurred N times, and then the lines are extracted in each layer using LSD. In the present application, the scale s and the number of layers of the pyramid are simplified, s=0.5, N=2. Next, if the alignment area points in the closed rectangle are less than the threshold, LSD sets the minimum density threshold to reject the line segment, and the minimum density threshold d can be set to 0.5 to speed up the processing speed. That is, the image scale parameter s and the minimum density threshold d can be adjusted according to s=0.5, d=0.5.

[0031] S12, setting a minimum length threshold for the line features, and filtering out line features with a length less than the minimum length threshold.

[0032] S13, respectively calculating the square of the pixel coordinate difference of the end points of the line features to be matched in the two frames of images, if less than or equal to the set threshold, further calculating the included angle of the line features to be matched, if less than or equal to the set threshold, establishing the matching relationship of the line features in the two frames of images.

[0033] Specifically, the square of the pixel coordinate difference of the end points of the line features to be matched in the two frames of images is calculated respectively, when the period is greater than the set threshold, it is considered that the matching is not passed; the included angle of the two line segments is calculated and compared, if the included angle is greater than the set threshold, the matching is not passed; if both the end point coordinates and the included angle are matched, the matching relationship of the line features in the two frames of images is established. The positional relationship between the two frames of images is established according to the matching relationship of the line features.

[0034] S2, performing semantic segmentation on the visual image by the off-line trained semantic segmentation model to obtain the semantic region.

[0035] In one specific embodiment of the present application, the visual image is segmented by the off-line trained semantic segmentation model to obtain the semantic region, specifically including the following steps S21-S23:

[0036] S21, using an encoder to extract the hierarchical features F low of the visual image, in which the encoder gradually reduces the channel number of the feature map while gradually increasing the spatial size of the feature map during operation.

[0037] Specifically, the encoder is responsible for extracting hierarchical features, and as the features change from low-dimensional to high-dimensional, the encoder gradually reduces the number of channels and expands the spatial dimensions of the features, achieving lightweight design.

[0038] S22, hierarchical features F are obtained by using a pyramid pooling module low After global average pooling, convolution, and up-sampling operations, the features are added and convolved to fuse the hierarchical features to generate fine features F up .

[0039] Specifically, the pyramid pooling module takes the deepest layer of the encoder as input, first performs global average pooling to obtain features of three spatial sizes, then reduces the spatial size of the features through 1x1 convolution and up-sampling operations, and then adds the three features together and performs a convolution to obtain fine features F up .

[0040] S23, the decoder is used to fuse the hierarchical features and the fine features, down-sample and up-sample the operations to output semantic regions.

[0041] In an embodiment of the present application, the decoder is used to fuse the hierarchical features and the fine features, specifically including: calculating the weight a through the spatial attention mechanism and the channel attention mechanism, multiplying and adding the hierarchical features and the fine features according to the weight a to generate fine features.

[0042] Specifically, the weight a is calculated by combining the spatial attention and the channel attention, so as to fully utilize the spatial and channel relationships of the input features. Then, the hierarchical features F low and the fine features F up are fused according to the weight coefficient a, for example, F out =F low ·a+F up ·(1-a) is fused, and F out is the output result after fusion. After down-sampling and up-sampling operations, the semantic regions are output.

[0043] The encoder gradually reduces the number of channels and expands the spatial dimensions of the features, and the decoder solves this problem by effectively strengthening the feature representation. The decoder first calculates the weight a through the attention module, and then fuses the input features with these weights. The pyramid pooling module (by reducing the number of intermediate and output channels, removing the shortcut connection, and replacing the splicing operation with the addition operation, the calculation cost is greatly reduced.

[0044] The architecture of the semantic segmentation model can be seen in Figure 2 .

[0045] S3, pre-integrate the IMU data, and initialize the visual-inertial navigation system through visual-inertial loose coupling.

[0046] Specifically, the IMU pre-integration is to integrate the data irrelevant to the state vector update in the IMU measurement data in a period of time as the IMU pre-integration item in the period of time, so as to avoid repeated integration of the data in the subsequent process, and finally achieve the purpose of simplifying the state estimation.

[0047] S4, screen the visual features of the semantic region, and tightly couple the screened visual features and the IMU data integration through a sliding window to optimize the visual sensor pose and the map landmarks.

[0048] Specifically, on the basis of the above feature extraction, point features, line features and other visual features can be obtained. In order to ensure the effectiveness of the visual features participating in the optimization between frames, the effective region of the visual features, i.e. the semantic region, is obtained through semantic segmentation. The visual features in the effective region can participate in the optimization in the sliding window, otherwise they are removed, so as to ensure the effectiveness of the visual features participating in the optimization, and further improve the robustness and positioning accuracy of the system.

[0049] S5, detect loop closure, and perform loop closure optimization to obtain a globally consistent map and pose when the loop closure is detected.

[0050] Specifically, the SLAM system continuously estimates its own pose (position and attitude) and simultaneously constructs an environment map during operation. Although the visual sensor (camera) and the IMU (inertial) sensor are complementary, they inevitably have measurement noise and errors. These small errors will gradually accumulate during long-term operation or large-scale movement, causing the estimated trajectory and map of the system to drift. That is, when the system actually returns to the starting point or a previously visited location, there will be a significant deviation (inconsistency) between the estimated position and the constructed map points and the previously recorded position and map points. Loop detection is a process of identifying that the device has returned to a previously visited location using current visual information. After detecting an effective loop, the system uses this powerful spatial and positional consistency constraint to significantly correct the drift error accumulated by the visual and inertial sensors during long-term operation / large-scale movement through back-end optimization, thereby restoring the global consistency of the trajectory and the map, which is an indispensable key technology to ensure the long-term robustness and accuracy of the SLAM system. The vision is responsible for "recognizing" the old location, the IMU assists in positioning and provides continuous motion constraints, and the combination of the two achieves drift correction through optimization, thereby obtaining a globally consistent map and pose.

[0051] In summary, according to the visual-inertial SLAM positioning method of the embodiment of the present application, for the indoor parking lot scene, the line features of the visual image are taken as the supplementary visual features, the problem of poor positioning accuracy and robustness caused by changes of light, texture and dynamic objects in the automatic driving scene is solved, the salient regions of the visual features are obtained through semantic segmentation, the features extracted from the unstable regions are excluded from the pose calculation and optimization, the visual features of the salient regions are screened, the vehicle pose is optimized through the tight coupling of the IMU measurement data, the positioning real-time performance can be ensured, the global pose is optimized through loop detection, the influence of the cumulative error is reduced, and thus the visual-inertial SLAM positioning performance is significantly improved.

[0052] Corresponding to the above-mentioned visual-inertial SLAM positioning method, the present application further provides a visual-inertial SLAM positioning device. Since the device embodiment of the present application corresponds to the above-mentioned method embodiment, the details not disclosed in the device embodiment can be referred to the above-mentioned method embodiment, which will not be described herein again.

[0053] Figure 3 is a block schematic diagram of the visual-inertial SLAM positioning device according to an embodiment of the present application, as shown in Figure 3 The device comprises an extraction module 1, a segmentation module 2, an initialization module 3, an updating module 4 and an optimization module 5.

[0054] The extraction module 1 is configured to extract the visual features of the visual image collected by the visual sensor, wherein the visual features comprise point features and line features; the segmentation module 2 is configured to perform semantic segmentation on the visual image through the offline trained semantic segmentation model to obtain semantic regions; the initialization module 3 is configured to pre-integrate the IMU data, and initialize the visual-inertial navigation system through the visual-inertial loose coupling manner; the updating module 4 is configured to screen the visual features of the semantic regions, and integrate the screened visual features and the IMU data through the sliding window tight coupling optimization to update the visual sensor pose and the map landmarks; and the optimization module 5 is configured to detect the loop, and when the loop is detected, the loop optimization is performed to obtain the globally consistent map and pose.

[0055] According to an embodiment of the present application, the extraction module 1 is specifically configured to extract the point features of the visual image using the DGTT algorithm from OpenCV.

[0056] According to one embodiment of the present application, the extraction module is specifically configured to: extract line features of the visual image using an LSD algorithm from OpenCV, and adjust image scale parameter s and minimum density threshold d according to the visual sensor parameters when extracting the line features; set a minimum length threshold for the line features, and filter out line features with a length less than the minimum length threshold; respectively calculate the square of the pixel coordinate difference of the end points of the line features to be matched in the two frames of images, and if the square is less than or equal to a set threshold, further calculate the included angle of the line features to be matched, and if the included angle is less than or equal to a set threshold, establish a matching relationship of the line features in the two frames of images.

[0057] According to one embodiment of the present application, the segmentation module 2 is specifically configured to: extract hierarchical features of the visual image using an encoder, wherein the encoder gradually reduces the number of channels of the feature map while gradually increasing the spatial size of the feature map during operation; perform global average pooling, convolution, and up-sampling operations using a pyramid pooling module, and then add and convolve the features to generate fine features by fusing the hierarchical features; and fuse, down-sample, and up-sample the hierarchical features and the fine features using a decoder to output semantic regions.

[0058] According to one embodiment of the present application, the segmentation module 2 is further configured to: calculate a weight a through a spatial attention mechanism and a channel attention mechanism, multiply the hierarchical features and the fine features according to the weight a, and add the multiplied results to generate fine features.

[0059] In summary, according to the visual-inertial SLAM positioning device of the embodiments of the present application, for the indoor parking lot scene, the line features of the visual image are used as supplementary visual features to solve the problem of poor positioning accuracy and robustness caused by changes in light, texture, and dynamic objects in the autonomous driving scene. Then, the semantic segmentation is performed to obtain the saliency regions of the visual features, and the features extracted from the unstable regions are excluded to calculate and optimize the pose. The visual features in the selected saliency regions are tightly coupled with the IMU measurement data to optimize the vehicle pose, which can ensure the real-time positioning. The global pose is optimized by detecting the loop to reduce the influence of cumulative errors, thereby significantly improving the performance of the visual-inertial SLAM positioning.

[0060] In the description of the present application, the terms "first", "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features. The meaning of "a plurality of" is two or more, unless otherwise specifically limited.

[0061] In the description of the specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the description of the specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the specification and the features of the different embodiments or examples, without contradiction. In the description of the specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the description of the specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the specification and the features of the different embodiments or examples, without contradiction.

[0062] Any process or method descriptions or descriptions of the flow diagrams in the flow charts described herein, or otherwise described in this specification, can be understood as representing the steps of a method or process, including one or more steps for implementing custom logic functions or processes, and the scope of the preferred embodiments of the present application includes additional implementation in which the steps are performed in a different order, including an order that is substantially simultaneous, or in reverse order, depending on the functionality involved, as will be understood by those skilled in the art.

[0063] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.

[0064] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0065] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0066] In addition, each function unit in each embodiment of the present application can be integrated in one processing module, or each unit can exist physically independently, or two or more units can be integrated in one module. The integrated module can be realized in the form of hardware, or in the form of software function module. When the integrated module is realized in the form of software function module and sold or used as an independent product, it can also be stored in a computer readable storage medium.

[0067] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it should be understood by those skilled in the art that the above embodiments are exemplary and should not be construed as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above embodiments within the scope of the present application.

[0068] Although the embodiments of the present application have been shown and described above, it should be understood by those skilled in the art that various changes, modifications, replacements and variations can be made to the embodiments without departing from the principles and spirits of the present application, and the scope of the present application is defined by the appended claims and their equivalents.

Claims

1. A visual inertial SLAM positioning method, characterized in that: The following steps are involved: Extracting visual features of a visual image collected by a visual sensor, wherein the visual features include point features and line features; Perform semantic segmentation on the visual image through the offline trained semantic segmentation model to obtain semantic regions; Pre-integrate the IMU data and initialize the visual-inertial navigation system through visual-inertial loose coupling; Filter the visual features of the semantic area, and optimize the filtered visual features and IMU data integration through a sliding window to update the visual sensor pose and map landmarks; Detect loops. When a loop is detected, perform loop optimization to obtain a globally consistent map and pose.

2. The visual-inertial SLAM positioning method according to claim 1, wherein Extract visual features of visual images captured by visual cameras, including: Extract point features from visual images using the DGTT algorithm from OpenCV.

3. The visual-inertial SLAM positioning method according to claim 1, wherein Extract visual features of visual images captured by visual cameras, including: Use the LSD algorithm from OpenCV to extract line features of visual images. When extracting line features, adjust the image scale parameter s and the minimum density threshold d according to the visual sensor parameters; Set the minimum length threshold of line features to filter out line features whose length is less than the minimum length threshold; The square of the pixel coordinate difference of the endpoints of the line features to be matched in the two frames of images is calculated respectively. If it is less than or equal to the set threshold, the angle of the line features to be matched is further calculated. If the angle is less than or equal to the set threshold, the matching relationship of the line features in the two frames of images is established.

4. The visual-inertial SLAM positioning method according to claim 1, wherein The offline trained semantic segmentation model is used to perform semantic segmentation on the visual image to obtain semantic regions, including: An encoder is used to extract hierarchical features of the visual image, wherein the encoder gradually increases the spatial size of the feature map while gradually reducing the number of channels of the feature map during operation; The pyramid pooling module is used to perform global average pooling, convolution, and upsampling operations on the hierarchical features, and then the features are added and convolved to fuse the hierarchical features into fine features; A decoder is used to fuse, downsample, and upsample the hierarchical features and the fine features to output a semantic region.

5. The visual-inertial SLAM positioning method according to claim 4, wherein The decoder is used to fuse the hierarchical features and the fine features, specifically including: The weight α is calculated through the spatial attention mechanism and the channel attention mechanism, and the hierarchical features and the fine features are multiplied and added according to the weight α to generate fine features.

6. A visual inertial SLAM positioning device, characterized in that: include: An extraction module, the extraction module is used to extract visual features of the visual image collected by the visual sensor, the visual features including point features and line features; A segmentation module, wherein the segmentation module is used to perform semantic segmentation on the visual image using an offline trained semantic segmentation model to obtain semantic regions; An initialization module, which is used to pre-integrate the IMU data and initialize the visual inertial navigation system through visual-inertial loose coupling; An update module is used to filter visual features of the semantic area, and optimize the filtered visual features and IMU data integration through a sliding window to update the visual sensor pose and map landmarks; The optimization module is used to detect loops. When a loop is detected, loop optimization is performed to obtain a globally consistent map and pose.

7. The visual-inertial SLAM positioning device according to claim 6, characterized in that: The extraction module is specifically used for: Extract point features from visual images using the DGTT algorithm from OpenCV.

8. The visual-inertial SLAM positioning device according to claim 6, wherein: The extraction module is specifically used for: Use the LSD algorithm from OpenCV to extract line features of visual images. When extracting line features, adjust the image scale parameter s and the minimum density threshold d according to the visual sensor parameters; Set the minimum length threshold of line features to filter out line features whose length is less than the minimum length threshold; The square of the pixel coordinate difference of the endpoints of the line features to be matched in the two frames of images is calculated respectively. If it is less than or equal to the set threshold, the angle of the line features to be matched is further calculated. If the angle is less than or equal to the set threshold, the matching relationship of the line features in the two frames of images is established.

9. The visual-inertial SLAM positioning device according to claim 6, characterized in that: The segmentation module is specifically used for: An encoder is used to extract hierarchical features of the visual image, wherein the encoder gradually increases the spatial size of the feature map while gradually reducing the number of channels of the feature map during operation; After using the pyramid pooling module to perform global average pooling, convolution, and upsampling operations, the features are then added and convolved to fuse the hierarchical features into fine features; A decoder is used to fuse, downsample, and upsample the hierarchical features and the fine features to output a semantic region.

10. The visual-inertial SLAM positioning device according to claim 9, characterized in that: The segmentation module is further configured to: The weight α is calculated through the spatial attention mechanism and the channel attention mechanism, and the hierarchical features and the fine features are multiplied and added according to the weight α to generate fine features.