Visual simultaneous localization and mapping method and device, equipment and medium
By introducing three-dimensional feature clustering and a semantic filter model of YOLOV8 neural network in the VSLAM system, low-level semantic regions in the visual sensor image are filtered, and the problem of insufficient positioning accuracy in the existing VSLAM technology is solved, and visual positioning and map construction with higher accuracy is achieved.
Patent Information
- Application Number
- CN202510688897.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-08-26
AI Technical Summary
When existing VSLAM technology visually positioning and navigation in dynamic scenes, it rarely studies the impact of scene semantic classification on positioning accuracy, resulting in insufficient positioning accuracy.
The semantic filter model built on three-dimensional feature clustering and YOLOV8 neural network is adopted to filter the low-level semantic regions in the visual sensor image, extract the target key images, and use grayscale values, gradient angles and grayscale histograms to cluster image features, and train the YOLOV8 neural network to improve positioning accuracy.
By filtering out low-level semantic areas, the accuracy of visual simultaneous positioning and map construction is improved, the cumulative positioning error is reduced, and the accuracy of the motion trajectory of the visual sensor is improved.
Smart Images

Figure CN120543779A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method, device, equipment and medium for simultaneous visual positioning and map construction. Background Art
[0002] Currently, VSLAM (Visual Simultaneous Localization and Mapping) technology uses visual sensors (monocular, binocular, and RGB-D cameras) to capture image or video data, constructing a real-time environmental map while simultaneously performing self-localization. It is widely used in areas such as autonomous driving, robot navigation, augmented reality (AR), surveying and modeling. With the decreasing cost of camera sensors and the increasing computing power, VSLAM is expected to have even greater potential for development compared to LIDAR-SLAM (Light Detection and Ranging).
[0003] Currently, VSLAM technology is primarily categorized into direct methods and feature point methods based on different visual odometry. Direct methods directly analyze image pixels. Feature point methods are more robust than direct methods, as pixels can vary dramatically under varying lighting conditions, making them the mainstream VSLAM approach. Feature point-based VSLAM methods primarily include the ORB (Oriented FAST and Rotated BRIEF) feature point method, dense VSLAM methods, and deep learning VSLAM methods. Since current mainstream methods primarily focus on visual positioning and navigation in dynamic scenes, little research has been conducted on the impact of scene semantic classification on VSLAM positioning accuracy.
[0004] In summary, how to improve the accuracy of simultaneous visual positioning and map construction is the current problem that needs to be solved. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a method, device, equipment and medium for simultaneous visual localization and mapping, which can improve the accuracy of simultaneous visual localization and mapping. The specific scheme is as follows:
[0006] In a first aspect, the present application discloses a method for simultaneous visual localization and mapping, which is applied to a system for simultaneous visual localization and mapping. The method comprises:
[0007] Acquire a plurality of initial images sent by a visual sensor, and extract a reference key image from the initial images;
[0008] A pre-built semantic filter model is used to filter the low-level semantic areas in the reference key image to obtain a target key image; wherein the semantic filter model is a filter model constructed based on three-dimensional feature clustering and YOLOV8 neural network; the three-dimensional features include grayscale value, gradient angle and grayscale histogram.
[0009] A target map including a motion trajectory of the visual sensor is obtained based on the target key image.
[0010] Optionally, before using the pre-built semantic filter model to filter the low-level semantic regions in the reference key image to obtain the target key image, the method further includes:
[0011] Acquire a temporary key image in a data set, extract an image feature block in the temporary key image, and determine a three-dimensional feature vector of the image feature block;
[0012] Performing three-dimensional feature clustering on the image feature blocks based on the three-dimensional feature vector to obtain feature block cluster sets of several different feature block categories;
[0013] The YOLOV8 neural network is trained based on the feature block clustering set to obtain a semantic filter model.
[0014] Optionally, the filtering of the low-level semantic regions in the reference key image by using a pre-built semantic filter model to obtain the target key image includes:
[0015] The pre-built semantic filter model is used to filter the image feature blocks of low-level semantics in the reference key image to obtain a target key image.
[0016] Optionally, extracting the image feature blocks from the temporary key image includes:
[0017] Extracting and matching ORB feature points of the temporary key images of two adjacent frames to obtain matching point pairs;
[0018] Screening the matched point pairs to eliminate point pairs whose distances from corresponding epipolar lines are greater than a predetermined threshold to obtain screened point pairs;
[0019] An image block of a predetermined size centered at an inlier point in the temporary key image is extracted as an image feature block in the temporary key image; the inlier point is a point corresponding to the temporary key image in the filtered point pair.
[0020] Optionally, performing ORB feature point extraction and matching on the temporary key images of two adjacent frames to obtain matched point pairs includes:
[0021] The ORB corner detector is written based on Python code to extract the ORB feature points of the temporary key image, and the brute force algorithm is used to match the ORB feature points of the temporary key images of two adjacent frames to obtain matched point pairs.
[0022] Optionally, screening the matched point pairs to remove point pairs whose distances from corresponding epipolar lines are greater than a predetermined threshold to obtain screened point pairs includes:
[0023] The matched point pairs are screened using a random sampling consensus algorithm to eliminate point pairs whose lengths from corresponding epipolar lines are greater than a predetermined threshold to obtain screened point pairs.
[0024] Optionally, performing three-dimensional feature clustering on the image feature blocks based on the three-dimensional feature vectors to obtain feature block cluster sets of several different feature block categories includes:
[0025] The K-means clustering algorithm is used to perform three-dimensional feature clustering on the image feature blocks based on the three-dimensional feature vectors to obtain feature block cluster sets of several different feature block categories.
[0026] In a second aspect, the present application discloses a visual simultaneous localization and mapping device, which is applied to a visual simultaneous localization and mapping system. The device includes:
[0027] An image extraction module is used to obtain a number of initial images sent by the visual sensor and extract a reference key image from the initial images;
[0028] A filtering module, configured to filter the low-level semantic regions in the reference key image using a pre-built semantic filter model to obtain a target key image; wherein the semantic filter model is a filter model built based on three-dimensional feature clustering and a YOLOV8 neural network; the three-dimensional features include grayscale value, gradient angle, and grayscale histogram;
[0029] A visual simultaneous positioning and map construction module is used to obtain a target map including the motion trajectory of the visual sensor based on the target key image.
[0030] In a third aspect, the present application discloses an electronic device, comprising:
[0031] Memory, used to store computer programs;
[0032] A processor is used to execute the computer program to implement the aforementioned disclosed method for simultaneous visual positioning and mapping.
[0033] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the aforementioned disclosed method of simultaneous visual positioning and mapping is implemented.
[0034] It can be seen that the present application obtains several initial images sent by the visual sensor, extracts a reference key image from the initial image; uses a pre-built semantic filter model to filter the low-level semantic areas in the reference key image to obtain a target key image; wherein, the semantic filter model is a filter model constructed based on three-dimensional feature clustering and YOLOV8 neural network; the three-dimensional features include grayscale value, gradient angle and grayscale histogram; based on the target key image, a target map including the motion trajectory of the visual sensor is obtained. It can be seen that the present application adds a semantic filter to filter out the low-level semantics in the reference key image to obtain the target key image, so as to improve the accuracy of subsequent visual simultaneous positioning and map construction; in addition, the semantic filter is constructed using three-dimensional feature clustering and YOLOV8 neural network, and the grayscale value, gradient angle and grayscale histogram fully consider the background, brightness, edge and texture feature information, give full play to the clustering effect, improve the ability to propose low-level semantics, and further improve the accuracy of subsequent visual simultaneous positioning and map construction. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0036] Figure 1 This is a flow chart of a method for simultaneous visual positioning and map construction disclosed in this application;
[0037] Figure 2 A schematic diagram of a visual simultaneous positioning and map construction process disclosed in this application;
[0038] Figure 3 A schematic diagram of a three-dimensional feature disclosed in this application;
[0039] Figure 4 A schematic diagram of a semantic filter construction process disclosed in this application;
[0040] Figure 5 This is a schematic diagram showing an image block disclosed in this application;
[0041] Figure 6 A trajectory comparison schematic diagram disclosed in this application;
[0042] Figure 7 This is a schematic diagram of absolute posture error comparison disclosed in this application;
[0043] Figure 8 This is a schematic diagram of the structure of a visual simultaneous positioning and map building device disclosed in this application;
[0044] Figure 9 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION
[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0046] Currently, VSLAM technology is mainly divided into direct methods and feature point methods based on different visual odometry. Direct methods directly analyze image pixels. Because pixels vary dramatically under changing lighting conditions, feature point methods are more robust than direct methods and therefore occupy the mainstream VSLAM methods. VSLAM based on feature point methods mainly includes ORB feature point methods, dense VSLAM methods, deep learning VSLAM methods, etc. Since the mainstream methods at this stage mainly focus on visual positioning and navigation in dynamic scenes, there is less research on the impact of scene semantic classification on VSLAM positioning accuracy.
[0047] To this end, an embodiment of the present application proposes a visual simultaneous positioning and map construction solution, which can improve the accuracy of visual simultaneous positioning and map construction.
[0048] The present application discloses a method for simultaneous visual positioning and map construction, see Figure 1 As shown, the method is applied to a visual simultaneous localization and mapping system, and includes:
[0049] Step S11: Acquire several initial images sent by the visual sensor, and extract a reference key image from the initial images.
[0050] In this embodiment, see Figure 2As shown, it is a schematic diagram of a visual simultaneous positioning and map construction process. In the figure, the image frame is first input, and after pre-processing, pose estimation, map tracking, key frame decision-making, and then the key frame is determined and saved, which is the reference key frame. After that, the reference key frame is filtered of low-level semantics and returned to the pre-processing step to obtain a new key frame, which is the target key frame. Subsequently, map points are updated, local optimization, local key frame removal, loop detection, closed-loop correction and global optimization are performed; among them, the main content of this application lies in the step of filtering low-level semantics, and other steps are not introduced in detail here. It should be pointed out that when the image frame is transmitted to the system from the visual sensor, the semantic filter model is first used to filter out the detected low-level semantic image block area, and then tracking, loop closure and mapping steps are performed. The ORB-SLAM3 system that applies this model can reduce the cumulative positioning error and improve the pose estimation and positioning accuracy.
[0051] In this embodiment, the reference key image may be a grayscale image.
[0052] Step S12: Filtering the low-level semantic regions in the reference key image using a pre-built semantic filter model to obtain a target key image; wherein the semantic filter model is a filter model built based on three-dimensional feature clustering and a YOLOV8 (You Only Look Once version 8) neural network; the three-dimensional features include grayscale value, gradient angle, and grayscale histogram.
[0053] In this embodiment, the semantic classification of scene images can be divided into specific and abstract semantic information based on their image representation. For deep learning methods, it is relatively easy to learn specific scene image semantic information, such as specific image representations of people, cars, trees, etc.; however, deep learning methods have not yet proposed a good solution for abstract scene image semantic information, such as low-semantic (low-quality matching area) image representation. However, this application uses degree value, gradient angle and grayscale histogram clustering to filter low-level semantics, which effectively solves the existing problems and helps improve accuracy.
[0054] In this embodiment, before using a pre-built semantic filter model to filter the low-level semantic regions in the reference key image to obtain the target key image, the method further includes: obtaining a temporary key image from a data set, extracting image feature blocks from the temporary key image, and determining a three-dimensional feature vector for the image feature block; performing three-dimensional feature clustering on the image feature block based on the three-dimensional feature vector to obtain a set of feature block clusters of several different feature block categories; and training a YOLOV8 neural network based on the set of feature block clusters to obtain a semantic filter model. Deep learning training is performed in the YOLOV8 neural network, and the generated network model is applied to the ORB-SLAM3 system.
[0055] See also Figure 3 The figure shows a schematic diagram of a three-dimensional feature. In the figure, weak semantic blocks (low semantic areas) are equal to edge information, background information, and corner information. Edge information corresponds to gradient angles (which can be used to describe edges), background information corresponds to grayscale values (which can be clustered or classified based on background similarity), and corner information corresponds to position information and gradient angles. These three elements form a three-dimensional vector for clustering. (Grayscale values represent various types of information in an image, making it easier to group similar backgrounds together. Grayscale histograms determine the brightness distribution of pixels in an image, measuring contrast differences. Gradient angles measure detailed features such as edges and textures, and can determine the corresponding information of corners.) Figure 3 In the edge information map, the white area represents edge information, the area framed by multiple points and their connecting lines in the background information map represents background information, and the positions framed by multiple circles in the corner information map represent corner information.
[0056] It should be pointed out that this application can first modify the file code of the ORB-SLAM3 system, add the function of automatically extracting and saving the key frame image to a local folder after the ORB-SLAM3 system ends running; use the ORB-SLAM3 system to run the data set (KITTI data set, TUM data set, and data sets collected in real scenes), run the modified code, and obtain the key frame image (RGB).
[0057] It should be pointed out that the method described in this application of using a pre-built semantic filter model to filter the low-level semantic areas in the reference key image to obtain the target key image includes: using a pre-built semantic filter model to filter the image feature blocks of the low-level semantics in the reference key image to obtain the target key image.
[0058] In this embodiment, the extracting of the image feature block in the temporary key image includes: performing ORB feature point extraction and matching on the temporary key images of two adjacent frames to obtain matched point pairs; filtering the matched point pairs to eliminate point pairs whose length values from the corresponding extreme lines are greater than a predetermined threshold to obtain filtered point pairs; extracting an image block of a predetermined size centered on an inlier in the temporary key image as the image feature block in the temporary key image; the inlier is a point in the filtered point pair corresponding to the temporary key image.
[0059] It should be noted that extracting and matching ORB feature points from two adjacent frames of the temporary key image to obtain a matched point pair includes: extracting ORB feature points from the temporary key image using an ORB corner detector written in Python code, and matching the ORB feature points from the two adjacent frames of the temporary key image to obtain a matched point pair using a brute force algorithm. The brute force algorithm is also known as the BF (Brute Force) algorithm.
[0060] It should be noted that screening the matched point pairs to remove point pairs whose lengths from the corresponding epipolar lines are greater than a predetermined threshold to obtain the screened point pairs includes: using a random sampling consensus algorithm (RANSAC algorithm) to screen the matched point pairs to remove point pairs whose lengths from the corresponding epipolar lines are greater than a predetermined threshold to obtain the screened point pairs. It should be noted that the screening primarily involves identifying inliers and outliers, with points whose lengths are greater than the predetermined threshold being considered outliers, and points whose lengths are less than the predetermined threshold being considered inliers.
[0061] It should be noted that the image feature block is an image block with a size of 32*32 (Pixel) centered on the inlier point.
[0062] It should be pointed out that this application performs ORB feature point extraction and brute force matching between two adjacent frames of continuous key frames. With the help of multi-view geometry principles, the RANSAC algorithm is used to screen the matched point pairs and remove point pairs whose point-to-polar line distance is greater than the set threshold.
[0063] In this embodiment, performing three-dimensional feature clustering on the image feature blocks based on the three-dimensional feature vectors to obtain feature block cluster sets of several different feature block categories includes: performing three-dimensional feature clustering on the image feature blocks based on the three-dimensional feature vectors using a K-means clustering algorithm to obtain feature block cluster sets of several different feature block categories. Specifically, the K-means method is used to calculate the Euclidean distance of the three-dimensional GHG (Gray value, Gray-level histogram, Gradient angle value) vectors of each image block. The third step is to classify and label the image based on the Euclidean distance to obtain feature block cluster sets. It should be noted that the number of cluster points can be 6-12, resulting in six YOLO format (feature block cluster set) sample data sets (label0-label5). GHG features are three-dimensional vectors composed of gray value, gradient angle, and gray histogram, and are capable of describing semantic information about similar gradient structures. The overall steps are as follows: First, flatten a 32*32 image block into a one-dimensional grayscale array, and the elements in the array are , i ranges from 1 to 1024, and then the grayscale histogram of this image is calculated, which is also expanded into a one-dimensional array with the elements being , i ranges from 0 to 255, and finally the gradient angle is calculated and expanded into a one-dimensional array, the elements are , where i ranges from 1 to 1024. These patches are then merged into a three-dimensional vector, and the corresponding Euclidean distance is calculated for each image. The K-means method is then used to find the cluster centers, and the semantic patches are divided into different categories based on their distance from the cluster center.
[0064] It's important to note that images can be divided into background, objects, and edges and corners. To extract useful information from an image, three image features are used: grayscale value (x), ranging from 0 to 255, gradient angle (y), and grayscale histogram (z). Grayscale value represents various types of information in an image, making it easier to group similar backgrounds together. The grayscale histogram determines the brightness distribution of pixels within the image, measuring contrast differences. The gradient angle measures details such as edges and textures, and can identify corner points. These are then expanded into one-dimensional vectors and merged into three-dimensional vectors. K-means is then used to calculate the distance between different image patches, with those closest in distance being grouped together. A total of six categories are defined, labels0-label5. Image patches are labeled based on their classification, and the training and validation sets are divided into a ratio of 8:2 to create a sample dataset in YOLO format.
[0065] In one specific embodiment, the K-means algorithm is often used to classify samples in a dataset into K distinct categories. The algorithm's main idea is to iteratively assign data points to K clusters, so that each data point is ultimately assigned to the cluster corresponding to its nearest cluster center. The specific steps and principles of the K-means algorithm for clustering GHG three-dimensional features are as follows:
[0066] First, initialization: There are three dimensions in total. First, use represents the data point, n represents the number of images, Represents the data center point. The data point U has three dimensional features. First, select the K clusters of pixel value x. , in this article, K is taken as 6, namely: Six clusters, including It is a cluster The mean of all data points in :
[0067] ;
[0068] Similarly, the gradient angle and grayscale histogram: and The total cluster center is ,include Three vectors, .
[0069] Second, calculate the distance: for each patch (data point) corresponding to , , Calculate its distance to each cluster center, such as To find the cluster center, the distance from all points to the center must be minimized, that is, the distance in each dimension must be minimized at the same time, and the coordinates of the center point can be obtained. :
[0070] ;
[0071] ;
[0072] ;
[0073] At the same time, when it is the smallest, the center of the first dimension (x dimension) is ; GHG characteristics correspond to three-dimensional vectors x, y, and z. Similarly, its center can be obtained as , . Then assign the patch (data point) n* to the nearest cluster center.
[0074] Assign: ;
[0075] Third, update the cluster center: for each cluster, calculate its mean and use the mean as the new cluster center. And assign the patch to the nearest center to form a new cluster. : ;
[0076] Fourth, iteration: Repeat the second and third steps until the cluster center no longer iterates and changes. When no more obvious changes occur, Enough hours. No longer changes significantly.
[0077] Fifth, output the results: get K cluster centers and the cluster label assigned to each data point.
[0078] Specific as Figure 4 As described above, a simple three-dimensional clustering visualization model is shown. The crosses represent the coordinates of the six center points, and different colors represent different categories.
[0079] In summary, see Figure 4 The figure shows a flowchart of the semantic filter construction process. In the figure, ORB corner point recognition is completed first, and then semantic blocks (image feature blocks) are extracted, and GHG three-dimensional features are used for classification and annotation. Finally, the YOLOV8 neural network is used to complete the training to obtain the semantic filter.
[0080] Step S13: obtaining a target map including the motion trajectory of the visual sensor based on the target key image.
[0081] In this embodiment, the training sample set prepared in the second step is put into the YOLOV8 neural network for deep learning training to obtain a semantic filter model, and the semantic filter is applied to the ORB-SLAM3 system.
[0082] In a specific embodiment, see Figure 5 As shown in the figure, it is a schematic diagram of an image block display; in the figure, the semantic filter detects the image blocks label0 to label5 in the image frame. The area framed by the square in the figure is the image block. The two values in the rectangular area next to the image block represent which cluster and the confidence level respectively. In the figure, 0, 1, 2, 3, 4, and 5 represent which cluster, and the following values such as 0.95 and 0.92 represent the confidence level. The confidence levels in the figure are all higher than 0.8.
[0083] In a specific embodiment, see Figure 6 The figure shows a trajectory comparison diagram. The ORB-SLAM3 system improved by the semantic filter model based on GHG features, after filtering out the label4 image semantic block, obtains a trajectory (label4) that is basically consistent with the gt (ground truth) true value.
[0084] In a specific embodiment, see Figure 7 As shown, this is a schematic diagram of absolute posture error comparison; Figure 7 This is a visualization of the APE (Absolute Pose Error) evaluation of the trajectory files obtained by running the ORB-SLAM3 system on the KITTI07 sequence, where keyframe is the trajectory file obtained by running the original ORB-SLAM3 system, and label0-label5 are trajectory files obtained by applying different categories of semantic filters. In the figure, the contents corresponding to max (maximum value), min (minimum value), std (Standard Deviation), median (median), mean (average value), and rmse (Root Mean Square Error) are in the order of d0.txt, d1.txt, d2.txt, d3.txt, d4.txt, d5.txt, and keyframe.txt from top to bottom.
[0085] As shown in Table 1, the RMSE (root mean square error) of the ATE (absolute trajectory error) of the trajectory obtained after filtering out the semantic image blocks corresponding to label0-5 levels and the original trajectory is shown. The ORB-SLAM3 system applies filters that filter out the corresponding levels of label4 and label5 and then runs the trajectory file obtained by ORB-SLAM3. Compared with the case when no filter is applied, the RMSE root mean square error is reduced, among which label4 is reduced the most, by about 27.03% (0.2703=(2.760-2.014) / 2.760).
[0086] Table 1
[0087] category RMSE (m) origin 2.760 Label0 2.697 Label1 2.926 Label2 3.151 Label3 2.333 Label4 2.014 Label5 2.103 Outlier0.8_0.9 2.314
[0088] As shown in Table 2, applying the semantic filter model without the label4 semantic block to other loop closure sequences in KITTI yields the largest improvement in the 02 sequence, approximately 25.08%. This result demonstrates that defining image semantic categories for different regions using GHG features and filtering low-level semantic regions can improve VSLAM system positioning accuracy and reduce cumulative positioning error.
[0089] Table 2
[0090] Loop sequence Original sequence precision (m) Remove the fourth category of accuracy (m) promote% 00 5.71 4.80 4.87% 02 25.69 16.66 25.08% 05 7.24 6.72 6.55% 06 15.33 13.45 12.26% 07 2.58 2.01 22.09% 08 58.35 54.40 6.77% 09 7.82 5.89 24.68%
[0091] It can be seen that the present application obtains several initial images sent by the visual sensor, extracts a reference key image from the initial image; uses a pre-built semantic filter model to filter the low-level semantic areas in the reference key image to obtain a target key image; wherein, the semantic filter model is a filter model constructed based on three-dimensional feature clustering and YOLOV8 neural network; the three-dimensional features include grayscale value, gradient angle and grayscale histogram; based on the target key image, a target map including the motion trajectory of the visual sensor is obtained. It can be seen that the present application adds a semantic filter to filter out the low-level semantics in the reference key image to obtain the target key image, so as to improve the accuracy of subsequent visual simultaneous positioning and map construction; in addition, the semantic filter is constructed using three-dimensional feature clustering and YOLOV8 neural network, and the grayscale value, gradient angle and grayscale histogram fully consider the background, brightness, edge and texture feature information, give full play to the clustering effect, improve the ability to propose low-level semantics, and further improve the accuracy of subsequent visual simultaneous positioning and map construction.
[0092] Correspondingly, the present application also discloses a visual simultaneous positioning and map construction device, see Figure 8 As shown, the device is applied to a visual simultaneous positioning and mapping system, and includes:
[0093] An image extraction module 11 is configured to obtain a plurality of initial images sent by a visual sensor and extract a reference key image from the initial images;
[0094] A filtering module 12 is configured to filter the low-level semantic regions in the reference key image using a pre-built semantic filter model to obtain a target key image; wherein the semantic filter model is a filter model built based on three-dimensional feature clustering and a YOLOV8 neural network; the three-dimensional features include grayscale value, gradient angle, and grayscale histogram;
[0095] The visual simultaneous positioning and map construction module 13 is used to obtain a target map including the motion trajectory of the visual sensor based on the target key image.
[0096] Among them, the more specific working processes of the above modules can refer to the corresponding contents disclosed in the above embodiments, which will not be repeated here.
[0097] It can be seen that the present application obtains several initial images sent by the visual sensor, extracts a reference key image from the initial image; uses a pre-built semantic filter model to filter the low-level semantic areas in the reference key image to obtain a target key image; wherein, the semantic filter model is a filter model constructed based on three-dimensional feature clustering and YOLOV8 neural network; the three-dimensional features include grayscale value, gradient angle and grayscale histogram; based on the target key image, a target map including the motion trajectory of the visual sensor is obtained. It can be seen that the present application adds a semantic filter to filter out the low-level semantics in the reference key image to obtain the target key image, so as to improve the accuracy of subsequent visual simultaneous positioning and map construction; in addition, the semantic filter is constructed using three-dimensional feature clustering and YOLOV8 neural network, and the grayscale value, gradient angle and grayscale histogram fully consider the background, brightness, edge and texture feature information, give full play to the clustering effect, improve the ability to propose low-level semantics, and further improve the accuracy of subsequent visual simultaneous positioning and map construction.
[0098] Furthermore, an embodiment of the present application also provides an electronic device. Figure 9 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram should not be considered as any limitation to the scope of application of the present application.
[0099] Figure 9 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a display 23, an input / output interface 24, a communication interface 25, a power supply 26, and a communication bus 27. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps of the visual simultaneous localization and mapping method disclosed in any of the aforementioned embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0100] In this embodiment, the power supply 26 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 25 can create a data transmission channel between the electronic device 20 and the external device. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface 24 is used to obtain external input data or output data to the outside world. Its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0101] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, magnetic disk, or optical disk, etc. The resources stored thereon can include a computer program 221, which can be stored in a temporary or permanent manner. In addition to including a computer program capable of implementing the visual simultaneous localization and mapping method performed by the electronic device 20 as disclosed in any of the aforementioned embodiments, the computer program 221 can further include a computer program capable of performing other specific tasks.
[0102] Furthermore, an embodiment of the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the aforementioned disclosed method of simultaneous visual positioning and mapping is implemented.
[0103] The specific steps of the method can be referred to the corresponding contents disclosed in the above embodiments, and will not be repeated here.
[0104] The various embodiments in this application are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method part.
[0105] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0106] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0107] Finally, it should be noted that, in this document, relational terms such as first and first are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.
[0108] The above is a detailed introduction to the visual simultaneous positioning and mapping method, device, equipment, and storage medium provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core ideas. At the same time, for those skilled in the art, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A method for simultaneous visual localization and map construction, characterized in that: Applied to a visual simultaneous localization and mapping system, the method includes: Acquire a plurality of initial images sent by a visual sensor, and extract a reference key image from the initial images; A pre-built semantic filter model is used to filter the low-level semantic regions in the reference key image to obtain a target key image; wherein the semantic filter model is a filter model built based on three-dimensional feature clustering and a YOLOV8 neural network; the three-dimensional features include grayscale value, gradient angle, and grayscale histogram; A target map including a motion trajectory of the visual sensor is obtained based on the target key image.
2. The method for simultaneous visual positioning and mapping according to claim 1, wherein: Before using the pre-built semantic filter model to filter the low-level semantic areas in the reference key image to obtain the target key image, the method further includes: Acquire a temporary key image in a data set, extract an image feature block in the temporary key image, and determine a three-dimensional feature vector of the image feature block; Performing three-dimensional feature clustering on the image feature blocks based on the three-dimensional feature vector to obtain feature block cluster sets of several different feature block categories; The YOLOV8 neural network is trained based on the feature block clustering set to obtain a semantic filter model.
3. The method for simultaneous visual positioning and mapping according to claim 2, wherein: The method of filtering the low-level semantic regions in the reference key image using the pre-built semantic filter model to obtain the target key image includes: The pre-built semantic filter model is used to filter the image feature blocks of low-level semantics in the reference key image to obtain a target key image.
4. The method for simultaneous visual positioning and mapping according to claim 2, wherein: The extracting of the image feature blocks from the temporary key image includes: Extracting and matching ORB feature points of the temporary key images of two adjacent frames to obtain matching point pairs; Screening the matched point pairs to eliminate point pairs whose distances from corresponding epipolar lines are greater than a predetermined threshold to obtain screened point pairs; An image block of a predetermined size centered at an inlier point in the temporary key image is extracted as an image feature block in the temporary key image; the inlier point is a point corresponding to the temporary key image in the filtered point pair.
5. The method for simultaneous visual positioning and mapping according to claim 4, wherein: The step of extracting and matching ORB feature points of the temporary key images of two adjacent frames to obtain matched point pairs includes: The ORB corner detector is written based on Python code to extract the ORB feature points of the temporary key image, and the brute force algorithm is used to match the ORB feature points of the temporary key images of two adjacent frames to obtain matched point pairs.
6. The method for simultaneous visual localization and mapping according to claim 4, wherein: The screening of the matched point pairs to eliminate point pairs whose distances from corresponding epipolar lines are greater than a predetermined threshold to obtain screened point pairs includes: The matched point pairs are screened using a random sampling consensus algorithm to eliminate point pairs whose lengths from corresponding epipolar lines are greater than a predetermined threshold to obtain screened point pairs.
7. The method for simultaneous visual localization and mapping according to any one of claims 2 to 6, wherein: The performing three-dimensional feature clustering on the image feature blocks based on the three-dimensional feature vector to obtain feature block cluster sets of several different feature block categories includes: The K-means clustering algorithm is used to perform three-dimensional feature clustering on the image feature blocks based on the three-dimensional feature vectors to obtain feature block cluster sets of several different feature block categories.
8. A visual simultaneous positioning and mapping device, characterized in that: Applied to a visual simultaneous positioning and mapping system, the device comprises: An image extraction module is used to obtain a number of initial images sent by the visual sensor and extract a reference key image from the initial images; A filtering module, configured to filter the low-level semantic regions in the reference key image using a pre-built semantic filter model to obtain a target key image; wherein the semantic filter model is a filter model built based on three-dimensional feature clustering and a YOLOV8 neural network; the three-dimensional features include grayscale value, gradient angle, and grayscale histogram; A visual simultaneous positioning and map construction module is used to obtain a target map including the motion trajectory of the visual sensor based on the target key image.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the visual simultaneous localization and mapping method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that Used to store a computer program; wherein, when the computer program is executed by a processor, the visual simultaneous localization and mapping method according to any one of claims 1 to 7 is implemented.