Simultaneous localization and mapping method and device, equipment and storage medium

By fusing the semantic information and depth information of the scene images collected by the camera at each continuous moment, accurately distinguishing dynamic and static feature points, the problem of low accuracy of simultaneous positioning and map construction in the prior art is solved, high-precision pose estimation and map construction are achieved, and computing resource consumption is reduced.

CN119991812AActive Publication Date: 2025-05-13SINOTRANS +1
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510437553.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-13
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

In the prior art, deep learning combined with geometric methods, the problem of low accuracy caused by simultaneous positioning and map construction after processing dynamic or static objects in the environment is not high.

Method used

By fusing the semantic information and depth information of the scene image collected by the camera at each continuous moment, and accurately distinguishing the dynamic feature points and static feature points in the current frame image, we can accurately eliminate the dynamic feature points, and then accurately estimate the camera's position and construct the scene map.

Benefits of technology

It achieves the improvement of camera pose estimation accuracy and accuracy and robustness of map construction, while reducing computing resource consumption and meeting the needs of different robot applications in complex dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991812A_ABST
    Figure CN119991812A_ABST
Patent Text Reader

Abstract

The invention provides a simultaneous localization and mapping method and device, equipment and a storage medium, which are applied to the technical field of image processing, and the method comprises the steps: obtaining a first image group collected by a camera for a scene at the current moment, and carrying out the semantic segmentation and depth estimation, obtaining a semantic segmentation result of each feature point in the first image group and a target depth image; the first image group comprises a first RGB image and a first initial depth image; determining a target category of each feature point in the first RGB image according to the semantic segmentation result and the target depth image of each feature point in the first image group and the semantic segmentation result and the target depth image of each feature point in each second image group; and removing feature points of which the target category belongs to a dynamic state in the first RGB image to obtain a target RGB image, carrying out simultaneous positioning and map construction, and determining a target pose and a target map of the camera. By adopting the technical scheme of the invention, the accuracy of simultaneous positioning and map construction can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a simultaneous positioning and map building method, device, equipment and storage medium. Background Art

[0002] With the rapid development of autonomous robots, augmented reality and drone technology, SLAM (Simultaneous Localization and Mapping) technology has become the key to the successful operation of intelligent mobile robots in unknown environments. SLAM needs to distinguish between static and dynamic objects in the environment during the simultaneous localization and mapping process, so that accurate positioning and higher-precision map construction can be achieved.

[0003] In related technologies, dynamic objects and / or static objects in the environment are generally processed by deep learning combined with geometric methods, thereby achieving simultaneous positioning and map construction.

[0004] However, the above technologies have the problem of low accuracy in simultaneous positioning and map construction. Summary of the invention

[0005] The present invention provides a simultaneous positioning and mapping method, device, equipment and storage medium, which are used to solve the defect of low accuracy caused by simultaneous positioning and mapping after processing dynamic or static objects in the environment through deep learning combined with geometric methods in the prior art. The method is achieved by fusing the semantic information and depth information of the scene image collected by the camera at each continuous moment, and accurately distinguishing the dynamic feature points and static feature points in the current frame image to accurately eliminate the dynamic feature points, thereby accurately estimating the camera's posture and accurately constructing the scene map, thereby improving the camera posture estimation accuracy and the accuracy and robustness of map construction, while reducing the consumption of computing resources, so as to meet the needs of different robot applications in complex dynamic environments.

[0006] The present invention provides a simultaneous positioning and map construction method, comprising: Obtain a first image group captured by the camera at the current moment of the scene, and perform semantic segmentation and depth estimation processing on the first image group to obtain the semantic segmentation results of each feature point in the first image group and a first target depth image; the first image group includes a first RGB image and a first initial depth image; Acquire multiple second image groups that are temporally correlated with the first image group and are captured by the camera, and obtain semantic segmentation results of each feature point in each second image group and a second target depth image, and determine the target category corresponding to each feature point in the first RGB image according to the semantic segmentation results of each feature point in the first image group and the first target depth image, and the semantic segmentation results of each feature point in each second image group and the second target depth image; the above target category is used to characterize whether the feature point belongs to a static feature point or a dynamic feature point, and each second image group includes a second RGB image and a second initial depth image; According to the target categories corresponding to the feature points in the first RGB image, the feature points of the target category in the first RGB image that are dynamic are eliminated to obtain a target RGB image corresponding to the first RGB image; Simultaneous positioning and mapping are performed based on the target RGB image to determine the target pose corresponding to the camera and the target map corresponding to the scene.

[0007] According to a simultaneous positioning and mapping method provided by the present invention, the semantic segmentation results of each feature point in the first image group include semantic category labels of each feature point in the first RGB image, and the first target depth image includes depth values ​​of each feature point in the first RGB image. The target category corresponding to each feature point in the first RGB image is determined according to the semantic segmentation results of each feature point in the first image group and the first target depth image, the semantic segmentation results of each feature point in each second image group and the second target depth image, including: For each feature point in the first RGB image, determine change information of the semantic category label of the feature point according to the semantic category label corresponding to the feature point in the first RGB image and the semantic category label corresponding to the feature point in each second RGB image; Determine change information of the spatial position of the feature point according to the depth value corresponding to the feature point in the first target depth image and the depth value corresponding to the feature point in each second target depth image; According to the change information of the semantic category label and the spatial position of the feature point, the target category corresponding to the feature point is determined.

[0008] According to a simultaneous positioning and mapping method provided by the present invention, the above-mentioned determining the change information of the semantic category label of the feature point according to the semantic category label corresponding to the feature point in the first RGB image and the semantic category label corresponding to the feature point in each second RGB image includes: Determine an average semantic category label corresponding to the feature point according to the semantic category label corresponding to the feature point in the first RGB image and the semantic category label corresponding to the feature point in each second RGB image; According to the semantic category labels of the feature points in each frame of RGB images and the average semantic category labels corresponding to the feature points, the change information of the semantic category labels of the feature points in each frame of RGB images is determined; the above-mentioned each frame of RGB images includes a first RGB image and each second RGB image.

[0009] According to a simultaneous positioning and mapping method provided by the present invention, the above-mentioned determination of change information of the spatial position of the feature point according to the depth value corresponding to the feature point in the first target depth image and the depth value corresponding to the feature point in each second target depth image includes: Determine the spatial position corresponding to the feature point in each frame of the RGB image according to the depth value corresponding to the feature point in the first target depth image, the depth value corresponding to the feature point in each second target depth image, the two-dimensional position corresponding to the feature point in the first RGB image, and the two-dimensional position corresponding to the feature point in each second RGB image; According to the spatial positions corresponding to the feature points in each frame of the RGB image, the average spatial positions corresponding to the feature points are determined; According to the spatial position of the feature point in each frame of RGB image and the average spatial position corresponding to the feature point, the change information of the spatial position of the feature point in each frame of RGB image is determined.

[0010] According to a simultaneous positioning and mapping method provided by the present invention, the change information of the semantic category label of the feature point includes the change information of the semantic category label of the feature point in each frame of RGB image, and the change information of the spatial position of the feature point includes the change information of the spatial position of the feature point in each frame of RGB image. The target category corresponding to the feature point is determined according to the change information of the semantic category label and the change information of the spatial position of the feature point, including: According to the change information of the semantic category label of the feature point in each frame of RGB image and the change information of the spatial position of the feature point in each frame of RGB image, the spatiotemporal consistency quantization value corresponding to the feature point is calculated; According to the spatiotemporal consistency quantization value corresponding to the feature point and a preset threshold, the target category corresponding to the feature point is determined.

[0011] According to a method for simultaneous positioning and mapping provided by the present invention, the above-mentioned simultaneous positioning and mapping based on the target RGB image is performed to determine the target posture corresponding to the camera and the target map corresponding to the scene, including: Obtain static feature points of the target category in the target RGB image, and calculate the corresponding projection positions of the static feature points on the image plane of the camera; Get the semantic category labels corresponding to each area in the target RGB image and get the current position of the camera; The current position of the camera is used as a node, the projection position of the static feature points and the semantic category labels of each area are used as the edges of the nodes, and a graph model and the objective function corresponding to the graph model are constructed; The objective function is iteratively solved to determine the target pose corresponding to the camera and the target map corresponding to the scene; the target map is composed of static feature points.

[0012] According to a simultaneous positioning and mapping method provided by the present invention, the semantic segmentation and depth estimation processing is performed on the first image group to obtain the semantic segmentation results of each feature point in the first image group and the first target depth image, including: Using a preset lightweight semantic segmentation and depth estimation joint network to perform semantic segmentation and depth estimation processing on the first image group, obtaining a semantic segmentation result of each feature point in the first image group and a first target depth image; Among them, in the lightweight semantic segmentation and depth estimation joint network, the semantic segmentation processing and the depth estimation processing share some convolutional layers, and the information in the two parts is linked across layers.

[0013] The present invention also provides a simultaneous positioning and map building device, comprising the following modules: A depth semantic perception module is used to obtain a first image group captured by the camera at the current moment of the scene, and perform semantic segmentation and depth estimation processing on the first image group to obtain the semantic segmentation results of each feature point in the first image group and a first target depth image; the first image group includes a first RGB image and a first initial depth image; A feature point category determination module is used to obtain multiple second image groups collected by the camera and having a time correlation with the first image group, and obtain the semantic segmentation results of each feature point in each second image group and the second target depth image, and determine the target category corresponding to each feature point in the first RGB image according to the semantic segmentation results of each feature point in the first image group and the first target depth image, the semantic segmentation results of each feature point in each second image group and the second target depth image; the above target category is used to characterize whether the feature point belongs to a static feature point or a dynamic feature point, and each second image group includes a second RGB image and a second initial depth image; A dynamic feature point elimination module is used to eliminate the feature points of the first RGB image whose target category is dynamic according to the target category corresponding to each feature point in the first RGB image, so as to obtain a target RGB image corresponding to the first RGB image; The positioning and map construction module is used to perform simultaneous positioning and map construction based on the target RGB image, determine the target pose corresponding to the camera and the target map corresponding to the scene.

[0014] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the simultaneous positioning and mapping method as described above is implemented.

[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the simultaneous positioning and mapping method as described in any one of the above is implemented.

[0016] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned simultaneous positioning and map building methods.

[0017] The simultaneous positioning and mapping method, device, equipment and storage medium provided by the present invention obtain a first image group including a first RGB image and a first initial depth image captured by a camera at a current moment of a scene, and perform semantic segmentation and depth estimation processing on the first image group to obtain semantic segmentation results of each feature point in the first image group and a first target depth image, and simultaneously obtain multiple second image groups that are time-correlated with the first image group captured by the camera and obtain semantic segmentation results of each feature point in each second image group and a second target depth image, and determine the target category of each feature point in the first RGB image based on the multiple semantic segmentation results and the multiple target depth images, and then eliminate the feature points in the first RGB image that belong to a dynamic target category to obtain a target RGB image, and then perform simultaneous positioning and mapping based on the target RGB image to determine the target posture of the camera and the target map of the scene; wherein the target category is used to characterize whether the feature point belongs to a static feature point or a dynamic feature point, and each second image group includes a second RGB image and a second initial depth image. In this method, the semantic segmentation results and depth information of the scene images collected by the camera at each consecutive moment can be fused, and the dynamic feature points and static feature points in the current frame image can be accurately distinguished to accurately eliminate the dynamic feature points, thereby accurately estimating the camera's pose and accurately constructing the scene map, thereby achieving the purpose of improving the accuracy of camera pose estimation and the accuracy and robustness of map construction. At the same time, the robot does not need to rely on the rigid motion assumption, so it can reduce computing resource consumption and meet the needs of different robot applications in complex dynamic environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0019] Figure 1 This is one of the flow charts of the simultaneous positioning and map building method provided by the present invention.

[0020] Figure 2 This is the second flow chart of the simultaneous positioning and map building method provided by the present invention.

[0021] Figure 3 This is the third flow chart of the simultaneous positioning and map building method provided by the present invention.

[0022] Figure 4 It is a detailed flow chart of the simultaneous positioning and map building method provided by the present invention.

[0023] Figure 5 It is a structural schematic diagram of the simultaneous positioning and map building device provided by the present invention.

[0024] Figure 6 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0025] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0026] In terms of simultaneous positioning and map construction, some technologies use deep learning technology to provide semantic information to achieve simultaneous positioning and map construction, but they cannot accurately judge the real-time motion state of objects. There are also some technologies that use deep learning combined with geometric methods and other methods to process dynamic objects and / or static objects in the environment to achieve simultaneous positioning and map construction. There are still many limitations, such as limited semantic understanding ability, resulting in low accuracy of simultaneous positioning and map construction, poor processing of complex dynamic scenes, and low computational efficiency. There are also some technologies that use ORB-SLAM3 (a SLAM system that supports vision, vision plus inertial navigation, and hybrid maps) for simultaneous positioning and map construction, which performs well in static environments, but is greatly affected by dynamic feature points in dynamic environments. There are also some technologies that use DynaSLAM (a visual SLAM system designed specifically for dynamic scenes) for simultaneous positioning and map construction, which relies on rigid motion assumptions and has high computational overhead. Based on this, the embodiments of the present invention provide a method, device, equipment, and storage medium for simultaneous positioning and map construction, which can solve this technical problem.

[0027] Combine the following Figure 1-Figure 4 The simultaneous localization and mapping method of the present invention is described.

[0028] It should be noted that the execution subject of the embodiment of the present invention may be a simultaneous positioning and mapping device, or an electronic device including the simultaneous positioning and mapping device, or other devices or equipment, such as a SLAM system. The electronic device here may be a terminal or a server. In the case of a terminal, the electronic device may be a robot or an electronic device in a robot. The following embodiments are described by taking an electronic device as an execution subject as an example.

[0029] Figure 1 is one of the flow charts of the simultaneous positioning and map building method provided by the present invention, such as Figure 1 As shown, the method comprises the following steps: Step 102, obtain a first image group captured by the camera at the current moment of the scene, and perform semantic segmentation and depth estimation on the first image group to obtain semantic segmentation results of each feature point in the first image group and a first target depth image; the first image group includes a first RGB image and a first initial depth image.

[0030] The camera is a camera installed on the robot, which can be one camera or multiple cameras. The camera can include a common camera that collects red, green, and blue (RGB) images of the scene, or a depth camera that collects depth data of objects in the scene. Alternatively, a laser radar can be installed on the robot to collect depth data of objects in the scene. The scene here can be a scene where the robot needs to perform simultaneous localization and mapping (SLAM), which can be an indoor or outdoor scene.

[0031] The camera on the robot can collect data on the scene and the objects therein at the current moment, and obtain the RGB image and depth data collected at the current moment. Here, the RGB image collected at the current moment is recorded as the first RGB image, and the depth data collected at the current moment is recorded as the first initial depth image (also recorded as D), which may include the depth information / depth value of each point in the scene. The first RGB image and the first initial depth image here can be recorded as the first image group. Objects in the scene can include dynamic objects (such as pedestrians, moving vehicles, etc.) and static objects (such as buildings, trees, etc.), and each object can include multiple points.

[0032] After obtaining the first RGB image and the first initial depth image in the first image group, the information on the two images can be fused for semantic segmentation and depth estimation. As an optional embodiment, a preset lightweight semantic segmentation and depth estimation joint network can be used to perform semantic segmentation and depth estimation on the first image group to obtain the semantic segmentation results of each feature point in the first image group and the first target depth image.

[0033] The lightweight semantic segmentation and depth estimation joint network can be recorded as Lightweight Semantic Segmentation and Depth Estimation Network, abbreviated as LSDEN. The lightweight semantic segmentation and depth estimation joint network can be a network improved based on a deep learning network framework, such as a network improved based on a deep learning network framework TensorFlow or PyTorch.

[0034] The lightweight semantic segmentation and depth estimation joint network is used to perform semantic segmentation processing and depth estimation processing on the input RGB image and the initial depth image at the same time. In the lightweight semantic segmentation and depth estimation joint network, the semantic segmentation processing and the depth estimation processing share some convolutional layers, and the information in the semantic segmentation processing and the depth estimation processing are linked across layers. That is to say, when the first RGB image and the first initial depth image are subjected to semantic segmentation processing and depth estimation processing, the two tasks of semantic segmentation processing and depth estimation processing can share a part of the underlying convolutional layer (for example, the two tasks share the convolutional layer for extracting features), which can reduce the amount of calculation to improve the efficiency of feature extraction. At the same time, during the processing of the two tasks, the semantic segmentation related information obtained in the middle will be fused with the depth information through multi-source information fusion in a cross-convolutional layer manner, so as to better estimate the depth information of the points in the scene through the semantic segmentation information, and at the same time, the depth information of the depth estimation can be used to better perform semantic segmentation on the points in the scene, thereby improving the accuracy of the semantic segmentation results and the accuracy of the depth estimation. In addition, the lightweight semantic segmentation and depth estimation joint network can divide the input RGB image and the initial depth image into multiple feature maps of different scales, perform semantic segmentation and depth estimation at different scales, and obtain the final semantic segmentation result and target depth image through upsampling and fusion operations.

[0035] For the above-mentioned lightweight semantic segmentation and depth estimation joint network, it can be pre-trained, and its training process can include: 1. Data preparation: collect a large number of RGB-D image groups / image datasets containing different scenes, objects and lighting conditions, and manually annotate semantic category labels and depth information, and divide the / image dataset into training sets, validation sets and test sets according to a certain ratio. 2. Network training: Use a deep learning framework (such as TensorFlow or PyTorch) to build an LSDEN network, and use the cross entropy loss function and mean square error loss function to train the semantic segmentation task and depth estimation task respectively. At the same time, during the training process, by adjusting parameters such as the learning rate, optimizer (such as Adam optimizer) and number of iterations, the performance of the network is continuously optimized to minimize the loss function value of the network on the validation set.

[0036] After the lightweight semantic segmentation and depth estimation joint network is trained, the first RGB image and the first initial depth image collected at the current moment can be input into the lightweight semantic segmentation and depth estimation joint network to perform semantic segmentation processing and depth estimation processing at the same time, and the semantic segmentation result and depth image corresponding to the first RGB image can be obtained. The semantic segmentation results of each feature point and the first target depth image can be obtained through the semantic segmentation results and depth image corresponding to the first RGB image.

[0037] Here, the first RGB image includes the RGB value of each pixel, and the first initial depth image includes the depth information / depth value corresponding to each pixel on the first RGB image.

[0038] The semantic segmentation result and the depth image corresponding to the first RGB image output by the lightweight semantic segmentation and depth estimation joint network are obtained to obtain the final semantic segmentation results of each feature point and the first target depth image, which may include the following methods: Method 1: The semantic segmentation result corresponding to the first RGB image output by the lightweight semantic segmentation and depth estimation joint network includes the semantic segmentation result of each pixel therein, and the depth image also includes the depth value corresponding to each pixel therein. The feature points in the first RGB image are extracted by using an improved feature extraction algorithm (such as an improved method based on FAST feature points and ORB descriptors), and the position and feature description information of the extracted feature points are recorded. The feature description information of the feature points may include information describing the local appearance of the feature points. Then, based on the position of the feature points, the semantic segmentation result corresponding to the feature points is found in the semantic segmentation result corresponding to the first RGB image, and the depth value corresponding to the feature points is found in the depth image, and the semantic segmentation results of each feature point in the first image group are obtained through the semantic segmentation results corresponding to the feature points, and the first target depth image is obtained through the depth values ​​corresponding to the feature points.

[0039] Method 2: The semantic segmentation result corresponding to the first RGB image output by the lightweight semantic segmentation and depth estimation joint network includes the semantic segmentation results of each feature point, and the depth image includes the depth value of each feature point. The semantic segmentation result corresponding to the first RGB image can be directly used as the semantic segmentation result of each feature point in the first image group, and the depth image output by the lightweight semantic segmentation and depth estimation joint network can be directly used as the first target depth image.

[0040] It should be noted that the above-mentioned feature points may be representative points in a scene or object, such as boundary points, corner points, etc. The number of feature points may be one or more, which is specifically set according to actual conditions.

[0041] In addition, the semantic segmentation result of each feature point may include the category label corresponding to the feature point, such as pedestrians, vehicles, buildings, etc. The first target depth image includes the depth value of each feature point, which combines the semantic information of the feature point, that is, combines the understanding information of the scene content, can distinguish objects / feature points of different category labels in the first RGB image, and distinguish the boundaries and structures between objects / feature points of different category labels in the first RGB image, which can assist in more accurate simultaneous positioning and map construction.

[0042] Step 104, obtain multiple second image groups captured by the camera and having a time correlation with the first image group, and obtain the semantic segmentation results of each feature point in each second image group and the second target depth image, and determine the target category corresponding to each feature point in the first RGB image based on the semantic segmentation results of each feature point in the first image group and the first target depth image, and the semantic segmentation results of each feature point in each second image group and the second target depth image.

[0043] In this step, each second image group includes a second RGB image and a second initial depth image, that is, each second image group includes an RGB image and an initial depth image collected by the camera at the same time. The acquisition time of each second image group can be a time that is continuous with the acquisition time of the first image group, that is, the camera can continuously collect multiple frames of RGB images and multiple frames of initial depth images, so as to accurately analyze the real-time motion state of the feature points in the first RGB image at the current moment. Here, the acquisition time of each second image group can be before the current moment and continuous with the current moment, that is, each second image group is continuously collected at multiple moments before the current moment, or the acquisition time of each second image group can be after the current moment and continuous with the current moment, that is, each second image group is continuously collected at a moment after the current moment, or the acquisition time of each second image group can be a part of the acquisition time before the current moment and another part of the acquisition time after the current moment, and both are continuous. The number of each second image group collected here can be set according to actual conditions, for example, it can be 5 to 10 consecutive second image groups.

[0044] After obtaining each second image group collected at multiple moments continuous with the current moment, semantic segmentation processing and depth estimation processing can be performed on each group of second image groups in the manner of the above step 102 to obtain semantic segmentation results and a second target depth image for each feature point corresponding to each second image group; or after extracting feature points in the above step 102, the extracted feature points can be tracked and semantic segmentation processing and depth estimation processing can be performed in each second image group to obtain semantic segmentation results and a second target depth image for each feature point corresponding to each second image group.

[0045] Afterwards, the semantic segmentation results of each feature point corresponding to the first image group at the current moment and the first target depth image, and the semantic segmentation results of each feature point in the second image group at each consecutive moment and the second target depth image can be used to analyze whether the semantic information of the feature point at multiple consecutive moments has changed and whether the depth information has changed, that is, whether the feature point has changed in time and space, and determine the target category of the feature point in the first image group based on the information of whether it has changed, that is, determine the target category of the feature point in the first RGB image. The above target category is used to characterize whether the feature point belongs to a static feature point or a dynamic feature point.

[0046] Optionally, if a feature point changes in at least one of time and space (for example, a change in the semantic category label of the same feature point at different times is a change in time, and a change in the depth value of the same feature point at different times is a change in space), the target category of the feature point is determined to be a dynamic feature point, that is, the feature point belongs to a dynamically changing point. If the feature point does not change in time and space, the target category of the feature point is determined to be a static feature point, that is, the feature point belongs to an unchanging / fixed point.

[0047] Step 106 , according to the target categories corresponding to the feature points in the first RGB image, the feature points of the target category belonging to dynamic in the first RGB image are eliminated to obtain a target RGB image corresponding to the first RGB image.

[0048] In this step, after obtaining the target category corresponding to each feature point in the first RGB image, the feature points with dynamic target category can be removed in the first RGB image, for example, the pixel values ​​corresponding to the dynamic feature points can be filled with the pixel values ​​corresponding to the background. After removing the dynamic feature points, the final RGB image at the current moment can be obtained, which is recorded as the target RGB image. The target RGB image only includes static feature points, eliminating the interference of dynamic feature points, so the camera can be more accurately positioned and mapped at the same time.

[0049] Step 108, performing simultaneous positioning and mapping based on the target RGB image to determine the target pose corresponding to the camera and the target map corresponding to the scene.

[0050] In this step, after obtaining the target RGB image, the current pose of the camera can be calculated through the pixel value and pixel position of each pixel in the target RGB image, which is recorded as the target pose. At the same time, the spatial position of each feature point in the scene can be calculated based on this, and a map of the scene can be constructed based on this. The final constructed map can be recorded as the target map.

[0051] From the above description, it can be seen that the simultaneous positioning and mapping method in this embodiment is a method for distinguishing dynamic and static feature points using multiple views collected by a camera, which is suitable for dynamic environments and can effectively distinguish static points and dynamic points in dynamic scenes, and the accuracy of simultaneous positioning and mapping is high. The system corresponding to this method can also be called an improved multi-view semantic SLAM system suitable for dynamic environments.

[0052] In this embodiment, a first image group including a first RGB image and a first initial depth image captured by the camera at the current moment of the scene is obtained, and semantic segmentation and depth estimation processing are performed on the first image group to obtain semantic segmentation results of each feature point in the first image group and a first target depth image, and at the same time, a plurality of second image groups that are temporally correlated with the first image group and captured by the camera are obtained, and semantic segmentation results of each feature point in each second image group and a second target depth image are obtained, and the target category of each feature point in the first RGB image is determined according to the plurality of semantic segmentation results and the plurality of target depth images, and then feature points in the first RGB image that belong to a dynamic target category are eliminated to obtain a target RGB image, and then simultaneous positioning and mapping are performed according to the target RGB image to determine the target posture of the camera and the target map of the scene; wherein the target category is used to characterize whether the feature point belongs to a static feature point or a dynamic feature point, and each second image group includes a second RGB image and a second initial depth image. In this method, the semantic segmentation results and depth information of the scene images collected by the camera at each consecutive moment can be fused, and the dynamic feature points and static feature points in the current frame image can be accurately distinguished to accurately eliminate the dynamic feature points, thereby accurately estimating the camera's pose and accurately constructing the scene map, thereby achieving the purpose of improving the accuracy of camera pose estimation and the accuracy and robustness of map construction. At the same time, the robot does not need to rely on the rigid motion assumption, so it can reduce computing resource consumption and meet the needs of different robot applications in complex dynamic environments.

[0053] The following embodiment describes a possible implementation method for determining the target category corresponding to each feature point in the first RGB image according to the semantic segmentation results of the RGB images and the initial depth image collected at multiple consecutive moments and the target depth image.

[0054] Figure 2 FIG. 2 is a flow chart of the simultaneous positioning and map building method provided by the present invention. Figure 2 As shown, in the above step 104, determining the target category corresponding to each feature point in the first RGB image according to the semantic segmentation result of each feature point in the first image group and the first target depth image, the semantic segmentation result of each feature point in each second image group and the second target depth image, may include the following steps: Step 202 : for each feature point in the first RGB image, determine change information of the semantic category label of the feature point according to the semantic category label corresponding to the feature point in the first RGB image and the semantic category label corresponding to the feature point in each second RGB image.

[0055] The semantic segmentation results of each feature point in the first image group include the semantic category labels of each feature point in the first RGB image, and the first target depth image includes the depth value of each feature point in the first RGB image. The semantic segmentation results of each feature point in the second image group include the semantic category labels of each feature point in the second RGB image, and the second target depth image includes the depth value of each feature point in the second RGB image.

[0056] In this embodiment, a spatio-temporal consistency based dynamic feature point filtering algorithm (STCDF) can be used to determine the target category of each feature point in the first RGB image, thereby distinguishing the dynamic feature points and the static feature points in the first RGB image.

[0057] Optionally, for the same feature point in the first RGB image, the average semantic category label corresponding to the feature point can be determined based on the semantic category label corresponding to the feature point in the first RGB image and the semantic category label corresponding to the feature point in each second RGB image; based on the semantic category label of the feature point in each frame of RGB image and the average semantic category label corresponding to the feature point, the change information of the semantic category label of the feature point in each frame of RGB image is determined; the above-mentioned frames of RGB images include the first RGB image and each second RGB image.

[0058] Among them, for the same feature point, the semantic category label corresponding to the feature point in the first RGB image and the semantic category label corresponding to each second RGB image can be directly summed and averaged to obtain the average semantic category label corresponding to the feature point; or corresponding weights can be set for the first RGB image and each second RGB image respectively, and then the semantic category label corresponding to the feature point in the first RGB image and the semantic category label corresponding to the second RGB image are weighted summed and averaged to obtain the average semantic category label corresponding to the feature point.

[0059] After obtaining the average semantic category label corresponding to each feature point, the absolute value of the difference between the semantic category label corresponding to the feature point in each frame of the RGB image and its average semantic category label can be calculated to obtain multiple semantic category label differences, and these semantic category label differences are used as the change information of the semantic category label of the feature point in each frame of the RGB image.

[0060] According to the above method, the change information of the semantic category labels of all feature points in the first RGB image in each frame of RGB image can be calculated, which will not be repeated here.

[0061] Step 204 : determining change information of the spatial position of the feature point according to the depth value corresponding to the feature point in the first target depth image and the depth value corresponding to the feature point in each second target depth image.

[0062] The first RGB image also includes the two-dimensional position / two-dimensional coordinate (x, y) of each feature point, and each second RGB image also includes the two-dimensional position / two-dimensional coordinate (x, y) of each feature point.

[0063] In this step, after obtaining the depth value corresponding to each feature point in the first RGB image in the first target depth image and the depth value corresponding to each second target depth image, as an option, for the same feature point in the first RGB image, the spatial position corresponding to the feature point in each frame of RGB image can be determined according to the depth value corresponding to the feature point in the first target depth image, the depth value corresponding to the feature point in each second target depth image, the two-dimensional position corresponding to the feature point in the first RGB image, and the two-dimensional position corresponding to the feature point in each second RGB image; according to the spatial position corresponding to the feature point in each frame of RGB image, the average spatial position corresponding to the feature point is determined; according to the spatial position of the feature point in each frame of RGB image and the average spatial position corresponding to the feature point, the change information of the spatial position of the feature point in each frame of RGB image is determined. Each frame of RGB image includes a first RGB image and each second RGB image.

[0064] Among them, for the same feature point, the depth value z corresponding to the feature point can be obtained through the first target depth image, and the two-dimensional position (x, y) corresponding to the feature point can be obtained through the first RGB image. Then, combined with the camera internal parameters, the two-dimensional position (x, y) of the feature point can be converted into a position (X, Y, Z) in three-dimensional space to obtain the spatial position of the feature point in the first RGB image. Similarly, the spatial position of the feature point in each second RGB image can be obtained in this way. Afterwards, the spatial position of the feature point in each frame of the RGB image can be directly summed and averaged along the axial direction to obtain the average spatial position of the feature point; or the spatial position of the feature point in each frame of the RGB image can be weighted summed and averaged along the axial direction to obtain the average spatial position of the feature point.

[0065] After obtaining the average spatial position of each feature point, the absolute value of the difference between the spatial position of the feature point in each frame of the RGB image and its corresponding average spatial position can be calculated to obtain multiple spatial position differences, and these spatial position differences are used as the change information of the spatial position of the feature point in each frame of the RGB image.

[0066] According to the above method, the change information of the spatial positions of all the feature points in the first RGB image in each frame of the RGB image can be calculated, which will not be repeated here.

[0067] Step 206 : determining the target category corresponding to the feature point according to the change information of the semantic category label and the change information of the spatial position of the feature point.

[0068] Among them, the change information of the semantic category label of the feature point includes the change information of the semantic category label of the feature point in each frame of RGB image, the change information of the spatial position of the feature point includes the change information of the spatial position of the feature point in each frame of RGB image, and each frame of RGB image includes a first RGB image and each second RGB image.

[0069] After obtaining the change information of the semantic category label of each feature point in each frame of RGB image and the change information of the spatial position of each feature point in each frame of RGB image, as an option, for the same feature point in the first RGB image, the spatiotemporal change information of the feature point in the same frame of RGB image can be determined according to the change information of the semantic category label and the change information of the spatial position of the feature point in the same frame of RGB image; the weight corresponding to each frame of RGB image is obtained; and the target category corresponding to the feature point is determined according to the weight corresponding to each frame of RGB image and the spatiotemporal change information in each frame of RGB image. The determining the target category corresponding to the feature point according to the weight corresponding to each frame of RGB image and the spatiotemporal change information in each frame of RGB image can include: for each feature point, according to the spatiotemporal change information of the feature point in the same frame of RGB image and the weight corresponding to each frame of RGB image, the spatiotemporal consistency quantization value corresponding to the feature point is calculated; according to the spatiotemporal consistency quantization value corresponding to the feature point and a preset threshold, the target category corresponding to the feature point is determined.

[0070] Alternatively, as an option, the spatiotemporal consistency quantization value corresponding to the feature point can be calculated directly based on the change information of the semantic category label of the feature point in each frame of RGB image and the change information of the spatial position of the feature point in each frame of RGB image; and the target category corresponding to the feature point can be determined based on the spatiotemporal consistency quantization value corresponding to the feature point and a preset threshold.

[0071] Specifically, for the same / each feature point, the change information of the semantic category label and the change information of the spatial position of the feature point in each frame of RGB image can be directly summed or weighted summed to obtain the spatiotemporal change information of the feature point in the same frame / each frame of RGB image; then the spatiotemporal change information of the feature point in the same frame / each frame of RGB image is multiplied by the weight corresponding to the corresponding frame of RGB image to obtain the weighted value of the feature point in each frame of RGB image, and then the weighted value of the feature point in each frame of RGB image is summed to obtain the change information and value of the feature point in each frame of RGB image, which is recorded as the spatiotemporal consistency quantization value. Then the spatiotemporal consistency quantization value of the feature point is compared with the preset threshold value. If the spatiotemporal consistency quantization value of the feature point is greater than the preset threshold value, it is determined that the target category corresponding to the feature point is a dynamic feature point. If the spatiotemporal consistency quantization value of the feature point is not greater than (less than or equal to) the preset threshold value, it is determined that the target category corresponding to the feature point is a static feature point. The above preset threshold is determined through experiments and data analysis.

[0072] As an option, the above-mentioned dynamic feature point screening algorithm STCDF based on spatiotemporal consistency can be expressed in the form of the following formula and its variation: .

[0073] Where i represents the i-th feature point in the first RGB image; t represents the t-th frame RGB image, T represents the total number of consecutive frames of RGB images; C i represents the spatiotemporal consistency quantization value of the i-th feature point; w t represents the weight corresponding to the RGB image of the tth frame; S i,t Indicates the semantic category label corresponding to the i-th feature point in the t-th frame RGB image; Represents the average semantic category label corresponding to the i-th feature point; represents the change information of the semantic category label of the i-th feature point in the t-th frame RGB image; P i,t Indicates the spatial position corresponding to the i-th feature point in the t-th frame RGB image; Represents the average spatial position corresponding to the i-th feature point; Represents the change information of the spatial position of the i-th feature point in the t-th frame RGB image; represents the parameter that balances the weights of semantic category labels and spatial positions.

[0074] In this embodiment, the semantic category label of the feature point in each frame of RGB image and the depth value in each depth image are used to determine the change information of the semantic category label of the feature point and the change information of the spatial position, and the dynamic and static categories of the feature point are determined accordingly, so that the dynamic points and static points in the first RGB image can be effectively distinguished, and the accuracy of subsequent simultaneous positioning and map construction can be improved. In addition, by calculating the average semantic category label of the feature point under each frame of RGB image, and determining the change information of the semantic category label of the feature point together with the semantic category label in each frame of RGB image, this method is simple and effective, so the efficiency and accuracy of determining the change information of the semantic category label of the feature point can be improved, thereby improving the accuracy of subsequent simultaneous positioning and map construction and reducing the amount of calculation for simultaneous positioning and map construction. At the same time, by calculating the average spatial position of the feature point under each frame of RGB image, and determining the change information of the spatial position of the feature point together with the spatial position under each frame of RGB image, this method is simple and effective, so the efficiency and accuracy of determining the change information of the spatial position of the feature point can be improved, thereby improving the accuracy of subsequent simultaneous positioning and map construction and reducing the amount of calculation for simultaneous positioning and map construction. Furthermore, the spatiotemporal consistency quantization value of the feature point is calculated through the change information of the semantic category label and the change information of the spatial position of the feature point, and the target category of the feature point is determined by comparing with the threshold. This can improve the efficiency and accuracy of distinguishing dynamic points from static points, thereby improving the accuracy of subsequent simultaneous positioning and map construction and reducing the computational complexity of simultaneous positioning and map construction.

[0075] The following example illustrates a possible implementation method of simultaneous positioning and map construction based on the target RGB image.

[0076] Figure 3 FIG. 3 is a flow chart of the simultaneous positioning and map building method provided by the present invention. Figure 3 As shown, in the above step 108, simultaneous positioning and mapping are performed according to the target RGB image to determine the target posture corresponding to the camera and the target map corresponding to the scene, which may include the following steps: Step 302, obtaining feature points of the target category belonging to static in the target RGB image, and calculating the corresponding projection positions of the static feature points on the image plane of the camera.

[0077] Step 304, obtaining the semantic category labels corresponding to each region in the target RGB image and obtaining the current position and posture of the camera.

[0078] Step 306 , taking the current position of the camera as a node, taking the projection position of the static feature points and the semantic category labels of each region as the edges of the node, and constructing a graph model and an objective function corresponding to the graph model.

[0079] Step 308, iteratively solve the objective function to determine the target pose corresponding to the camera and the target map corresponding to the scene; the target map is composed of static feature points.

[0080] In this embodiment, an improved multi-view geometry optimization module is used to perform simultaneous positioning and map construction, and in the processing of the improved multi-view geometry optimization module, a multi-view geometry optimization method based on graph optimization (Graph-based Optimization for Multiview Geometry, GOMG) is introduced to perform simultaneous positioning and map construction.

[0081] Among them, after obtaining the target RGB image, the static feature points and the corresponding position information of the static feature points in the target RGB image can be obtained. At the same time, when each camera collects the RGB image and the initial depth image at the current moment, the image plane corresponding to each camera will be formed, and then the corresponding projection position of the static feature points on the image plane of each camera can be obtained through the corresponding position information of the static feature points in the target RGB image, which is recorded as the observation point. At the same time, when each camera collects images, the current posture of the camera can be obtained, which is recorded as the current posture or posture transformation matrix (including parameters such as the translation and rotation of the camera).

[0082] In addition, the target RGB image can be divided into multiple regions in advance, each region can include one or more feature points, or each region can include one or more pixels, and then the semantic category label corresponding to each region can be obtained through the semantic category label of the feature point in each region. At the same time, a deep learning network that is pre-trained to identify regional semantic category labels can be used to identify each region in the target RGB image to obtain the predicted semantic category label of each region.

[0083] Afterwards, during the operation of the robot / SLAM system, the current position of the camera can be used as a node, and the projection position of the static feature points and the semantic category labels and predicted semantic category labels (i.e., semantic constraints) of each region can be used as the edges of the node to construct a graph model and an objective function corresponding to the graph model. The objective function can be an objective function that minimizes the reprojection error and the semantic consistency error. The objective function of the graph model constructed above is shown in the following formula: .

[0084] Where E is the total error of the objective function, ρ is the robust kernel function, π is the projection function, and T j is the pose transformation matrix of the jth camera, P i is the world coordinate of the ith feature point (which can be calculated using the method for calculating the spatial position of the feature point in the previous embodiment), p i,j is the observation point / projection position of the i-th feature point on the image plane of the j-th camera, S k is the semantic category label of the kth region in the target RGB image, is the predicted semantic label of the kth region in the target RGB image, and µ is the error weight of the semantic category label.

[0085] After determining the objective function, a nonlinear optimization algorithm (such as the Levenberg-Marquardt algorithm) can be used to iteratively optimize the objective function of the graph model to continuously update the camera pose and map until convergence (for example, if the change in the objective function value E after multiple consecutive iterations is less than a preset threshold, it converges), and finally obtain the optimal target pose of the camera at the current moment and the map composed of each three-dimensional point in the scene.

[0086] In this embodiment, by constructing a graph model including camera pose, feature points and semantic information, and using the constraint relationship between nodes for joint optimization, the accuracy of camera pose estimation and map construction can be improved. Compared with the traditional method based only on geometric constraints, the method based on GOMG in this embodiment can better utilize semantic information to constrain the optimization process, reduce error accumulation, and effectively improve the accuracy of simultaneous positioning and map construction.

[0087] A detailed embodiment is given below to illustrate the simultaneous positioning and map construction method of the present invention. Figure 4 The detailed flow chart of the simultaneous positioning and mapping method provided by the present invention is shown in FIG. The simultaneous positioning and mapping method in the embodiment of the present invention can be implemented by using a SLAM system. The SLAM system can include a depth semantic perception module, a dynamic feature point screening module, and a multi-view geometry optimization module. Based on these modules, the method can include: The cameras on the robot collect RGB images and initial depth images at the current moment, and then the deep semantic perception module extracts features through the LSDEN network to obtain feature maps, and performs semantic segmentation and depth estimation on the feature maps to obtain the semantic segmentation results and target depth images at the current moment. At the same time, the multiple cameras on the robot can collect multiple frames of RGB images and multiple frames of initial depth images at multiple moments consecutive to the current moment, and use the LSDEN network to obtain the semantic segmentation results and target depth images at each moment.

[0088] Afterwards, the dynamic feature point screening module uses the STCDF algorithm to calculate the spatiotemporal consistency quantization value of the feature points based on the semantic segmentation results and the target depth image at the current moment and the semantic segmentation results and the target depth image at each subsequent moment, and performs threshold screening through the calculated spatiotemporal consistency quantization value of the feature points to eliminate the dynamic points in the RGB image at the current moment and obtain the target RGB image that only includes static feature points.

[0089] Then, the multi-view geometry optimization module adopts the GOMG method to perform pose estimation and map construction based on the target RGB image, and continuously iterates and optimizes the objective function of the constructed graph model to finally obtain the precise pose of the camera and the constructed map.

[0090] It can be seen from the above description that the technical solution of this embodiment has the following technical effects: Through the innovative design of the deep semantic perception module, the semantic and depth information of the environment can be obtained more accurately, which improves the richness and accuracy of information compared to traditional methods and provides a better foundation for subsequent processing. The spatiotemporal consistency algorithm of the dynamic feature point screening module effectively utilizes the information of multiple frames of images, can more accurately identify dynamic feature points, and reduce misjudgment, thereby significantly improving the accuracy of camera pose estimation. The positioning error in complex dynamic scenes is significantly reduced compared to traditional technologies. The improved multi-view geometry optimization module uses graph optimization and semantic information fusion to further improve the accuracy and robustness of map construction, and the generated map performs better in terms of details and consistency. While ensuring performance improvement, the overall system of this embodiment reduces computing resource consumption through the optimization of network structure and algorithm, making it more suitable for resource-constrained platforms, such as small mobile robots, and has a wider range of application scenarios.

[0091] The simultaneous positioning and mapping apparatus provided by the present invention is described below. The simultaneous positioning and mapping apparatus described below and the simultaneous positioning and mapping method described above can be referred to each other.

[0092] Figure 5 is a structural diagram of the simultaneous positioning and map building device provided by the present invention, see Figure 5 As shown, the device may include: The depth semantic perception module 510 is used to obtain a first image group captured by the camera at the current moment of the scene, and perform semantic segmentation and depth estimation processing on the first image group to obtain the semantic segmentation results of each feature point in the first image group and a first target depth image; the first image group includes a first RGB image and a first initial depth image; The feature point category determination module 520 is used to obtain multiple second image groups collected by the camera and having a time correlation with the first image group, and obtain the semantic segmentation results of each feature point in each second image group and the second target depth image, and determine the target category corresponding to each feature point in the first RGB image according to the semantic segmentation results of each feature point in the first image group and the first target depth image, the semantic segmentation results of each feature point in each second image group and the second target depth image; the above target category is used to characterize whether the feature point belongs to a static feature point or a dynamic feature point, and each second image group includes a second RGB image and a second initial depth image; The dynamic feature point elimination module 530 is used to eliminate the feature points of the first RGB image whose target category is dynamic according to the target category corresponding to each feature point in the first RGB image, so as to obtain a target RGB image corresponding to the first RGB image; The positioning and map construction module 540 is used to perform simultaneous positioning and map construction based on the target RGB image, and determine the target posture corresponding to the camera and the target map corresponding to the scene.

[0093] In some embodiments, the semantic segmentation result of each feature point in the first image group includes a semantic category label of each feature point in the first RGB image, the first target depth image includes a depth value of each feature point in the first RGB image, and the feature point category determination module 520 includes: A semantic label change information determining unit, configured to determine, for each feature point in the first RGB image, change information of the semantic category label of the feature point according to the semantic category label corresponding to the feature point in the first RGB image and the semantic category label corresponding to the feature point in each second RGB image; A spatial position change information determining unit, configured to determine change information of the spatial position of the feature point according to a depth value corresponding to the feature point in the first target depth image and a depth value corresponding to the feature point in each second target depth image; The feature point category determination unit is used to determine the target category corresponding to the feature point according to the change information of the semantic category label and the change information of the spatial position of the feature point.

[0094] Optionally, the above-mentioned semantic label change information determination unit is specifically used to determine the average semantic category label corresponding to the feature point based on the semantic category label corresponding to the feature point in the first RGB image and the semantic category label corresponding to the feature point in each second RGB image; determine the change information of the semantic category label of the feature point in each frame of RGB image based on the semantic category label of the feature point in each frame of RGB image and the average semantic category label corresponding to the feature point; the above-mentioned RGB image frames include the first RGB image and each second RGB image.

[0095] Optionally, the above-mentioned spatial position change information determination unit is specifically used to determine the spatial position corresponding to the feature point in each frame of RGB image according to the depth value corresponding to the feature point in the first target depth image, the depth value corresponding to the feature point in each second target depth image, the two-dimensional position corresponding to the feature point in the first RGB image, and the two-dimensional position corresponding to the feature point in each second RGB image; determine the average spatial position corresponding to the feature point according to the spatial position corresponding to the feature point in each frame of RGB image; determine the change information of the spatial position of the feature point in each frame of RGB image according to the spatial position of the feature point in each frame of RGB image and the average spatial position corresponding to the feature point.

[0096] Optionally, the change information of the semantic category label of the feature point includes the change information of the semantic category label of the feature point in each frame of RGB image, and the change information of the spatial position of the feature point includes the change information of the spatial position of the feature point in each frame of RGB image. The feature point category determination unit is specifically used to calculate the spatiotemporal consistency quantization value corresponding to the feature point based on the change information of the semantic category label of the feature point in each frame of RGB image and the change information of the spatial position of the feature point in each frame of RGB image; determine the target category corresponding to the feature point based on the spatiotemporal consistency quantization value corresponding to the feature point and a preset threshold.

[0097] In some embodiments, the positioning and map construction module 540 is specifically used to obtain static feature points belonging to the target category in the target RGB image, and calculate the corresponding projection positions of the static feature points on the image plane of the camera; obtain the semantic category labels corresponding to each area in the target RGB image and obtain the current posture of the camera; use the current posture of the camera as a node, the projection positions of the static feature points and the semantic category labels of each area as the edges of the node, and construct a graph model and an objective function corresponding to the graph model; iteratively solve the objective function to determine the target posture corresponding to the camera and the target map corresponding to the scene; the above target map is composed of static feature points.

[0098] In some embodiments, the above-mentioned depth semantic perception module 510 is specifically used to perform semantic segmentation and depth estimation processing on the first image group using a preset lightweight semantic segmentation and depth estimation joint network to obtain the semantic segmentation results of each feature point in the first image group and the first target depth image; wherein, in the lightweight semantic segmentation and depth estimation joint network, the semantic segmentation processing and the depth estimation processing share some convolutional layers, and the information in the semantic segmentation processing and the depth estimation processing are linked across layers.

[0099] It should be noted here that the above-mentioned device provided in the embodiment of the present invention can implement all the method steps implemented in the above-mentioned method embodiment, and can achieve the same technical effect. The parts and beneficial effects that are the same as the method embodiment in this embodiment will not be described in detail here.

[0100] Figure 6 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 6As shown, the electronic device may include: a processor 610, a communications interface 620, a memory 630 and a communication bus 640, wherein the processor 610, the communications interface 620 and the memory 630 communicate with each other via the communications bus 640. The processor 610 may call the logic instructions in the memory 630 to execute a simultaneous positioning and map building method, the method comprising: obtaining a first image group captured by the camera of the scene at the current moment, and performing semantic segmentation and depth estimation processing on the first image group to obtain the semantic segmentation results of each feature point in the first image group and a first target depth image; the first image group includes a first RGB image and a first initial depth image; obtaining a plurality of second image groups captured by the camera and having a time correlation with the first image group, and obtaining the semantic segmentation results of each feature point in each second image group and a second target depth image, and according to the semantic segmentation results of each feature point in the first image group and the first The target depth image, the semantic segmentation results of each feature point in each second image group and the second target depth image are used to determine the target category corresponding to each feature point in the first RGB image; the above target category is used to characterize whether the feature point belongs to a static feature point or a dynamic feature point, and each second image group includes a second RGB image and a second initial depth image; according to the target category corresponding to each feature point in the first RGB image, the feature points whose target category in the first RGB image is dynamic are eliminated to obtain the target RGB image corresponding to the first RGB image; simultaneous positioning and map construction are performed based on the target RGB image to determine the target posture corresponding to the camera and the target map corresponding to the scene.

[0101] In addition, the logic instructions in the above-mentioned memory 630 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0102] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the simultaneous positioning and mapping method provided by the above methods, which includes: obtaining a first image group captured by the camera of the scene at the current moment, and performing semantic segmentation and depth estimation on the first image group to obtain the semantic segmentation results of each feature point in the first image group and a first target depth image; the first image group includes a first RGB image and a first initial depth image; obtaining multiple second image groups captured by the camera and having time correlation with the first image group, and obtaining the semantic segmentation results of each feature point in each second image group and the first target depth image; Two target depth images, and determine the target category corresponding to each feature point in the first RGB image according to the semantic segmentation result of each feature point in the first image group and the first target depth image, the semantic segmentation result of each feature point in each second image group and the second target depth image; the above target category is used to characterize whether the feature point belongs to a static feature point or a dynamic feature point, and each second image group includes a second RGB image and a second initial depth image; according to the target category corresponding to each feature point in the first RGB image, the feature points whose target category in the first RGB image is dynamic are eliminated to obtain the target RGB image corresponding to the first RGB image; simultaneous positioning and map construction are performed according to the target RGB image to determine the target posture corresponding to the camera and the target map corresponding to the scene.

[0103] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the simultaneous positioning and mapping method provided by the above-mentioned methods, the method comprising: obtaining a first image group captured by the camera of the scene at the current moment, and performing semantic segmentation and depth estimation processing on the first image group to obtain the semantic segmentation results of each feature point in the first image group and a first target depth image; the first image group comprises a first RGB image and a first initial depth image; obtaining a plurality of second image groups captured by the camera and having a time correlation with the first image group, and obtaining the semantic segmentation results of each feature point in each second image group and a second target depth image, and according to the first image The semantic segmentation results of each feature point in the group and the first target depth image, the semantic segmentation results of each feature point in each second image group and the second target depth image, determine the target category corresponding to each feature point in the first RGB image; the above target category is used to characterize whether the feature point belongs to a static feature point or a dynamic feature point, and each second image group includes a second RGB image and a second initial depth image; according to the target category corresponding to each feature point in the first RGB image, the feature points whose target category in the first RGB image is dynamic are eliminated to obtain the target RGB image corresponding to the first RGB image; simultaneous positioning and map construction are performed according to the target RGB image to determine the target posture corresponding to the camera and the target map corresponding to the scene.

[0104] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0105] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for simultaneous positioning and map construction, characterized in that: include: Acquire a first image group captured by the camera at the current moment of the scene, and perform semantic segmentation and depth estimation processing on the first image group to obtain a semantic segmentation result of each feature point in the first image group and a first target depth image; The first image group includes a first RGB image and a first initial depth image; Acquire a plurality of second image groups captured by the camera and having a temporal correlation with the first image group, and acquire a semantic segmentation result of each feature point in each of the second image groups and a second target depth image, and determine a target category corresponding to each feature point in the first RGB image according to the semantic segmentation result of each feature point in the first image group and the first target depth image, and the semantic segmentation result of each feature point in each of the second image groups and the second target depth image; the target category is used to characterize whether the feature point belongs to a static feature point or a dynamic feature point, and each of the second image groups includes a second RGB image and a second initial depth image; According to the target categories corresponding to the feature points in the first RGB image, the feature points of the target category in the first RGB image are removed to obtain a target RGB image corresponding to the first RGB image; Simultaneous positioning and mapping are performed according to the target RGB image to determine the target posture corresponding to the camera and the target map corresponding to the scene.

2. The method for simultaneous positioning and mapping according to claim 1, characterized in that: The semantic segmentation result of each feature point in the first image group includes a semantic category label of each feature point in the first RGB image, and the first target depth image includes a depth value of each feature point in the first RGB image. The determining, according to the semantic segmentation result of each feature point in the first image group and the first target depth image, the semantic segmentation result of each feature point in each second image group and the second target depth image, the target category corresponding to each feature point in the first RGB image includes: For each feature point in the first RGB image, determine change information of the semantic category label of the feature point according to the semantic category label corresponding to the feature point in the first RGB image and the semantic category label corresponding to the feature point in each of the second RGB images; Determine change information of the spatial position of the feature point according to the depth value corresponding to the feature point in the first target depth image and the depth value corresponding to the feature point in each of the second target depth images; The target category corresponding to the feature point is determined according to the change information of the semantic category label and the change information of the spatial position of the feature point.

3. The method for simultaneous positioning and mapping according to claim 2, characterized in that: The determining, according to the semantic category label corresponding to the feature point in the first RGB image and the semantic category label corresponding to the feature point in each of the second RGB images, change information of the semantic category label of the feature point includes: Determine an average semantic category label corresponding to the feature point according to the semantic category label corresponding to the feature point in the first RGB image and the semantic category label corresponding to the feature point in each of the second RGB images; According to the semantic category label of the feature point in each frame of RGB image and the average semantic category label corresponding to the feature point, the change information of the semantic category label of the feature point in each frame of RGB image is determined; each frame of RGB image includes the first RGB image and each second RGB image.

4. The method for simultaneous positioning and mapping according to claim 2, characterized in that: The determining, according to the depth value corresponding to the feature point in the first target depth image and the depth value corresponding to the feature point in each of the second target depth images, the change information of the spatial position of the feature point includes: Determine the spatial position corresponding to the feature point in each frame of the RGB image according to the depth value corresponding to the feature point in the first target depth image, the depth value corresponding to the feature point in each of the second target depth images, the two-dimensional position corresponding to the feature point in the first RGB image, and the two-dimensional position corresponding to the feature point in each of the second RGB images; Determine the average spatial position corresponding to the feature point according to the spatial position corresponding to the feature point in each frame of the RGB image; According to the spatial position of the feature point in each frame of RGB image and the average spatial position corresponding to the feature point, the change information of the spatial position of the feature point in each frame of RGB image is determined.

5. The method for simultaneous positioning and mapping according to claim 2, characterized in that: The change information of the semantic category label of the feature point includes the change information of the semantic category label of the feature point in each frame of RGB image, the change information of the spatial position of the feature point includes the change information of the spatial position of the feature point in each frame of RGB image, and determining the target category corresponding to the feature point according to the change information of the semantic category label and the change information of the spatial position of the feature point includes: Calculate the spatiotemporal consistency quantization value corresponding to the feature point according to the change information of the semantic category label of the feature point in each frame of RGB image and the change information of the spatial position of the feature point in each frame of RGB image; The target category corresponding to the feature point is determined according to the spatiotemporal consistency quantization value corresponding to the feature point and a preset threshold.

6. The simultaneous positioning and mapping method according to any one of claims 1 to 5, characterized in that: The simultaneous positioning and mapping according to the target RGB image to determine the target posture corresponding to the camera and the target map corresponding to the scene includes: Obtaining static feature points of the target category in the target RGB image, and calculating the corresponding projection positions of the static feature points on the image plane of the camera; Obtaining the semantic category labels corresponding to each region in the target RGB image and obtaining the current position and posture of the camera; Taking the current position of the camera as a node, taking the projection position of the static feature point and the semantic category label of each area as the edge of the node, constructing a graph model and an objective function corresponding to the graph model; The objective function is iteratively solved to determine the target posture corresponding to the camera and the target map corresponding to the scene; the target map is composed of the static feature points.

7. The method for simultaneous positioning and mapping according to any one of claims 1 to 5, characterized in that: The performing semantic segmentation and depth estimation processing on the first image group to obtain a semantic segmentation result of each feature point in the first image group and a first target depth image includes: Using a preset lightweight semantic segmentation and depth estimation joint network to perform semantic segmentation and depth estimation processing on the first image group, to obtain a semantic segmentation result of each feature point in the first image group and a first target depth image; Among them, in the lightweight semantic segmentation and depth estimation joint network, the semantic segmentation processing and the depth estimation processing share some convolutional layers, and the information in the semantic segmentation processing and the depth estimation processing are linked across layers.

8. A simultaneous positioning and map building device, characterized in that: include: A depth semantic perception module is used to obtain a first image group captured by the camera at the current moment of the scene, and perform semantic segmentation and depth estimation processing on the first image group to obtain a semantic segmentation result of each feature point in the first image group and a first target depth image; The first image group includes a first RGB image and a first initial depth image; A feature point category determination module, used to obtain a plurality of second image groups captured by the camera and having a time correlation with the first image group, and to obtain a semantic segmentation result of each feature point in each of the second image groups and a second target depth image, and to determine a target category corresponding to each feature point in the first RGB image according to the semantic segmentation result of each feature point in the first image group and the first target depth image, and the semantic segmentation result of each feature point in each of the second image groups and the second target depth image; the target category is used to characterize whether the feature point belongs to a static feature point or a dynamic feature point, and each of the second image groups includes a second RGB image and a second initial depth image; A dynamic feature point elimination module, used to eliminate feature points belonging to dynamic target categories in the first RGB image according to target categories corresponding to each feature point in the first RGB image, so as to obtain a target RGB image corresponding to the first RGB image; The positioning and map construction module is used to perform simultaneous positioning and map construction according to the target RGB image, and determine the target posture corresponding to the camera and the target map corresponding to the scene.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the simultaneous positioning and mapping method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the simultaneous positioning and mapping method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Dynamic environment camera pose estimation and semantic map construction method based on semantic SLAM

    CN111402336A

  • Image processing method and device, shooting device and movable platform

    CN111837158A

  • Outdoor monocular synchronous mapping and positioning method fusing scene semantics

    CN112734845A

  • Three-dimensional map construction method and device, electronic equipment and storage medium

    CN113674416A

  • Semantic-based positioning and mapping method and system and intelligent robot

    CN114926536A