A semantically-integrated unsupervised depth estimation and visual odometry method and system
By introducing semantic segmentation and adaptive semantic masks into the visual odometer method, the impact of illumination changes and dynamic objects on reprojection losses is solved, and the performance of visual odometers in outdoor scenes is improved.
Patent Information
- Application Number
- CN202410070464.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-18
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-01-18
AI Technical Summary
The utility of reprojection loss is affected when the existing visual odometry method faces light changes and dynamic objects, especially in outdoor scenes, where the model's pose estimation subnet cannot be effectively constrained.
Unsupervised depth estimation and visual odometry method of fused semantics are used to semantic segmentation subnetworks on the target image and the reconstructed target image, construct semantic masks and semantic reprojection losses, reduce the impact of illumination changes, and avoid the influence of dynamic objects and infinity regions through adaptive semantic masks.
It effectively reduces the impact of light changes and dynamic objects on reprojection loss, enhances the effectiveness of reprojection loss, and improves the training effect of the model in outdoor scenes.
Smart Images

Figure CN118052841B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a semantically integrated unsupervised depth estimation and visual odometer method and system. Background Art
[0002] Real-time Localization and Mapping (SLAM) aims to enable mobile robots to achieve self-localization and map construction without a preset environmental map. The SLAM system is mainly divided into two parts: the front-end and the back-end. The front-end visual odometry aims to estimate the movement of the robot through visual data, providing the necessary data for the back-end data optimization and mapping. Researchers first used feature extraction methods such as ORB and SIFT to implement the visual odometry task. However, these traditional feature extraction methods can only obtain low-dimensional features of the image. As deep learning can obtain high-dimensional and high-representation features of images through continuous training iterations, a visual odometry calculation method based on deep learning has been proposed.
[0003] Existing visual odometry calculation methods generally use reprojection loss as their constraint means. Assume that the image at time t is called the target image, and the image at time t+1 is called the source image. Reprojection can use the depth of the target image and the pose information between the target image and the source image to obtain the reconstructed target image. The reconstructed target image and the target image should be consistent in theory. The reprojection loss can be constructed using the difference in photometric values between the two. In order to obtain the depth of the target image and the pose information between the target image and the source image, the existing prediction model contains two subnetworks, namely the depth estimation subnetwork and the pose estimation subnetwork. The input of the depth estimation subnetwork is the target image, and the output is the corresponding depth map; the input of the pose estimation subnetwork is the target image and the source image, and the output is the relative pose between the two.
[0004] When there are illumination changes in the scene, the correct corresponding points between the target image and the reconstructed target image may have different luminosity values due to the illumination changes, resulting in large reprojection errors. When there are dynamic objects in the scene, the static scene assumption of the reprojection loss is not met. For outdoor scenes, there are areas such as the end of the road and the sky where the depth tends to infinity. When the reprojection loss constraint model is trained for such areas, the pose estimation subnetwork in the model cannot be well constrained. Although the prior art has proposed relevant masking strategies to solve the problem of failure of dynamic objects and infinite reprojection loss, this masking strategy relies on the difference between the network's estimated depth and the depth obtained by reprojection. In addition, although the prior art has masks constructed using image semantic categories, it needs to rely on experience to judge the categories that need to be covered. Summary of the invention
[0005] The technical problem to be solved by the present invention is to provide a semantically integrated unsupervised depth estimation and visual odometer method and system, which effectively reduces the impact of illumination changes and dynamic objects on reprojection loss in visual odometer and enhances the effectiveness of reprojection loss.
[0006] The technical solution adopted by the present invention to solve the technical problem is: to provide a semantically fused unsupervised depth estimation and visual odometer method, comprising the following steps:
[0007] receiving a video sequence;
[0008] The video sequence is input into a prediction model to obtain the pose between each two adjacent frames in the video sequence and the depth map of each frame image; wherein the prediction model includes:
[0009] The depth estimation subnetwork is used to estimate the depth map corresponding to the input target image;
[0010] The pose estimation subnetwork is used to estimate the relative pose between the input target image and the source image;
[0011] The semantic segmentation subnetwork is used to perform semantic segmentation on the input target image and obtain the semantic segmentation result of the target image;
[0012] The semantic mask module is used to construct a semantic mask based on the semantic segmentation result of the target image;
[0013] A reprojection module is used to reproject the source image according to the depth map, relative pose and semantic mask to obtain a reconstructed target image;
[0014] The semantic segmentation subnetwork is also used to perform semantic segmentation on the reconstructed target image to obtain a semantic segmentation result of the reconstructed target image; the target image semantic segmentation result and the reconstructed target image semantic segmentation result are used to construct a semantic reprojection loss.
[0015] The semantic reprojection loss is: Among them, L sem is the semantic reprojection loss, dis() represents the distance change function, and f sem () is the semantic segmentation sub-network, I t is the target image, To reconstruct the target image.
[0016] The semantic mask module adds n hyperparameters involved in training, where n is the number of categories of the semantic segmentation results; the semantic mask module assigns different weights to each category in the semantic segmentation results, where the different weights assigned to each category serve as semantic masks.
[0017] The semantic mask module uses regularization loss during training, and the regularization loss is expressed as: L reg =-log(Sem mask ), where L reg is the regularization loss, Sem mask is a semantic mask.
[0018] The technical solution adopted by the present invention to solve the technical problem is: to provide an unsupervised depth estimation and visual odometer system integrating semantics, including:
[0019] A receiving module, used for receiving a video sequence;
[0020] A prediction module is used to input the video sequence into a prediction model to obtain the pose between each two adjacent frames in the video sequence and the depth map of each frame image; wherein the prediction model includes:
[0021] The depth estimation subnetwork is used to estimate the depth map corresponding to the input target image;
[0022] The pose estimation subnetwork is used to estimate the relative pose between the input target image and the source image;
[0023] The semantic segmentation subnetwork is used to perform semantic segmentation on the input target image and obtain the semantic segmentation result of the target image;
[0024] The semantic mask module is used to construct a semantic mask based on the semantic segmentation result of the target image;
[0025] A reprojection module is used to reproject the source image according to the depth map, relative pose and semantic mask to obtain a reconstructed target image;
[0026] The semantic segmentation subnetwork is also used to perform semantic segmentation on the reconstructed target image to obtain a semantic segmentation result of the reconstructed target image; the target image semantic segmentation result and the reconstructed target image semantic segmentation result are used to construct a semantic reprojection loss.
[0027] The semantic reprojection loss is: Among them, L sem is the semantic reprojection loss, dis() represents the distance change function, and f sem () is the semantic segmentation sub-network, I t is the target image, To reconstruct the target image.
[0028] The semantic mask module adds n hyperparameters involved in training, where n is the number of categories of the semantic segmentation results; the semantic mask module assigns different weights to each category in the semantic segmentation results, where the different weights assigned to each category serve as semantic masks.
[0029] The semantic mask module uses regularization loss during training, and the regularization loss is expressed as: L reg = -log(Sem mask ), where L reg is the regularization loss, Sem mask is a semantic mask.
[0030] The technical solution adopted by the present invention to solve its technical problem is: to provide an electronic device, including a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the above-mentioned semantically fused unsupervised depth estimation and visual odometer method are implemented.
[0031] The technical solution adopted by the present invention to solve its technical problem is: providing a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned unsupervised depth estimation and visual odometer method with fusion semantics are implemented.
[0032] Beneficial Effects
[0033] Due to the adoption of the above-mentioned technical scheme, the present invention has the following advantages and positive effects compared with the prior art: the present invention designs a semantic reprojection loss, trains the prediction model according to the semantic category constraints between the target image and the reconstructed target image, and reduces the influence of illumination changes in the scene on the reprojection loss. The present invention also designs an adaptive semantic mask, which can well avoid the influence of dynamic objects and infinitely distant areas on the model training process without relying on the experience of engineers. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is a flow chart of a method for unsupervised depth estimation and visual odometer integrating semantics in a first embodiment of the present invention;
[0035] Figure 2 It is a schematic diagram of the structure of the prediction model in the first embodiment of the present invention. DETAILED DESCRIPTION
[0036] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not intended to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms fall within the scope limited by the appended claims of the application equally.
[0037] The first embodiment of the present invention relates to a semantically fused unsupervised depth estimation and visual odometer method, such as Figure 1As shown, the following steps are included:
[0038] Step 1: Receive a monocular RGB video sequence.
[0039] Step 2: Input the monocular RGB video sequence into the prediction model to obtain the pose between each two adjacent frames in the monocular RGB video sequence and the depth map of each frame image.
[0040] like Figure 2 As shown, the prediction model in this embodiment includes:
[0041] The depth estimation subnetwork is used to estimate the depth map corresponding to the input target image. In this embodiment, the input of the depth estimation subnetwork is the image at time t in the monocular RGB video sequence, that is, the target image, and the output is the depth map corresponding to the image.
[0042] The pose estimation subnetwork is used to estimate the relative pose between the input target image and the source image. In this embodiment, the input of the pose estimation subnetwork is the image at time t and the image at time t+1 in the monocular RGB video sequence, that is, the target image and the source image, and the output is the relative pose between the two images.
[0043] The semantic segmentation subnetwork is used to perform semantic segmentation on the input target image to obtain the semantic segmentation result of the target image. The semantic segmentation subnetwork in this embodiment can use a pre-trained model, including but not limited to OneFormer and DeepLab series algorithms, and the model parameters do not participate in the training.
[0044] The semantic mask module is used to construct a semantic mask based on the semantic segmentation result of the target image. The semantic mask can be used to cover distant objects and dynamic objects when processing outdoor scenes, thereby enhancing the effectiveness of the photometric reprojection loss.
[0045] The reprojection module is used to reproject the source image according to the depth map, relative pose and semantic mask to obtain a reconstructed target image.
[0046] The semantic segmentation subnetwork is also used to perform semantic segmentation on the reconstructed target image to obtain a semantic segmentation result of the reconstructed target image; the target image semantic segmentation result and the reconstructed target image semantic segmentation result are used to construct a semantic reprojection loss.
[0047] The training of unsupervised visual odometry mainly relies on reprojection loss as the main constraint. However, when there are illumination changes in the scene, there may still be large luminosity differences between the correct corresponding points of the target image and the reconstructed target image. In addition, the above situation may also occur in two images taken at different angles at different times. To this end, this embodiment introduces a semantic segmentation subnetwork, and the target image and the reconstructed target image are input into the semantic segmentation subnetwork to obtain the semantic segmentation result of the target image and the semantic segmentation result of the reconstructed target image. The semantic segmentation result determines the category information corresponding to each pixel point. The correct corresponding points between the target image and the reconstructed target image should have the same semantic category, and the semantic category information is not affected by illumination factors. However, since the semantic category information is binary information, this embodiment further processes it by applying a distance change function to obtain the final semantic reprojection loss. The semantic reprojection loss in this embodiment is shown as follows:
[0048]
[0049] Among them, L sem is the semantic reprojection loss, dis() represents the distance change function, and f sem () is the semantic segmentation sub-network, I t is the target image, To reconstruct the target image. This embodiment uses the semantic segmentation result to construct the reprojection loss, which can reduce the impact of illumination changes in the scene on the reprojection loss.
[0050] The reprojection loss needs to be based on the assumption of a static scene. The pose estimation subnetwork in the prediction model predicts the relative pose between the cameras of adjacent frames. If there are objects in the scene that move relative to the camera, the reprojection will not be able to obtain the corresponding corresponding points. At the same time, in outdoor scenes, the infinitely far area cannot use the pose information of adjacent frames in the reprojection calculation, resulting in the area being unable to constrain the pose estimation subnetwork in the model for training. For this reason, this embodiment proposes a semantic mask. When the image passes through the semantic segmentation subnetwork, the category information of each pixel is obtained. Some categories, such as vehicles and pedestrians, generally exist in the scene in relative motion with the camera, and other categories, such as the sky, are generally in areas with infinite depth. Covering these categories will help the training of the model. Therefore, the semantic mask module in this embodiment adds n hyperparameters involved in the training, where n is the number of categories of the semantic segmentation results. During the training process of the prediction model, the semantic mask module will assign different weights to each category in the semantic segmentation result according to the influence of different categories, and use it as a semantic mask. In addition, to ensure that the model does not set all hyperparameters to 0, this implementation also adds a regularization loss, which is expressed as:
[0051] Lreg =-log(Sem mask )
[0052] Among them, L reg is the regularization loss, Sem mask is the semantic mask, that is, the weights assigned to different categories.
[0053] To verify the feasibility of this implementation, experiments were conducted on the KITTIOdometry dataset. By adding semantic loss to SC-Depth, the translation error of the present invention was reduced by 6.43% and 5.90% in the 09 sequence and 10 sequence, and the rotation error was reduced by 34.43% and 43.67%. By adding semantic mask to SC-Depth, the translation error of the present invention was reduced by 15.87% and 3.47% in the 09 sequence and 10 sequence, and the rotation error was reduced by 41.64% and 23.47%.
[0054] It is not difficult to find that the present invention designs a semantic reprojection loss, trains the prediction model according to the semantic category constraints between the target image and the reconstructed target image, and reduces the impact of illumination changes in the scene on the reprojection loss. The present invention also designs an adaptive semantic mask, which can effectively avoid the impact of dynamic objects and infinitely distant areas on the model training process without relying on the experience of engineers.
[0055] The second embodiment of the present invention relates to an unsupervised depth estimation and visual odometer system integrating semantics, comprising:
[0056] A receiving module, used for receiving a video sequence;
[0057] A prediction module is used to input the video sequence into a prediction model to obtain the pose between each two adjacent frames in the video sequence and the depth map of each frame image; wherein the prediction model includes:
[0058] The depth estimation subnetwork is used to estimate the depth map corresponding to the input target image;
[0059] The pose estimation subnetwork is used to estimate the relative pose between the input target image and the source image;
[0060] The semantic segmentation subnetwork is used to perform semantic segmentation on the input target image and obtain the semantic segmentation result of the target image;
[0061] The semantic mask module is used to construct a semantic mask based on the semantic segmentation result of the target image;
[0062] A reprojection module is used to reproject the source image according to the depth map, relative pose and semantic mask to obtain a reconstructed target image;
[0063] The semantic segmentation subnetwork is also used to perform semantic segmentation on the reconstructed target image to obtain a semantic segmentation result of the reconstructed target image; the target image semantic segmentation result and the reconstructed target image semantic segmentation result are used to construct a semantic reprojection loss.
[0064] The semantic reprojection loss is: Among them, L sem is the semantic reprojection loss, dis() represents the distance change function, and f sem () is the semantic segmentation sub-network, I t is the target image, To reconstruct the target image.
[0065] The semantic mask module adds n hyperparameters involved in training, where n is the number of categories of the semantic segmentation results; the semantic mask module assigns different weights to each category in the semantic segmentation results, where the different weights assigned to each category serve as semantic masks.
[0066] The semantic mask module uses regularization loss during training, and the regularization loss is expressed as: L reg = -log(Sem mask ), where L reg is the regularization loss, Sem mask is a semantic mask.
[0067] A third embodiment of the present invention relates to an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the semantically fused unsupervised depth estimation and visual odometer method of the first embodiment when executing the computer program.
[0068] A fourth embodiment of the present invention relates to a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the semantically fused unsupervised depth estimation and visual odometer method of the first embodiment are implemented.
[0069] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) that contain computer-usable program code.
[0070] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0071] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction method, which is implemented in the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0072] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0073] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A semantically-integrated unsupervised depth estimation and visual odometry method, characterized in that: The following steps are involved: receiving a video sequence; The video sequence is input into a prediction model to obtain the pose between each two adjacent frames in the video sequence and the depth map of each frame image; wherein the prediction model includes: The depth estimation subnetwork is used to estimate the depth map corresponding to the input target image; The pose estimation subnetwork is used to estimate the relative pose between the input target image and the source image; The semantic segmentation subnetwork is used to perform semantic segmentation on the input target image and obtain the semantic segmentation result of the target image; The semantic mask module is used to construct a semantic mask based on the semantic segmentation result of the target image; A reprojection module is used to reproject the source image according to the depth map, relative pose and semantic mask to obtain a reconstructed target image; The semantic segmentation subnetwork is also used to perform semantic segmentation on the reconstructed target image to obtain a semantic segmentation result of the reconstructed target image; the semantic segmentation result of the target image and the semantic segmentation result of the reconstructed target image are used to construct a semantic reprojection loss; the semantic reprojection loss is: Among them, L sem is the semantic reprojection loss, dis() represents the distance change function, and f sem () is the semantic segmentation sub-network, I t is the target image, To reconstruct the target image.
2. The semantically fused unsupervised depth estimation and visual odometer method according to claim 1, characterized in that: The semantic mask module adds n hyperparameters involved in training, where n is the number of categories of the semantic segmentation results; the semantic mask module assigns different weights to each category in the semantic segmentation results, where the different weights assigned to each category serve as semantic masks.
3. The semantically fused unsupervised depth estimation and visual odometer method according to claim 1, characterized in that: The semantic mask module uses regularization loss during training, and the regularization loss is expressed as: L reg = -log(Sem mask ), where L reg is the regularization loss, Sem mask is a semantic mask.
4. A semantically fused unsupervised depth estimation and visual odometry system, characterized in that: include: A receiving module, used for receiving a video sequence; A prediction module is used to input the video sequence into a prediction model to obtain the pose between each two adjacent frames in the video sequence and the depth map of each frame image; wherein the prediction model includes: The depth estimation subnetwork is used to estimate the depth map corresponding to the input target image; The pose estimation subnetwork is used to estimate the relative pose between the input target image and the source image; The semantic segmentation subnetwork is used to perform semantic segmentation on the input target image and obtain the semantic segmentation result of the target image; The semantic mask module is used to construct a semantic mask based on the semantic segmentation result of the target image; A reprojection module is used to reproject the source image according to the depth map, relative pose and semantic mask to obtain a reconstructed target image; The semantic segmentation subnetwork is also used to perform semantic segmentation on the reconstructed target image to obtain a semantic segmentation result of the reconstructed target image; the semantic segmentation result of the target image and the semantic segmentation result of the reconstructed target image are used to construct a semantic reprojection loss; the semantic reprojection loss is: Among them, L sem is the semantic reprojection loss, dis() represents the distance change function, and f sem () is the semantic segmentation sub-network, I t is the target image, To reconstruct the target image.
5. The semantically fused unsupervised depth estimation and visual odometer system according to claim 4, characterized in that: The semantic mask module adds n hyperparameters involved in training, where n is the number of categories of the semantic segmentation results; the semantic mask module assigns different weights to each category in the semantic segmentation results, where the different weights assigned to each category serve as semantic masks.
6. The semantically fused unsupervised depth estimation and visual odometer system according to claim 4, characterized in that: The semantic mask module uses regularization loss during training, and the regularization loss is expressed as: L reg = -log(Sem mask ), where L reg is the regularization loss, Sem mask is a semantic mask.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the semantically fused unsupervised depth estimation and visual odometer method as described in any one of claims 1 to 3 are implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the semantically fused unsupervised depth estimation and visual odometer method as claimed in any one of claims 1 to 3 are implemented.
Citation Information
Patent Citations
Semantic segmentation and visual SLAM tight coupling method for dynamic environment
CN110827305A
Unsupervised monocular depth estimation method fusing full-scale and adjacent frame feature information
CN116071412A