Monocular depth estimation method based on multi-modal information fusion
Through multimodal information fusion and depth estimation model, the problems of depth edge blur and insufficient adaptability in monocular depth estimation technology are solved, and more efficient and accurate depth information generation is achieved, suitable for naked-eye 3D display and autonomous driving.
Patent Information
- Application Number
- CN202510563308.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing monocular depth estimation technology has blurring when dealing with depth edges, lack of adaptability to multiple scenes, and it is difficult to balance computing efficiency and accuracy, which limits its application in fields such as naked-eye 3D display and autonomous driving.
By collecting monocular images and laser point cloud data, a monocular depth estimation model of multi-scale convolutional layer, residual connection layer, semantic segmentation layer and detail optimizer is constructed. Combined with Laplace edge detection and local ternary mode texture analysis, the accuracy and scene adaptability of depth estimation are improved.
It improves the accuracy and calculation efficiency of monocular depth estimation, generates more accurate depth information, solves the problems of blurred depth edges, poor adaptability of multiple scenes, and difficult to balance calculation efficiency and accuracy. It is suitable for scenarios such as naked-eye 3D display and autonomous driving.
Smart Images

Figure CN120471977A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a monocular depth estimation method based on multimodal information fusion. Background Art
[0002] In today's computer vision field, monocular depth estimation technology is dedicated to accurately inferring the three-dimensional depth information of a scene from a two-dimensional image captured by a monocular camera. This technology plays an indispensable and key role in many cutting-edge application fields. In the field of naked-eye 3D display, accurate monocular depth estimation is the core element for presenting extremely realistic stereoscopic visual effects, allowing users to immerse themselves in lifelike three-dimensional scenes without the need for any additional equipment. In the field of autonomous driving, monocular depth estimation helps vehicles accurately perceive the distance to objects in the surrounding environment, providing a core basis for safe driving decisions, which is directly related to driving safety and traffic efficiency. In the field of robot visual navigation, monocular depth information helps robots fully understand the surrounding spatial layout, thereby planning reasonable and efficient movement paths and realizing intelligent operations.
[0003] However, current monocular depth estimation technology faces a series of severe challenges. The problems of depth edge blurring and detail loss are extremely prominent. Most existing models are prone to blurring when processing depth edges due to inherent algorithmic limitations. This leads to blurred object outlines and the inability to effectively retain a large amount of key detail information. This problem seriously undermines the realism and naturalness of stereoscopic vision in naked-eye 3D display scenarios, which have almost stringent requirements for visual effects, and greatly reduces the user experience. Secondly, it lacks adaptability to multiple scenes. Real scenes are rich and varied. Different lighting conditions such as direct strong light, low light, complex light and shadow, object materials such as metal, plastic, rubber, glass, and texture features such as smooth surfaces, complex textures, and periodic textures will seriously interfere with the accuracy of monocular depth estimation. Therefore, existing monocular depth estimation methods are difficult to fully adapt to such complex and diverse scenes, and cannot meet the application requirements of naked-eye 3D display, autonomous driving, and other scenarios in various scenarios, greatly limiting their widespread application. In addition, it is difficult to balance computational efficiency and accuracy. Some monocular depth estimation models that pursue high precision use complex network structures and massive parameters, resulting in an exponential increase in computational complexity and extremely high requirements for hardware performance. They are difficult to run in real time on resource-constrained platforms such as mobile devices and embedded devices. Although some lightweight models have a fast computational speed, the accuracy of monocular depth estimation is far from meeting the actual application standards. For example, in naked-eye 3D real-time rendering and autonomous driving real-time perception scenarios, they cannot provide depth information in a timely and accurate manner, which seriously affects the real-time interactive experience and driving safety, and limits their development in more fields. Summary of the Invention
[0004] In response to the problems existing in the prior art, the present invention provides a monocular depth estimation method based on multimodal information fusion, which can simultaneously take into account image edge detection capability, scene adaptability and computational efficiency, and improve the accuracy and precision of monocular depth estimation.
[0005] To achieve the above technical objectives, the present invention adopts the following technical solution: a monocular depth estimation method based on multimodal information fusion, comprising the following steps:
[0006] Step S1: collecting multimodal data consisting of monocular images and laser point cloud data, and fusing the multimodal data;
[0007] Step S2: construct a monocular depth estimation model consisting of a multi-scale convolutional layer, a residual connection layer, an upsampling layer, a semantic segmentation layer, and a detail optimizer;
[0008] Step S3: Input the fused multimodal data into the multi-scale convolution layer and semantic segmentation layer of the monocular depth estimation model respectively, use the multi-scale convolution layer to extract multi-scale features, use the semantic segmentation layer to identify the object category in the monocular image, use the residual connection layer to pass the multi-scale features to the upsampling layer, determine the preliminary monocular depth of the monocular image, combine the identified object category, use the detail optimizer to adjust the preliminary monocular depth, and predict the monocular depth of the monocular image.
[0009] Furthermore, step S1 includes the following sub-steps:
[0010] Step S1.1: Collect monocular images from the video stream and use the optical flow method combined with the convolutional neural network to determine the motion state of the object in the monocular image;
[0011] Step S1.2, obtaining laser point cloud data corresponding to the monocular image, and converting the laser point cloud data into a depth map;
[0012] Step S1.3, performing scene environment data enhancement on the monocular image based on the generative adversarial network;
[0013] Step S1.4: Perform multimodal data fusion on the object motion state in the monocular image, the enhanced scene environment data, and the depth map.
[0014] Furthermore, the specific process of step S1.1 is as follows: use the TV-L1 optical flow algorithm to calculate the optical flow field between adjacent monocular images in the video stream, perform channel splicing on the optical flow field and the current monocular image, input the convolutional neural network based on ResNet-18, perform feature extraction through the convolution layer and the residual block, reduce the dimension of the extracted features through the pooling layer, and output the reduced dimension features through the fully connected layer to output the motion state of the object on the current monocular image.
[0015] Furthermore, when enhancing scene environment data with changing illumination, the input of the generator in the generative adversarial network is a monocular image and the corresponding illumination change parameters; when enhancing scene environment data with different materials, the input of the generator in the generative adversarial network is a monocular image and random noise.
[0016] Furthermore, the multi-scale convolutional layer includes several parallel convolutional layers, and a successively increasing expansion rate is set for each convolutional layer. The fused multimodal data is used to generate feature maps of different scales using the parallel convolutional layers. The channel attention mechanism is used to generate channel attention weights for the feature maps of each scale, and the feature maps of different scales are weighted using the corresponding channel attention weights to obtain the fusion features of the multi-scale feature maps.
[0017] Furthermore, the residual connection layer adds a learnable weight matrix to the residual connection path, optimizes the weight matrix through back propagation, and dynamically adjusts the contribution of the residual signal η = W r ×identity, where W r Represents a learnable weight matrix, initialized using Kaiming W r ; identity represents the residual signal.
[0018] Furthermore, the semantic segmentation layer includes: a backbone network composed of ResNeXt-101, a region proposal network RPN, a mask prediction branch network and a classification prediction branch network. The backbone network is used to extract semantic features from the fused multimodal data; the region proposal network RPN generates candidate regions containing objects based on semantic features; the mask prediction branch network performs pixel-level segmentation on each candidate region to generate a mask map of the object; the classification prediction branch network predicts the object category of the corresponding candidate region based on the mask map of the object.
[0019] Furthermore, the detail optimizer adjusts the preliminary monocular depth of the monocular image using Laplacian edge detection and local ternary pattern texture analysis based on the object category predicted in the monocular image.
[0020] Furthermore,
[0021] i. For each object category area in the monocular image, use the Laplacian operator to extract the object's edge map as a mask and perform edge sharpening on the corresponding area in the preliminary monocular depth of the monocular image:
[0022]
[0023] Among them, (x, y) represents the pixel coordinates on the monocular image, D initial(x,y) represents the initial monocular depth of the pixel (x,y), α represents the sharpening coefficient, E(x,y) represents the mask of the pixel (x,y), represents the mask threshold, D edge-refined (x,y) represents the monocular depth of the coordinate point (x,y) for edge sharpening;
[0024] ii. For each object category area in the monocular image, the local ternary pattern texture is used to extract texture features, and the corresponding area in the edge-sharpened monocular depth is corrected according to the intensity of the texture features:
[0025] D(x,y)=w·D initial (x,y)+(1-w)·D edge-refined (x,y)
[0026] Among them, w represents the correction weight set according to the strength of the texture feature, and D(x,y) represents the corrected monocular depth.
[0027] Furthermore, for each pixel in the monocular image, the corresponding disparity offset is calculated based on the predicted monocular depth The monocular image is rearranged according to the calculated disparity offset to generate the disparity map of the left and right views, where: and β are constants determined according to the parameters of the naked-eye 3D display device and the image resolution, respectively.
[0028] Compared with the prior art, the present invention has the following beneficial effects: the monocular depth estimation method based on multimodal information fusion of the present invention comprehensively obtains information that can reflect the characteristics of the monocular image by collecting multimodal data, providing an accurate and reliable data basis for subsequent monocular depth estimation; the monocular depth estimation method of the present invention adopts a monocular depth estimation model with a semantic segmentation layer and a detail optimizer, thereby further correcting the monocular depth according to the object category on the monocular image, greatly improving the accuracy, scene adaptability and computational efficiency of monocular depth estimation, and being able to generate more accurate and actual scene depth information, effectively solving the problems of depth edge blur, poor multi-scene adaptability, difficulty in balancing computational efficiency and accuracy, and insufficient application adaptability existing in traditional monocular depth estimation technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 Flowchart of the monocular depth estimation method based on multimodal information fusion of the present invention;
[0030] Figure 2 Schematic diagram of the monocular depth estimation model in the present invention. DETAILED DESCRIPTION
[0031] The technical solution of the present invention will be further explained below with reference to the accompanying drawings.
[0032] like Figure 1 The present invention is a flow chart of a monocular depth estimation method based on multimodal information fusion, which includes the following steps:
[0033] Step S1: Collect multimodal data consisting of monocular images and laser point cloud data, fuse the multimodal data, and comprehensively obtain information that can reflect the characteristics of the monocular image, providing an accurate and reliable data basis for subsequent monocular depth estimation; including the following sub-steps:
[0034] Step S1.1. Collect monocular images in the video stream and use the optical flow method combined with the convolutional neural network to determine the motion state of the object in the monocular image. Specifically, use the TV-L1 optical flow algorithm to calculate the optical flow field between adjacent monocular images in the video stream, which directly reflects the object's motion trajectory, is sensitive to small movements, and can improve the accuracy of subsequent object motion estimation. The optical flow field is channel-joined with the current monocular image and input into the convolutional neural network based on ResNet-18. Feature extraction is performed through the convolutional layer and the residual block. The extracted features are reduced in dimensionality through the pooling layer, and the reduced dimensionality features are output through the fully connected layer to the object motion state on the current monocular image.
[0035] Step S1.2: Obtain the laser point cloud data corresponding to the monocular image and convert the laser point cloud data into a depth map.
[0036] Step S1.3: Enhance the monocular image with scene environment data based on a generative adversarial network. Specifically, when enhancing scene environment data with varying illumination, the input to the generator in the generative adversarial network is the monocular image and the corresponding illumination variation parameters, including illumination intensity and color temperature. The monocular image can retain the basic visual information of the scene, and the illumination variation parameters can accurately describe the illumination variation, providing a basis for the generator to adjust the image color and contrast, generating an enhanced image that adapts to illumination variations, thereby reducing the interference of illumination on depth estimation. When enhancing scene environment data of different materials, the input to the generator in the generative adversarial network is the monocular image and random noise. The random noise is used to introduce diversity, helping the generator learn the changes in the reflective characteristics of the corresponding material under different illumination and viewing angles, prompting the generator to generate rich and diverse feature maps.
[0037] Step S1.4: Perform multimodal data fusion on the object motion state in the monocular image, the enhanced scene environment data, and the depth map.
[0038] Step S2: Figure 2, constructing a monocular depth estimation model consisting of a multi-scale convolutional layer, a residual connection layer, an upsampling layer, a semantic segmentation layer and a detail optimizer, which can further correct the monocular depth according to the object category in the monocular image, greatly improving the accuracy, scene adaptability and computational efficiency of monocular depth estimation, and can generate more accurate and actual scene depth information, effectively solving the problems of traditional monocular depth estimation technology such as blurred depth edges, poor adaptability to multiple scenes, difficulty in balancing computational efficiency and accuracy, and insufficient application adaptability.
[0039] The multi-scale convolution layer includes several parallel convolution layers. An increasing expansion rate is set for each convolution layer. The fused multimodal data is used to generate feature maps of different scales using parallel convolution layers. The multi-scale convolution layer can extract multi-scale features of each modality in parallel, avoiding the insufficient expression ability of single-scale features for complex data, and can reduce the number of parameters and computational overhead; the channel attention mechanism is used to generate channel attention weights for the feature maps of each scale, and the feature maps of different scales are weighted using the corresponding channel attention weights to obtain the fusion features of the multi-scale feature maps.
[0040] The residual connection layer adds a learnable weight matrix to the residual connection path, optimizes the weight matrix through back propagation, and dynamically adjusts the contribution of the residual signal η = W r ×identity, effectively enhances feature propagation and avoids the gradient disappearance problem, where W r Represents a learnable weight matrix, initialized using Kaiming W r ; identity represents the residual signal.
[0041] The semantic segmentation layer includes: a backbone network composed of ResNeXt-101, a region proposal network RPN, a mask prediction branch network and a classification prediction branch network. The backbone network is used to extract semantic features from the fused multimodal data. Specifically, ResNeXt-101 enhances the expressiveness of semantic features while maintaining computational efficiency through grouped convolution and residual connection; the region proposal network RPN generates candidate regions containing objects based on semantic features. Specifically, it slides a window on the semantic feature map extracted by the backbone network and predicts a set of anchor points of different scales and aspect ratios for each position through convolution operations. For each anchor point, RPN outputs its attribute. The probability of being an object or background, as well as the corresponding bounding box regression parameters, is used to screen out high-quality region proposals using the non-maximum suppression algorithm. These region proposals will serve as candidate regions for subsequent target detection and semantic segmentation; the mask prediction branch network performs pixel-level segmentation on each candidate region and generates a binary mask map of the object, which represents the specific shape and position of the object in the region. The mask prediction branch network adopts the structure of a fully convolutional network to gradually restore the resolution of the feature map and finally outputs a mask map of the same size as the input image. Each pixel corresponds to the probability value of the object or background; the classification prediction branch network predicts the object category corresponding to the candidate region based on the object mask map.
[0042] Since the initial monocular depth map may be blurred at the edges of objects and the depth estimation in areas with complex textures is inaccurate, the present invention uses a detail optimizer to adjust the initial monocular depth of the monocular image based on the object category predicted in the monocular image using Laplacian edge detection and local ternary pattern texture analysis to improve the accuracy of depth estimation and edge consistency. Specifically:
[0043] i. For each object category region in the monocular image, use the Laplacian operator to extract the object's edge map. Use this as a mask to sharpen the edges of the corresponding region in the preliminary monocular depth of the monocular image to enhance the edge consistency of the monocular depth and align the depth mutation region with the edge of the monocular image.
[0044] The process of edge sharpening the corresponding area in the preliminary monocular depth of the monocular image in the present invention is as follows:
[0045]
[0046] Among them, (x, y) represents the pixel coordinates on the monocular image, D initial (x,y) represents the initial monocular depth of the pixel (x,y), α represents the sharpening coefficient, E(x,y) represents the mask of the pixel (x,y), represents the mask threshold, D edge-refined (x,y) represents the monocular depth of the coordinate point (x,y) for edge sharpening;
[0047] ii. For each object category area in the monocular image, the local ternary pattern texture is used to extract texture features, and the corresponding area in the edge-sharpened monocular depth is corrected according to the intensity of the texture features to correct blurred or erroneous areas.
[0048] The process of correcting the corresponding area in the edge-sharpened monocular depth in the present invention is:
[0049] D(x,y)=w·D initial (x,y)+(1-w)·D edge-refined (x,y)
[0050] Among them, w represents the correction weight set according to the strength of the texture feature, and D(x,y) represents the corrected monocular depth.
[0051] Step S3: input the fused multimodal data into the multi-scale convolution layer and semantic segmentation layer of the monocular depth estimation model respectively, use the multi-scale convolution layer to perform multi-scale feature extraction, use the semantic segmentation layer to identify the object category in the monocular image, use the residual connection layer to pass the multi-scale features to the upsampling layer, determine the preliminary monocular depth of the monocular image, combine the identified object category, use the detail optimizer to adjust the preliminary monocular depth, and predict the monocular depth of the monocular image. This can take into account the image edge detection capability, scene adaptability and computational efficiency at the same time, and improve the accuracy and precision of the monocular depth estimation.
[0052] In one technical solution of the present invention, for each pixel in the monocular image, the corresponding disparity offset is calculated according to the predicted monocular depth. The monocular image is rearranged according to the calculated disparity offset to generate the disparity map of the left and right views, where: and β are constants determined by the parameters of the naked-eye 3D display device and the image resolution, respectively. During the rearrangement process, bilinear interpolation is used to fill the holes caused by the offset, ensuring the integrity and continuity of the image and achieving a high-quality stereoscopic display effect.
[0053] In one technical solution of the present invention, a computer-readable storage medium is further provided, storing a computer program, wherein the computer program enables a computer to execute a monocular depth estimation method based on multimodal information fusion.
[0054] In one technical solution of the present invention, an electronic device is also provided, including: a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, a monocular depth estimation method based on multimodal information fusion is implemented.
[0055] In the embodiments disclosed herein, computer storage media can be tangible media that can contain or store programs for use by or in conjunction with an instruction execution system, device, or apparatus. Computer storage media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. More specific examples of computer storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0056] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0057] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions based on the principles of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should be considered within the scope of protection of the present invention.
Claims
1. A monocular depth estimation method based on multimodal information fusion, characterized in that: The steps include: Step S1: Collect multimodal data consisting of monocular images and laser point cloud data, and fuse the multimodal data; Step S2: Construct a monocular depth estimation model consisting of a multi-scale convolutional layer, a residual connection layer, an upsampling layer, a semantic segmentation layer, and a detail optimizer; Step S3: The fused multimodal data is input into the multi-scale convolution layer and semantic segmentation layer of the monocular depth estimation model respectively. The multi-scale convolution layer is used to extract multi-scale features. The semantic segmentation layer is used to identify the object category in the monocular image. The multi-scale features are transferred to the upsampling layer using the residual connection layer to determine the preliminary monocular depth of the monocular image. Combined with the identified object category, the detail optimizer is used to adjust the preliminary monocular depth and predict the monocular depth of the monocular image.
2. The monocular depth estimation method based on multimodal information fusion according to claim 1, characterized in that: Step S1 includes the following sub-steps: Step S1.1: Collect monocular images from the video stream and use the optical flow method combined with the convolutional neural network to determine the motion state of the object in the monocular image; Step S1.2: Obtain laser point cloud data corresponding to the monocular image and convert the laser point cloud data into a depth map; Step S1.3: Perform scene environment data enhancement on the monocular image based on the generative adversarial network; Step S1.4: Perform multimodal data fusion on the object motion state in the monocular image, the enhanced scene environment data, and the depth map.
3. The monocular depth estimation method based on multimodal information fusion according to claim 2, characterized in that: The specific process of step S1.1 is as follows: use the TV-L1 optical flow algorithm to calculate the optical flow field between adjacent monocular images in the video stream, perform channel splicing on the optical flow field and the current monocular image, input the convolutional neural network based on ResNet-18, perform feature extraction through the convolution layer and residual block, reduce the dimension of the extracted features through the pooling layer, and output the reduced dimension features through the fully connected layer to output the motion status of the object in the current monocular image.
4. The monocular depth estimation method based on multimodal information fusion according to claim 2, characterized in that: When enhancing scene environment data with changing illumination, the input of the generator in the generative adversarial network is a monocular image and the corresponding illumination change parameters; when enhancing scene environment data with different materials, the input of the generator in the generative adversarial network is a monocular image and random noise.
5. The monocular depth estimation method based on multimodal information fusion according to claim 1, characterized in that: The multi-scale convolutional layer includes several parallel convolutional layers, and a sequentially increasing expansion rate is set for each convolutional layer. The fused multimodal data is used to generate feature maps of different scales using the parallel convolutional layers. The channel attention mechanism is used to generate channel attention weights for the feature maps of each scale. The feature maps of different scales are weighted using the corresponding channel attention weights to obtain the fusion features of the multi-scale feature maps.
6. The monocular depth estimation method based on multimodal information fusion according to claim 5, characterized in that: The residual connection layer adds a learnable weight matrix to the residual connection path, optimizes the weight matrix through back propagation, and dynamically adjusts the contribution of the residual signal η = W r ×identity, where W r Represents a learnable weight matrix, initialized using Kaiming W r ; identity represents the residual signal.
7. The monocular depth estimation method based on multimodal information fusion according to claim 6, characterized in that: The semantic segmentation layer includes: a backbone network composed of ResNeXt-101, a region proposal network (RPN), a mask prediction branch network, and a classification prediction branch network. The backbone network is used to extract semantic features from fused multimodal data; the region proposal network (RPN) generates candidate regions containing objects based on semantic features; the mask prediction branch network performs pixel-level segmentation on each candidate region to generate a mask map of the object; and the classification prediction branch network predicts the object category of the corresponding candidate region based on the object mask map.
8. The monocular depth estimation method based on multimodal information fusion according to claim 7, characterized in that: The detail optimizer adjusts the preliminary monocular depth of the monocular image using Laplacian edge detection and local ternary pattern texture analysis based on the object category predicted on the monocular image.
9. The monocular depth estimation method based on multimodal information fusion according to claim 8, characterized in that: i. For each object category area in the monocular image, use the Laplacian operator to extract the object's edge map as a mask and perform edge sharpening on the corresponding area in the preliminary monocular depth of the monocular image: Among them, (x, y) represents the pixel coordinates on the monocular image, D initial (x,y) represents the initial monocular depth of the pixel (x,y), α represents the sharpening coefficient, E(x,y) represents the mask of the pixel (x,y), represents the mask threshold, D edge-refined (x,y) represents the monocular depth of the coordinate point (x,y) for edge sharpening; ii. For each object category area in the monocular image, the local ternary pattern texture is used to extract texture features, and the corresponding area in the edge-sharpened monocular depth is corrected according to the intensity of the texture features: D(x,y)=w·D initial (x,y)+(1-w)·D edge-refined (x,y) Among them, w represents the correction weight set according to the strength of the texture feature, and D(x,y) represents the corrected monocular depth.
10. The monocular depth estimation method based on multimodal information fusion according to any one of claims 1 to 9, characterized in that: For each pixel in the monocular image, the corresponding disparity offset is calculated based on the predicted monocular depth The monocular image is rearranged according to the calculated disparity offset to generate the disparity map of the left and right views, where: and β are constants determined according to the parameters of the naked-eye 3D display device and the image resolution, respectively.
Citation Information
Cited By
Method and system for quickly positioning vertex of suspension arm based on laser radar point cloud calibration monocular depth estimation
CN121685631A