Virtual and real occlusion processing method based on lightweight monocular depth estimation
By adopting lightweight network architecture and deep learning optimization technology in the monocular depth estimation algorithm, the problem of insufficient accuracy and real-time performance in complex environments is solved, and virtual and real occlusion processing with high precision and low power consumption is achieved.
Patent Information
- Application Number
- CN202510232868.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-17
AI Technical Summary
Existing monocular depth estimation algorithms based on deep learning are difficult to achieve high accuracy, real-time and robustness in complex industrial environments, especially in dynamic scenes and weak texture areas.
A lightweight monocular depth estimation network based on deep learning is adopted, and a lightweight convolutional neural network and a small visual Transformer layer is combined, and a model optimization is used using unsupervised learning and pixel-level contrast loss function, and real-time occlusion processing is performed in combination with OpenGL.
In dynamic scenarios, the depth map generation speed is increased by 14.7 times, the depth error is reduced to 8.7%, the occlusion judgment accuracy is increased to 93.2%, and the GPU power consumption is reduced by 42%, meeting the real-time needs of mobile terminals.
Smart Images

Figure CN120164078A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of augmented reality, and in particular to a virtual-real occlusion processing method based on lightweight monocular depth estimation. Background Art
[0002] Augmented reality is a technology that cleverly integrates virtual information with the real world. It can effectively enhance the seamless integration between the real environment and the virtual environment, and ultimately achieve natural, realistic, and harmonious human-computer interaction. The core function of augmented reality is to add virtual objects made by computers or other systems and some information such as text and patterns as prompts to real scenes. Augmented reality combines virtual information with real scenes to make real scenes richer and more meaningful. Since augmented reality has the characteristics of enhancing the real environment, augmented reality can be applied to many fields in people's lives. The use of augmented reality can make people's lives more convenient and can also facilitate scientific research in various disciplines.
[0003] In augmented reality, virtual-real occlusion technology is an important part, among which the monocular depth estimation algorithm based on deep learning is an important content. Each algorithm has its own advantages and disadvantages and its scope of application. How to improve the accuracy, real-time and robustness of the algorithm is the key issue of this technology. Although the monocular depth estimation algorithm has made great progress, in actual industrial applications, due to the complexity of the industrial environment, the depth estimation accuracy based on deep learning has never achieved the expected effect. For example, when virtual-real occlusion is applied to the pod assembly scene, the assembly positions are close and concentrated in a small space, and the assemblers need to move frequently to obtain a more convenient assembly posture. Traditional virtual-real occlusion processing methods are often difficult to maintain reliability in the face of complex environmental conditions. These complex conditions include light changes when the assemblers move, motion blur, and mutual occlusion of objects to be assembled.
[0004] After searching, the application publication number CN112365516B is a method for processing virtual and real occlusion in augmented reality, including a method for determining virtual and real occlusion and a method for rendering virtual and real object occlusion. Compared with the prior art, the present invention first determines the occlusion relationship between virtual and real objects through the SFM algorithm to separate the occluders; then, the occlusion rendering of the virtual and real objects is performed by establishing a mask, so that the augmented reality system can achieve a better occlusion edge and a higher accuracy occlusion effect.
[0005] "The CN112365516B patent uses the SFM (Structure from Motion) algorithm for depth estimation, and its defects are as follows: it relies on multi-frame image matching and is prone to failure in dynamic scenes; the computational complexity reaches O(n^3), making it difficult to meet the real-time requirements of mobile devices; the depth estimation error in weak texture areas is > 35%. The present invention replaces SFM with a monocular depth estimation network based on deep learning (inference time < 15 ms / frame). In actual device measurements: the depth error in dynamic scenes is reduced to 8.7%, the occlusion judgment accuracy is increased to 93.2%, and the GPU power consumption is reduced by 42%. " Summary of the Invention
[0006] The present invention aims to solve the above problems of the prior art. A virtual-real occlusion processing method based on lightweight monocular depth estimation is proposed. The technical solution of the present invention is as follows:
[0007] A virtual-real occlusion processing method based on lightweight monocular depth estimation, comprising the following steps:
[0008] Step S1, using a pre-collected dataset containing monocular images and corresponding depth maps for model training;
[0009] Step S2, using a lightweight convolutional neural network combined with a small vision Transformer layer as a network model for feature extraction and depth prediction;
[0010] Step S3, through an unsupervised learning method, using an autoencoder architecture to enable the network model to predict the corresponding depth map through the input monocular image;
[0011] Step S4, applying a pixel-level contrast loss function to optimize the model;
[0012] Step S5, based on a depth sorting-based virtual-real occlusion algorithm, using depth information to judge the occlusion relationship between virtual objects and real-world objects;
[0013] Step S6, using OpenGL for real-time occlusion processing to ensure the correct display of virtual objects in the real world;
[0014] Step S7, using a model optimization tool to accelerate the model;
[0015] Step S8, deploying the system on a mobile device, including further performance optimization using the Android NDK or iOS Core ML.
[0016] Further, in the step S2, the lightweight convolutional neural network is selected from MobileNet or ShuffleNet.
[0017] Further, in the step S2, the small vision Transformer layer is configured as Mini-ViT.
[0018] Further, in the step S3, the unsupervised learning autoencoder architecture includes pre-training a network model using transfer learning technology. The unsupervised learning autoencoder architecture uses MobileNetv2 pre-trained on ImageNet as the encoder, freezes the shallow convolutional parameters and fine-tunes the deep parameters, and improves the depth prediction accuracy through cross-domain feature transfer.
[0019] Further, in the step S4, the pixel-level contrast loss function is the L1 loss; the L1 loss is used to optimize the training of the depth prediction model, and its formula is expressed as:
[0020] [L=\frac{1}{N}\sum_{i=1}^{N}|p_i - g_i|]
[0021] where L represents the loss value, N represents the total number of pixels in the image, p_i represents the depth value of the i-th pixel in the predicted depth map, and g_i represents the depth value of the i-th pixel in the true depth map. The goal of this function is to minimize the total absolute error between the predicted depth map and the true depth map.
[0022] Further, in the step S6, the OpenGL real-time occlusion processing further includes dynamically adjusting the display of virtual and real objects to ensure that virtual objects are correctly occluded visually.
[0023] Further, in the step S7, the model optimization tool uses TensorRT or ONNX Runtime to perform model acceleration optimization, and the optimized inference time can be expressed as:
[0024] [t_{optimized}=\frac{t_{original}}{S}]
[0025] where t_{optimized} represents the inference time of the optimized model, t_{original} represents the inference time of the model before optimization, and S represents the acceleration ratio.
[0026] Further, in the step S8, the system deployment further includes using corresponding localization development tools on the Android or iOS platform for the efficient operation of the model on mobile devices.
[0027] A computer device, which includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor is caused to execute the virtual-real occlusion processing method based on lightweight monocular depth estimation as described in any one of the claims.
[0028] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the processor is caused to execute the virtual-real occlusion processing method based on lightweight monocular depth estimation as described in any one of the claims.
[0029] The advantages and beneficial effects of the present invention are as follows:
[0030] The present invention adopts a monocular depth estimation algorithm based on deep learning for virtual-real occlusion processing to meet the requirement of high precision. At the same time, the algorithm is optimized and the model is compressed through lightweight operations, reducing the number of model parameters, simplifying the network structure, and using model compression technology to reduce the volume and computational amount of the deep learning model. In this way, even under the limited computing power of AR devices, a lightweight model can still be loaded and run, thus realizing real-time virtual-real occlusion processing and meeting the requirement of high real-time performance.
[0031] Through the collaborative design of algorithm architecture innovation and hardware-level optimization, the present invention has achieved a double breakthrough in the accuracy and speed of mobile AR virtual-real occlusion processing, specifically manifested as:
[0032] 1. High-precision depth estimation (corresponding to claims 1-4)
[0033] Using MobileNetv2 pre-trained on ImageNet as the encoder backbone network (claim 4), through the cross-domain transfer strategy of freezing the parameters of shallow convolutional layers and fine-tuning deep features, in device actual measurements:
[0034] The depth estimation error in weak texture areas is reduced from 35% of the traditional SFM algorithm to 8.7% (compared with CN112365516B). The depth map generation speed in dynamic scenes reaches 62FPS, which is 14.7 times higher than that of traditional multi-view geometry methods.
[0035] 2. Low-power real-time inference (corresponding to claims 5-7)
[0036] The network is compressed and optimized through the FP16 quantization and layer fusion technology of TensorRT (claim 7): the number of model parameters is compressed from 12.5M to 3.8M (compression rate 69.6%). The GPU memory occupancy is reduced to 86MB, and the power consumption is reduced by 42%.
[0037] 3. Precise dynamic occlusion processing (corresponding to claims 8-9)
[0038] Combined with the hardware acceleration features of OpenGL ES 3.0 (Claim 9), a dual-channel Z-Buffer update mechanism is designed: the occlusion judgment accuracy is increased to 93.2%, which is 31.5% higher than the traditional template matching method, the depth buffer update latency ≤ 2ms, and 60Hz AR rendering without frame tearing is supported.
[0039] 4. Full-scene compatibility on mobile devices (corresponding to Claims 5-7)
[0040] Achieved through structured model compression technology (Claim 6): support for computing power adaptation from flagship models to mid-range devices, and in low-light (<50lux) and motion blur (angular velocity > 3rad / s) scenarios, the stability error of depth estimation < 12%. Description of the Drawings
[0041] Figure 1 It is a schematic flowchart of a virtual-real occlusion processing method based on lightweight monocular depth estimation according to a preferred embodiment provided by the present invention. Detailed Description of the Invention
[0042] Next, the technical solutions in the embodiments of the present invention will be clearly and detailedly described in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention.
[0043] The technical solution of the present invention to solve the above technical problems is:
[0044] Please refer to Figure 1 , the present invention provides a technical solution:
[0045] A virtual-real occlusion processing method based on lightweight monocular depth estimation, comprising the following steps:
[0046] Step S1, using a pre-collected dataset containing monocular images and corresponding depth maps for model training;
[0047] Step S2, using a lightweight convolutional neural network combined with a small vision Transformer layer as a network model for feature extraction and depth prediction;
[0048] Preferably, the lightweight convolutional neural network is selected from MobileNet or ShuffleNet;
[0049] Further, the small vision Transformer layer is configured as Mini-ViT;
[0050] In this step, a lightweight convolutional neural network and a vision Transformer layer are combined. The purpose of this structure is to extract effective features and perform depth prediction. Its processing flow includes:
[0051] Lightweight convolutional networks (such as MobileNet or ShuffleNet):
[0052] [FeatureCNN = CNNLayers(InputImage)]
[0053] where (Feature_CNN) are the features extracted from the input image;
[0054] Vision Transformer layer (Mini-ViT):
[0055] [FeatureViT = TransformerEncoder(FeatureCNN)]
[0057] where (Feature_ViT) further extracts global information from the features output by the convolutional network.
[0058] Step S3, through an unsupervised learning method, using an autoencoder architecture, enabling the network model to predict the corresponding depth map from the input monocular image;
[0059] Preferably, the unsupervised learning autoencoder architecture further includes pre-training the network model using transfer learning techniques to improve the speed and accuracy of depth prediction;
[0060] In this step, depth prediction is performed using an autoencoder architecture through an unsupervised learning method. The autoencoder consists of an encoder and a decoder, and its formula is:
[0061] [Encoded = Encoder(Input Image)]
[0062] [Predicted Depth = Decoder(Encoded)].
[0063] Step S4, applying a pixel-level contrast loss function to optimize the model and improve the accuracy of depth prediction;
[0064] Preferably, the pixel-level contrast loss function is the L1 loss; the L1 loss is used to optimize the training of the depth prediction model, and its formula is expressed as:
[0065] [L=\frac{1}{N}\sum_{i = 1}^{N}|p_i - g_i|]
[0066] Among them, (L) represents the loss value, (N) represents the total number of pixels in the image, (p_i) represents the depth value of the i-th pixel in the predicted depth map, and (g_i) represents the depth value of the i-th pixel in the ground truth depth map. The objective of this function is to minimize the total absolute error between the predicted depth map and the ground truth depth map, which is used for regression tasks in deep learning.
[0067] Step S5, based on the depth sorting-based virtual-real occlusion algorithm, use depth information to determine the occlusion relationship between virtual objects and real-world objects;
[0068] Preferably, collect the depth information (D(x, y)) of each object, where ((x, y)) is the pixel coordinate. The depth information is directly obtained through a depth camera or calculated from the disparity (\text{disp}(x, y)) of stereo vision:
[0069] [D(x, y)=\frac{f\cdot B}{\text{disp}(x, y)}]
[0070] Here, (f) is the focal length of the camera, and (B) is the baseline distance between the cameras;
[0071] Furthermore, for each pixel, compare the depth values of the real-world object and the virtual object:
[0072] If (D_{\text{virtual}}(x, y)<D_{\text{real}}(x, y)), the virtual object is in front;
[0073] Otherwise, the real object is in front, and the Z-Buffer algorithm is used to manage occlusion:
[0074] [Z(x, y)=\min(D_{\text{virtual}}(x, y),D_{\text{real}}(x, y))]
[0075] Render the image according to the (Z) value to ensure the correct layer order.
[0076] Step S6, use OpenGL for real-time occlusion processing to ensure that virtual objects are correctly displayed in the real world;
[0077] Preferably, the OpenGL real-time occlusion processing further includes dynamically adjusting the display of virtual and real objects to ensure that virtual objects are correctly occluded visually.
[0078] Step S7: Use a model optimization tool to accelerate the model to meet the operating requirements of edge devices;
[0079] Preferably, the model optimization tool uses TensorRT or ONNX Runtime for model acceleration optimization. The performance calculation formula depends on the used tool, TensorRT or ONNX Runtime. When optimizing, the reduction in inference time will be concerned. The optimized inference time can be expressed as:
[0080] [t_{optimized}=\frac{t_{original}}{S}]
[0081] Where, (t_{optimized}) represents the inference time of the optimized model, (t_{original}) represents the inference time of the model before optimization, and (S) represents the acceleration ratio.
[0082] Step S8: Deploy the system on a mobile device, including using the NDK of Android or Core ML of iOS for further performance optimization;
[0083] Preferably, the system deployment further includes using corresponding localization development tools on the Android or iOS platform for the efficient operation of the model on the mobile device.
[0084] A computer device includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes any one of the virtual-real occlusion processing methods based on lightweight monocular depth estimation.
[0085] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the processor executes any one of the virtual-real occlusion processing methods based on lightweight monocular depth estimation.
[0086] In this embodiment, the complete steps from the input image to depth estimation and then to occlusion processing are as follows:
[0087] Step 1: The input monocular image first passes through a lightweight convolutional neural network to extract basic features:
[0088] [\text{Feature}\text{CNN}=\text{CNN}\text{Layers}(\text{InputImage})]
[0089] Step 2, perform global feature extraction on the output of Step 1 through a small Vision Transformer layer: [Feature_ViT = Transformer_Encoder(Feature_CN
[0090] N})]
[0091] Step 3, use the decoder part in the autoencoder architecture to predict the depth map from the enhanced features:
[0092] [Predicted_Depth = Decoder(Feature_ViT)]
[0093] Step 4, optimize the predicted depth map through a pixel-level L1 loss function to make it close to the real depth map:
[0094] [L = \frac{1}{N}\sum_{i = 1}^{N}|p_i - g_i|]
[0095] Here, (p_i) is the predicted depth value, and (g_i) is the actual depth value.
[0096] Step 5, calculate the depth information of each pixel and use this information to judge the occlusion relationship between the virtual object and the real-world object:
[0097] [D(x,y) = \frac{f\cdot B}{disp(x,y)}]
[0098] [Z(x,y) = \min(D_{\text{virtual}}(x,y),D_{\text{real}}(x,y))]
[0099] Here, (Z(x,y)) determines the final rendering depth;
[0100] Step 6, use OpenGL technology to render the final image based on the result of (Z(x,y)), ensuring the correct occlusion relationship;
[0101] Step 7, use TensorRT or ONNX Runtime to accelerate the model to adapt to mobile devices:
[0102] [t_{optimized} = \frac{t_{original}}{S}]
[0103] Then, deploy the accelerated model on the Android or iOS platform.
[0104] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions.
[0105] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0106] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.
[0107] The above embodiments should be understood as being only for illustrative purposes of the present invention and not for limiting the scope of protection of the present invention. After reading the content described in the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.
Claims
1. A virtual and real occlusion processing method based on lightweight monocular depth estimation, characterized in that: The following steps are involved: Step S1, using a pre-collected data set containing a monocular image and a corresponding depth map to perform model training; Step S2, using a lightweight convolutional neural network combined with a small visual Transformer layer as a network model for feature extraction and depth prediction; Step S3, using an unsupervised learning method and an autoencoder architecture, so that the network model can predict the corresponding depth map through the input monocular image; Step S4, applying a pixel-level contrast loss function to optimize the model; Step S5, a virtual-real occlusion algorithm based on depth sorting, using depth information to determine the occlusion relationship between the virtual object and the real-world object; Step S6, using OpenGL to perform real-time occlusion processing to ensure that the virtual object is correctly displayed in the real world; Step S7, using a model optimization tool to accelerate the model; Step S8, deploying the system on a mobile device, including using Android's NDK or iOS's Core ML for further performance optimization.
2. The virtual and real occlusion processing method based on lightweight monocular depth estimation according to claim 1 is characterized in that , in step S2, the lightweight convolutional neural network is selected from MobileNet or ShuffleNet.
3. The virtual and real occlusion processing method based on lightweight monocular depth estimation according to claim 1, characterized in that: In step S2, the small visual Transformer layer is configured as Mini-ViT.
4. The virtual and real occlusion processing method based on lightweight monocular depth estimation according to claim 1 is characterized in that: In step S3, the unsupervised learning autoencoder architecture uses MobileNetv2 pre-trained based on ImageNet as an encoder, freezes shallow convolution parameters and fine-tunes deep parameters, and improves depth prediction accuracy through cross-domain feature migration.
5. The virtual and real occlusion processing method based on lightweight monocular depth estimation according to claim 1 is characterized in that: In step S4, the pixel-level contrast loss function is L1 loss; the L1 loss is used to optimize the training of the depth prediction model, and its formula is expressed as: [L=\frac{1}{N}\sum_{i=1}^{N}|p_i-g_i|] Among them, L represents the loss value, N represents the total number of pixels in the image, p_i represents the depth value of the i-th pixel in the predicted depth map, g_i represents the depth value of the i-th pixel in the real depth map, and the goal of this function is to minimize the total absolute error between the predicted depth map and the real depth map.
6. The virtual and real occlusion processing method based on lightweight monocular depth estimation according to claim 1, characterized in that: In step S6, the OpenGL real-time occlusion processing further includes dynamically adjusting the display of virtual and real objects to ensure that the virtual object is visually correctly occluded.
7. The virtual and real occlusion processing method based on lightweight monocular depth estimation according to claim 1, characterized in that: In step S7, the model optimization tool uses TensorRT or ONNX Runtime to perform model acceleration optimization, and the optimized reasoning time can be expressed as: [t_{optimized}=\frac{t_{original}}{S}] Among them, t_{optimized} represents the optimized model inference time, t_{original} represents the model inference time before optimization, and S represents the acceleration ratio.
8. The virtual and real occlusion processing method based on lightweight monocular depth estimation according to claim 1, characterized in that: In step S8, the system deployment further includes using corresponding localized development tools on the Android or iOS platform for efficient operation of the model on the mobile device.
9. A computer device, characterized in that: It includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the virtual and real occlusion processing method based on lightweight monocular depth estimation as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a processor, the processor executes the virtual and real occlusion processing method based on lightweight monocular depth estimation as described in any one of claims 1 to 8.
Citation Information
Patent Citations
A method for handling virtual and real occlusion in augmented reality
CN112365516B