An AI deep information enhancement-based naked-eye 3D display method, device and equipment
Patent Information
- Application Number
- CN202610621923.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-08
- Publication Date
- 2026-09-22
AI Technical Summary
但是裸眼3D显示设备需要输入两路分别捕捉自略微不同视角的两幅场景图像(模拟人类双目之间的间隔) 视频流才能有效,而两路视频的捕捉需要专门的双目腹腔镜镜头,对于单目的镜头无法使用
[0068](1)本发明利用人工智能深度估计技术,学习大量带有真实深度标签的图像数据,挖掘图像中的深度线索(如物体大小、遮挡关系、纹理梯度、透视、光影、已知物体的先验知识等),建立从像素外观到深度值的复杂映射模型,通过映射模型在手术视频图像中生成精准的深度图,从而实现从单目视频到双目视频的端到端自动化合成,生成的立体视频流具有良好的立体感和时空连续性,满足裸眼3D显示的严格要求;
Smart Images

Figure CN122802666A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of 3D display technology, and in particular relates to a naked-eye 3D display method, device and equipment based on AI depth information enhancement. Background Technology
[0002] Endoscopic surgery (such as thoracoscopic surgery and laparoscopic surgery) is a type of surgical procedure performed through tiny incisions in the body surface or natural cavities, inserting instruments equipped with cameras and light sources. Compared to traditional open surgery, its core advantages include smaller incisions, less bleeding, less postoperative pain, faster recovery time, shorter hospital stays, smaller and more aesthetically pleasing scars, and a reduced risk of complications such as wound infection. Furthermore, the high-definition magnified imaging provided by endoscopes offers surgeons a clearer field of vision and more precise manipulation. It is increasingly being used in surgical procedures.
[0003] However, the single-lens two-dimensional imaging system commonly used in endoscopic surgery transforms the natural binocular vision formed by direct visualization of the surgical organs in traditional open surgery into monocular vision of direct visualization of a two-dimensional screen. This makes it difficult for surgeons to perceive depth and intuitively judge the anterior-posterior distances and layered relationships between different tissues, trachea, or surgical instruments, such as the depth of tumor burial or the relative position of blood vessels and organs. This significantly reduces the safety of the surgery.
[0004] Glasses-free 3D technology projects different images to the observer's left and right eyes separately using grating hardware (such as parallax barriers or cylindrical lenses). It then utilizes the brain's visual synthesis mechanism to form a three-dimensional image, thereby reconstructing depth information in the field of view. However, glasses-free 3D display devices require two video streams—one captured from slightly different perspectives (simulating the distance between human binoculars)—to be effective. Capturing these two video streams requires specialized binocular laparoscopic lenses; monocular lenses cannot be used. Furthermore, even if binocular vision can be fully reconstructed using dual-channel video and glasses-free 3D display devices under laparoscopic visualization, the limited operating space makes it difficult to accurately perceive the depth relationships between key anatomical tissues using only binocular parallax. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of the existing technology and provide a naked-eye 3D display method, device and equipment based on AI depth information enhancement. It realizes end-to-end automated synthesis from monocular video to binocular video. The generated stereoscopic video stream has good stereoscopic effect and spatiotemporal continuity, meeting the strict requirements of naked-eye 3D display.
[0006] This invention is achieved through the following technical solution:
[0007] The first aspect of this invention discloses a glasses-free 3D display method based on AI depth information enhancement, comprising:
[0008] A training dataset is generated based on the left and right views captured by a binocular camera.
[0009] Construct a depth estimation model, and train the depth estimation model based on training samples;
[0010] Each frame of the monocular video is used as a reference image and input into the trained depth estimation model to obtain a depth prediction map and an uncertainty estimation map.
[0011] Based on the depth prediction map, the pixel coordinate mapping relationship from the reference image to the target camera's view image is calculated through three-dimensional back projection, coordinate system transformation, and reprojection. The target camera is a virtual camera.
[0012] Based on the pixel coordinate mapping relationship, the pixel values of each pixel in the reference image are assigned to the target image plane to generate another view image corresponding to the reference image;
[0013] Combine all image frames of a monocular video and their corresponding images from another perspective in chronological order to form a binocular video, and output the binocular video for display;
[0014] The depth estimation model includes:
[0015] An encoder is used to encode an input image through a multi-layer Transformer structure and output multi-scale features;
[0016] The attention fusion module is used to weight and fuse the multi-scale features using a spatial attention weight map;
[0017] The decoder is used to upsample the weighted fused multi-scale features to restore the spatial resolution of the input image and output a dense feature map.
[0018] The uncertainty estimation module is used to generate a depth prediction map and a corresponding uncertainty map based on the dense feature map through parallel branches.
[0019] Furthermore, a training dataset is generated based on the left and right views acquired by the stereo camera, including:
[0020] The stereo camera is calibrated to obtain its internal and external parameters.
[0021] The epipolar correction method is used to correct the left and right views acquired by the binocular camera so that the projections of the same spatial point on the left and right views are in the same row.
[0022] Stereo matching is performed on the left and right views to calculate disparity values and generate a disparity map;
[0023] Calculate the depth value of the pixel based on the focal length, baseline length, and parallax value;
[0024] Generate a depth map based on the depth value;
[0025] Training sample pairs are constructed using the left view as the input image and the corresponding depth map as the supervision label.
[0026] Furthermore, the encoder is specifically used for:
[0027] The image patch sequence is mapped to a dimensional representation vector through a linear projection layer, and then positional encoding is added to obtain the initial sequence;
[0028] The initial sequence is input into an encoding structure consisting of multiple Transformer structures;
[0029] Features are extracted from different depth layers of the coding structure to form multi-scale features.
[0030] Furthermore, the training process of the encoder includes: initializing the student encoder and the teacher encoder, with the student encoder and the teacher encoder having the same initial parameters; executing the training steps of the student encoder and the teacher encoder until the preset training termination condition is met, and then determining the teacher encoder as the trained encoder.
[0031] The training steps for the student encoder and the teacher encoder include:
[0032] Data augmentation is performed on the input image to generate two different augmented views;
[0033] Each enhanced view is divided into multiple image blocks to form an image block sequence;
[0034] Calculate the similarity between corresponding image patches in two enhanced views, and construct a similarity binary matrix based on a preset similarity threshold;
[0035] The augmented views are input into the student encoder and the teacher encoder respectively to generate the corresponding feature matrices;
[0036] Based on the similarity binary matrix and the feature matrix, the image patch similarity regularization loss and self-distillation loss are calculated, and then summed to obtain the total training loss;
[0037] Based on the total training loss, the parameters of the student encoder are updated using the backpropagation algorithm;
[0038] The parameters of the teacher encoder are updated using an exponential moving average method based on the updated parameters of the student encoder.
[0039] Furthermore, the attention fusion module is specifically used for:
[0040] Each feature representation in the multi-scale features output by the encoder is reconstructed into a two-dimensional spatial feature map.
[0041] Calculate the attention weight map for each two-dimensional spatial feature map;
[0042] The corresponding attention weight maps are applied to weight each two-dimensional spatial feature map.
[0043] Furthermore, the encoder is a ViT-B / 16 model, and the decoder is a UNet model.
[0044] Furthermore, the total loss function of the depth estimation model is:
[0045] in, These are weighting coefficients. It's a depth tag. It is a prediction of depth. This represents the uncertainty value of the i-th pixel.
[0046] Furthermore, based on the depth prediction map, the pixel coordinate mapping relationship from the reference image to the target camera's viewpoint image is calculated through 3D backprojection, coordinate system transformation, and reprojection, including:
[0047] Based on the predicted depth map and the intrinsic parameter matrix of the reference camera, the pixels in the reference image are transformed into three-dimensional points in the coordinate system of the reference camera through three-dimensional back projection. The reference camera is a camera that captures monocular video.
[0048] Based on the preset external parameters of the target camera relative to the reference camera, the three-dimensional points are transformed from the reference camera coordinate system to the target camera coordinate system;
[0049] Based on the intrinsic parameter matrix of the target camera, the three-dimensional points after coordinate system transformation are reprojected onto the two-dimensional image plane of the target camera to obtain the pixel coordinates of the target image;
[0050] Calculate the pixel coordinate mapping relationship from the reference image to the target camera's viewpoint image.
[0051] A second aspect of the present invention discloses a naked-eye 3D display device based on AI depth information enhancement, comprising:
[0052] The training set generation module is used to generate training datasets based on the left and right views captured by the binocular camera.
[0053] A model building module is used to build a depth estimation model and train the depth estimation model based on training sample pairs;
[0054] The image prediction module is used to input each frame of the monocular video as a reference image into the trained depth estimation model to obtain a depth prediction map and an uncertainty estimation map.
[0055] The mapping relationship calculation module is used to calculate the pixel coordinate mapping relationship from the reference image to the target camera's view image based on the depth prediction map through three-dimensional back projection, coordinate system transformation and reprojection, wherein the target camera is a virtual camera;
[0056] The image generation module is used to assign the pixel values of each pixel in the reference image to the target image plane based on the pixel coordinate mapping relationship, thereby generating another view image corresponding to the reference image.
[0057] The video generation module is used to combine all image frames of a monocular video and their corresponding images from another perspective in chronological order to form a stereo video, and to output and display the stereo video.
[0058] The depth estimation model includes:
[0059] An encoder is used to encode an input image through a multi-layer Transformer structure and output multi-scale features;
[0060] The attention fusion module is used to weight and fuse the multi-scale features using a spatial attention weight map;
[0061] The decoder is used to upsample the weighted fused multi-scale features to restore the spatial resolution of the input image and output a dense feature map.
[0062] The uncertainty estimation module is used to generate a depth prediction map and a corresponding uncertainty map based on the dense feature map through parallel branches.
[0063] A third aspect of the present invention discloses an electronic device comprising:
[0064] At least one processor;
[0065] A memory that is communicatively connected to the at least one processor;
[0066] The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the naked-eye 3D display method according to the first aspect of the present invention.
[0067] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0068] (1) This invention utilizes artificial intelligence depth estimation technology to learn a large amount of image data with real depth labels, mine depth clues in the image (such as object size, occlusion relationship, texture gradient, perspective, light and shadow, prior knowledge of known objects, etc.), establish a complex mapping model from pixel appearance to depth value, generate accurate depth map in surgical video image through the mapping model, thereby realizing end-to-end automated synthesis from monocular video to binocular video, the generated stereo video stream has good stereo sense and spatiotemporal continuity, and meets the strict requirements of naked-eye 3D display;
[0069] (2) By introducing a spatial attention fusion module, the present invention enables the network to adaptively focus on key anatomical structures such as blood vessels, trachea, and cardia, which significantly improves the spatial accuracy and boundary consistency of depth estimation.
[0070] (3) The present invention employs an uncertainty estimation module, which enables the network to automatically identify and reduce the confidence of unreliable areas in complex intraoperative scenarios (such as tissue reflection, liquid occlusion, and smoke interference). Through uncertainty weighted loss function and post-processing optimization, the robustness of the system in environments such as low texture and abnormal lighting is greatly enhanced.
[0071] (4) In this invention, the encoder uses a dense feature regularization module when performing representation learning pre-training on endoscopic images. This allows the encoder to focus more on local details, with different attention heads [CLS] focusing on and understanding different semantic objects (blood vessels, surgical instruments, cavities, etc.) in the image.
[0072] (5) Existing technologies usually require joint training of the encoder and decoder, which is slow, unstable, and has poor generalization ability. In this invention, the encoder is pre-trained through self-distillation. During the training of the deep estimation model, the encoder parameters are frozen and only the decoder parameters are updated. Due to the self-supervised pre-training and freezing mechanism of the encoder, the generated features are more stable and general, and the training convergence speed of the deep estimation model is faster. Attached Figure Description
[0073] The accompanying drawings, which are included to provide a further understanding of embodiments of the invention and form part of this application, do not constitute a limitation thereof. In the drawings:
[0074] Figure 1 This is a flowchart illustrating one method of naked-eye 3D display in this invention.
[0075] Figure 2 This is a flowchart illustrating the training dataset generation method in this invention;
[0076] Figure 3This is a schematic flowchart of the encoder training method in this invention;
[0077] Figure 4 This is a schematic diagram of a naked-eye 3D display device according to the present invention;
[0078] Figure 5 This is a schematic diagram of the structure of the electronic device provided in this invention. Detailed Implementation
[0079] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0080] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual pictures. They should not be construed as limiting the invention. To better illustrate the embodiments of the invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0081] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "inner," and "outer" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the accompanying drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.
[0082] In the description of this invention, unless otherwise explicitly specified and limited, the term "connection" or similar designation indicating a connection between components should be interpreted broadly. For example, it can refer to a fixed connection, a detachable connection, or an integral part; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can refer to the internal communication between two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0083] like Figures 1 to 5 As shown in the figure, this embodiment discloses a naked-eye 3D display method, device and equipment based on AI depth information enhancement.
[0084] The first aspect of this embodiment discloses a naked-eye 3D display method based on AI depth information enhancement, such as... Figure 1As shown, the naked-eye 3D display method includes steps S100 to S600.
[0085] Step S100. Generate a training dataset based on the left and right views acquired by the binocular camera.
[0086] In some implementations of this embodiment, such as Figure 2 As shown, a training dataset is generated based on the left and right views acquired by the binocular camera, including steps S110 to S160.
[0087] Step S110. Calibrate the stereo camera and obtain its internal and external parameters.
[0088] The internal parameters of the binocular camera include focal length.
[0089] The external parameters of the binocular camera are used to determine the baseline length and relative pose between the left and right cameras.
[0090] In these implementations, the geometric accuracy of subsequent depth calculations can be guaranteed by calibrating the binocular camera.
[0091] Step S120. Use the epipolar correction method to correct the left and right views acquired by the binocular camera so that the projections of the same spatial point on the left and right views are in the same row.
[0092] In these implementations, the epipolar correction method is used to correct the left and right views, simplifying the matching search to a single horizontal dimension, thereby improving the accuracy and computational efficiency of disparity calculation.
[0093] Step S130. Perform stereo matching on the left and right views to calculate the disparity value and generate a disparity map.
[0094] For example, taking the left view as a reference, for each pixel in the left view, search for the corresponding pixel in the same row of the right view; calculate the difference in the horizontal coordinate between the pixel in the left view and the corresponding pixel in the right view as the disparity value; map and store the disparity value corresponding to each pixel in the left view as the pixel value of the corresponding position to generate a disparity map of the same size as the left view.
[0095] The magnitude of parallax is inversely proportional to the depth of the target point; closer objects correspond to larger parallax, while distant objects correspond to smaller parallax.
[0096] Step S140. Calculate the depth value of the pixel based on the focal length, baseline length, and disparity value.
[0097] The formula for calculating the depth value of the pixel is:
[0098]
[0099] In the formula, M represents the depth value, f represents the focal length of the camera, B represents the baseline length of the camera, and d represents the parallax value.
[0100] Step S150. Generate a depth map based on the depth values.
[0101] Step S160. Construct training sample pairs using the left view as the input image and the corresponding depth map as the supervision label.
[0102] Each training sample pair includes a left view and the corresponding depth map.
[0103] In some embodiments, a training dataset covering different lighting, textures, and endoscopic scenarios is formed by acquiring a large number of intraoperative videos.
[0104] Step S200. Construct a depth estimation model and train the depth estimation model based on training sample pairs.
[0105] The depth estimation model includes: an encoder for encoding the input image through a multi-layer Transformer structure and outputting multi-scale features; an attention fusion module for weighted fusion of the multi-scale features using a spatial attention weight map; a decoder for upsampling the weighted fused multi-scale features to restore the spatial resolution of the input image and outputting a dense feature map; and an uncertainty estimation module for generating a depth prediction map and a corresponding uncertainty map based on the dense feature map through parallel branches.
[0106] In some embodiments of this example, the encoding includes a multi-layer Transformer structure, and the encoder is specifically used to: map the image patch sequence to a dimensional representation vector through a linear projection layer, and then add positional encoding to obtain an initial sequence; input the initial sequence into an encoding structure composed of a multi-layer Transformer structure; and extract features from different depth layers of the encoding structure to form multi-scale features.
[0107] In some embodiments of this example, the encoder is a ViT-B / 16 model, which consists of 12 layers of standard Transformer stacks.
[0108] For example, the input to the encoder is a two-dimensional image. Where H represents the image height, W represents the image width, and C represents the number of image channels, typically 3, corresponding to RGB channels. The encoder is used to divide a two-dimensional image into a sequence of non-overlapping image blocks. ,in, Indicates the length of the image patch sequence. Represents the resolution of an image patch, for example, = This indicates that the resolution of the image block is 16×16.
[0109] The hidden dimension of the ViT-B / 16 model is D. The ViT-B / 16 model maps the image patch sequence to a representation vector of dimension D through a linear projection layer and adds positional encoding to obtain the initial sequence:
[0110]
[0111] in, Represents the linear projection matrix. Represents the position encoding matrix, This indicates a category token.
[0112] In these implementations, the order of the image patch sequence is determined by adding positional encoding to the vector.
[0113] In these implementations, the first step is calculated using multi-head self-attention (MSA) and multilayer perceptron (MLP). The output of the first layer of Transformer, The representation vector output by the Transformer layer is:
[0114]
[0115] Here, MSA represents multi-head self-attention layer, MLP represents feedforward network layer, and Norm represents layer normalization operation.
[0116] In some implementations of this embodiment, such as Figure 3 As shown, the training process of the encoder includes steps S210 to S290.
[0117] Step S210. Initialize the student encoder and the teacher encoder. The initial parameters of the student encoder and the teacher encoder are the same.
[0118] For example, initialize a ViT-B / 16 encoder network, and create two instances of the ViT-B / 16 encoder network as a student encoder and a teacher encoder, respectively, with the student encoder and the teacher encoder having the same initial parameters.
[0119] Step S220. Perform data augmentation on the input image to generate two different augmented views.
[0120] For example, two random enhancements are applied to the input image to generate two different global enhanced views, u and v, with dimensions C×H×W.
[0121] Step S230. Divide each enhanced view into multiple image blocks to form an image block sequence.
[0122] For example, dividing the enhanced views u and v into N non-overlapping image blocks (e.g., image block size is 16×16) yields an image block sequence. and .
[0123] Step S240. Calculate the similarity between corresponding image patches in two enhanced views, and construct a similarity binary matrix based on a preset similarity threshold.
[0124] For example, the cosine similarity between corresponding image patches is calculated, and a binary similarity matrix B and its inverse matrix are constructed based on a preset similarity threshold T. If the cosine similarity between two image patches is greater than the similarity threshold, the two image patches are considered similar; otherwise, the two image patches are considered dissimilar.
[0125] The similarity binary matrix is as follows:
[0126]
[0127] Where sim(,) represents the cosine similarity. express The i-th row vector The j-th row vector, i,j∈[0, N).
[0128] Inverting the similarity binary matrix bitwise yields its inverse matrix:
[0129] .
[0130] Step S250. Input the enhanced view into the student encoder and the teacher encoder respectively to generate the corresponding feature matrix.
[0131] Specifically, the augmented view is input into the student encoder and the teacher encoder respectively to generate the corresponding feature matrices, including: inputting the augmented view into the student encoder to generate the student feature matrix, and inputting the augmented view into the teacher encoder to generate the teacher feature matrix.
[0132] Step S260. Based on the similarity binary matrix and the feature matrix, calculate the image patch similarity regularization loss and the self-distillation loss, and then sum them to obtain the total training loss.
[0133] In some implementations of this embodiment, based on the similarity binary matrix and the feature matrix, the image patch similarity regularization loss and the self-distillation loss are calculated, and then summed to obtain the total training loss. This includes: calculating the image patch similarity regularization loss based on the similarity binary matrix and the feature matrix; calculating the self-distillation loss between the student encoder and the teacher encoder; and summing the image patch similarity regularization loss and the self-distillation loss to obtain the total training loss.
[0134] In these implementations, the similarity regularization loss function forces the student encoder to learn representations of local features, enhancing the correlation of features in similar image patches and reducing the correlation of features in dissimilar image patches, thereby enhancing the ability to perform dense prediction tasks, such as depth estimation.
[0135] For example, the formula for calculating the similarity regularization loss is:
[0136]
[0137]
[0138]
[0139] in, This represents the feature matrix output by the student encoder to the augmented view v. This represents the feature matrix output by the student encoder for the augmented view u. This represents the feature matrix output by the teacher encoder to the enhanced view v. d represents the feature matrix output by the teacher encoder for the augmented view u, where d is the embedding dimension of the model.
[0140] The formula for calculating the total training loss is:
[0141]
[0142] in, This represents the loss due to similarity regularization. This indicates loss due to self-distillation.
[0143] Step S270. Update the parameters of the student encoder using the backpropagation algorithm based on the total training loss.
[0144] Step S280. Update the parameters of the teacher encoder using an exponential moving average method based on the updated parameters of the student encoder.
[0145] Step S290. Determine whether the preset training termination condition is met. If yes, then determine the teacher encoder as the trained encoder; otherwise, proceed to step S220.
[0146] In some embodiments of this example, the attention fusion module is specifically used to: reshape each feature representation in the multi-scale features output by the encoder into a two-dimensional spatial feature map; calculate an attention weight map for each two-dimensional spatial feature map; and apply the corresponding attention weight map to weight each two-dimensional spatial feature map respectively.
[0147] For example, the multi-scale features output by the encoder Each feature in the representation Reconstructed into a two-dimensional spatial feature map ,in, and These represent the height and width of the two-dimensional spatial feature map, respectively.
[0148] The formula for calculating the attention weight map is:
[0149]
[0150] Where Conv represents the convolution operation, and These represent convolution operations with kernel sizes of 1×1 and 3×3, respectively; ReLU is a non-linear activation function. This is the Sigmoid function.
[0151] The weighted calculation formula for the two-dimensional spatial feature map is as follows:
[0152]
[0153] in, This indicates element-wise multiplication.
[0154] In these implementations, the attention fusion module enables the network to automatically focus on regions with significant texture and depth variations through a learnable attention weight map, thereby improving the structural clarity and boundary consistency of the predictions and enhancing the spatial sensitivity of the depth estimation model to key anatomical structures (such as blood vessels, trachea, and cardia).
[0155] The attention fusion module will fuse the multi-scale features. It is fed into the corresponding layer in the decoder to complete layer-by-layer decoding and depth reconstruction.
[0156] In some embodiments of this example, the decoder is a UNet model.
[0157] In these implementations, the decoder uses the deconvolution part of the UNet architecture, and performs upsampling and feature reconstruction layer by layer by merging with multi-scale features after attention fusion through skip connections.
[0158] In some implementations of this embodiment, the uncertainty estimation module consists of two parallel convolutional branches: the first convolutional branch is used to generate a depth prediction map; the second convolutional branch is used to generate an uncertainty map, which is then used to calculate a Gaussian-weighted loss function. .
[0159] In some embodiments of this example, the total loss function of the depth estimation model is:
[0160]
[0161] in, These are weighting coefficients. This represents the mean squared error loss function. This represents the Gaussian weighted loss function.
[0162] The mean squared error loss function is:
[0163]
[0164] in, It's a depth tag. It is the predicted depth.
[0165] The Gaussian weighted loss function is:
[0166]
[0167] in, This represents the uncertainty value of the i-th pixel.
[0168] Step S300. Input each frame of the monocular video as a reference image into the trained depth estimation model to obtain a depth prediction map and an uncertainty estimation map.
[0169] Step S400. Based on the depth prediction map, the pixel coordinate mapping relationship from the reference image to the target camera's viewpoint image is calculated through three-dimensional back projection, coordinate system transformation, and reprojection. The target camera is a virtual camera.
[0170] In some implementations of this embodiment, based on the depth prediction map, the pixel coordinate mapping relationship from the reference image to the target camera view image is calculated through three-dimensional back projection, coordinate system transformation and reprojection, including steps S410 to S440.
[0171] Step S410. Based on the predicted depth map and the intrinsic parameter matrix of the reference camera, the pixels in the reference image are transformed into three-dimensional points in the coordinate system of the reference camera through three-dimensional back projection. The reference camera is a camera that captures monocular video.
[0172] For example, for any pixel in the reference image, its pixel coordinates are (u1, v1), and its corresponding depth value is Z(u1, v1). Then, the three-dimensional coordinates of this pixel in the reference camera coordinate system can be obtained by the following formula:
[0173]
[0174]
[0175]
[0176] in, This indicates the position of the camera's principal point on the pixel plane, usually close to the image center; , Z(u1,v1) represents the focal length of the camera in the horizontal and vertical directions, respectively, in pixels; Z(u1,v1) is the depth value, representing the distance from the corresponding 3D scene point to the camera's optical center.
[0177] Using matrix form, the above relationship can be simplified to:
[0178] ,
[0179]
[0180] in, Represents the homogeneous coordinates of a pixel; This indicates the three-dimensional position of a pixel in the reference camera coordinate system.
[0181] Step S420. Based on the preset external parameters of the target camera relative to the reference camera, transform the three-dimensional point from the reference camera coordinate system to the target camera coordinate system.
[0182] For example, suppose the extrinsic parameters of the target camera are determined by the rotation matrix. Translation vector If the composition is such that the pixel is represented in the target camera coordinate system, then:
[0183]
[0184] Wherein, rotation matrix It is a 3×3 orthogonal matrix used to describe the orientation transformation between the reference camera coordinate system and the target camera coordinate system; the translation vector... It is a 3×1 vector used to describe the relative position between the origins of the two camera coordinate systems.
[0185] Step S430. Based on the intrinsic parameter matrix of the target camera, reproject the three-dimensional points after coordinate system transformation onto the two-dimensional image plane of the target camera to obtain the pixel coordinates of the target image.
[0186] The intrinsic parameter matrix of the target camera is the same as that of the reference camera.
[0187] For example, when a pixel is projected back onto the pixel plane of the target camera, its pixel coordinates on the target image are obtained as follows:
[0188] ,
[0189] in, It is the representation of the pixel in homogeneous coordinates.
[0190] The coordinates are then normalized to obtain the actual pixel positions on the target camera image plane:
[0191] .
[0192] Step S440. Calculate the pixel coordinate mapping relationship from the reference image to the target camera view image.
[0193] Step S500. Based on the pixel coordinate mapping relationship, assign the pixel values of each pixel in the reference image to the target image plane to generate another view image corresponding to the reference image.
[0194] For example, after obtaining the pixel position of the target image, the pixel value in the reference image is used. By copying the past, we obtain an image from another camera perspective:
[0195]
[0196] Thus, the pixel colors in the reference camera image can be transferred to the target camera plane according to spatial geometric relationships, ultimately generating a complete image frame from another perspective corresponding to the reference image, realizing the synthetic extension from monocular image to binocular image.
[0197] Step S600. Combine all image frames of the monocular video and their corresponding other viewpoint images in chronological order to form a binocular video, and output the binocular video for display.
[0198] For example, the generated binocular video can be output to a corresponding glasses-free 3D display device for display.
[0199] The second aspect of this embodiment discloses a naked-eye 3D display device based on AI depth information enhancement, such as... Figure 4 As shown, the naked-eye 3D display device includes a training set generation module, a model building module, an image prediction module, a mapping relationship calculation module, an image generation module, and a video generation module.
[0200] The training set generation module is used to generate a training dataset based on the left and right views captured by the binocular camera.
[0201] The model building module is used to build a depth estimation model and train the depth estimation model based on training samples.
[0202] The image prediction module is used to input each frame of the monocular video as a reference image into the trained depth estimation model to obtain a depth prediction map and an uncertainty estimation map.
[0203] The mapping relationship calculation module is used to calculate the pixel coordinate mapping relationship from the reference image to the target camera's view image based on the depth prediction map through three-dimensional back projection, coordinate system transformation and reprojection, wherein the target camera is a virtual camera.
[0204] The image generation module is used to assign the pixel values of each pixel in the reference image to the target image plane based on the pixel coordinate mapping relationship, thereby generating another perspective image corresponding to the reference image.
[0205] The video generation module is used to combine all image frames of a monocular video and their corresponding images from another perspective in chronological order to form a binocular video, and to output and display the binocular video.
[0206] The depth estimation model includes an encoder, an attention fusion module, a decoder, and an uncertainty estimation module.
[0207] The encoder is used to encode the input image through a multi-layer Transformer structure and output multi-scale features; the attention fusion module is used to perform weighted fusion of the multi-scale features through a spatial attention weight map; the decoder is used to upsample the weighted fused multi-scale features to restore the spatial resolution of the input image and output a dense feature map; the uncertainty estimation module is used to generate a depth prediction map and a corresponding uncertainty map based on the dense feature map through parallel branches.
[0208] It should be noted that those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system or device described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0209] The third aspect of this embodiment discloses an electronic device, such as... Figure 5 As shown, the electronic device includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus.
[0210] The memory stores a computer program, which, when executed by the processor, causes the processor to perform the steps of the naked-eye 3D display method described in the first aspect of this embodiment.
[0211] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0212] The communication interface is used for communication between the aforementioned electronic devices and other devices.
[0213] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0214] The processors mentioned above can be general-purpose processors, including central processing units, network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits, field-programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0215] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A naked-eye 3D display method based on AI depth information enhancement, characterized in that, include: A training dataset is generated based on the left and right views captured by a binocular camera. Construct a depth estimation model, and train the depth estimation model based on training samples; Each frame of the monocular video is used as a reference image and input into the trained depth estimation model to obtain a depth prediction map and an uncertainty estimation map. Based on the depth prediction map, the pixel coordinate mapping relationship from the reference image to the target camera's view image is calculated through three-dimensional back projection, coordinate system transformation, and reprojection. The target camera is a virtual camera. Based on the pixel coordinate mapping relationship, the pixel values of each pixel in the reference image are assigned to the target image plane to generate another view image corresponding to the reference image; Combine all image frames of a monocular video and their corresponding images from another perspective in chronological order to form a binocular video, and output the binocular video for display; The depth estimation model includes: An encoder is used to encode an input image through a multi-layer Transformer structure and output multi-scale features; The attention fusion module is used to weight and fuse the multi-scale features using a spatial attention weight map; The decoder is used to upsample the weighted fused multi-scale features to restore the spatial resolution of the input image and output a dense feature map. The uncertainty estimation module is used to generate a depth prediction map and a corresponding uncertainty map based on the dense feature map through parallel branches.
2. The naked-eye 3D display method based on AI depth information enhancement according to claim 1, characterized in that, Training datasets are generated based on left and right views acquired by binocular cameras, including: The stereo camera is calibrated to obtain its internal and external parameters. The epipolar correction method is used to correct the left and right views acquired by the binocular camera so that the projections of the same spatial point on the left and right views are in the same row. Stereo matching is performed on the left and right views to calculate disparity values and generate a disparity map; Calculate the depth value of the pixel based on the focal length, baseline length, and parallax value; Generate a depth map based on the depth value; Training sample pairs are constructed using the left view as the input image and the corresponding depth map as the supervision label.
3. The naked-eye 3D display method based on AI depth information enhancement according to claim 1, characterized in that, The encoder is specifically used for: The image patch sequence is mapped to a dimensional representation vector through a linear projection layer, and then positional encoding is added to obtain the initial sequence; The initial sequence is input into an encoding structure consisting of multiple Transformer structures; Features are extracted from different depth layers of the coding structure to form multi-scale features.
4. The naked-eye 3D display method based on AI depth information enhancement according to claim 1, characterized in that, The training process of the encoder includes: initializing the student encoder and the teacher encoder, with the same initial parameters for the student encoder and the teacher encoder; executing the training steps for the student encoder and the teacher encoder until the preset training termination condition is met, and then determining the teacher encoder as the trained encoder. The training steps for the student encoder and the teacher encoder include: Data augmentation is performed on the input image to generate two different augmented views; Each enhanced view is divided into multiple image blocks to form an image block sequence; Calculate the similarity between corresponding image patches in two enhanced views, and construct a similarity binary matrix based on a preset similarity threshold; The augmented views are input into the student encoder and the teacher encoder respectively to generate the corresponding feature matrices; Based on the similarity binary matrix and the feature matrix, the image patch similarity regularization loss and self-distillation loss are calculated, and then summed to obtain the total training loss; Based on the total training loss, the parameters of the student encoder are updated using the backpropagation algorithm; The parameters of the teacher encoder are updated using an exponential moving average method based on the updated parameters of the student encoder.
5. The naked-eye 3D display method based on AI depth information enhancement according to claim 1, characterized in that, The attention fusion module is specifically used for: Each feature representation in the multi-scale features output by the encoder is reconstructed into a two-dimensional spatial feature map. Calculate the attention weight map for each two-dimensional spatial feature map; The corresponding attention weight maps are applied to weight each two-dimensional spatial feature map.
6. The naked-eye 3D display method based on AI depth information enhancement according to claim 1, characterized in that, The encoder is a ViT-B / 16 model, and the decoder is a UNet model.
7. The naked-eye 3D display method based on AI depth information enhancement according to claim 1, characterized in that, The total loss function of the depth estimation model is: ; ; ; in, These are weighting coefficients. It's a depth tag. It is a prediction of depth. This represents the uncertainty value of the i-th pixel.
8. The naked-eye 3D display method based on AI depth information enhancement according to claim 1, characterized in that, Based on the depth prediction map, the pixel coordinate mapping relationship from the reference image to the target camera's viewpoint image is calculated through 3D backprojection, coordinate system transformation, and reprojection, including: Based on the predicted depth map and the intrinsic parameter matrix of the reference camera, the pixels in the reference image are transformed into three-dimensional points in the coordinate system of the reference camera through three-dimensional back projection. The reference camera is a camera that captures monocular video. Based on the preset external parameters of the target camera relative to the reference camera, the three-dimensional points are transformed from the reference camera coordinate system to the target camera coordinate system; Based on the intrinsic parameter matrix of the target camera, the three-dimensional points after coordinate system transformation are reprojected onto the two-dimensional image plane of the target camera to obtain the pixel coordinates of the target image; Calculate the pixel coordinate mapping relationship from the reference image to the target camera's viewpoint image.
9. A naked-eye 3D display device based on AI depth information enhancement, characterized in that, include: The training set generation module is used to generate training datasets based on the left and right views captured by the binocular camera. A model building module is used to build a depth estimation model and train the depth estimation model based on training sample pairs; The image prediction module is used to input each frame of the monocular video as a reference image into the trained depth estimation model to obtain a depth prediction map and an uncertainty estimation map. The mapping relationship calculation module is used to calculate the pixel coordinate mapping relationship from the reference image to the target camera's view image based on the depth prediction map through three-dimensional back projection, coordinate system transformation and reprojection, wherein the target camera is a virtual camera; The image generation module is used to assign the pixel values of each pixel in the reference image to the target image plane based on the pixel coordinate mapping relationship, thereby generating another view image corresponding to the reference image. The video generation module is used to combine all image frames of a monocular video and their corresponding images from another perspective in chronological order to form a stereo video, and to output and display the stereo video. The depth estimation model includes: An encoder is used to encode an input image through a multi-layer Transformer structure and output multi-scale features; The attention fusion module is used to weight and fuse the multi-scale features using a spatial attention weight map; The decoder is used to upsample the weighted fused multi-scale features to restore the spatial resolution of the input image and output a dense feature map. The uncertainty estimation module is used to generate a depth prediction map and a corresponding uncertainty map based on the dense feature map through parallel branches.
10. An electronic device, characterized in that, include: At least one processor; A memory that is communicatively connected to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the naked-eye 3D display method according to any one of claims 1-8.