Depth map generation method and device based on large model, three-dimensional reconstruction method and device, electronic equipment and storage medium
Through a large-model-based depth map generation method and the fusion technology of visual encoding and pre-trained large language models, the low accuracy problem of monocular depth estimation in complex scenes is solved, and efficient and high-precision depth map generation and three-dimensional reconstruction are achieved, which is suitable for fields such as autonomous driving, robot navigation and augmented reality.
Patent Information
- Application Number
- CN202510830781.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-19
AI Technical Summary
Existing monocular depth estimation technology has the disadvantages of high cost, complex deployment, high power consumption or environmental restrictions. In addition, the monocular depth estimation model has low accuracy in complex scenarios, making it difficult to meet the actual needs of fields such as autonomous driving, robot navigation and augmented reality.
A large-model-based depth map generation method is adopted. By visually encoding the monocular image and fusing it with a pre-trained large language model, global guided features are generated. Noise is added to the color image for denoising, and implicit features matching the joint semantic information are generated, ultimately generating a high-precision depth map.
It improves the accuracy and efficiency of monocular depth estimation, can generate high-quality depth maps, support high-precision 3D reconstruction, and meet application requirements in fields such as autonomous driving, robot navigation, and augmented reality.
Smart Images

Figure CN120672926A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the fields of artificial intelligence technology, particularly computer vision, deep learning, and large-scale models. It can be applied to scenarios such as real-time road scene depth perception, three-dimensional environmental reconstruction and obstacle avoidance, and virtual-real scene fusion. More specifically, this disclosure provides a large-scale model-based depth map generation method, three-dimensional reconstruction method, apparatus, electronic device, and storage medium. Background Art
[0002] With the rapid development of autonomous driving, robotic navigation, augmented reality (AR), and virtual reality (VR), acquiring depth information has become a core requirement for environmental perception, spatial interaction, and intelligent decision-making. However, depth perception technologies (such as LiDAR, binocular cameras, or multi-sensor fusion) suffer from high cost, complex deployment, high power consumption, and environmental constraints. Monocular depth estimation, a lightweight solution that predicts 3D depth information from a single 2D image, has gradually become a research hotspot. Summary of the Invention
[0003] The present disclosure provides a large model-based depth map generation method, a three-dimensional reconstruction method, an apparatus, an electronic device, and a storage medium.
[0004] According to one aspect of the present disclosure, a large-model-based depth map generation method is provided, comprising: visually encoding a monocular image to obtain an encoded image; fusing the encoded image and target text into a pre-trained large language model to obtain fused features; generating a global guiding feature based on the fused feature, the global guiding feature including joint semantic information of visual features and text features; adding noise to a color image of the monocular image to obtain a noise feature sequence; denoising the noise feature sequence based on the global guiding feature to generate implicit features matching the joint semantic information; and generating a depth map based on the implicit features.
[0005] According to another aspect of the present disclosure, a three-dimensional reconstruction method is provided, comprising: acquiring a monocular image acquired by an image acquisition device; converting a depth map corresponding to the monocular image into point cloud data based on posture information of the image acquisition device; performing three-dimensional reconstruction based on the point cloud data to obtain a three-dimensional reconstruction result; wherein the depth map is generated according to the large model-based depth map generation method of the present disclosure.
[0006] According to another aspect of the present disclosure, a large-model-based depth map generation device is provided, comprising: a visual encoding module for visually encoding a monocular image to obtain an encoded image; a fusion module for fusing the encoded image and target text input into a pre-trained large language model to obtain a fusion feature; a first generation module for generating a global guide feature based on the fusion feature, the global guide feature including joint semantic information of visual features and text features; a noise addition module for adding noise to a color image of the monocular image to obtain a noise feature sequence; a denoising module for denoising the noise feature sequence based on the global guide feature to generate an implicit feature matching the joint semantic information; and a second generation module for generating a depth map based on the implicit feature.
[0007] According to another aspect of the present disclosure, a three-dimensional reconstruction device is provided, comprising: an acquisition module for acquiring a monocular image acquired by an image acquisition device; a conversion module for converting a depth map corresponding to a monocular image sequence into point cloud data based on posture information of the image acquisition device; and a reconstruction module for performing three-dimensional reconstruction based on the point cloud data to obtain a three-dimensional reconstruction result; wherein the depth map is generated by the large model-based depth map generation device of the present disclosure.
[0008] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided according to the present disclosure.
[0009] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided. The computer instructions are used to cause a computer to execute the method provided according to the present disclosure.
[0010] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements the method provided according to the present disclosure when executed by a processor.
[0011] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention.
[0013] Figure 1 1 is a schematic diagram of an exemplary system architecture to which a large model-based depth map generation method and apparatus, and a 3D reconstruction method and apparatus can be applied according to an embodiment of the present disclosure;
[0014] Figure 2 is a flowchart of a method for generating a depth map based on a large model according to an embodiment of the present disclosure;
[0015] Figure 3 is a flowchart of visual encoding of a monocular image according to one embodiment of the present disclosure;
[0016] Figure 4 is a flow chart of gradually denoising a noise feature sequence according to one embodiment of the present disclosure;
[0017] Figure 5 is a flow chart of a three-dimensional reconstruction method according to an embodiment of the present disclosure;
[0018] Figure 6 is a flowchart for realizing autonomous driving based on a three-dimensional reconstruction method according to one embodiment of the present disclosure;
[0019] Figure 7 is a block diagram of a depth map generation apparatus based on a large model according to an embodiment of the present disclosure;
[0020] Figure 8 is a block diagram of a three-dimensional reconstruction apparatus according to an embodiment of the present disclosure; and
[0021] Figure 9 A schematic block diagram of an example electronic device 900 is shown, which may be used to implement embodiments of the present disclosure. DETAILED DESCRIPTION
[0022] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0023] Monocular estimation methods can include two methods: the first method is based on depth annotation data, using a pure visual backbone network to learn geometric features from visual data to predict depth information; the second method is a monocular depth estimation method that uses text and image fusion.
[0024] However, the first method requires large-scale labeled data, which has high labeling costs. When the scene changes, the data needs to be re-labeled and the model needs to be trained, which has poor domain generalization and high development costs. It is prone to errors and has poor results for complex scenes with transparent / reflective objects. The model is also difficult to interpret.
[0025] The second method is based on the simple fusion of text and image for monocular estimation. However, the cross-modal information fusion is insufficient, resulting in low estimation accuracy and poor quality of the generated depth map, which in turn affects the actual application effect of the depth map.
[0026] In view of this, an embodiment of the present disclosure provides a method and device for generating a depth map based on a large model, the method comprising: visually encoding a monocular image to obtain an encoded image. The encoded image and the target text are input into a pre-trained large language model for fusion to obtain fusion features. Based on the fusion features, a global guidance feature is generated, and the global guidance feature includes joint semantic information of visual features and text features. Noise is added to the color image of the monocular image to obtain a noise feature sequence. Based on the global guidance feature, the noise feature sequence is denoised to generate implicit features that match the joint semantic information. A depth map is generated based on the implicit features. The depth map obtained in this way can achieve high-precision three-dimensional reconstruction, and thus can better meet the practical applications in fields such as autonomous driving, robot navigation, and augmented reality / virtual reality.
[0027] Figure 1 This is a schematic diagram of an exemplary system architecture that can be applied to a large model-based depth map generation method and device, and a 3D reconstruction method and device according to an embodiment of the present disclosure. It should be noted that: Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure, but do not mean that the embodiments of the present disclosure may not be used in other devices, systems, environments or scenarios.
[0028] like Figure 1 As shown, the system architecture 100 according to this embodiment may include image acquisition devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium for providing communication links between the image acquisition devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0029] Image acquisition devices 101, 102, and 103 can be various electronic devices with image acquisition capabilities, including but not limited to smartphones, AR glasses, cameras, and webcams. The images captured by image acquisition devices 101, 102, and 103 can be monocular images. Image acquisition devices 101, 102, and 103 can interact with server 105 via network 104 to receive or send information, etc.
[0030] Server 105 can be a server that provides various services, such as a background management server that performs depth estimation on monocular images obtained by image acquisition devices 101, 102, and 103 (for example only). The background management server can perform depth estimation on the monocular images transmitted by image acquisition devices 101, 102, and 103 via network 104 to generate a depth map. The background management server can also perform 3D reconstruction based on the depth map to obtain a 3D reconstruction result.
[0031] Image acquisition devices 101, 102, and 103 may also be electronic devices with display screens, which can receive and display the depth map and 3D reconstruction results generated by server 105 via the network. For example, if image acquisition devices 101, 102, and 103 are AR glasses, they can perform 3D reconstruction based on monocular images on server 105, synthesize virtual information into the real scene, achieve virtual-real fusion, and finally present it to the user through the display screen of the AR glasses.
[0032] It should be noted that the large model-based depth map generation method and the three-dimensional reconstruction method provided in the embodiments of the present disclosure can generally be executed by the server 105. Accordingly, the large model-based depth map generation device and the three-dimensional reconstruction device provided in the embodiments of the present disclosure can generally be set in the server 105. The large model-based depth map generation method and the three-dimensional reconstruction method provided in the embodiments of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the image acquisition devices 101, 102, 103 and / or the server 105. Accordingly, the large model-based depth map generation device and the three-dimensional reconstruction device provided in the embodiments of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the image acquisition devices 101, 102, 103 and / or the server 105.
[0033] I understand. Figure 1 The number and type of image acquisition devices, networks, and servers in the embodiment are merely illustrative. Any number of image acquisition devices, networks, and servers may be used as needed.
[0034] It is understood that the system architecture of the present disclosure is described above, and the method of the present disclosure will be described below. It is also understood that the sequence numbers of the various operations in the following method are merely used to indicate the operation for the purpose of description and should not be regarded as indicating the order in which the various operations must be performed. Unless explicitly stated, the method does not need to be executed in the exact order shown.
[0035] In the technical solutions disclosed herein, all information and data involved (including but not limited to data used for analysis, stored data, displayed data, etc.) are authorized information and data, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, adopt necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for the choice of authorization or rejection.
[0036] Figure 2 4 is a flowchart of a method for generating a depth map based on a large model according to an embodiment of the present disclosure.
[0037] like Figure 2 As shown, the method 200 may include operations S210 to S240.
[0038] In operation S210 , visual encoding is performed on the monocular image to obtain an encoded image.
[0039] According to embodiments of the present disclosure, a monocular image may be information acquired at a single time and is a two-dimensional image that can provide two-dimensional information, such as width and height. Visual encoding of a monocular image can be understood as mapping or converting the two-dimensional information in the monocular image into a visual representation that can be processed by a computer or understood by humans, such as converting the object "dog" in the image into visual features that can be processed by a computer.
[0040] In operation S220 , the encoded image and the target text are input into a pre-trained large language model for fusion to obtain fusion features.
[0041] According to embodiments of the present disclosure, the target text can be a requirement text input based on an actual application scenario. For example, the target text description may be: Generate a depth map where the foreground object is close and the background is far away; or another example is: Generate a relative depth map of such an image. The present disclosure does not limit the specific content of the target text.
[0042] According to an embodiment of the present disclosure, the pre-trained large language model can be obtained by cross-modal training based on deep learning using text datasets and image datasets.
[0043] According to an embodiment of the present disclosure, fused features may include semantic association information obtained by capturing the linguistic consistency between the visual features of an image and the text features. For example, the semantic association between "a golden retriever running on the grass" in the image and the text "a dog running on the grass" is performed, and the visual features (such as the dog's shape and the color of the grass) and the text features (such as "dog", "grass", and "running") are aligned through a model to form a shared semantic space. Fusion features may also include fusing visual features and text features to form a joint representation containing information from both modalities. For example, the visual features of "dog" in the image and the word embedding of "dog" in the text are close in the feature space. Fusion features may also include contextual scene information. For example, "beach" in the image and "sunny afternoon" in the text jointly describe a scene, and fusion features will integrate this information.
[0044] In operation S230 , a global guiding feature is generated based on the fused features, where the global guiding feature includes joint semantic information of the visual feature and the text feature.
[0045] According to the embodiments of the present disclosure, the global guidance feature can provide a global perspective for the model, guide the task processing process, and improve the model's ability to understand complex inputs.
[0046] According to embodiments of the present disclosure, joint semantic information may include, for example, joint semantic information between objects and concepts, joint semantic information between actions and events, joint semantic information between attributes and descriptions, and joint semantic information between scenes and contexts. For example, joint semantic information between objects and concepts can be associated in feature space between an object in an image (e.g., "dog") and a concept in text (e.g., "pet" or "animal"), forming a cross-modal object-concept representation. For example, joint semantic information between actions and events can be associated with an action in an image (e.g., "running") and an event description in text (e.g., "dog running on the grass"), forming an action-event semantic representation. For example, joint semantic information between attributes and descriptions can be associated with attributes in an image (e.g., "red" or "round") and a description in text (e.g., "red ball"), forming an attribute-description semantic representation. For example, joint semantic information between scenes and contexts can be associated with a scene in an image (e.g., "beach") and a context in text (e.g., "sunny afternoon"), forming a scene-context semantic representation.
[0047] In operation S240 , noise is added to the color image of the monocular image to obtain a noise feature sequence.
[0048] According to an embodiment of the present disclosure, the color image is denoised in order to subsequently recover data from the noise without directly modeling the entire data distribution. The noise feature sequence contains both color image information and the added noise.
[0049] In operation S250 , the noise feature sequence is denoised based on the global guide feature to generate implicit features that match the joint semantic information.
[0050] According to the embodiments of the present disclosure, since the global guiding features can include the joint semantic information of body and concept, the joint semantic information of action and event, the joint semantic information of attribute and description, the joint semantic information of scene and context, etc., when processing the color image after adding noise, it can guide the extraction of features that are closely related to the text features contained in the target text, that is, the global guiding features can extract implicit features that match the target text.
[0051] In operation S260 , a depth map is generated based on the implicit features.
[0052] The large-model-based depth map generation method of the disclosed embodiment, combined with the reasoning capabilities of visual encoding and a pre-trained large language model, can more accurately capture the semantic relationship between text features and visual features. On this basis, the features obtained from the pre-trained large language model are used as initialization features for denoising and de-noising, which can reduce the number of denoising and de-noising steps and improve the efficiency and accuracy of monocular depth estimation. Global guidance features provide global context and semantic guidance for the denoising and de-noising processes, enhancing semantic understanding capabilities, thereby achieving more accurate and context-aware depth estimation and obtaining a more accurate depth map.
[0053] In an embodiment of the present disclosure, visually encoding a monocular image to obtain an encoded image may include:
[0054] The multiple image blocks obtained by dividing the monocular image are visually encoded respectively and the spatial information of the multiple image blocks is added to obtain a visual feature sequence.
[0055] A text feature space is constructed based on multiple text features of the target text, and the text features correspond to the dimensions of the text feature space.
[0056] The visual feature sequence is mapped to the text feature space to obtain multiple encoded images.
[0057] Figure 3 The present invention is a flowchart of visual encoding of a monocular image according to an embodiment of the present invention.
[0058] like Figure 3As shown, for example, a visual language model 310 can be used to visually encode a monocular image. The visual language model uses a multimodal pre-training model (Contrastive Language–Image Pretraining, CLIP). A monocular image (e.g., 224×224 pixels) is input into an image encoder 311, which divides the monocular image into M image patches. Each image patch can be, for example, a 14×14 grid of 16×16 pixels. A learnable positional encoding can be added to the M image patches to preserve spatial information, thereby generating a visual feature sequence.
[0059] The target text is input into the text encoder 312, and the text encoder 312 extracts multiple text features (tokens) from the target text. Text features can be words, phrases, topics, etc. The text feature space can be a high-dimensional vector space, where each dimension corresponds to a text feature. For example, if the target text is "The dog sat on the mat," the word segmentation results in {the, cat, sat, on, mat}, where each word corresponds to a dimension.
[0060] Image encoder 311 inputs the visual feature sequence into a multi-layer perceptron (MLP) 313. MLP 313 maps each visual feature in the visual feature sequence to a text feature space, generating M encoded images (tokens). For example, MLP 313 may have two layers.
[0061] Through the large-model-based depth map generation method of the embodiment of the present disclosure, by first dividing the monocular image into blocks and then encoding it and adding spatial information, a visual feature sequence that retains local and global information can be obtained, and then the visual features are mapped to the text feature space. Contrastive learning improves the model's generalization ability, so that the subsequent pre-trained large model can better understand the semantics of the image and text at the same time, improve processing accuracy and efficiency, and thereby improve the accuracy and generation efficiency of the depth map.
[0062] In an embodiment of the present disclosure, the encoded image and the target text are input into a pre-trained large model for fusion to obtain fusion features, which may include:
[0063] A first association relationship between a plurality of text words in a text word sequence, a second association relationship between coded images in a plurality of coded images, and a third association relationship between a plurality of text words and the coded images are obtained.
[0064] A single forward propagation is adopted to fuse multiple encoded images with text word sequences based on the first association relationship, the second association relationship and the third association relationship to obtain a fused feature.
[0065] According to an embodiment of the present disclosure, a text word sequence can be obtained by segmenting and encoding the target text.
[0066] Exemplarily, a pre-trained large language model can employ a transformer architecture with multiple decoders, each of which can include an attention mechanism and a feedforward neural network. The input to the pre-trained large language model can be the output of the visual language model (M images) and the target text. By segmenting and encoding the target text, a self-attention mechanism is used to capture the semantic or grammatical relationships between multiple text tokens to obtain a first association relationship. A self-attention mechanism is used to capture the spatial or semantic relationships between encoded images in multiple encoded images to obtain a second association relationship. A cross-attention mechanism is used to cause text tokens to focus on image features, or vice versa. A single forward pass of the feedforward neural network is then used to complete the calculation of all association relationships and feature fusion in a single forward pass to generate a fused feature.
[0067] Through the large model-based depth map generation method of the embodiment of the present disclosure, the pre-trained large language model can fuse visual features and text features at one time using a single forward propagation on the basis of multimodal fusion, which not only ensures the comprehensiveness of the fusion features, but also improves the fusion efficiency, thereby improving the accuracy and generation efficiency of the depth map.
[0068] In an embodiment of the present disclosure, generating a global guiding feature based on the fused feature may include:
[0069] A linear transformation is performed on the fused features to adjust the dimension of the fused features, and the visual features and text features in the fused features are mapped to the same feature space in order to learn the alignment relationship between the visual features and the text features and generate global guided features.
[0070] A connector can be used to linearly transform the fusion features output by the pre-trained large language model. The connector can include a linear layer, which can be used to transform the fusion features Projecting to a low-dimensional space reduces the dimensionality of the fused features and generates a compact global representation. A linear layer can be used to map the visual and textual features in the fused features to the same feature space. Based on contrastive learning, using contrastive loss, the similarity of aligned visual-text feature pairs is calculated, bringing aligned visual-text feature pairs closer together and pushing misaligned visual-text feature pairs further apart. This leads to an alignment between visual and textual features, resulting in a global guided feature.
[0071] The large-model-based depth map generation method of the disclosed embodiments adjusts dimensionality through linear transformation, reducing the computational effort of subsequent layers and improving model efficiency. Furthermore, linear transformation captures global information from both visual and textual data, avoiding the limitations of local information. By learning the alignment relationships between multiple modalities, the back-diffusion model possesses improved generalization capabilities, thereby improving depth map accuracy.
[0072] In an embodiment of the present disclosure, gradually adding noise to a color image of a monocular image to obtain a noise feature sequence may include:
[0073] Encode the color image and obtain the latent space features.
[0074] A plurality of target residual features are determined according to the initial residual features between the latent space features and the global guided features and a plurality of scaling factors.
[0075] Based on multiple target residual features and multiple random Gaussian noises, noise is added to the latent space features to generate a noise feature sequence.
[0076] According to embodiments of the present disclosure, noise addition and denoising can be implemented based on a diffusion model. This model uses a random process of gradually adding noise (forward diffusion) and denoising (reverse reconstruction) to a color image, leveraging a neural network to learn data distribution patterns, ultimately generating a high-quality depth map from random noise. The forward diffusion process can involve gradually applying noise to the original data (e.g., an image) to approximate a Gaussian distribution. The reverse diffusion process begins with a noisy state, gradually removes the noise, and reconstructs the original data to produce a depth map.
[0077] In the diffusion model, target and condition are the two core concepts that define the generated depth map. The global guided feature can be used as the condition of the diffusion model, and the latent space feature of the color image can be used as the target to generate implicit features.
[0078] According to embodiments of the present disclosure, noise can be added gradually. This can be understood as adding noise based on a time step. Noise is added to the latent space features at intervals of one time step. A time step is a unit of time in a discrete time series, representing a discrete point in the time dimension of the sequence data. During the noise addition process, the time step can be a parameter that controls the degree of noise addition, used to gradually convert the latent space features into pure noise.
[0079] According to an embodiment of the present disclosure, as the time step increases, the scaling factor gradually increases.
[0080] For example, the product of the initial residual features and multiple scaling factors can be used to determine multiple target residual features. As the scaling factors gradually increase, the noise added to the latent space features gradually increases, tending to pure noise. During the noise addition process, a hyperparameter that controls the noise variance can also be introduced to control the noise addition.
[0081] Through the large-model-based depth map generation method of the embodiment of the present disclosure, noise is added by gradually adjusting the residual based on the time step, which can guide the latent space features to change in a specific direction, that is, to approach the joint semantic information. In this way, the diffusion model can better utilize the joint semantic information and generate more qualified implicit features, thereby improving the quality and controllability of the generated depth map.
[0082] In an embodiment of the present disclosure, encoding a color image of a monocular image may include:
[0083] The color image is mapped to the latent space so that the features of the color image are converted into the distribution parameters of the latent space.
[0084] Feature reconstruction is performed based on distribution parameters to obtain latent space features.
[0085] According to an embodiment of the present disclosure, the latent space can be a low-dimensional potential representation space to which the deep generative model is mapped. The latent space captures the essential features of the data (such as shape, color, and semantics) and discards redundant information (such as noise and background details), so it can be used for subsequent denoising processing.
[0086] According to embodiments of the present disclosure, a variational autoencoder (VAE) can be used, for example, to map a color image to a latent space, generating parameters (mean and variance) of the latent space distribution. This distribution can be, for example, a normal distribution. Latent space features are then sampled and reconstructed from the latent space distribution. The color image can be, for example, an RGB image.
[0087] The large-model-based depth map generation method of the disclosed embodiment uses multimodal monocular depth estimation, which is based on the continuity and structuring of the latent space, and is conducive to semantically guided operations, thereby improving the accuracy of the depth map.
[0088] In an embodiment of the present disclosure, stepwise denoising of a noise feature sequence based on a global guiding feature may include:
[0089] The encoder in the encoder-decoder structure is used to gradually downsample the noise feature sequence to extract multi-scale features.
[0090] The decoder is used to gradually upsample the multi-scale features and gradually restore the spatial resolution.
[0091] Figure 4The present invention is a flowchart of gradually denoising a noise feature sequence according to an embodiment of the present disclosure.
[0092] like Figure 4 As shown, in an embodiment of the present disclosure, an encoder-decoder structure 400 can be used to gradually denoise the noise feature sequence. Multiple encoders 410 connected in series can constitute a downsampling path, and multiple decoders 430 connected in series can constitute an upsampling path. The encoder can be connected to the decoder of the corresponding layer through a skip connection. For example, the output of the first encoder 410 can be connected to the input of the last decoder 430, the output of the second encoder 410 can be connected to the input of the second-to-last decoder 430, and so on.
[0093] During the denoising process, the global guidance features can be input into the decoder 430 based on the skip connection to help the decoder maintain consistency with the global target when recovering spatial details. For example, during the upsampling process of the decoder 430, the global guidance features are combined with the low-level features of the encoder (transmitted through the skip connection) to ensure that the generated details (such as edges and textures) conform to the global semantics.
[0094] Since the deep features obtained by the encoder-decoder structure contain more global semantic information, the global guidance features can also be input into the intermediate layer 420 (deep features) of the encoder-decoder structure. Through cross-attention or feature concatenation, the global guidance features are fused with the intermediate features of the encoder-decoder structure to make the generated depth map more consistent with the text description.
[0095] The global guided features can also be input into each layer in the encoder-decoder structure together with the time step to dynamically adjust the denoising direction at different time steps, ensuring that the generated implicit features always meet the global objectives.
[0096] Through the large model-based depth map generation method of the embodiment of the present disclosure, denoising based on the encoder-decoder structure can fuse multi-scale features in the process of gradual denoising, and the introduction of global guided features can generate implicit features that are more in line with global goals in the denoising process, thereby improving the overall performance of the diffusion model and efficiently generating high-precision depth maps.
[0097] The following will be combined Figure 5 A three-dimensional reconstruction method according to an embodiment of the present disclosure is schematically described. Figure 5 is a flowchart of a 3D reconstruction method according to an embodiment of the present disclosure.
[0098] like Figure 5 As shown, the method 500 may include operations S510 to S530.
[0099] In operation S510 , a monocular image captured by an image capturing device is acquired.
[0100] In operation S520 , the depth map corresponding to the monocular image is converted into point cloud data based on the pose information of the image acquisition device.
[0101] In operation S530 , three-dimensional reconstruction is performed based on the point cloud data to obtain a three-dimensional reconstruction result.
[0102] According to embodiments of the present disclosure, an image acquisition device may be, for example, a smartphone, AR glasses, a camera, or a webcam. Pose information may be parameters related to the image acquisition device. For example, for a camera, pose information may include camera intrinsic parameters and camera extrinsic parameters. Camera intrinsic parameters may include focal length, principal point coordinates, etc., which are used to convert pixel coordinates into 3D points in the camera coordinate system; camera extrinsic parameters include the camera's rotation matrix and translation vector, which are used to convert points in the camera coordinate system into the world coordinate system.
[0103] Taking the camera as an example, since each pixel value in the depth map represents the distance from the point to the camera, for each pixel in the depth map, the pixel coordinates can be converted into a 3D point in the camera coordinate system based on the camera's focal length, the principal point coordinates, and the depth value of the pixel. Then, the 3D points corresponding to all pixels are combined to form point cloud data.
[0104] Based on the 3D reconstruction results obtained from the above 3D reconstruction, real-time road scene depth perception, environmental 3D reconstruction and obstacle avoidance, and virtual and real scene fusion can be performed.
[0105] For example, in real-time road depth perception scenarios, the 3D reconstruction method described above can analyze monocular images captured by a camera to determine the distance to obstacles ahead, enabling timely adjustments to driving speed or direction, enabling autonomous driving. Furthermore, by estimating the depth of vehicles or pedestrians ahead, the system can proactively make decisions to slow down or avoid them. In tunnels or urban areas with densely populated buildings, the 3D reconstruction method can supplement GPS signals for more accurate positioning.
[0106] For example, in the scenario of three-dimensional reconstruction and obstacle avoidance of the environment: in an unknown environment, the robot can use the images captured by the monocular camera to estimate the depth of the surrounding environment in real time, thereby planning a safe path, realizing robot navigation, and preventing the robot from colliding with obstacles.
[0107] For example, in the fusion of virtual and real scenes, in AR games, virtual characters or objects can be accurately placed on the ground or tabletop based on the depth information of the real scene. By estimating the scene's depth information, VR systems can generate more realistic 3D models, enhancing the user's sense of immersion. In photography post-processing, the above 3D reconstruction method can help photographers achieve a more natural background blur effect.
[0108] Figure 6 This is a flowchart of realizing autonomous driving based on a three-dimensional reconstruction method according to an embodiment of the present disclosure.
[0109] like Figure 6 As shown, the vehicle's camera captures a single image including the road and surrounding objects, and inputs the monocular image into the vehicle's controller, which can be integrated with a visual language model 610, a pre-trained large language model 620, a connector 630, a variational autoencoder 640, a diffusion model 650, and a variational autodecoder 660.
[0110] The visual language model 610 encodes the monocular image captured by the camera to produce an encoded image. The visual language model 610 inputs the encoded image into the pre-trained large language model 620. The pre-trained large language model 620 also receives the target text (to generate a depth map of road obstacles) and fuses the encoded image and target text to produce fused features. The pre-trained large oracle model 620 inputs the fused features into the connector 630, which generates global guiding features based on the fused features (for example, the alignment between the text "obstacle" and the image showing vehicles, animals, etc. in the middle of the road). The variational autoencoder 640 compresses and encodes the RGB image of the monocular image to produce latent space features. The connector 630 inputs the global guiding features into the diffusion model 650 as conditions. The variational autoencoder 640 inputs the latent space features into the diffusion model 650 as targets. The diffusion model 650 performs denoising and de-noising to produce latent features. The latent features are then input into the variational autodecoder 660, which generates a depth map based on the latent features. The controller performs three-dimensional reconstruction based on the depth map, which can determine the distance of obstacles ahead and adjust the driving speed or direction in time to achieve autonomous driving.
[0111] It should be understood that Figure 6 The various models shown are intended to more clearly illustrate the application of the three-dimensional reconstruction method in real-time road scene depth perception, environmental three-dimensional reconstruction and obstacle avoidance, virtual and real scene fusion, etc., and do not limit the present disclosure.
[0112] Since the 3D reconstruction method of the disclosed embodiment utilizes the depth map generated by the method for generating a depth map based on a large model described in the aforementioned embodiment, the specific implementation details can be found in the embodiment section on generating a depth map based on a large model and will not be repeated here. The depth map obtained in this manner enables high-precision 3D reconstruction, which in turn can better meet practical applications in areas such as autonomous driving, robotic navigation, and augmented reality / virtual reality.
[0113] The following will be combined Figure 7 A large model-based depth map generation device according to an embodiment of the present disclosure is schematically described. Figure 7 4 is a block diagram of a depth map generation apparatus based on a large model according to an embodiment of the present disclosure.
[0114] like Figure 7 As shown, the large model-based depth map generation device 700 may include a visual encoding module 710, a fusion module 720, a first generation module 730, a noise addition module 740, a denoising module 750 and a second generation module 760.
[0115] The visual encoding module 710 is configured to perform visual encoding on the monocular image to obtain an encoded image. In one embodiment, the visual encoding module 710 may be configured to perform the operation S210 described above, which will not be described in detail herein.
[0116] The fusion module 720 is used to input the encoded image and the target text into the pre-trained large language model for fusion to obtain fusion features. In one embodiment, the fusion module 720 can be used to perform the operation S220 described above, which will not be repeated here.
[0117] The first generating module 730 is configured to generate a global guiding feature based on the fused feature, wherein the global guiding feature includes the combined semantic information of the visual feature and the text feature. In one embodiment, the first generating module 730 may be configured to perform the operation S230 described above, which will not be described in detail here.
[0118] The noise adding module 740 is used to add noise to the color image of the monocular image to obtain a noise feature sequence. In one embodiment, the noise adding module 740 can be used to perform the operation S240 described above, which will not be repeated here.
[0119] The denoising module 750 is used to denoise the noise feature sequence based on the global guiding feature to generate implicit features that match the joint semantic information. In one embodiment, the denoising module 750 can be used to perform the operation S250 described above, which will not be repeated here.
[0120] The second generating module 760 is configured to generate a depth map based on the implicit features. In one embodiment, the second generating module 760 may be configured to perform the operation S260 described above, which will not be described in detail here.
[0121] According to an embodiment of the present disclosure, the visual encoding module 710 performs visual encoding on a monocular image to obtain an encoded image, which may include:
[0122] The multiple image blocks obtained by dividing a monocular image are visually encoded and their spatial information is added to generate a visual feature sequence. A text feature space is constructed based on multiple text features of the target text, where the text features correspond to the dimensions of the text feature space. The visual feature sequence is mapped to the text feature space to generate multiple encoded images.
[0123] According to an embodiment of the present disclosure, the fusion module 720 inputs the encoded image and the target text into a pre-trained large language model for fusion to obtain fusion features, which may include:
[0124] A first association relationship between multiple text grammars in a text grammar sequence, a second association relationship between encoded images in multiple encoded images, and a third association relationship between multiple text grammars and encoded images are obtained. The text grammar sequence is obtained by segmenting and encoding the target text. A single forward propagation is used to fuse the multiple encoded images with the text grammar sequence based on the first association relationship, the second association relationship, and the third association relationship to obtain a fused feature.
[0125] According to an embodiment of the present disclosure, the first generating module 730 generates a global guiding feature based on the fused feature, which may include:
[0126] A linear transformation is performed on the fused features to adjust the dimension of the fused features, and the visual features and text features in the fused features are mapped to the same feature space in order to learn the alignment relationship between the visual features and the text features and generate global guided features.
[0127] According to an embodiment of the present disclosure, the noise adding module 740 adds noise to the color image of the monocular image to obtain a noise feature sequence, which may include:
[0128] The color image is encoded to obtain latent space features. Multiple target residual features are determined based on the initial residual features between the latent space features and the global guided features and multiple scaling factors. Based on the multiple target residual features and multiple random Gaussian noises, noise is added to the latent space features to generate a noise feature sequence. The scaling factors gradually increase with increasing time steps.
[0129] According to an embodiment of the present disclosure, the denoising module 740 may encode the color image of the monocular image, which may include:
[0130] The color image is mapped to the latent space so that the features of the color image are converted into the distribution parameters of the latent space. The features are reconstructed based on the distribution parameters to obtain the latent space features.
[0131] According to an embodiment of the present disclosure, the denoising module 750 denoises the noise feature sequence based on the global guide feature, which may include:
[0132] The encoder in the encoder-decoder structure is used to gradually downsample the noise feature sequence to extract multi-scale features. The decoder is used to gradually upsample the multi-scale features to gradually restore the spatial resolution.
[0133] It should be noted that the details of other embodiments of the depth map generation device based on a large model and the technical effects brought about are the same or similar to the details of the embodiments of the depth map generation method based on a large model and the technical effects brought about, and will not be repeated here.
[0134] The following will be combined Figure 8 A three-dimensional reconstruction device according to an embodiment of the present disclosure is schematically described. Figure 8 FIG. 4 is a block diagram of a 3D reconstruction apparatus according to an embodiment of the present disclosure.
[0135] like Figure 8 As shown, the 3D reconstruction device 800 may include an acquisition module 810 , a conversion module 820 and a reconstruction module 830 .
[0136] The acquisition module 810 is configured to acquire a monocular image acquired by an image acquisition device. In one embodiment, the acquisition module 810 may be configured to execute the operation S510 described above, which will not be described in detail herein.
[0137] The conversion module 820 is used to convert the depth map corresponding to the monocular image into point cloud data based on the pose information of the image acquisition device. In one embodiment, the conversion module 820 can be used to perform the operation S5200 described above, which will not be repeated here.
[0138] The reconstruction module 830 is configured to perform 3D reconstruction based on the point cloud data to obtain a 3D reconstruction result. In one embodiment, the reconstruction module 830 may be configured to perform the operation S530 described above, which will not be described in detail herein.
[0139] According to an embodiment of the present disclosure, the depth map of the monocular image captured by the image acquisition device can be used to Figure 7 The depth map generation device 700 based on the large model is shown.
[0140] It should be noted that the details of other embodiments of the three-dimensional reconstruction device and the technical effects brought about are the same or similar to the details of the embodiments of the three-dimensional reconstruction method and the technical effects brought about, and will not be repeated here.
[0141] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0142] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0143] like Figure 9 As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. RAM 903 may also store various programs and data required for the operation of device 900. Computing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to bus 904.
[0144] Various components in the device 900 are connected to the I / O interface 905, including an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0145] The computing unit 901 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as the text processing method and / or the method for deploying a deep learning framework. For example, in some embodiments, the text processing method and / or the method for deploying a deep learning framework can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into RAM 903 and executed by computing unit 901, one or more steps of the text processing method and / or deep learning framework deployment method described above may be performed. Alternatively, in other embodiments, computing unit 901 may be configured to perform the text processing method and / or deep learning framework deployment method in any other suitable manner (e.g., via firmware).
[0146] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0147] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable large-model-based depth map generation device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0148] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM) or flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0149] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) display or a liquid crystal display (LCD)) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0150] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0151] Computer systems may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other.
[0152] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0153] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for generating a depth map based on a large model, comprising: Perform visual encoding on the monocular image to obtain an encoded image; Inputting the encoded image and the target text into a pre-trained large language model for fusion to obtain fusion features; generating a global guiding feature based on the fused feature, wherein the global guiding feature includes joint semantic information of visual features and text features; adding noise to the color image of the monocular image to obtain a noise feature sequence; denoising the noise feature sequence based on the global guide feature to generate implicit features that match the joint semantic information; and A depth map is generated based on the implicit features.
2. The method according to claim 1, wherein The visual encoding of the monocular image to obtain the encoded image includes: Performing visual encoding on each of the multiple image blocks obtained by dividing the monocular image and adding spatial information of each of the multiple image blocks to obtain a visual feature sequence; Constructing a text feature space based on a plurality of text features of the target text, wherein the text features correspond to dimensions of the text feature space; and The visual feature sequence is mapped to the text feature space to obtain a plurality of the encoded images.
3. The method according to claim 2, wherein: The step of inputting the encoded image and the target text into a pre-trained large language model for fusion to obtain fusion features includes: Obtaining a first association relationship between a plurality of text words in a text word sequence, a second association relationship between encoded images in a plurality of the encoded images, and a third association relationship between the plurality of text words and the encoded images, wherein the text word sequence is obtained by segmenting and encoding the target text; and A single forward propagation is adopted to fuse the plurality of encoded images with the text word sequence based on the first association relationship, the second association relationship, and the third association relationship to obtain the fused feature.
4. The method according to claim 1, wherein The generating of the global guiding feature based on the fusion feature includes: A linear transformation is performed on the fused features to adjust the dimension of the fused features, and the visual features and text features in the fused features are mapped to the same feature space so as to learn the alignment relationship between the visual features and the text features, thereby generating the global guiding features.
5. The method according to claim 1, wherein Adding noise to the color image of the monocular image to obtain a noise feature sequence includes: Encoding the color image to obtain latent space features; determining a plurality of target residual features based on an initial residual feature between the latent space feature and the global guided feature and a plurality of scaling factors; and Based on multiple target residual features and multiple random Gaussian noises, noise is added to the latent space features respectively to generate the noise feature sequence; Wherein, as the time step increases, the scaling factor gradually increases.
6. The method according to claim 5, wherein: The color image encoding of the monocular image includes: Mapping the color image to a latent space so as to convert features of the color image into distribution parameters of the latent space; and Feature reconstruction is performed based on the distribution parameters to obtain the latent space features.
7. The method according to claim 1, wherein denoising the noise feature sequence based on the global guiding feature comprises: The encoder in the encoder-decoder structure is used to gradually downsample the noise feature sequence to extract multi-scale features; A decoder is used to gradually upsample the multi-scale features to gradually restore the spatial resolution.
8. A three-dimensional reconstruction method, comprising: Acquiring a monocular image acquired by an image acquisition device; Based on the pose information of the image acquisition device, converting the depth map corresponding to the monocular image into point cloud data; Performing three-dimensional reconstruction based on the point cloud data to obtain a three-dimensional reconstruction result; Wherein, the depth map is generated according to the method according to any one of claims 1 to 7.
9. A depth map generation device based on a large model, comprising: A visual encoding module is used to visually encode a monocular image to obtain an encoded image; A fusion module, configured to input the encoded image and target text into a pre-trained large language model for fusion to obtain fusion features; A first generation module is used to generate a global guiding feature based on the fused feature, wherein the global guiding feature includes joint semantic information of the visual feature and the text feature; a noise adding module, configured to add noise to the color image of the monocular image to obtain a noise feature sequence; a denoising module, configured to denoise the noise feature sequence based on the global guiding feature, and generate an implicit feature that matches the joint semantic information; as well as The second generating module is configured to generate a depth map based on the implicit features.
10. A three-dimensional reconstruction device comprising: An acquisition module is used to acquire a monocular image sequence acquired by an image acquisition device; A conversion module, configured to convert a depth map corresponding to each monocular image in a monocular image sequence into point cloud data based on the pose information of the image acquisition device; A reconstruction module, configured to perform three-dimensional reconstruction based on the point cloud data to obtain a three-dimensional reconstruction result; Wherein, the depth map is generated using the apparatus described in claim 9.
11. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 8.
13. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.
Citation Information
Cited By
Monocular depth estimation method and device, electronic equipment and storage medium
CN121121768A
Image reconstruction method and device, storage medium and electronic equipment
CN121414590A
Robot, operation method thereof, operation device, storage medium, and program product
CN121424349A
FMRI video nerve decoding method based on visual perception and semantic consistency
CN121814971A
Text-to-touch signal controllable generation method and system based on perceptual decoupling
CN122219778A