Three-dimensional model generation method and device, equipment and storage medium

By generating multi-view image sequences through a symmetric causal three-dimensional network, the problems of high computing resource consumption and unstable generation results in the existing technology are solved, and efficient and stable three-dimensional model generation is achieved, which is suitable for real-time scenarios.

CN120707773AActive Publication Date: 2025-09-26BEIJING ZHIXIANG FUTURE TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510741286.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-26
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

Existing image-to-3D model generation methods rely on initial views, consume a lot of computing resources, and the generated results are multifaceted and geometrically distorted, making them difficult to apply to real-time scenarios.

Method used

Multi-view image sequences are generated through a symmetric causal 3D network, and temporal attention layers and 3D convolutional layers are used to gradually generate multi-view image sequences, and a real-time neural rendering method is used to reconstruct the 3D model.

Benefits of technology

It achieves efficient generation of high-quality three-dimensional models, gets rid of the high dependence on two-dimensional images, improves the geometric consistency and texture coherence between generated images, and is suitable for real-time scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707773A_ABST
    Figure CN120707773A_ABST
Patent Text Reader

Abstract

The invention provides a three-dimensional model generation method and device, equipment and a storage medium. The method comprises the steps that a two-dimensional image is acquired; generating a multi-view image sequence through a symmetric causal relationship three-dimensional network based on the two-dimensional image; wherein the symmetric causal relationship three-dimensional network comprises a time attention layer and a three-dimensional convolutional layer; and reconstructing a three-dimensional model based on the multi-view image sequence. According to the method, the multi-view image sequence of the two-dimensional image is generated through the symmetric causal relationship three-dimensional network, and the three-dimensional model is reconstructed based on the multi-view image sequence, so that the high dependence of the three-dimensional model on the two-dimensional image is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a three-dimensional model generation method, apparatus, device, and storage medium. Background Art

[0002] Existing methods for generating 3D models from images are primarily based on optimization. These methods do not rely on paired images and 3D data, but instead leverage large-scale pre-trained 2D diffusion models to introduce image priors during the optimization of 3D implicit representations, enabling the generation of 3D content from images.

[0003] For example, through a two-stage generation from coarse to fine, the first stage optimizes the neural radiation field to obtain coarse geometry, and the second stage uses a differentiable mesh representation to refine details and textures, and combines two-dimensional and three-dimensional diffusion priors to guide view supervision.

[0004] Although existing optimization-based methods have demonstrated strong representation capabilities to a certain extent, their optimization process is highly dependent on the initial view, consumes a lot of computing resources, and usually requires tens of minutes or even hours of inversion time. Summary of the Invention

[0005] In order to solve one of the above technical defects, the present application provides a three-dimensional model generation method, device, equipment, and storage medium.

[0006] In a first aspect, the present application provides a method for generating a three-dimensional model, the method comprising:

[0007] Acquire a two-dimensional image;

[0008] Based on a 2D image, a multi-view image sequence is generated through a symmetric causal 3D network; wherein the symmetric causal 3D network includes a temporal attention layer and a 3D convolutional layer;

[0009] Reconstruct 3D models based on multi-view image sequences.

[0010] Optionally, based on the two-dimensional image, a multi-view image sequence is generated by a symmetric causal three-dimensional network, including:

[0011] Based on 2D images, a multi-view image sequence is generated in multiple steps from near to far through a symmetric causal 3D network.

[0012] The last step generates a frame of view image, and each of the remaining steps simultaneously generates a pair of two frame of view image pairs that are left-right symmetrical about the central view; the viewing angles of each frame of view image in the multi-view image sequence are different.

[0013] Optionally, the attention mechanism of the temporal attention layer is

[0014] Where Q, K, and V are different vectors obtained according to (H×W)×N×C, H is the height of the two-dimensional image, W is the width of the two-dimensional image, C is the number of channels of the two-dimensional image, and N is the number of viewing angles; d is the scaling factor, T is the transpose operator, Softmax[·] is the normalization function, and M is the causal mask.

[0015] Optionally, M is an N×N matrix,

[0016] Among them, a is the row identifier, b is the column identifier, and m (a,b) is the value of the element in row a and column b in M.

[0017] Optionally, the filling strategy of the 3D convolutional layer is: fill in the front of the time dimension (k t -1) frames;

[0018] Among them, k t is the convolution kernel size of the time dimension of the 3D convolutional layer.

[0019] Optionally, reconstructing a three-dimensional model based on the multi-view image sequence includes:

[0020] A real-time neural rendering method is used to recover implicit representations from multi-view image sequences and camera parameters.

[0021] Reconstruct a 3D model from the implicit representation.

[0022] Optionally, the loss function used in training the symmetric causal three-dimensional network is:

[0023]

[0024] in, is the expected function, z0 is the potential space encoding of the training sample two-dimensional image, x is the set of training sample two-dimensional images, e is the pitch angle of the camera, ∈ is random noise, t is the time step identifier, ∈ θ (z t ; t,x,e) is the prediction noise, z t is the encoding obtained after adding noise to z0 at time step t.

[0025] In a second aspect of the present application, a three-dimensional model generation device is provided, the device comprising:

[0026] An acquisition module, used for acquiring a two-dimensional image;

[0027] A generation module is configured to generate a multi-view image sequence based on the two-dimensional image acquired by the acquisition module through a symmetric causal three-dimensional network; wherein the symmetric causal three-dimensional network includes a temporal attention layer and a three-dimensional convolution layer;

[0028] The reconstruction module is used to reconstruct a three-dimensional model based on the multi-view image sequence generated by the generation module.

[0029] In a third aspect of the present application, an electronic device is provided, comprising:

[0030] Memory;

[0031] processor; and

[0032] computer programs;

[0033] The computer program is stored in the memory and is configured to be executed by the processor to implement the method as described in the first aspect above.

[0034] In a fourth aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored; the computer program is executed by a processor to implement the method described in the first aspect above.

[0035] This application provides a method, apparatus, device, and storage medium for generating a three-dimensional model. The method comprises: acquiring a two-dimensional image; generating a multi-view image sequence based on the two-dimensional image using a symmetric causal three-dimensional network; wherein the symmetric causal three-dimensional network includes a temporal attention layer and a three-dimensional convolution layer; and reconstructing a three-dimensional model based on the multi-view image sequence. This application generates a multi-view image sequence of the two-dimensional image using a symmetric causal three-dimensional network, and then reconstructs the three-dimensional model based on the multi-view image sequence, eliminating the three-dimensional model's high dependence on the two-dimensional image. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0037] Figure 1 A schematic diagram of a process for generating a three-dimensional model provided in an embodiment of the present application;

[0038] Figure 2 A schematic diagram for generating an existing multi-view image sequence;

[0039] Figure 3 A schematic diagram of generating a multi-view image sequence provided in an embodiment of the present application;

[0040] Figure 4A schematic diagram of a symmetrical causal three-dimensional network provided in an embodiment of the present application;

[0041] Figure 5 is a schematic diagram of the existing mask;

[0042] Figure 6 A schematic diagram of a causal mask provided in an embodiment of the present application;

[0043] Figure 7 A schematic diagram of the symmetric causal relationship three-dimensional network training process provided in an embodiment of the present application;

[0044] Figure 8 A schematic diagram illustrating the implementation principle of the three-dimensional model generation method provided in an embodiment of the present application;

[0045] Figure 9 A schematic diagram comparing the effects of a three-dimensional model generation method provided in an embodiment of the present application and an existing method in terms of perspective production;

[0046] Figure 10 A schematic diagram comparing the effects of a three-dimensional model generation method provided in an embodiment of the present application and an existing method on three-dimensional models;

[0047] Figure 11 A schematic diagram of the structure of a three-dimensional model generation device provided in an embodiment of the present application;

[0048] Figure 12 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0049] In order to make the technical solutions and advantages of the embodiments of the present application more clearly understood, the exemplary embodiments of the present application are further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, and are not an exhaustive list of all the embodiments. It should be noted that the embodiments and features in the embodiments of the present application can be combined with each other unless they conflict.

[0050] During the development of this application, the inventors discovered that existing methods for converting images into 3D models are primarily based on optimization. These methods do not rely on paired images and 3D data, but instead utilize large-scale pre-trained 2D diffusion models to introduce image priors during the optimization of 3D implicit representations, enabling the generation of 3D content from images.

[0051] Although existing optimization-based methods have demonstrated strong representation capabilities to a certain extent, their optimization process is highly dependent on the initial view and consumes a lot of computing resources. The inversion time is usually tens of minutes or even hours, and the generated results often have multi-faceted and geometric distortion problems, making them difficult to apply to real-time scenarios.

[0052] To address the above issues, embodiments of the present application provide a method, apparatus, device, and storage medium for generating a three-dimensional model. The method comprises: acquiring a two-dimensional image; generating a multi-view image sequence based on the two-dimensional image using a symmetric causal three-dimensional network; wherein the symmetric causal three-dimensional network includes a temporal attention layer and a three-dimensional convolution layer; and reconstructing a three-dimensional model based on the multi-view image sequence. The present application generates a multi-view image sequence of two-dimensional images using a symmetric causal three-dimensional network, and then reconstructs a three-dimensional model based on the multi-view image sequence, eliminating the three-dimensional model's high dependence on the two-dimensional image.

[0053] See also Figure 1 This embodiment provides a three-dimensional model generation method, and the implementation process of the method is as follows:

[0054] 101, acquire a two-dimensional image.

[0055] This step can be implemented using existing solutions, such as obtaining a two-dimensional image through an input interface, which will not be explained in detail here.

[0056] This 2D image is used to generate a 3D model.

[0057] 102, based on 2D images, generating multi-view image sequences through a symmetric causal 3D network.

[0058] In step 102 , a multi-view image sequence (ie, multi-view) may be generated based on the two-dimensional image (ie, single view) acquired in step 101 .

[0059] The existing method will generate them in order from left to right, such as Figure 2 As shown, Figure 2 The input is a two-dimensional image. If the multi-view image sequence includes images from 16 perspectives, then 15 steps are used to generate images 01 to 15, a total of 15 frames of images from different perspectives. Together with the input two-dimensional image (i.e., input), a total of 16 perspective images are generated, thus forming a multi-view image sequence from 16 perspectives. However, for a 360° object observation sequence, this approach has two problems: (1) The farthest back view (the most difficult to predict) is in the middle of the sequence, which may lead to error concentration and fail to ensure structural consistency between long-distance views. (2) The naturally existing left-right symmetric structure in the sequence is not utilized.

[0060] Unlike existing methods, step 102 employs a "next symmetric view prediction" strategy to generate a multi-view image sequence (i.e., multiple views) from the two-dimensional image (i.e., single view) acquired in step 101 using an autoregressive approach. Specifically, in step 102, the multi-view image sequence is generated in multiple steps, starting from the two-dimensional image and proceeding from near to far, using a symmetric causal three-dimensional network. The final step generates a single view image, while each of the remaining steps simultaneously generates a pair of two view images that are symmetrical about the central view. Each view image in the multi-view image sequence has a different perspective.

[0061] That is, in step 102, each step generates a pair of images that are symmetrical about the central view (e.g., in any step u, the generated image pair is (p u ,p N-u )), and proceed in the order of "from near to far", such as Figure 3 shown.

[0062] The multi-view image sequence has N frames of images with different viewing angles, which are Step 1 generates all multi-view image sequences, i.e. exist Each step will generate two symmetrical frames of the central view, corresponding to the u-th frame image and the Nu-th frame image in the multi-view image sequence. Step 1 generates a frame of view image, i.e. Frame image, such as Figure 3 shown.

[0063] exist Figure 3 The input is a two-dimensional image, N = 16. Then, in step 1, images 01 and 15 are generated, which are bilaterally symmetrical in the central view. In step 2, images 02 and 14 are generated, which are bilaterally symmetrical in the central view. In step 3, images 03 and 13 are generated, which are bilaterally symmetrical in the central view. In step 4, images 04 and 12 are generated, which are bilaterally symmetrical in the central view. In step 5, images 05 and 11 are generated, which are bilaterally symmetrical in the central view. In step 6, images 06 and 10 are generated, which are bilaterally symmetrical in the central view. In step 7, images 07 and 09 are generated, which are bilaterally symmetrical in the central view. In step 8, a frame of image 08 is generated. So far, 15 frames of images from different perspectives have been generated through 8 steps. Together with the input two-dimensional image (i.e., input), a total of 16 perspectives are generated, thus forming a multi-view image sequence of 16 perspectives.

[0064] The “next symmetric view prediction” strategy of step 102 is implemented by a SymmetricCausal 3D U-Net, which includes a temporal attention layer and a 3D convolutional layer.

[0065] In specific implementation, the Symmetric Causal 3D U-Net can be Figure 4 As shown, it includes a spatial layer ( Figure 4 spatiallayer) and temporal layer ( Figure 4 In the temporal layer ( Figure 4 The temporal layer in Figure 4 The temporal attention layer in ) and the 3D convolutional layer ( Figure 4 Conv3D in

[15] .

[0066] 1. Temporal Attention Layer ( Figure 4 temporal attention layer in

[0067] Temporal Attention Layer ( Figure 4 The attention mechanism of temporal attention layer in is

[0068] Where Q, K, and V are vectors derived from the formula (H × W) × B × C, H is the height of the 2D image, W is the width of the 2D image, C is the number of channels of the 2D image, and N is the number of viewpoints (in a specific implementation, N is also the number of frames in a multi-view image sequence). d is a scaling factor, T is the transpose operator, Softmax[·] is the normalization function, and M is the causal mask.

[0069] That is, the temporal attention layer ( Figure 4 The attention mechanism of the temporal attention layer in the image is the causal attention mechanism, and the core of the causal attention mechanism is the causal mask. That is, for a feature map of size N×H×W×C If the spatial dimension is expanded into a batch, the feature sequence of the multi-view image sequence of (H×W)×N×C is obtained, and its Q, K, V vectors are calculated respectively. After we introduce the causal mask M, the causal attention mechanism is

[0070] Existing masks such as Figure 5 As shown, if a and b in the mask index different frames in the multi-view image sequence respectively. For each frame a, based on Figure 5The existing mask attention mechanism shown in shields its future frames (i.e., b>a) in the attention weight, thereby ensuring that each frame can only focus on its past frames and "cannot see" future frames (i.e., ensuring that the a-th frame can only rely on the previous a frames and will not leak future information). Figure 5 Different from the existing mask shown in FIG, the causal mask M used in this embodiment is a symmetric causal mask, such as Figure 6 As shown, M is an N×N matrix,

[0071] Among them, a is the row identifier, b is the column identifier, and m (a,b) is the value of the element in row a and column b in M.

[0072] Right now Figure 6 The element value of the white background is -∞, and the element value of the gray background is 0.

[0073] The causal mask M used in this embodiment can ensure that each frame can access not only its historical frames but also its left-right symmetric frames, while shielding distant future frames, thereby ensuring the generation of the next symmetric view.

[0074] 2. 3D convolutional layer ( Figure 4 Conv3D in

[0075] 3D convolutional layer ( Figure 4 The padding strategy (i.e., padding strategy) of Conv3D in the time dimension is: fill in the front of the time dimension (k t -1) frames.

[0076] Among them, k t is the convolution kernel size of the time dimension of the 3D convolutional layer.

[0077] If the 3D convolutional layer ( Figure 4 The convolution kernel size of Conv3D in is (k t ,k h ,k w ), where k t The convolution kernel size in the time dimension, k h is the high-dimensional convolution kernel size, k w is the convolution kernel size in the wide dimension.

[0078] If the existing front and back symmetrical filling is used This method will cause each frame to depend on both past and future frames (i.e. the current frame sees the next frame), and it is impossible to generate autoregressive images that only rely on past frames. Therefore, the three-dimensional convolution layer ( Figure 4 Conv3D in the time dimension) by padding in front of the time dimension (k t-1) frames, so that each output frame depends only on its current and past information, preventing future information from interfering with the generation of the current frame, and realizing a "symmetric accessibility" convolutional architecture. This ensures that the output frame depends only on current and past information, preventing future information from interfering with the generation of the current frame, and achieving autoregressive generation.

[0079] This completes the description of the Symmetric Causal 3D U-Net structure.

[0080] In a specific implementation, the Symmetric Causal 3DU-Net used in step 102 is a trained Symmetric Causal 3DU-Net. The loss function used in training the Symmetric Causal 3D U-Net is:

[0081]

[0082] in, is the expected function, z0 is the potential space encoding of the training sample two-dimensional image, x is the set of training sample two-dimensional images, e is the pitch angle of the camera, ∈ is random noise, t is the time step identifier, ∈ θ (z t ; t,x,e) is the prediction noise, z t is the encoding obtained after adding noise to z0 at time step t.

[0083] For example, the training process is as follows Figure 7 As shown:

[0084] 1) Obtain the training sample two-dimensional image x and the corresponding multi-view image sequence.

[0085] 2) The training sample two-dimensional image x is encoded into the latent space through the Variational Autoencoder (VAE) to obtain the latent code z0.

[0086] 3) Apply different degrees of Gaussian noise to z0 through multiple time steps. For example, after adding noise ∈ at time step t, the noisy latent code z is obtained. t .

[0087] This step enables the Symmetric Causal 3D U-Net to make correct predictions in a mixed state where some frames are denoised and some still have noise.

[0088] 4) Compare z0 with the noisy latent code zt The concatenated features are then stitched together and used as input to the SymmetricCausal 3D U-Net. The concatenated CLIP (Contrastive Language-Image Pre-Training) feature embedding is introduced into the SymmetricCausal 3D U-Net and fused using the attention mechanism.

[0089] In each Transformer module, the CLIP feature matrix serves as the Q, K, and V of the attention mechanism. This processing can effectively transfer the high-level semantic information of the training sample 2D image x to the denoising Symmetric Causal 3D U-Net.

[0090] Considering that different batches of multi-view image sequences may have different elevations, the camera’s elevation angle e is input into the Symmetric Causal3D U-Net as additional conditional information. For example, the camera’s elevation angle e is first embedded using sinusoidal position encoding, and then input together with the diffusion time step t into the prediction noise ∈ θ middle.

[0091] The final training goal is:

[0092] After the above training process, the symmetric causal 3D U-Net based on the above structure will obtain a trained symmetric causal 3D U-Net. The trained symmetric causal 3D U-Net can generate a multi-view image sequence in an autoregressive manner. Therefore, in step 102, a multi-view image sequence is generated based on the two-dimensional image through the trained symmetric causal 3D U-Net. Figure 8 As shown, the two-dimensional image is Figure 8 The input image in is latently encoded and predicted by the trained SymmetricCausal 3D U-Net according to the "symmetric generation order". Two left-right symmetrical views are generated at each step until a complete 16-view image sequence is synthesized.

[0093] in, Figure 8 Where t is the time step, t∈{1,2,…,T}, and T is the maximum time step. Time step t represents the number of diffusion steps uniformly sampled from the diffusion time axis {1,2,…,T} and is used to construct the noisy samples during training. It controls the degree of noise in the latent variable. U(T) is the process of sampling from the set of diffusion steps according to a uniform distribution, denoted as t~U(T). It is used during training to select samples from different diffusion stages to improve the robustness of the model.

[0094] 103, reconstructing a three-dimensional model based on a multi-view image sequence.

[0095] In step 103, a real-time neural rendering method can be used to recover an implicit representation from the multi-view image sequence and camera parameters, and to reconstruct a 3D model based on the implicit representation. For example, a real-time neural rendering method can be used to recover a differentiable implicit representation from the complete multi-view image sequence and its corresponding camera parameters, and a high-quality 3D model can be reconstructed based on the implicit representation.

[0096] In specific implementation, step 103 can be implemented in an existing way, such as through the open source framework instant-nsr-pl, which is built on Instant-NGP and introduces SDF (Signed Distance Function) field representation to achieve an efficient and differentiable 3D reconstruction process. Figure 8 As shown in the SDF-based Reconstruction in step 102, the multi-view image sequence (ie Figure 8 The generated multi-view images (generated multi-view images) and their corresponding camera extrinsics (poses) are fed into the instant-nsr-pl framework. This framework optimizes an implicit function (SDF) represented by a sparse voxel hash structure to ensure that the synthesized image is as consistent as possible with the input image under multi-view supervision. Differentiable rendering techniques can also be used in this process to backpropagate image reconstruction errors and further adjust the implicit function parameters. Finally, a triangular mesh is extracted from the optimized SDF representation and combined with color information to generate a textured 3D model.

[0097] The above methods are not only efficient, but also perform well in reconstructing the geometric structure and details of objects.

[0098] A schematic diagram comparing the effects of a three-dimensional model generation method provided in this embodiment and an existing method in terms of perspective production is shown in FIG. Figure 9 As shown, Figure 9 In , the input image is a two-dimensional image.

[0099] A schematic diagram comparing the effects of a 3D model generation method provided in this embodiment and an existing method in terms of 3D models is shown in FIG. Figure 10 As shown, Figure 10 In , the input image is a two-dimensional image.

[0100] A large number of experimental results show that the three-dimensional model generation method provided by this embodiment has the advantages of view synthesis and single Figure 3 In the three-dimensional reconstruction tasks, it achieved better performance than the existing methods, verifying its advancedness and practical value in the field of image to three-dimensional model generation.

[0101] The three-dimensional model generation method provided in this embodiment can achieve high-quality single-image three-dimensional reconstruction. Its core idea is to model multi-view generation as a "next symmetric view prediction" task. This approach gets rid of the problem that existing methods rely on input images to generate all perspective images in parallel.

[0102] The Symmetric Causal 3D U-Net used in this example is causal, thus supporting frame-by-frame autoregressive modeling. Its temporal attention layer and 3D convolutional layer capture the relationship between frames in the temporal dimension, performing autoregression based solely on past frames to generate multi-view image sequences.

[0103] The 3D model generation method provided in this embodiment is a novel image-to-3D generation solution. It is based on the autoregressive modeling paradigm of "next-symmetric view prediction" and can achieve next-symmetric perspective prediction. This method gradually generates a 360° orbital view sequence around the input image, effectively improving the geometric consistency and texture coherence between the generated images.

[0104] In the "next symmetric view prediction" strategy, each step of autoregression generates a pair of left and right symmetric perspectives, which are then expanded layer by layer from near to far, effectively alleviating the problems of distortion and cumulative error in far-angle perspectives.

[0105] The three-dimensional model generation method provided in this embodiment complies with the natural symmetry properties of multi-view sequences, enabling the model to better capture geometric structures and enhance long-distance consistency.

[0106] This embodiment provides a 3D model generation method that obtains a 2D image; generates a multi-view image sequence based on the 2D image using a symmetric causal 3D network; the symmetric causal 3D network includes a temporal attention layer and a 3D convolution layer; and reconstructs a 3D model based on the multi-view image sequence. This method generates a multi-view image sequence of the 2D image using the symmetric causal 3D network, and then reconstructs the 3D model based on the multi-view image sequence, eliminating the 3D model's high dependence on the 2D image.

[0107] Based on the same inventive concept of the three-dimensional model generation method, this embodiment provides a three-dimensional model generation device, see Figure 11 , the device comprises:

[0108] The acquisition module 1101 is used to acquire a two-dimensional image.

[0109] The generating module 1102 is configured to generate a multi-view image sequence through a symmetric causal three-dimensional network based on the two-dimensional image acquired by the acquiring module 1101. The symmetric causal three-dimensional network includes a temporal attention layer and a three-dimensional convolution layer.

[0110] The reconstruction module 1103 is configured to reconstruct a three-dimensional model based on the multi-view image sequence generated by the generation module 1102 .

[0111] The generating module 1102 is configured to generate a multi-view image sequence in multiple steps in a near-to-far order based on a two-dimensional image through a symmetrical causal three-dimensional network.

[0112] The last step generates a frame of view image, and each of the remaining steps simultaneously generates a pair of two frame view images that are symmetrical about the central view. The viewing angles of each frame view image in the multi-view image sequence are different.

[0113] Among them, the attention mechanism of the temporal attention layer is

[0114] Where Q, K, and V are vectors derived from the (H×W)×N×C matrix, H is the height of the 2D image, W is the width of the 2D image, C is the number of channels, and N is the number of viewpoints. d is the scaling factor, T is the transpose operator, Softmax[·] is the normalization function, and M is the causal mask.

[0115] Where M is an N×N matrix,

[0116] Among them, a is the row identifier, b is the column identifier, and m (a,b) is the value of the element in row a and column b in M.

[0117] Among them, the filling strategy of the 3D convolutional layer is: fill in the front of the time dimension (k t -1) frames.

[0118] Among them, k t is the convolution kernel size of the time dimension of the 3D convolutional layer.

[0119] The reconstruction module 1103 is configured to recover the implicit representation from the multi-view image sequence and camera parameters using a real-time neural rendering method, and reconstruct the 3D model based on the implicit representation.

[0120] Among them, the loss function used in the training of the symmetric causal three-dimensional network is:

[0121]

[0122] in, is the expected function, z0 is the potential space encoding of the training sample two-dimensional image, x is the set of training sample two-dimensional images, e is the pitch angle of the camera, ∈ is random noise, t is the time step identifier, ∈ θ (z t ; t,x,e) is the prediction noise, z t is the encoding obtained after adding noise to z0 at time step t.

[0123] The device provided in this embodiment generates a multi-view image sequence of two-dimensional images through a symmetric causal three-dimensional network, and then reconstructs a three-dimensional model based on the multi-view image sequence, thus eliminating the three-dimensional model's high dependence on two-dimensional images.

[0124] Based on the same inventive concept of the three-dimensional model generation method, this embodiment provides an electronic device, such as Figure 12 As shown, it includes: a memory 1201, a processor 1202, and a computer program.

[0125] The computer program is stored in the memory 1201 and is configured to be executed by the processor 1202 to implement the above-mentioned three-dimensional model generation method.

[0126] Specifically,

[0127] Acquire a 2D image.

[0128] Based on a 2D image, a multi-view image sequence is generated through a symmetric causal 3D network, which consists of a temporal attention layer and a 3D convolutional layer.

[0129] Reconstruct 3D models based on multi-view image sequences.

[0130] Among them, based on the two-dimensional image, a multi-view image sequence is generated through a symmetric causal three-dimensional network, including:

[0131] Based on two-dimensional images, a multi-view image sequence is generated in multiple steps in the order from near to far through a symmetric causal three-dimensional network.

[0132] The last step generates a frame of view image, and each of the remaining steps simultaneously generates a pair of two frame view images that are symmetrical about the central view. The viewing angles of each frame view image in the multi-view image sequence are different.

[0133] Among them, the attention mechanism of the temporal attention layer is

[0134] Where Q, K, and V are vectors derived from the (H×W)×N×C matrix, h is the height of the 2D image, W is the width of the 2D image, C is the number of channels, and N is the number of viewpoints. d is the scaling factor, T is the transpose operator, Softmax[·] is the normalization function, and M is the causal mask.

[0135] Where M is an N×N matrix,

[0136] Among them, a is the row identifier, b is the column identifier, and m (a,b) is the value of the element in row a and column b in M.

[0137] Among them, the filling strategy of the 3D convolutional layer is: fill in the front of the time dimension (k t -1) frames.

[0138] Among them, k t is the convolution kernel size of the time dimension of the 3D convolutional layer.

[0139] The three-dimensional model is reconstructed based on the multi-view image sequence, including:

[0140] A real-time neural rendering method is used to recover implicit representations from multi-view image sequences and camera parameters.

[0141] Reconstruct a 3D model from the implicit representation.

[0142] Among them, the loss function used in the training of the symmetric causal three-dimensional network is:

[0143]

[0144] in, is the expected function, z0 is the potential space encoding of the training sample two-dimensional image, x is the set of training sample two-dimensional images, e is the pitch angle of the camera, ∈ is random noise, t is the time step identifier, ∈ θ (z t ; t,x,e) is the prediction noise, z t is the encoding obtained after adding noise to z0 at time step t.

[0145] The electronic device provided by this embodiment has a computer program executed by a processor to generate a multi-view image sequence of two-dimensional images through a symmetric causal three-dimensional network, and then reconstruct a three-dimensional model based on the multi-view image sequence, thereby eliminating the three-dimensional model's high dependence on two-dimensional images.

[0146] Based on the same inventive concept of the three-dimensional model generation method, this embodiment provides a computer-readable storage medium having a computer program stored thereon. The computer program is executed by a processor to implement the three-dimensional model generation method.

[0147] Specifically,

[0148] Acquire a 2D image.

[0149] Based on a 2D image, a multi-view image sequence is generated through a symmetric causal 3D network, which consists of a temporal attention layer and a 3D convolutional layer.

[0150] Reconstruct 3D models based on multi-view image sequences.

[0151] Among them, based on the two-dimensional image, a multi-view image sequence is generated through a symmetric causal three-dimensional network, including:

[0152] Based on two-dimensional images, a multi-view image sequence is generated in multiple steps in the order from near to far through a symmetric causal three-dimensional network.

[0153] The last step generates a frame of view image, and each of the remaining steps simultaneously generates a pair of two frame view images that are symmetrical about the central view. The viewing angles of each frame view image in the multi-view image sequence are different.

[0154] Among them, the attention mechanism of the temporal attention layer is

[0155] Where Q, K, and V are vectors derived from the (H×W)×N×C matrix, h is the height of the 2D image, W is the width of the 2D image, C is the number of channels, and N is the number of viewpoints. d is the scaling factor, T is the transpose operator, Softmax[·] is the normalization function, and M is the causal mask.

[0156] Where M is an N×N matrix,

[0157] Among them, a is the row identifier, b is the column identifier, and m (a,b) is the value of the element in row a and column b in M.

[0158] Among them, the filling strategy of the 3D convolutional layer is: fill in the front of the time dimension (k t -1) frames.

[0159] Among them, k t is the convolution kernel size of the time dimension of the 3D convolutional layer.

[0160] The three-dimensional model is reconstructed based on the multi-view image sequence, including:

[0161] A real-time neural rendering method is used to recover implicit representations from multi-view image sequences and camera parameters.

[0162] Reconstruct a 3D model from the implicit representation.

[0163] Among them, the loss function used in the training of the symmetric causal three-dimensional network is:

[0164]

[0165] in, is the expected function, z0 is the potential space encoding of the training sample two-dimensional image, x is the set of training sample two-dimensional images, e is the pitch angle of the camera, ∈ is random noise, t is the time step identifier, ∈ θ (z t ; t,x,e) is the prediction noise, z t is the encoding obtained after adding noise to z0 at time step t.

[0166] The computer-readable storage medium provided in this embodiment has a computer program on which a processor executes to generate a multi-view image sequence of two-dimensional images through a symmetric causal three-dimensional network, and then reconstruct a three-dimensional model based on the multi-view image sequence, thereby eliminating the three-dimensional model's high dependence on two-dimensional images.

[0167] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The scheme in the embodiment of the present application can be implemented in various computer languages, for example, object-oriented programming language Java and literal translation scripting language JavaScript, etc.

[0168] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0169] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0170] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0171] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0172] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A three-dimensional model generation method, characterized in that: The method comprises: Acquire a two-dimensional image; Based on the two-dimensional image, a multi-view image sequence is generated through a symmetric causal three-dimensional network; wherein the symmetric causal three-dimensional network includes a temporal attention layer and a three-dimensional convolution layer; A three-dimensional model is reconstructed based on the multi-view image sequence.

2. The method according to claim 1, characterized in that The method of generating a multi-view image sequence based on the two-dimensional image through a symmetric causal three-dimensional network includes: Based on the two-dimensional image, a multi-view image sequence is generated in multiple steps in a near-to-far order through a symmetric causal three-dimensional network; The last step generates a frame of view image, and each of the remaining steps simultaneously generates a pair of two frame of view image pairs that are left-right symmetrical about the central view; the viewing angles of each frame of view image in the multi-view image sequence are different.

3. The method according to claim 1, characterized in that The attention mechanism of the temporal attention layer is Where Q, K, and V are different vectors obtained according to (H×W)×N×C, H is the height of the two-dimensional image, W is the width of the two-dimensional image, C is the number of channels of the two-dimensional image, and N is the number of viewing angles; d is the scaling factor, T is the transpose operator, Softmax[·] is the normalization function, and M is the causal mask.

4. The method according to claim 3, characterized in that The M is an N×N matrix, Among them, a is the row identifier, b is the column identifier, and m (a,b) is the value of the element in row a and column b in M.

5. The method according to claim 1, wherein The filling strategy of the three-dimensional convolutional layer is: fill in the front of the time dimension (k t -1) frames; Among them, k t is the convolution kernel size of the time dimension of the 3D convolutional layer.

6. The method according to claim 1, characterized in that The reconstructing of a three-dimensional model based on the multi-view image sequence comprises: Recovering implicit representations from the multi-view image sequence and camera parameters using a real-time neural rendering method; A three-dimensional model is reconstructed according to the implicit representation.

7. The method according to claim 1, characterized in that The loss function used in the training of the symmetric causal three-dimensional network is: in, is the expected function, z0 is the potential space encoding of the training sample two-dimensional image, x is the set of training sample two-dimensional images, e is the pitch angle of the camera, ∈ is random noise, t is the time step identifier, ∈ θ (z t ; t,x,e) is the prediction noise, z t is the encoding obtained after adding noise to z0 at time step t.

8. A three-dimensional model generating device, characterized in that: The device comprises: An acquisition module, used for acquiring a two-dimensional image; A generation module, configured to generate a multi-view image sequence based on the two-dimensional image acquired by the acquisition module through a symmetric causal three-dimensional network; wherein the symmetric causal three-dimensional network includes a temporal attention layer and a three-dimensional convolution layer; The reconstruction module is used to reconstruct a three-dimensional model based on the multi-view image sequence generated by the generation module.

9. An electronic device, characterized in that: include: Memory; processor; as well as computer programs; The computer program is stored in the memory and configured to be executed by the processor to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that A computer program is stored thereon; the computer program is executed by a processor to implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Action prediction method based on human skeleton sequence

    CN114582024A

  • Three-dimensional reconstruction method based on symmetric fusion and attention mechanism

    CN116030193A

  • Audio and video generation method and device, equipment and storage medium

    CN117373455A

  • Multi-dimensional feature and three-plane representation enhanced three-dimensional grid reconstruction method and device

    CN119540497A

  • Arc screen multi-view 3D image generation method and device

    CN119671885A