Three-dimensional reconstruction method and device, equipment and storage medium
By optimizing the image rendering and generation model of the monocular image reconstruction, the image is generated and repaired, and the three-dimensional reconstruction is carried out again, the problem of low quality of monocular image reconstruction is solved, and a higher quality three-dimensional model generation and convenient operation process is achieved.
Patent Information
- Application Number
- CN202510335829.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-20
AI Technical Summary
In the prior art, the three-dimensional model of monocular image reconstruction is not of high quality, and limb distortion and breakage are prone to problems, and users need to provide multi-view information, which is inconvenient to operate.
By performing three-dimensional reconstruction based on monocular images, an initial three-dimensional model is generated, and then image rendering is rendered to generate a rendered image. The generated model is used to optimize the rendered image to generate a repaired image, and finally, a three-dimensional reconstruction is carried out again based on the initial image and the repaired image to generate a higher quality three-dimensional model.
The quality of the three-dimensional model generated by monocular images is improved, and the limb distortion and breakage are avoided. The user does not need to provide multi-view information, making it easy to operate.
Smart Images

Figure CN120182499A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to artificial intelligence technology and computer vision technology, and in particular to a three-dimensional reconstruction method, apparatus, device, and storage medium. Background Art
[0002] With the continuous development of artificial intelligence (AI) technology, significant progress has been made in AI generation technology in recent years. Three-dimensional reconstruction is an important task in the field of computer vision, and among them, human-oriented reconstruction is an important research direction in the field of three-dimensional reconstruction.
[0003] The methods of human body model reconstruction can be divided into multi-view reconstruction and monocular (single-image) reconstruction. Multi-view reconstruction requires providing images of at least two views, and usually requires providing comprehensive information of each view of the human body, and optimizing through representation methods such as 3-Dimensional Gaussian Splatting (3DGS), Neural Radiance Fields (NeRF), Mesh, etc. to obtain three-dimensional model data. Due to the lack of information on invisible views of the human body, monocular reconstruction usually needs to use a generative model to supplement views, such as generating some images or videos of other views based on a single monocular image, and then using the generated images and videos for three-dimensional reconstruction.
[0004] However, multi-view reconstruction requires inputting information of multiple views, which is not convenient for users. The method of first generating and then reconstructing used in monocular reconstruction heavily depends on the appearance consistency between the generated images. If the geometric shape and appearance of the target human body are inconsistent, it will lead to situations such as limb distortion and fragmentation in the generated three-dimensional model. Summary of the Invention
[0005] Embodiments of the present disclosure provide a three-dimensional reconstruction method, apparatus, device, and storage medium, which can improve the quality of the three-dimensional model generated using a single monocular image.
[0006] An aspect of the embodiments of the present disclosure provides a three-dimensional reconstruction method, including:
[0007] Performing three-dimensional reconstruction based on the obtained initial image to generate a first three-dimensional model of the target object, where the initial image is a monocular image including the target object;
[0008] Performing image rendering on the first three-dimensional model to obtain at least one rendered image, where the image acquisition view of the at least one rendered image is different from the image acquisition view of the initial image;
[0009] Input the at least one rendered image into a generation model, and optimize the at least one rendered image through the generation model to generate at least one repaired image;
[0010] Perform 3D reconstruction based on the initial image and the at least one repaired image to generate a second 3D model of the target object.
[0011] Optionally, the step of inputting the at least one rendered image into a generation model, and optimizing the at least one rendered image through the generation model to generate at least one repaired image includes:
[0012] Encode the at least one rendered image through an encoder to obtain image features of the at least one rendered image;
[0013] Concatenate the image features with noise features to obtain a feature pair;
[0014] Input the feature pair and the facial features of the target object into the generation model, and perform inverse diffusion processing on the noise features based on the image features and the facial features through the generation model to obtain target image features;
[0015] Decode the target image features through a decoder to obtain the at least one repaired image.
[0016] Optionally, the step of performing image rendering on the first 3D model to obtain at least one rendered image includes:
[0017] Collect and render images around the first 3D model at preset angular intervals to obtain a first sequence of rendered images;
[0018] Copy the first frame image in the first sequence of rendered images, and paste the copied image after the last frame image in the first sequence of rendered images to obtain a second sequence of rendered images.
[0019] Optionally, the generation model includes a cross-attention network, a self-attention network, and an inverse diffusion network;
[0020] The step of inputting the feature pair and the facial features of the target object into the generation model, and performing inverse diffusion processing on the noise features based on the image features and the facial features through the generation model to obtain target image features includes:
[0021] Input the sequence of feature pairs corresponding to the second sequence of rendered images and the facial features into the cross-attention network for cross-attention calculation to obtain a first sequence of fused features;
[0022] Input the first fusion feature sequence into the temporal self-attention network for temporal self-attention calculation to obtain a second fusion feature sequence. The preset attention unit in the temporal self-attention network has been pre-masked. The preset attention unit includes the self-attention unit of the first frame image in the second rendered image sequence, the self-attention unit of the last frame image in the second rendered image sequence, and the attention unit between the first frame image and the last frame image in the second rendered image sequence;
[0023] Input the second fusion feature sequence into the inverse diffusion network for inverse diffusion processing to obtain the target image feature.
[0024] Optionally, the method further includes:
[0025] Render the first sample model to obtain at least one first sample image;
[0026] Extract parameters from the first sample model to obtain a sample parametric representation model, which is used to represent the shape of the first sample model;
[0027] Perform 3D reconstruction based on the first sample image from a preset acquisition perspective and the human body parametric representation model to obtain a second sample model;
[0028] Render the second sample model based on the image acquisition perspective of the at least one first sample image to obtain at least one second sample image;
[0029] Train the generation model using the at least one first sample image and the at least one second sample image.
[0030] Optionally, the training of the generation model using the at least one first sample image and the at least one second sample image includes:
[0031] Encode the first sample image and the second sample image respectively to obtain a first image feature and a second image feature;
[0032] Add noise to the first image feature using a sample noise feature to obtain a noisy image feature;
[0033] Input the noisy image feature and the second image feature into the generation model, and through the generation model, perform inverse diffusion processing on the noisy image feature based on the second image feature to generate a predicted noise feature;
[0034] Using a loss function, calculate the function value of the loss function based on the sample noise feature and the predicted noise feature, and update the parameters of the generation model based on the function value.
[0035] Optionally, the three-dimensional reconstruction based on the obtained initial image to generate the first three-dimensional model of the target object includes:
[0036] Extract parameters from the initial image to obtain a target parametric representation model of the target object, where the target parametric representation model is used to represent the shape of the target object;
[0037] Perform target recognition on the initial image to obtain the image region corresponding to the target object;
[0038] Use a preset three-dimensional reconstruction technique to generate the first three-dimensional model based on the target parametric representation model and the image region corresponding to the target object.
[0039] On the other hand, an embodiment of the present disclosure provides a three-dimensional reconstruction device, including:
[0040] A first reconstruction module for performing three-dimensional reconstruction based on the obtained initial image to generate a first three-dimensional model of the target object, where the initial image is a monocular image including the target object;
[0041] A first rendering module for performing image rendering on the first three-dimensional model to obtain at least one rendered image, where the image acquisition perspective of the at least one rendered image is different from the image acquisition perspective of the initial image;
[0042] A repair module for inputting the at least one rendered image into a generation model, and performing image optimization on the at least one rendered image through the generation model to generate at least one repaired image;
[0043] A second reconstruction module for performing three-dimensional reconstruction based on the initial image and the at least one repaired image to generate a second three-dimensional model of the target object.
[0044] Optionally, the repair module is further configured to:
[0045] Encode the at least one rendered image through an encoder to obtain the image features of the at least one rendered image;
[0046] Concatenate the image features with noise features to obtain a feature pair;
[0047] Input the feature pair and the facial features of the target object into the generation model, and perform inverse diffusion processing on the noise features based on the image features and the facial features through the generation model to obtain target image features;
[0048] The decoder decodes the target image features to obtain the at least one restored image.
[0049] Optionally, the first rendering module is further configured to:
[0050] Collect and render images around the first 3D model at a preset angular interval to obtain a first rendered image sequence;
[0051] Copy the first frame image in the first rendered image sequence, and paste the copied image after the last frame image in the first rendered image sequence to obtain a second rendered image sequence.
[0052] Optionally, the generation model includes a cross-attention network, a self-attention network, and an inverse diffusion network;
[0053] The restoration module is further configured to:
[0054] Input the feature pair sequence corresponding to the second rendered image sequence and the facial features into the cross-attention network for cross-attention calculation to obtain a first fused feature sequence;
[0055] Input the first fused feature sequence into the temporal self-attention network for temporal self-attention calculation to obtain a second fused feature sequence, where the preset attention units in the temporal self-attention network are pre-masked, and the preset attention units include the self-attention unit of the first frame image in the second rendered image sequence, the self-attention unit of the last frame image in the second rendered image sequence, and the attention unit between the first frame image and the last frame image in the second rendered image sequence;
[0056] Input the second fused feature sequence into the inverse diffusion network for inverse diffusion processing to obtain the target image features.
[0057] On the other hand, an embodiment of the present disclosure provides an electronic device, including:
[0058] A memory for storing a computer program;
[0059] A processor for executing the computer program stored in the memory, and when the computer program is executed, implementing the method described in the above aspect.
[0060] On the other hand, an embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, implementing the method described in the above aspect.
[0061] Another aspect of the embodiments of the present disclosure provides a computer program including computer program instructions, which when executed by a processor implement the method described in the above aspect.
[0062] Based on the embodiments of the present disclosure, first, monocular images are used for 3D reconstruction to initially obtain a rough first 3D model. Then, multi-view image rendering is performed on the first 3D model to obtain images from other perspectives, and a generative model is used to perform image inpainting on the rendered images so that the obtained inpainted images contain clear target objects with consistent geometric shapes and appearances. Then, 3D reconstruction is performed again using the inpainted images and the initial images to generate a second 3D model with higher quality. The method of using the generative model to inpaint the rendered images to obtain multi-view images improves the quality of the 3D model generated based on monocular images and eliminates the need for users to collect multi-view images, making the operation convenient.
[0063] The technical solutions of the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings
[0064] The drawings forming a part of the specification depict the embodiments of the present disclosure and, together with the description, are used to explain the principles of the present disclosure.
[0065] Referring to the accompanying drawings, the present disclosure can be more clearly understood according to the following detailed description, where:
[0066] Figure 1 is a flowchart of an embodiment of the 3D reconstruction method of the present disclosure;
[0067] Figure 2 is a flowchart of another embodiment of the 3D reconstruction method of the present disclosure;
[0068] Figure 3 is a schematic diagram of the spatio-temporal self-attention matrix shown in an embodiment of the 3D reconstruction method of the present disclosure;
[0069] Figure 4 is a flowchart of another embodiment of the 3D reconstruction method of the present disclosure;
[0070] Figure 5 is a schematic diagram of the training process of the generative model shown in an embodiment of the 3D reconstruction method of the present disclosure;
[0071] Figure 6 is a schematic diagram of the process of generating a 3D human body model based on monocular images shown in an embodiment of the 3D reconstruction method of the present disclosure;
[0072] Figure 7 is a schematic structural diagram of an embodiment of the 3D reconstruction device of the present disclosure;
[0073] Figure 8Schematic structural diagram of another embodiment of the 3D reconstruction device of the present disclosure;
[0074] Figure 9 Schematic structural diagram of an application embodiment of an electronic device of the present disclosure. Detailed implementation manners
[0075] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present disclosure.
[0076] Those skilled in the art can understand that terms such as "first" and "second" in the embodiments of the present disclosure are only used to distinguish different steps, devices, or modules, etc., and neither represent any specific technical meaning nor indicate an inevitable logical order between them.
[0077] It should also be understood that in the embodiments of the present disclosure, "a plurality of" may refer to two or more, and "at least one" may refer to one, two, or more.
[0078] It should also be understood that for any component, data, or structure mentioned in the embodiments of the present disclosure, without clear limitation or contrary indication in the context, it can generally be understood as one or more.
[0079] In addition, the term "and / or" in the present disclosure is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present disclosure generally represents an "or" relationship between the associated objects before and after.
[0080] It should also be understood that the present disclosure emphasizes the differences between various embodiments, and the same or similar parts thereof can be referred to each other. For the sake of brevity, they will not be described one by one.
[0081] At the same time, it should be understood that for the sake of description, the dimensions of each part shown in the drawings are not drawn according to the actual proportional relationship.
[0082] The following description of at least one exemplary embodiment is actually merely illustrative and in no way a limitation on the present disclosure and its application or use.
[0083] Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and devices should be regarded as part of the specification.
[0084] It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0085] Figure 1 The flowchart of the 3D reconstruction method provided for an exemplary embodiment of the present disclosure. The 3D reconstruction method of the embodiments of the present disclosure can be implemented by an electronic device deployed with a generation model.
[0086] As Figure 1 shown, the method includes the following steps:
[0087] Step 101, perform 3D reconstruction based on the acquired initial image to generate a first 3D model of the target object, where the initial image is a monocular image including the target object.
[0088] Optionally, the initial image can be obtained by directly collecting an image of the target object, or can be transmitted by other devices (such as cameras, smartphones, laptops, etc.).
[0089] The initial image is a monocular image (i.e., a single image) including the target object. Schematically, the initial image can be a single image taken from the front view of the target object. In addition, the initial image can also be an image in the form of an image generated by a drawing tool, a video screenshot, etc. The embodiments of the present disclosure do not limit the acquisition method of the initial image. The target object can be a human body, an animal or an item, and the embodiments of the present disclosure do not limit the type of the target object.
[0090] In a possible implementation manner, the initial image can be used to preliminarily perform 3D reconstruction to obtain a relatively rough 3D model of the target object (hereinafter referred to as the first 3D model). Schematically, 3D reconstruction can be performed using 3D expressions such as 3DGS, NeRF, Mesh, etc. to generate the first 3D model of the target object.
[0091] Step 102, perform image rendering on the first 3D model to obtain at least one rendered image.
[0092] Among them, the image acquisition perspective of at least one rendered image is different from the image acquisition perspective of the initial image.
[0093] In a possible implementation, after generating the first three-dimensional model, an image rendering tool can be used to collect and render images of the first three-dimensional model from different image acquisition perspectives, where the image acquisition perspectives of each rendered image are different and the image acquisition perspective of each rendered image is different from that of the initial image. For example, the image acquisition perspective of the initial image is the front of the target object, the image acquisition perspective of one rendered image is the side of the target object, and the image acquisition perspective of another rendered image is the back of the target object. Alternatively, video can be collected around the first three-dimensional model, and the obtained video frames can be used as rendered images.
[0094] Schematically, tools such as Open Graphics Library (OpenGL), Blender, and Web Graphics Library (WebGL) can be used to render the first three-dimensional model to obtain at least one rendered image.
[0095] Step 103: Input at least one rendered image into a generation model, and optimize the at least one rendered image through the generation model to generate at least one repaired image.
[0096] The generation model can be implemented using Diffusion Models, such as the U-Net neural network model. The model inference principle of the generation model is to perform inverse diffusion processing on random noise, gradually denoise, and finally obtain the image features of a clear image. In the embodiments of the present disclosure, at least one rendered image can be input into a pre-trained generation model, so that the generation model performs inverse diffusion processing on random noise based on the rough shape of the first three-dimensional model in the rendered image, improves the geometric consistency and appearance consistency of the first three-dimensional model between the rendered images, eliminates phenomena such as limb distortion and fragmentation, and thus obtains multi-perspective images of the target object with higher quality, that is, at least one repaired image.
[0097] Step 104: Perform three-dimensional reconstruction based on the initial image and at least one repaired image to generate a second three-dimensional model of the target object.
[0098] In a possible implementation, the initial image and at least one repaired image can be used to perform three-dimensional reconstruction again. Since the basis for this three-dimensional reconstruction includes multi-perspective images, a three-dimensional model of the target object with higher quality (hereinafter referred to as the second three-dimensional model) can be obtained. Schematically, 3DGS, NeRF, Mesh, and other 3D representations can be used to perform three-dimensional reconstruction to generate the second three-dimensional model of the target object.
[0099] Based on the embodiments of the present disclosure, first, 3D reconstruction is performed using a monocular image to initially obtain a rough first 3D model. Then, multi-view image rendering is performed on the first 3D model to obtain images from other perspectives, and a generative model is used to repair the rendered images so that the obtained repaired images contain clear target objects with consistent geometric shapes and appearances. Then, 3D reconstruction is performed again using the repaired images and the initial images to generate a second 3D model with higher quality. The method of using the generative model to repair the rendered images to obtain multi-view images improves the quality of the 3D model generated based on the monocular image, and there is no need for the user to collect multi-view images, which is convenient to operate.
[0100] In a possible implementation manner, in order to further improve the quality of the second 3D model, when performing image rendering on the first 3D model, multi-angle image acquisition and rendering can be performed around the first 3D model at a preset angular interval to provide sufficient conditions for subsequent 3D reconstruction. The above step 102 may specifically include the following steps:
[0101] Step 102a, perform image acquisition and rendering around the first 3D model at a preset angular interval to obtain a first sequence of rendered images.
[0102] Step 102b, copy the first frame image in the first sequence of rendered images, and paste the copied image after the last frame image in the first sequence of rendered images to obtain a second sequence of rendered images.
[0103] Optionally, image acquisition and rendering can be performed on the first 3D model at a preset angular interval from each perspective to obtain a first sequence of rendered images, or video acquisition can be performed by uniformly rotating around the first 3D model, and then video frames can be extracted at a preset video frame interval to obtain a first sequence of rendered images.
[0104] Illustratively, starting from the front of the first 3D model, rotate in a clockwise direction around the first 3D model, and perform image rendering every 20° rotation to obtain a first sequence of rendered images containing 18 rendered images. Then, copy the first frame image in the first sequence of rendered images and paste the copied image after the 18th frame image to obtain a second sequence of rendered images containing 19 rendered images.
[0105] Based on the embodiments of the present disclosure, by performing multi-angle image acquisition and rendering around the first 3D model at preset angular intervals, sufficient conditions can be provided for subsequent 3D reconstruction, thereby improving the quality of the second 3D model. Moreover, after obtaining the first rendered image sequence, the first frame image in the first rendered image sequence is copied and pasted after the last frame image to obtain the second rendered image sequence, enabling the attention mechanism of the generation model to focus on the correlation between the first frame image and the last frame image, preventing the last few rendered images from having a large gap with the first few rendered images in the second rendered image sequence, further improving the consistency of all rendered images, and thus improving the quality of the second 3D model.
[0106] In a possible implementation manner, as Figure 2 shown, step 103 above may specifically include the following steps:
[0107] Step 201, encoding at least one rendered image through an encoder to obtain the image features of at least one rendered image.
[0108] Optionally, an image encoder is also deployed in the electronic device, and at least one rendered image can be encoded through the image encoder to obtain the image features of at least one rendered image, so as to input the image features into the generation model for inverse diffusion processing.
[0109] Step 202, splicing the image features and the noise features to obtain a feature pair.
[0110] For the image features of each rendered image, a frame of noise feature is correspondingly generated to obtain a feature pair corresponding to each rendered image. Among them, the noise feature can be obtained by performing feature extraction on random noise.
[0111] Step 203, inputting the feature pair and the facial features of the target object into the generation model, and the generation model performs inverse diffusion processing on the noise feature based on the image features and the facial features to obtain the target image features.
[0112] Among them, when the target object is a human body, the facial features of the target object can be obtained by performing face recognition and feature extraction on the initial image through a face recognition model. In the embodiments of the present disclosure, inputting the facial features of the target object into the generation model to assist in inverse diffusion processing can prevent randomly generating other faces during the inverse diffusion process, ensure the consistency of the face in the repaired image and the initial image, and further ensure the consistency of the facial features of the generated second 3D model and the facial features of the target object.
[0113] In a possible implementation, the generation model includes a cross-attention network, a self-attention network, and an inverse diffusion network. The self-attention network is used to calculate the relationship between the image features of each frame of the rendered image and the image features of other rendered images, so as to capture long-range dependencies. When at least one rendered image input to the generation model is the second rendered image sequence generated in step 102b above, since the last frame of the rendered image is obtained by copying the first frame of the rendered image, when the generation model performs self-attention calculation, the correlation between the last frame of the rendered image and the first frame of the rendered image is too high, which will cause the attention of the model to focus on the last frame of the rendered image and the first frame of the rendered image and ignore the intermediate rendered images. Therefore, by masking the attention network, the model can be made to focus on the intermediate rendered images that need to be repaired, thereby improving the image quality of the repaired image. Step 203 may specifically include the following steps:
[0114] Step 203a: Input the feature pair sequence corresponding to the second rendered image sequence and the facial features into the cross-attention network for cross-attention calculation to obtain a first fused feature sequence.
[0115] Step 203b: Input the first fused feature sequence into the temporal self-attention network for temporal self-attention calculation to obtain a second fused feature sequence.
[0116] Among them, the preset attention units in the temporal self-attention network are pre-masked. The preset attention units include the self-attention unit of the first frame image in the second rendered image sequence, the self-attention unit of the last frame image in the second rendered image sequence, and the attention unit between the first frame image in the second rendered image sequence and the last frame image in the second rendered image sequence.
[0117] Step 203c: Input the second fused feature sequence into the inverse diffusion network for inverse diffusion processing to obtain the target image features.
[0118] The temporal self-attention network is used to perform temporal self-attention calculation to generate an attention weight matrix composed of attention units, where each attention unit is the attention weight between two corresponding frames of images. The generation model can perform subsequent inverse diffusion processing based on this attention weight matrix. By masking the preset attention units, the corresponding attention weights are reduced to 0, which can make the generation model pay more attention to the correlation between other images.
[0119] Schematically, Figure 3An attention weight matrix is shown, where the left side is the attention weight matrix without masking, and the right side is the attention weight matrix of the preset attention unit after masking. Each square in the attention weight matrix represents an attention unit. For example, the attention unit in the first row and the first column is the attention weight of the first frame image itself, and the attention unit in the first row and the second column is the attention weight between the first frame image and the second frame image. It can be seen that in the attention matrix on the left without masking, the attention weights of the first frame image itself, the last frame image itself, and the attention weight between the first frame image and the last frame image are relatively high, and the attention weights corresponding to other attention units are relatively low. While in the attention matrix on the right after masking, the attention weights corresponding to the attention units in the middle part are significantly increased, which can make the generation model pay more attention to the intermediate rendering images that need to be repaired.
[0120] Step 204, decode the target image features through a decoder to obtain at least one repaired image.
[0121] Optionally, an image decoder is also deployed in the electronic device, and each target image feature can be decoded through the image decoder to obtain at least one repaired image. The image encoder and the image decoder can be pre-trained.
[0122] In a possible implementation manner, the generation model in the above embodiment is pre-trained. As Figure 4 shown, the method provided by the embodiments of the present disclosure may further include the following steps:
[0123] Step 401, perform image rendering on the first sample model to obtain at least one first sample image.
[0124] Among them, the first sample model is a three-dimensional model obtained in advance. The first sample model can be a high-quality three-dimensional model obtained through methods such as point cloud scanning and multi-view image generation.
[0125] Optionally, an image rendering tool can be used to perform image rendering on the first sample model from different image acquisition perspectives, where the image acquisition perspectives of each first sample image are different. For example, video acquisition can be performed around the first sample model, and the obtained video frames can be used as the first sample images, or the first sample model can be surrounded and image rendering can be performed at preset angular intervals to obtain at least one first sample image.
[0126] Schematically, tools such as OpenGL, Blender, and WebGL can be used to perform image rendering on the first sample model to obtain at least one first sample image. The at least one first sample image can be used as an optimization condition for training the generation model.
[0127] Step 402: Extract parameters from the first sample model to obtain a sample parametric representation model.
[0128] Among them, the sample parametric representation model is used to represent the shape of the first sample model. The parametric representation model (Skinned Multi-Person Linear Model, SMPL) is a human parametric representation model based on skin vertices, which can accurately represent the body shape in various natural human postures. Schematically, a sample SMPL can be estimated from the first sample image using Neural Localizer Fields (NLF).
[0129] Step 403: Perform 3D reconstruction based on the first sample image with a preset acquisition perspective and the human parametric representation model to obtain a second sample model.
[0130] Among them, the first sample image with a preset acquisition perspective refers to a first sample image whose acquisition perspective is the same as the preset acquisition perspective, or a first sample image whose acquisition perspective is closest to the preset acquisition perspective. For example, as Figure 5 shown, a first sample image with an acquisition perspective closest to the front view of the first sample model can be selected, and 3D reconstruction can be performed in combination with the sample SMPL to obtain a second sample model.
[0131] Performing 3D reconstruction using the first sample image with a preset acquisition perspective and the human parametric representation model results in a relatively rough second sample model.
[0132] Step 404: Render the second sample model based on the image acquisition perspectives of at least one first sample image to obtain at least one second sample image.
[0133] Optionally, the image acquisition perspectives of at least one second sample image correspond one-to-one with the image acquisition perspectives of at least one first sample image, so as to splice the first sample image and the second sample image for model training. For example, when rendering the first sample image, if starting from the front of the first sample model and clockwise around the first sample model, a first sample image is collected and rendered every 20°, then starting from the front of the second sample model and clockwise around the second sample model, a second sample image can be collected and rendered every 20°.
[0134] Step 405: Train the generation model using at least one first sample image and at least one second sample image.
[0135] Using at least one first sample image and at least one second sample image, adding noise to the first sample image so that the generative model can restore the first sample image based on the second sample image as much as possible, and updating the model parameters of the generative model, so that the trained generative model has the ability to repair the rendered image based on the rough model into a clear image with the same geometric shape and appearance of the target object. In a possible implementation manner, step 405 above may specifically include the following steps:
[0136] Step 405a, encoding the first sample image and the second sample image respectively to obtain a first image feature and a second image feature.
[0137] Optionally, a pre-trained image encoder can be used to encode the first sample image and the second sample image respectively to obtain the image feature corresponding to the first sample image (hereinafter referred to as the first image feature) and the image feature corresponding to the second sample image (hereinafter referred to as the second image feature).
[0138] Step 405b, adding noise to the first image feature by using the sample noise feature to obtain a noisy image feature.
[0139] Optionally, the sample noise feature can be superimposed on the first image feature to obtain a noisy image feature, where the dimension of the sample noise feature is the same as that of the first image feature.
[0140] Step 405c, inputting the noisy image feature and the second image feature into the generative model, and performing inverse diffusion processing on the noisy image feature based on the second image feature through the generative model to generate a predicted noise feature.
[0141] Optionally, when the first sample model is a three-dimensional human body model, face recognition and feature extraction can also be performed on the first sample image through a face recognition model to obtain sample facial features, and the sample facial features are used as a constraint condition and input into the generative model, so that the generative model refers to the sample facial features to perform inverse diffusion processing on the noisy image feature to generate a predicted noise feature.
[0142] Step 405e, using a loss function to calculate the function value of the loss function based on the sample noise feature and the predicted noise feature, and updating the parameters of the generative model based on the function value.
[0143] Among them, the function value of the loss function is used to characterize the difference between the sample noise feature and the predicted noise feature. After updating the parameters of the generative model based on the function value, the next round of training can be performed based on the generative model with updated parameters until the preset training conditions are met. The preset training conditions may include, for example, but are not limited to at least one of the following conditions: the number of iterative training reaches a preset number, the training duration reaches a preset duration, and the function value of the loss function is less than a preset threshold.
[0144] Schematically, the loss function value is the average of the norms corresponding to the noise feature differences, where the noise feature differences are the differences between the sample noise features and the predicted noise features, and the expression is as follows:
[0145] L = E[||ε - ε θ (z′, c)|| 2 (1)
[0146] where E represents taking the mean, "||·||" represents the norm, ε represents the sample noise features, and ε θ (z′, c) is the predicted noise feature, z′ represents the noisy image feature, and c represents the preset control condition, which can include, for example, the facial features of the target object and the structural information of the target object (such as the key point positions). The facial features can be obtained by extracting the facial features of the initial image through a face recognition model, and the key point positions can be obtained based on the received key point annotation operations.
[0147] Based on the embodiments of the present disclosure, at least one first sample image with clear, consistent geometric shape and appearance is rendered using a high-quality first sample model, and then a second sample model with lower quality is reconstructed using one first sample image (monocular image) and image rendering is performed from the same perspective. The training data can be formed by pairing the first sample image with the second sample image with rough image content.
[0148] In a possible implementation manner, the target parametric representation model of the target object and the image region corresponding to the target object can be extracted from the initial image for three-dimensional reconstruction. The above step 101 can specifically include the following steps:
[0149] Step 101a: Extract parameters from the initial image to obtain the target parametric representation model of the target object, which is used to represent the shape of the target object.
[0150] Among them, the target parametric representation model can accurately represent the body shape in various natural human postures. Schematically, the target SMPL can be estimated from the initial image using Neural Localizer Fields (NLF).
[0151] Step 101b: Perform target recognition on the initial image to obtain the image region corresponding to the target object.
[0152] Schematically, image segmentation models such as Segment Anything, Mask2Former, and Swin Transfomer can be used to perform target recognition and image segmentation on the initial image to obtain the image region (mask) corresponding to the target object.
[0153] Step 101c: Using a preset 3D reconstruction technique, generate a first 3D model based on the target parametric representation model and the image region corresponding to the target object.
[0154] Schematically, the target parametric representation model (target SMPL) and the image region (mask) corresponding to the target object can be input into models such as 3DGS, NeRF, and Mesh for 3D reconstruction to generate the first 3D model of the target object.
[0155] Correspondingly, after obtaining the repaired image, the second 3D model can also be generated in the above manner.
[0156] Combining the above method embodiments, Figure 6 shows a schematic diagram of the process of generating a human 3D model based on a monocular image using the method provided in the embodiments of the present disclosure. After obtaining the initial image (a frontal human image), the initial image is input into a 3D reconstruction model to obtain a rough first 3D model. The first 3D model is rendered at multiple angles to obtain at least one rendered image. The at least one rendered image is input into an encoder to obtain image features. Then, the image features are concatenated with noise features and input into a generation model together with the facial features extracted from the initial image. After attention calculation and inverse diffusion processing, target image features are obtained. The decoder decodes the target image features to generate at least one repaired image. By inputting the repaired image into the 3D reconstruction model, a high-quality second 3D model can be obtained.
[0157] Please refer to Figure 7 , which shows a structural block diagram of a 3D reconstruction device provided in an exemplary embodiment of the present disclosure. The 3D reconstruction device provided in this embodiment includes:
[0158] A first reconstruction module 701, configured to perform 3D reconstruction based on the obtained initial image to generate a first 3D model of the target object, where the initial image is a monocular image including the target object;
[0159] A first rendering module 702, configured to render the first 3D model generated by the first reconstruction module 701 to obtain at least one rendered image, where the image acquisition perspective of the at least one rendered image is different from the image acquisition perspective of the initial image;
[0160] A repair module 703, configured to input the at least one rendered image rendered by the first rendering module 702 into a generation model, and optimize the at least one rendered image through the generation model to generate at least one repaired image;
[0161] The second reconstruction module 704 is configured to perform three-dimensional reconstruction based on the initial image and at least one repaired image obtained by the repair module 703, and generate a second three-dimensional model of the target object.
[0162] Optionally, in a possible implementation, the above-mentioned repair module 703 can also be used to:
[0163] Encode at least one rendered image through an encoder to obtain image features of the at least one rendered image;
[0164] Concatenate the image features with noise features to obtain a feature pair;
[0165] Input the feature pair and the facial features of the target object into a generation model, and the generation model performs inverse diffusion processing on the noise features based on the image features and facial features to obtain target image features;
[0166] Decode the target image features through a decoder to obtain at least one repaired image.
[0167] Optionally, in a possible implementation, the above-mentioned first rendering module 702 can also be used to:
[0168] Collect and render images around the first three-dimensional model at a preset angular interval to obtain a first rendered image sequence;
[0169] Copy the first frame image in the first rendered image sequence, and paste the copied image after the last frame image in the first rendered image sequence to obtain a second rendered image sequence.
[0170] Optionally, in a possible implementation, the generation model includes a cross-attention network, a self-attention network, and an inverse diffusion network;
[0171] The above-mentioned repair module 703 can also be used to:
[0172] Input the feature pair sequence corresponding to the second rendered image sequence and the facial features into the cross-attention network for cross-attention calculation to obtain a first fused feature sequence;
[0173] Input the first fused feature sequence into a temporal self-attention network for temporal self-attention calculation to obtain a second fused feature sequence, where the preset attention unit in the temporal self-attention network is pre-masked, and the preset attention unit includes the self-attention unit of the first frame image in the second rendered image sequence, the self-attention unit of the last frame image in the second rendered image sequence, and the attention unit between the first frame image and the last frame image in the second rendered image sequence;
[0174] Input the second fusion feature sequence into the inverse diffusion network for inverse diffusion processing to obtain the target image feature.
[0175] Optionally, in a possible implementation, as Figure 8 shown, the device provided by the embodiments of the present disclosure may further include:
[0176] A second rendering module 801, configured to perform image rendering on the first sample model to obtain at least one first sample image;
[0177] An extraction module 802, configured to extract parameters from the first sample model to obtain a sample parametric representation model, where the sample parametric representation model is used to represent the shape of the first sample model;
[0178] A third reconstruction module 803, configured to perform three-dimensional reconstruction based on the first sample image at a preset acquisition perspective rendered by the second rendering module 801 and the human body parametric representation model extracted by the extraction module 802 to obtain a second sample model;
[0179] A third rendering module 804, configured to perform image rendering on the second sample model generated by the third reconstruction module 803 based on the image acquisition perspective of at least one first sample image to obtain at least one second sample image;
[0180] A training module 805, configured to train the generation model by using at least one first sample image rendered by the second rendering module 801 and at least one second sample image rendered by the third rendering module 804.
[0181] Optionally, in a possible implementation, the training module 805 is further configured to:
[0182] Encode the first sample image and the second sample image respectively to obtain a first image feature and a second image feature;
[0183] Perform noise addition processing on the first image feature by using the sample noise feature to obtain a noise-added image feature;
[0184] Input the noise-added image feature and the second image feature into the generation model, and through the generation model, perform inverse diffusion processing on the noise-added image feature based on the second image feature to generate a predicted noise feature;
[0185] Use the loss function to calculate the function value of the loss function based on the sample noise feature and the predicted noise feature, and update the parameters of the generation model based on the function value.
[0186] Optionally, in a possible implementation, the first reconstruction module 701 is further configured to:
[0187] Extract parameters from the initial image to obtain a target parametric representation model of the target object, where the target parametric representation model is used to represent the shape of the target object;
[0188] Perform target recognition on the initial image to obtain the image region corresponding to the target object;
[0189] Using a preset 3D reconstruction technique, generate a first 3D model based on the target parametric representation model and the image region corresponding to the target object.
[0190] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is the difference from other embodiments. For the same, similar or corresponding parts among the embodiments, reference can be made to each other. Since the method, apparatus, and device embodiments basically correspond, reference can be made to the corresponding parts of the description for the relevant parts. The methods, apparatuses, and devices of the embodiments of the present disclosure also correspond to each other in specific implementation and beneficial technical effects. The relevant content can be referred to each other and will not be elaborated here.
[0191] In addition, the embodiments of the present disclosure also provide an electronic device, including:
[0192] A memory for storing a computer program;
[0193] A processor for executing the computer program stored in the memory, and when the computer program is executed, implementing the 3D reconstruction method described in any one of the above embodiments of the present disclosure.
[0194] Figure 9 This is a schematic structural diagram of an application embodiment of the electronic device of the present disclosure. Next, refer to Figure 9 to describe the electronic device according to the embodiments of the present disclosure. The electronic device can be any one or both of the first device and the second device, or a stand-alone device independent of them. The stand-alone device can communicate with the first device and the second device to receive the input signals collected from them.
[0195] As Figure 9 shown, the electronic device includes one or more processors and a memory.
[0196] The processor can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device to perform desired functions.
[0197] The memory may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media, and the processor may run the program instructions to implement the three-dimensional reconstruction method of various embodiments of the present disclosure described above and / or other desired functions.
[0198] In one example, the electronic device may further include: an input device and an output device, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0199] In addition, the input device may further include, for example, a keyboard, a mouse, and so on.
[0200] The output device may output various information to the outside, including the determined distance information, direction information, etc. The output device may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, and so on.
[0201] Of course, for simplicity, Figure 9 only some of the components related to the present disclosure in the electronic device are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device may further include any other appropriate components.
[0202] In addition to the above methods and devices, an embodiment of the present disclosure may also be a computer program product, which includes computer program instructions, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the three-dimensional reconstruction method according to various embodiments of the present disclosure described in the above part of this specification.
[0203] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages, such as Java, C++, etc., and also include conventional procedural programming languages, such as the "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0204] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium storing computer program instructions, which, when run by a processor, cause the processor to execute the steps in the three-dimensional reconstruction method according to various embodiments of the present disclosure described in the foregoing part of this specification.
[0205] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0206] Those of ordinary skill in the art can understand that all or part of the steps for implementing the foregoing method embodiments may be completed by hardware related to program instructions. The foregoing program may be stored in a computer-readable storage medium, and when executed, it performs the steps including the foregoing method embodiments; and the foregoing storage medium includes: ROM, RAM, magnetic disk, or optical disk and other media that can store program codes.
[0207] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above-disclosed specific details are only for the purposes of illustration and facilitating understanding, and are not limitations. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.
[0208] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference may be made to each other. For the system embodiment, since it basically corresponds to the method embodiment, the description is relatively simple. For the relevant parts, reference may be made to the partial description of the method embodiment.
[0209] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including," "comprising," "having," etc. are open-ended terms, meaning "including but not limited to," and can be used interchangeably with each other. The words "or" and "and" used herein refer to the word "and / or" and can be used interchangeably with it, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with it.
[0210] The methods and apparatuses of the present disclosure can be implemented in many ways. For example, the methods and apparatuses of the present disclosure can be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of the steps for the methods is for illustrative purposes only, and the steps of the methods of the present disclosure are not limited to the specific order described above, unless otherwise specifically stated. In addition, in some embodiments, the present disclosure can also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the methods according to the present disclosure. Thus, the present disclosure also covers a recording medium storing a program for executing the methods according to the present disclosure.
[0211] It should also be noted that in the apparatuses, equipment, and methods of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present disclosure.
[0212] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0213] The above description has been given for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions, and sub-combinations thereof.
Claims
1. A three-dimensional reconstruction method, characterized in that: include: Performing three-dimensional reconstruction based on the acquired initial image to generate a first three-dimensional model of the target object, wherein the initial image is a monocular image including the target object; Performing image rendering on the first three-dimensional model to obtain at least one rendered image, wherein an image acquisition perspective of the at least one rendered image is different from an image acquisition perspective of the initial image; Inputting the at least one rendered image into a generation model, performing image optimization on the at least one rendered image through the generation model, and generating at least one repaired image; Perform three-dimensional reconstruction based on the initial image and the at least one repaired image to generate a second three-dimensional model of the target object.
2. The method according to claim 1, characterized in that The step of inputting the at least one rendered image into a generation model, and performing image optimization on the at least one rendered image through the generation model to generate at least one repaired image comprises: Encoding the at least one rendered image by an encoder to obtain image features of the at least one rendered image; Concatenating the image feature with the noise feature to obtain a feature pair; Inputting the feature pair and the facial features of the target object into the generative model, and performing inverse diffusion processing on the noise features based on the image features and the facial features through the generative model to obtain target image features; The target image features are decoded by a decoder to obtain the at least one repaired image.
3. The method according to claim 2, characterized in that The performing image rendering on the first three-dimensional model to obtain at least one rendered image includes: Performing image acquisition and rendering around the first three-dimensional model at preset angle intervals to obtain a first rendered image sequence; The first frame image in the first rendered image sequence is copied, and the copied image is pasted after the last frame image in the first rendered image sequence to obtain a second rendered image sequence.
4. The method according to claim 3, characterized in that: The generation model includes a cross attention network, a self-attention network and a reverse diffusion network; The step of inputting the feature pair and the facial features of the target object into the generation model, and performing inverse diffusion processing on the noise features based on the image features and the facial features through the generation model to obtain target image features includes: Inputting the feature pair sequence corresponding to the second rendered image sequence and the facial features into the cross attention network to perform cross attention calculation to obtain a first fused feature sequence; Inputting the first fused feature sequence into the temporal self-attention network for temporal self-attention calculation to obtain a second fused feature sequence, wherein the preset attention units in the temporal self-attention network are pre-masked, and the preset attention units include the self-attention unit of the first frame image in the second rendered image sequence, the self-attention unit of the last frame image in the second rendered image sequence, and the attention unit between the first frame image in the second rendered image sequence and the last frame image in the second rendered image sequence; The second fused feature sequence is input into the inverse diffusion network for inverse diffusion processing to obtain the target image feature.
5. The method according to any one of claims 2 to 4, characterized in that: The method further comprises: Performing image rendering on the first sample model to obtain at least one first sample image; Extracting parameters from the first sample model to obtain a sample parameterized representation model, where the sample parameterized representation model is used to represent the shape of the first sample model; Performing three-dimensional reconstruction based on the first sample image of the preset acquisition angle and the human body parameterized representation model to obtain a second sample model; Performing image rendering on the second sample model based on the image acquisition perspective of the at least one first sample image to obtain at least one second sample image; The generation model is trained using the at least one first sample image and the at least one second sample image.
6. The method according to claim 5, characterized in that The using the at least one first sample image and the at least one second sample image to train the generation model comprises: Encoding the first sample image and the second sample image respectively to obtain a first image feature and a second image feature; Performing noise processing on the first image feature using the sample noise feature to obtain a noisy image feature; Inputting the noisy image feature and the second image feature into the generation model, and performing inverse diffusion processing on the noisy image feature based on the second image feature through the generation model to generate a predicted noise feature; The loss function is used to calculate a function value of the loss function based on the sample noise feature and the predicted noise feature, and the parameters of the generation model are updated based on the function value.
7. The method according to any one of claims 1 to 4, characterized in that: The three-dimensional reconstruction based on the acquired initial image to generate a first three-dimensional model of the target object includes: Extracting parameters from the initial image to obtain a target parameterized representation model of the target object, wherein the target parameterized representation model is used to represent the shape of the target object; Performing target recognition on the initial image to obtain an image area corresponding to the target object; The first three-dimensional model is generated based on the target parameterized representation model and the image area corresponding to the target object by using a preset three-dimensional reconstruction technology.
8. A three-dimensional reconstruction device, characterized in that: include: A first reconstruction module, configured to perform three-dimensional reconstruction based on an acquired initial image to generate a first three-dimensional model of a target object, wherein the initial image is a monocular image containing the target object; A first rendering module, configured to perform image rendering on the first three-dimensional model to obtain at least one rendered image, wherein an image acquisition perspective of the at least one rendered image is different from an image acquisition perspective of the initial image; A restoration module, used for inputting the at least one rendered image into a generation model, performing image optimization on the at least one rendered image through the generation model, and generating at least one restored image; The second reconstruction module is used to perform three-dimensional reconstruction based on the initial image and the at least one repaired image to generate a second three-dimensional model of the target object.
9. The device according to claim 8, characterized in that The repair module is further used for: Encoding the at least one rendered image by an encoder to obtain image features of the at least one rendered image; Concatenating the image feature with the noise feature to obtain a feature pair; Inputting the feature pair and the facial features of the target object into the generative model, and performing inverse diffusion processing on the noise features based on the image features and the facial features through the generative model to obtain target image features; The target image features are decoded by a decoder to obtain the at least one repaired image.
10. The device according to claim 9, characterized in that The first rendering module is further used for: Performing image acquisition and rendering around the first three-dimensional model at preset angle intervals to obtain a first rendered image sequence; The first frame image in the first rendered image sequence is copied, and the copied image is pasted after the last frame image in the first rendered image sequence to obtain a second rendered image sequence.
11. The device according to claim 10, characterized in that The generation model includes a cross attention network, a self-attention network and a reverse diffusion network; The repair module is further used for: Inputting the feature pair sequence corresponding to the second rendered image sequence and the facial features into the cross attention network to perform cross attention calculation to obtain a first fused feature sequence; Inputting the first fused feature sequence into the temporal self-attention network for temporal self-attention calculation to obtain a second fused feature sequence, wherein the preset attention units in the temporal self-attention network are pre-masked, and the preset attention units include the self-attention unit of the first frame image in the second rendered image sequence, the self-attention unit of the last frame image in the second rendered image sequence, and the attention unit between the first frame image in the second rendered image sequence and the last frame image in the second rendered image sequence; The second fused feature sequence is input into the inverse diffusion network for inverse diffusion processing to obtain the target image feature.
12. An electronic device, characterized in that: include: Memory for storing computer programs; A processor is used to execute the computer program stored in the memory, and when the computer program is executed, the method described in any one of claims 1 to 7 is implemented.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method described in any one of claims 1 to 7 is implemented.
14. A computer program comprising computer program instructions, characterized in that When the computer program instructions are executed by a processor, the method described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Single-view three-dimensional object reconstruction method, device and equipment and storage medium
CN116843832A
Map generation method and device, equipment, computer program product and storage medium
CN118135114A
Transition video generation method and system
CN119211642A
Map generation method and device, medium, equipment and computer program product
CN119251373A
Cited By
Video generation method and device, electronic equipment, storage medium and program product
CN120434373A
Video generation methods, apparatuses, electronic devices, storage media, and software products
CN120434373B
Three-dimensional model generation method and device, computer equipment and storage medium
CN120747356A
Three-dimensional model generation method and apparatus, computer device, and storage medium
CN120747356B
Object three-dimensional reconstruction method and electronic equipment
CN121353503A