Three-dimensional data generation method and device, electronic equipment and storage medium

By diffusion networks, fusion of images and describing text features, and using neural radiation fields to adjust losses, the cumbersome problems of the traditional three-dimensional content creation process are solved, and the effect of quickly and efficiently generating high-quality three-dimensional data is achieved.

CN120070721AActive Publication Date: 2025-05-30BEIJING ZITIAO NETWORK TECH CO LTD

Patent Information

Application Number
CN202311619292.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-29
Publication Date
2025-05-30
Estimated Expiration
2043-11-29

AI Technical Summary

Technical Problem

The traditional three-dimensional content creation process is complicated and complicated, making it difficult to quickly and efficiently generate high-quality three-dimensional data.

Method used

The image features, description text features and noise image features of the target object are fused through a diffusion network to generate a second image, and the neural radiation field is adjusted through fractional distillation loss to determine the three-dimensional model of the target object.

Benefits of technology

It realizes the rapid and efficient generation of high-quality three-dimensional data based on the image and description text of the target object, and simplifies the three-dimensional content creation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070721A_ABST
    Figure CN120070721A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a three-dimensional data generation method and device, electronic equipment and a storage medium. The method comprises the steps that a first image of a target object and a description text corresponding to the first image are acquired; obtaining a noise image; wherein the noise image comprises a rendering image added with noise; the rendered image comprises an image rendered by the neural radiation field at a preset view angle; performing cross attention fusion on the features of the first image, the features of the description text and the features of the noise image through a diffusion network to generate a second image; wherein the second image comprises an image of the target object at a preset view angle; determining fractional distillation loss according to the second image, the noise image and the added noise, and adjusting a neural radiation field according to the fractional distillation loss; and determining a three-dimensional model of the target object according to the adjusted neural radiation field. And high-quality three-dimensional data can be quickly and efficiently generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, and in particular, to a method, apparatus, electronic device, and storage medium for generating three-dimensional data. Background Art

[0002] The traditional three-dimensional (3D) content creation process is cumbersome and complex, requiring a large amount of human and material costs. How to quickly and efficiently generate high-quality 3D data is a technical problem to be solved urgently. Summary of the Invention

[0003] Embodiments of the present disclosure provide a method, apparatus, electronic device, and storage medium for generating three-dimensional data, which can quickly and efficiently generate high-quality three-dimensional data.

[0004] In a first aspect, an embodiment of the present disclosure provides a method for generating three-dimensional data, including:

[0005] Obtaining a first image of a target object and a description text corresponding to the first image;

[0006] Obtaining a noise image; wherein the noise image includes a rendered image with added noise; the rendered image includes an image rendered by a neural radiance field at a preset viewing angle;

[0007] Through a diffusion network, cross-attention fusion is performed on the features of the first image, the features of the description text, and the features of the noise image to generate a second image; wherein the second image includes an image of the target object at the preset viewing angle;

[0008] Determining a score distillation loss according to the second image, the noise image, and the added noise, and adjusting the neural radiance field according to the score distillation loss;

[0009] Determining a three-dimensional model of the target object according to the adjusted neural radiance field.

[0010] In a second aspect, an embodiment of the present disclosure further provides a device for generating three-dimensional data, including:

[0011] A first acquisition module, configured to acquire a first image of a target object and a description text corresponding to the first image;

[0012] A second acquisition module, configured to acquire a noise image; wherein the noise image includes a rendered image with added noise; the rendered image includes an image rendered by a neural radiance field at a preset viewing angle;

[0013] Diffusion module, configured to perform cross-attention fusion on the features of the first image, the features of the description text, and the features of the noise image through a diffusion network to generate a second image; wherein, the second image includes an image of the target object at the preset perspective;

[0014] 3D representation module, configured to determine a score distillation loss according to the second image, the noise image, and the added noise, and adjust the neural radiance field according to the score distillation loss;

[0015] Model generation module, configured to determine a 3D model of the target object according to the adjusted neural radiance field.

[0016] In a third aspect, an embodiment of the present disclosure further provides an electronic device, which includes:

[0017] One or more processors;

[0018] A storage device, configured to store one or more programs,

[0019] When the one or more programs are executed by the one or more processors, the one or more processors implement the method for generating 3D data as described in any one of the embodiments of the present disclosure.

[0020] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the method for generating 3D data as described in any one of the embodiments of the present disclosure when executed by a computer processor.

[0021] The technical solution of the embodiment of the present disclosure is to obtain a first image of a target object and a description text corresponding to the first image; obtain a noise image; wherein, the noise image includes a rendered image with added noise; the rendered image includes an image rendered by a neural radiance field at a preset perspective; perform cross-attention fusion on the features of the first image, the features of the description text, and the features of the noise image through a diffusion network to generate a second image; wherein, the second image includes an image of the target object at the preset perspective; determine a score distillation loss according to the second image, the noise image, and the added noise, and adjust the neural radiance field according to the score distillation loss; determine a 3D model of the target object according to the adjusted neural radiance field.

[0022] By performing cross-attention fusion on the features of the first image, the features of the descriptive text, and the features of the noise image based on a diffusion network, it is possible to generate a second image based on the noise image with the first image and the descriptive text as control conditions. By determining the rendered image at a preset viewing angle based on a neural radiance field, generating a noise image for input to the diffusion network based on the rendered image in a way of adding noise, and constructing the neural radiance field through a score distillation loss, it is possible to make the three-dimensional model expressed based on the neural radiance field have the same feature distribution as the second image generated based on the first image and the descriptive text. Thus, it is possible to quickly and efficiently generate high-quality three-dimensional data according to the first image and the descriptive text. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the accompanying drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the original components and elements are not necessarily drawn to scale.

[0024] Figure 1 It is a flowchart of a method for generating three-dimensional data provided by an embodiment of the present disclosure;

[0025] Figure 2 It is a framework diagram of a method for generating three-dimensional data provided by an embodiment of the present disclosure;

[0026] Figure 3 It is a schematic diagram of feature cross-attention fusion in a method for generating three-dimensional data provided by an embodiment of the present disclosure;

[0027] Figure 4 It is a flowchart of a method for generating three-dimensional data provided by an embodiment of the present disclosure;

[0028] Figure 5 It is a schematic diagram of a third image in a method for generating three-dimensional data provided by an embodiment of the present disclosure;

[0029] Figure 6 It is a flowchart of a method for generating three-dimensional data provided by an embodiment of the present disclosure;

[0030] Figure 7 It is a structural diagram of a device for generating three-dimensional data provided by an embodiment of the present disclosure;

[0031] Figure 8 It is a structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0033] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0034] As used herein, the term "including" and its variations are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0035] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.

[0036] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly stated in the context, it should be understood as "one or more".

[0037] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0038] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.

[0039] Figure 1 It is a schematic flowchart of a method for generating three-dimensional data provided for the embodiments of the present disclosure. The embodiments of the present disclosure are applicable to the situation of generating a similar three-dimensional model according to graphics and text. This method can be executed by a device for generating three-dimensional data, and this device can be implemented in the form of software and / or hardware, and can be configured in an electronic device, such as configured in a computer.

[0040] As Figure 1As shown in the figure, the method for generating three-dimensional data provided in this embodiment may include:

[0041] S110. Obtain a first image of a target object and a description text corresponding to the first image.

[0042] In the embodiments of the present disclosure, the target object may be an object or a biological object, etc. The first image includes an image containing the target object, and the description text can be considered as the text describing the target object in the first image, and the first image and the description text have content consistency. Among them, the first image and the description text can be used as control conditions for generating a three-dimensional model of the target object, so that the generated three-dimensional model is similar to the first image and the description text.

[0043] Among them, the first image and the description text can be obtained by means of user input, or by reading a preset storage space, etc., and no exhaustive listing is made here.

[0044] S120. Obtain a noise image; wherein, the noise image includes a rendered image with added noise; the rendered image includes an image rendered from a neural radiance field at a preset viewing angle.

[0045] In the embodiments of the present disclosure, a neural radiance field (NeRF) can be randomly initialized, and camera parameters at preset viewing angles can be sampled in spherical coordinates, and point light sources can be sampled around the camera. Then, the NeRF field can be rendered at the preset viewing angle to obtain a rendered image.

[0046] After obtaining the rendered image of the NeRF at the preset viewing angle, noise (such as Gaussian noise, etc.) can be added to the rendered image to obtain a noise image. Among them, the preset viewing angle may include at least one viewing angle; correspondingly, the rendered image may include at least one image, and the noise image may also include at least one image.

[0047] Exemplarily, Figure 2 is a schematic framework diagram of a method for generating three-dimensional data provided in the embodiments of the present disclosure. Refer to Figure 2 , rendering the NeRF field at the preset viewing angle can obtain a rendered image a; adding noise b to the rendered image can obtain a noise image c.

[0048] S130. Through a diffusion network, cross-attention fusion is performed on the features of the first image, the features of the description text, and the features of the noise image to generate a second image; wherein, the second image includes an image of the target object at a preset viewing angle.

[0049] In the embodiments of the present disclosure, the diffusion network may include a two-dimensional diffusion network, and the diffusion network may be a pre-constructed network model (i.e., the parameters in the diffusion network in this embodiment are fixed). Referring again to Figure 2 , the diffusion network can perform inverse diffusion on the input noise image c to generate a second image e. Since the noise image c is generated based on the rendered image a under a preset perspective, the second image e also belongs to the image under the preset perspective.

[0050] Figure 2 In , during the process of the diffusion network generating the second image, the features of the extracted noise image c can be cross-attention fused with the features of the first image d and the features of the description text t, so that the generated second image is similar to the first image and the description text.

[0051] Exemplarily, Figure 3 is a schematic diagram of feature cross-attention fusion in a method for generating three-dimensional data provided by an embodiment of the present disclosure. Referring to Figure 3 , in some alternative implementation manners, cross-attention fusing the features of the first image, the features of the description text, and the features of the noise image through the diffusion network may include: performing feature processing on the noise image through each denoising sub-network in the diffusion network; performing feature processing on the first image through a mirror network having the same structure as the diffusion network; extracting the features of the description text through a text feature extraction network; and cross-attention fusing the features of the noise image processed by the denoising sub-network in the diffusion network, the features of the first image processed by the corresponding denoising sub-network in the mirror network, and the features of the description text through the cross-attention sub-network in the diffusion network.

[0052] Referring to Figure 3 , the diffusion network may include multiple denoising sub-networks (such as the U-Net network in Figure 3 ), and the input noise image can be subjected to feature processing based on the denoising sub-network. Figure 3In this case, the first image can be feature - processed through a mirror network with the same structure as the diffusion network to obtain the processed features of the first image; the description text can be feature - extracted based on an existing text feature extraction network to obtain the features of the description text. In some optional implementation manners, the feature - processing of the noise image by each denoising sub - network in the diffusion network may include: the noise image is feature - extracted by an image encoder to obtain a noise feature map, and then the noise feature map is feature - processed by each denoising sub - network in the diffusion network; the feature - processing of the first image by a mirror network with the same structure as the diffusion network may include: the first image is feature - extracted by an image encoder to obtain a first feature map, and then the first feature map is feature - processed by a mirror network with the same structure as the diffusion network. It should be noted that the mirror network and the diffusion network having the same structure may mean that they include the same number of denoising sub - networks.

[0053] Figure 3 In this case, the diffusion network may further include a cross - attention sub - network; through the cross - attention sub - network, the features of the processed noise image in each denoising sub - network can be cross - attention fused with the features of the first image processed by the corresponding denoising sub - network in the mirror network and the features of the description text. That is, the features of the noise image processed in the i - th denoising sub - network in the diffusion network can be cross - attention fused with the features of the first image processed in the i - th denoising sub - network in the mirror network and the features of the description text.

[0054] In some optional implementation manners, the cross - attention fusion of the features of the noise image processed by the denoising sub - network in the diffusion network, the features of the first image processed by the corresponding denoising sub - network in the mirror network, and the features of the description text may include: cross - attention fusing the features of the noise image processed by at least some of the denoising sub - networks in the diffusion network with the features of the first image processed by the corresponding denoising sub - network in the mirror network and the features of the description text. And, in the diffusion network, for each denoising sub - network that needs to perform cross - attention fusion, a cross - attention sub - network may be correspondingly deployed.

[0055] In some optional implementation manners, the cross - attention fusion of the features of the noise image processed by the denoising sub - network in the diffusion network, the features of the first image processed by the corresponding denoising sub - network in the mirror network, and the features of the description text may include: cross - attention fusing the features of the noise image processed by at least some network layers of the denoising sub - network in the diffusion network with the features of the first image processed by the corresponding network layers of the corresponding denoising sub - network in the mirror network and the features of the description text.

[0056] In the embodiments of the present disclosure, the first image is processed by using a mirror network consistent with the diffusion network, and intermediate features of the noisy image processed by at least part of the denoising sub-networks in the mirror network are fused into the corresponding denoising sub-networks in the diffusion network through a cross-attention sub-network. At the same time, the cross-attention sub-network can also fuse the features of the description text into the corresponding denoising sub-networks in the diffusion network, so that the generated second image has a high similarity with the first image and the description text.

[0057] Moreover, at least part of the denoising sub-networks that need to perform cross-attention fusion and at least one network layer in the denoising sub-networks that need to perform cross-attention fusion can be set according to the actual application scenario, so as to save network computing amount to a certain extent and improve the generation efficiency of the second image while ensuring the similarity between the second image and the first image and the description text.

[0058] S140. Determine a score distillation loss according to the second image, the noisy image, and the added noise, and adjust the neural radiance field according to the score distillation loss.

[0059] Refer to again Figure 2 , the second image e and the noisy image c can be compared to determine the added noise f predicted by the diffusion network. The score distillation loss (Score Distillation Sampling, SDS) can be determined according to the predicted added noise f and the actually added noise b. Among them, the closer the feature distribution of the picture rendered based on NeRF is to the second image, the closer the noise predicted by the diffusion network is to the actually added noise. Therefore, the parameters in NeRF can be adjusted based on the score distillation loss, so that the rendering result of NeRF has a similarity with the feature distribution of the second image. That is, an implicit expression of the three-dimensional model can be constructed based on the loss of the two-dimensional image.

[0060] S150. Determine a three-dimensional model of the target object according to the adjusted neural radiance field.

[0061] In the embodiments of the present disclosure, the three-dimensional model may include a three-dimensional mesh model. An initial three-dimensional model can be extracted based on the constructed NeRF based on existing algorithms. For example, an initial three-dimensional model can be extracted based on the Deep Marching Tetrahedra (DMTet) algorithm based on the constructed NeRF. In addition, a high-resolution three-dimensional model can also be generated based on existing algorithms on the basis of the initial three-dimensional model. The generated three-dimensional model of the target object can have a high similarity with the first image and the description text. Compared with the generation of traditional three-dimensional models, the technical solution of the embodiments of the present disclosure can quickly and efficiently generate high-quality three-dimensional models according to the first image and the description text.

[0062] The technical solution of the embodiment of the present disclosure is to obtain a first image of a target object and a description text corresponding to the first image; obtain a noise image; wherein, the noise image includes a rendered image with added noise; the rendered image includes an image rendered by a neural radiance field at a preset viewing angle; through a diffusion network, cross-attention fusion is performed on the features of the first image, the features of the description text, and the features of the noise image to generate a second image; wherein, the second image includes an image of the target object at a preset viewing angle; determine a score distillation loss according to the second image, the noise image, and the added noise, and adjust the neural radiance field according to the score distillation loss; determine a three-dimensional model of the target object according to the adjusted neural radiance field.

[0063] By performing cross-attention fusion on the features of the first image, the features of the description text, and the features of the noise image based on the diffusion network, it is possible to generate a second image based on the noise image with the first image and the description text as control conditions. By determining the rendered image at a preset viewing angle based on the neural radiance field, generating a noise image for input to the diffusion network from the rendered image in a way of adding noise, and constructing the neural radiance field through the score distillation loss, it is possible to make the three-dimensional model expressed based on the neural radiance field have the same feature distribution as the second image generated based on the first image and the description text. Thus, it is possible to quickly and efficiently generate high-quality three-dimensional data according to the first image and the description text.

[0064] The embodiments of the present disclosure can be combined with various alternative solutions in the method for generating three-dimensional data provided in the above embodiments. In the method for generating three-dimensional data provided in this embodiment, when the target object includes a target human object, the control conditions for generating the second image are described in detail. By adding a third image in a preset pose as a control condition for generating the second image, it is possible to make the generated three-dimensional model also in the preset pose, which is beneficial to subsequent further processing of the three-dimensional model, such as beneficial to bone binding and other processing.

[0065] Figure 4 It is a flowchart of a method for generating three-dimensional data provided by an embodiment of the present disclosure. As Figure 4 shown, the method for generating three-dimensional data provided in this embodiment, when the target object includes a target human object, may include:

[0066] S410. Obtain a first image of the target human object and a description text corresponding to the first image.

[0067] S420. Obtain a noise image; wherein, the noise image includes a rendered image with added noise; the rendered image includes an image rendered by a neural radiance field at a preset viewing angle.

[0068] S430. Obtain a third image of a preset three-dimensional model from a preset perspective; wherein, the preset three-dimensional model includes a preset human body model in a preset pose.

[0069] In this embodiment, the preset three-dimensional model may include a three-dimensional mesh model. Among them, a preset three-dimensional model can be generated based on an existing neural network conditional on pose control parameters. For example, through the (Skinned Multi-Person Linear eXpressive, SMPL-X) model, according to the pose control parameters of each preset region, the pre-configured standard model can be deformed to generate a preset human body model in a preset pose.

[0070] Among them, the preset three-dimensional model can be rendered from each preset perspective to obtain each third image; wherein, each third image can present the entire surface of the preset human body. Exemplarily, Figure 5 is a schematic diagram of the third image in a method for generating three-dimensional data provided by an embodiment of the present disclosure. Refer to Figure 5 , the preset three-dimensional model is in a natural standing pose, and third images of the preset three-dimensional model from the front perspective and the back perspective can be obtained.

[0071] S440. Based on the features of the first image, the features of the descriptive text, the features of the noise image, and the features of the third image, generate a second image through a diffusion network; wherein, the second image includes an image of the target human body object in a preset pose from a preset perspective.

[0072] In this embodiment, while cross-attention fusing the features of the first image, the features of the descriptive text, and the features of the noise image through the diffusion network, the features of the third image can also be extracted and introduced into the denoising sub-network of the diffusion network. Thus, the generated second image can, while maintaining similarity with the first image and the text description, also have the same preset pose as the preset human body model in the third image, thereby enabling adjustment of the pose presentation of the three-dimensional model.

[0073] S450. Determine a score distillation loss based on the second image, the noise image, and the added noise, and adjust the neural radiance field according to the score distillation loss.

[0074] S460. Determine the three-dimensional model of the target human body object according to the adjusted neural radiance field.

[0075] In some alternative implementation manners, after determining the three-dimensional model of the target human body object, it may further include: binding the three-dimensional model of the target human body object to a skeletal model in a preset pose; wherein, the three-dimensional model of the bound skeletal model is driven based on pose control parameters.

[0076] In these alternative implementations, the preset pose can be, for example, the pose presented by an existing skeletal model. Since the three-dimensional model presents the same pose as the skeletal model, it is very convenient to align the standard key points of the three-dimensional model with the skeletal model. Among them, an existing binding tool can be used to bind the three-dimensional model and the skeletal model, and then the skeletal model can be driven by pose control parameters to generate a dynamic three-dimensional model.

[0077] In the technical solution of the embodiment of the present disclosure, when the target object includes a target human object, the control conditions for generating the second image are described in detail. By adding a third image presenting a preset pose as the control condition for generating the second image, the generated three-dimensional model can also present the preset pose, which is beneficial to subsequent further processing of the three-dimensional model, such as skeletal binding and other processing. The method for generating three-dimensional data provided by the embodiment of the present disclosure and the method for generating three-dimensional data provided by the above embodiment belong to the same general inventive concept. Technical details not described in detail in this embodiment can be referred to in the above embodiment, and the same technical features have the same beneficial effects in this embodiment and the above embodiment.

[0078] The various alternative solutions in the method for generating three-dimensional data provided by the embodiment of the present disclosure and the above embodiment can be combined. The method for generating three-dimensional data provided by this embodiment describes the application scenario in detail. By detecting and segmenting human objects and scene objects in the target video to obtain a segmented image, the segmented image can be used as a control condition for generating a three-dimensional model, so that the generated three-dimensional model has similarity with the human object and scene object in the target video. Based on this, the construction and rendering of a three-dimensional scene similar to the target video can be realized.

[0079] Figure 6 It is a schematic flowchart of a method for generating three-dimensional data provided by an embodiment of the present disclosure. As Figure 6 shown, the method for generating three-dimensional data provided by this embodiment may include:

[0080] S610. Detect objects in the target video to determine video frames containing the target object.

[0081] In this embodiment, based on existing image processing algorithms, objects (such as scene objects and human objects, etc.) in each frame of the target video can be detected and tracked. When a new object is detected, the detected object instance can be saved, and at the same time, the video frame containing the object instance can be recorded. Among them, at least one of the detected objects can be determined as the target object. Thus, the video frames containing the target object in the target video can be determined.

[0082] S620. Determine the target video frame from the video frames containing the target object.

[0083] Among them, based on a preset selection strategy, a target video frame can be determined from each video containing the target object. For example, the video frame with the highest proportion of the target object area in each video frame can be determined as the target video frame; or, the video frame with the highest clarity of the target object in each video frame can be determined as the target video frame. In addition, other selection strategies can also be applied here, and they will not be enumerated one by one.

[0084] S630. Segment the target video frame to obtain the first image of the target object.

[0085] In this embodiment, based on an existing image segmentation algorithm, the target object area in the target video frame can be segmented to obtain the first image of the target object.

[0086] S640. Through an image description network, based on the first image of the target object, generate a description text corresponding to the first image.

[0087] In this embodiment, the image description network may include an existing image annotation generation network. For example, it may include a model based on Transformer, such as a Bidirectional Encoder Representations from Transformers (BERT) based on Transformer, etc. The first image can be input into the image description network so that it outputs the corresponding description text. Thus, the first image and the description text of the target object in the target video frame can be obtained.

[0088] S650. Obtain a noise image; among them, the noise image includes a rendered image with added noise; the rendered image includes an image rendered from a neural radiance field at a preset viewing angle.

[0089] S660. Through a diffusion network, perform cross-attention fusion on the features of the first image, the features of the description text, and the features of the noise image to generate a second image; among them, the second image includes an image of the target object at a preset viewing angle.

[0090] Among them, when the target object includes a target human object, the method may further include: obtaining a third image of a preset three-dimensional model at a preset viewing angle; among them, the preset three-dimensional model includes a preset human model in a preset pose;

[0091] Correspondingly, through the diffusion network, cross-attention fusion of the features of the first image, the features of the descriptive text, and the features of the noise image is performed to generate a second image, which may include: generating a second image through the diffusion network based on the features of the first image, the features of the descriptive text, the features of the noise image, and the features of a third image; wherein, the second image includes an image of the target human object in a preset pose under a preset perspective.

[0092] S670. Determine a score distillation loss according to the second image, the noise image, and the added noise, and adjust the neural radiance field according to the score distillation loss.

[0093] S680. Determine a three-dimensional model of the target object according to the adjusted neural radiance field.

[0094] Among them, the corresponding neural radiance field can be adjusted for each object to determine the three-dimensional models corresponding to the respective target objects.

[0095] S690. Generate a three-dimensional scene corresponding to the target video according to the three-dimensional model of the target object.

[0096] Among them, the three-dimensional models corresponding to the target objects in the target video can be used to build a three-dimensional scene according to the spatial positions, viewing directions, etc. of the target objects in the target video, so as to generate a three-dimensional scene with a high similarity to the target video. Thus, it is possible to quickly build a corresponding three-dimensional scene based on a two-dimensional video, with a wide range of application prospects and meeting the user's needs for three-dimensional scene building.

[0097] In addition, preset three-dimensional model materials can also be placed in the built three-dimensional scene to enrich the three-dimensional scene. It is also possible to perform three-dimensional scene rendering with different perspectives, lighting, etc. from the target video to further enrich the presentation of the three-dimensional scene.

[0098] In some alternative implementation manners, in the case where the target object includes a target human object, the method for generating three-dimensional data may further include: performing motion estimation on the target human object in the target video to determine a sequence of pose control parameters; driving the three-dimensional model of the target human object according to the sequence of pose control parameters.

[0099] Among them, based on an existing human pose estimation model (such as a dense pose model, etc.), motion estimation can be performed on the target human object in the video frame corresponding to the target human object to determine the pose parameters of each joint point of the target human object in each frame (such as position, angle, etc. parameters), so as to obtain a sequence of pose control parameters for each joint point.

[0100] After determining the attitude control parameter sequence, the attitude control sequences of each joint point can be loaded into the corresponding 3D model, and then the driving of the 3D model of the target human object can be realized. It can realize the movement of the 3D model of the target human object in the 3D scene, which has high similarity with the target video. Thus, the construction of a dynamic scene can be realized, and the user experience can be improved.

[0101] The technical solution of the embodiment of the present disclosure describes the application scenario in detail. By detecting and segmenting the human object and the scene object in the target video to obtain the segmented image, the segmented image can be used as the control condition for generating the 3D model, so that the generated 3D model has similarity with the human object and the scene object in the target video. Based on this, the construction and rendering of a 3D scene similar to the target video can be realized. The method for generating 3D data provided by the embodiment of the present disclosure and the method for generating 3D data provided by the above embodiment belong to the same general concept. The technical details not described in detail in this embodiment can be referred to the above embodiment, and the same technical features have the same beneficial effects in this embodiment and the above embodiment.

[0102] Figure 7 It is a schematic structural diagram of a device for generating 3D data provided by an embodiment of the present disclosure. The device for generating 3D data provided by this embodiment is applicable to the situation of generating a similar 3D model according to graphics and text.

[0103] As Figure 7 shown, the device for generating 3D data provided by the embodiment of the present disclosure may include:

[0104] The first acquisition module 710 is configured to acquire the first image of the target object and the description text corresponding to the first image;

[0105] The second acquisition module 720 is configured to acquire a noise image; wherein, the noise image includes a rendered image with added noise; the rendered image includes an image rendered by the neural radiance field at a preset viewing angle;

[0106] The diffusion module 730 is configured to perform cross-attention fusion on the features of the first image, the features of the description text, and the features of the noise image through a diffusion network to generate a second image; wherein, the second image includes an image of the target object at a preset viewing angle;

[0107] The 3D representation module 740 is configured to determine a score distillation loss according to the second image, the noise image, and the added noise, and adjust the neural radiance field according to the score distillation loss;

[0108] The model generation module 750 is configured to determine the 3D model of the target object according to the adjusted neural radiance field.

[0109] In some alternative implementation manners, the diffusion module can be used for:

[0110] Performing feature processing on a noisy image through each denoising sub-network in the diffusion network;

[0111] Performing feature processing on a first image through a mirror network having the same structure as the diffusion network;

[0112] Performing feature extraction on the description text through a text feature extraction network;

[0113] Performing cross-attention fusion on the features of the noisy image processed by the denoising sub-network in the diffusion network, the features of the first image processed by the corresponding denoising sub-network in the mirror network, and the features of the description text through the cross-attention sub-network in the diffusion network.

[0114] In some alternative implementation manners, the diffusion module can be used for:

[0115] Performing cross-attention fusion on the features of the noisy image processed by at least part of the denoising sub-networks in the diffusion network, the features of the first image processed by the corresponding denoising sub-networks in the mirror network, and the features of the description text.

[0116] In some alternative implementation manners, the diffusion module can be used for:

[0117] Performing cross-attention fusion on the features of the noisy image processed by at least part of the network layers in the denoising sub-network, the features of the first image processed by the corresponding network layers in the corresponding denoising sub-network, and the features of the description text.

[0118] In some alternative implementation manners, when the target object includes a target human object, the three-dimensional data generation device may further include:

[0119] A third acquisition module, configured to acquire a third image of a preset three-dimensional model from a preset perspective; wherein, the preset three-dimensional model includes a preset human model in a preset pose;

[0120] Correspondingly, the diffusion module can be used for:

[0121] Generating a second image through the diffusion network based on the features of the first image, the features of the description text, the features of the noisy image, and the features of the third image; wherein, the second image includes an image of the target human object in the preset pose from the preset perspective.

[0122] In some alternative implementation manners, the three-dimensional data generation device may further include:

[0123] The model-driven module is used to bind the three-dimensional model of the target human object to the skeleton model in a preset pose after determining the three-dimensional model of the target human object;

[0124] Among them, the three-dimensional model of the bound skeleton model is driven based on the pose control parameters.

[0125] In some optional implementation manners, the first acquisition module can be used to:

[0126] Detect the object in the target video to determine the video frame containing the target object;

[0127] Determine the target video frame from the video frames containing the target object;

[0128] Segment the target video frame to obtain the first image of the target object.

[0129] In some optional implementation manners, the first acquisition module can also be used to:

[0130] Generate the description text corresponding to the first image based on the first image of the target object through the image description network.

[0131] In some optional implementation manners, the model-driven module can be used to:

[0132] Perform motion estimation on the target human object in the target video to determine the pose control parameter sequence;

[0133] Drive the three-dimensional model of the target human object according to the pose control parameter sequence.

[0134] In some optional implementation manners, the model generation module can also be used to:

[0135] Generate a three-dimensional scene corresponding to the target video according to the three-dimensional model of the target object.

[0136] The three-dimensional data generation device provided by the embodiments of the present disclosure can execute the three-dimensional data generation method provided by any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects for executing the method.

[0137] It should be noted that the various units and modules included in the above device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the embodiments of the present disclosure.

[0138] Next, refer to Figure 8 , which shows an electronic device suitable for implementing the embodiments of the present disclosure (for example Figure 8Schematic diagram of the structure of the terminal device or server) 800 in. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The electronic device shown is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present disclosure.

[0139] As Figure 8 shown, the electronic device 800 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 801, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 802 or the program loaded from the storage device 808 into the random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 are also stored. The processing device 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. The input / output (I / O) interface 805 is also connected to the bus 804.

[0140] Generally, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 8 the electronic device 800 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be implemented or had alternatively.

[0141] Particularly, according to the embodiments of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, the embodiments of the present disclosure include a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above-mentioned functions defined in the method for generating three-dimensional data of the embodiments of the present disclosure are executed.

[0142] The electronic device provided by an embodiment of the present disclosure and the method for generating three-dimensional data provided by the above embodiment belong to the same general inventive concept. Technical details not described in detail in this embodiment can be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0143] An embodiment of the present disclosure provides a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method for generating three-dimensional data provided by the above embodiment.

[0144] It should be noted that the computer-readable medium of the present disclosure described above can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), or a flash memory (FLASH), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0145] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0146] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.

[0147] The above computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to:

[0148] Obtain a first image of a target object and the description text corresponding to the first image; obtain a noise image; wherein the noise image includes a rendered image with added noise; the rendered image includes an image rendered from a neural radiance field at a preset viewing angle; through a diffusion network, perform cross-attention fusion on the features of the first image, the features of the description text, and the features of the noise image to generate a second image; wherein the second image includes an image of the target object at a preset viewing angle; determine a score distillation loss based on the second image, the noise image, and the added noise, and adjust the neural radiance field according to the score distillation loss; determine a three-dimensional model of the target object based on the adjusted neural radiance field.

[0149] Computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof. The above programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).

[0150] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0151] The units involved in the embodiments described in the present disclosure can be implemented in software or in hardware. Among them, the names of the units and modules do not, in some cases, constitute a limitation on the units and modules themselves.

[0152] The functions described above in this document can be at least partially performed by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.

[0153] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0154] According to one or more embodiments of the present disclosure, a method for generating three-dimensional data is provided, the method comprising:

[0155] Obtaining a first image of a target object and a description text corresponding to the first image;

[0156] Obtaining a noise image; wherein the noise image includes a rendered image with added noise; the rendered image includes an image rendered from a neural radiance field at a preset viewing angle;

[0157] Through a diffusion network, cross-attention fusion is performed on the features of the first image, the features of the description text, and the features of the noise image to generate a second image; wherein the second image includes an image of the target object at the preset viewing angle;

[0158] Determining a score distillation loss based on the second image, the noise image, and the added noise, and adjusting the neural radiance field according to the score distillation loss;

[0159] Determining a three-dimensional model of the target object according to the adjusted neural radiance field.

[0160] According to one or more embodiments of the present disclosure, a method for generating three-dimensional data further comprises:

[0161] In some alternative implementations, the performing cross-attention fusion on the features of the first image, the features of the description text, and the features of the noise image through a diffusion network includes:

[0162] Performing feature processing on the noise image through each denoising sub-network in the diffusion network;

[0163] Performing feature processing on the first image through a mirror network having the same diffusion network structure;

[0164] Performing feature extraction on the description text through a text feature extraction network;

[0165] Performing cross-attention fusion on the features of the noise image processed by the denoising sub-network in the diffusion network, the features of the first image processed by the corresponding denoising sub-network in the mirror network, and the features of the description text through the cross-attention sub-network in the diffusion network.

[0166] According to one or more embodiments of the present disclosure, a method for generating three-dimensional data is provided, further including:

[0167] In some alternative implementation manners, the performing cross-attention fusion on the features of the noise image processed by the denoising sub-network in the diffusion network, the features of the first image processed by the corresponding denoising sub-network in the mirror network, and the features of the description text includes:

[0168] Performing cross-attention fusion on the features of at least part of the noise images processed by the denoising sub-networks in the diffusion network, the features of the first images processed by the corresponding denoising sub-networks in the mirror network, and the features of the description text.

[0169] According to one or more embodiments of the present disclosure, a method for generating three-dimensional data is provided, further including:

[0170] In some alternative implementation manners, the performing cross-attention fusion on the features of the noise image processed by the denoising sub-network in the diffusion network, the features of the first image processed by the corresponding denoising sub-network in the mirror network, and the features of the description text includes:

[0171] Performing cross-attention fusion on the features of the noise images processed by at least part of the network layers in the denoising sub-network, the features of the first images processed by the corresponding network layers in the corresponding denoising sub-network, and the features of the description text.

[0172] According to one or more embodiments of the present disclosure, a method for generating three-dimensional data is provided, further including:

[0173] In some alternative implementation manners, when the target object includes a target human object, the method further includes:

[0174] Obtaining a third image of a preset three-dimensional model at the preset viewing angle; wherein, the preset three-dimensional model includes a preset human model in a preset pose;

[0175] Correspondingly, the cross-attention fusion of the features of the first image, the features of the description text, and the features of the noise image through the diffusion network to generate a second image includes:

[0176] Generating a second image through a diffusion network based on the features of the first image, the features of the description text, the features of the noise image, and the features of the third image; wherein, the second image includes an image of the target human object in the preset pose from the preset perspective.

[0177] According to one or more embodiments of the present disclosure, a method for generating three-dimensional data is provided, further including:

[0178] In some optional implementation manners, after determining the three-dimensional model of the target human object, it further includes:

[0179] Binding the three-dimensional model of the target human object to the bone model in the preset pose;

[0180] Wherein, the three-dimensional model of the bound bone model is driven based on pose control parameters.

[0181] According to one or more embodiments of the present disclosure, a method for generating three-dimensional data is provided, further including:

[0182] In some optional implementation manners, the obtaining of the first image of the target object includes:

[0183] Detecting the object in the target video to determine the video frame containing the target object;

[0184] Determining the target video frame from the video frames containing the target object;

[0185] Segmenting the target video frame to obtain the first image of the target object.

[0186] According to one or more embodiments of the present disclosure, a method for generating three-dimensional data is provided, further including:

[0187] In some optional implementation manners, obtaining the description text corresponding to the first image includes:

[0188] Generating the description text corresponding to the first image through an image description network based on the first image of the target object.

[0189] According to one or more embodiments of the present disclosure, a method for generating three-dimensional data is provided, further including:

[0190] In some optional implementation manners, when the target object includes a target human object, the method further includes:

[0191] Perform motion estimation on the target human object in the target video to determine a sequence of pose control parameters;

[0192] Drive the 3D model of the target human object according to the sequence of pose control parameters.

[0193] According to one or more embodiments of the present disclosure, a method for generating 3D data is provided, which further includes:

[0194] In some alternative implementation manners, after determining the 3D model of the target object, it further includes:

[0195] Generate a 3D scene corresponding to the target video according to the 3D model of the target object.

[0196] According to one or more embodiments of the present disclosure, a device for generating 3D data is provided, and the device includes:

[0197] A first acquisition module, configured to acquire a first image of a target object and a description text corresponding to the first image;

[0198] A second acquisition module, configured to acquire a noise image; wherein, the noise image includes a rendered image with added noise; the rendered image includes an image rendered by a neural radiance field at a preset viewing angle;

[0199] A diffusion module, configured to perform cross-attention fusion on the features of the first image, the features of the description text, and the features of the noise image through a diffusion network to generate a second image; wherein, the second image includes an image of the target object at the preset viewing angle;

[0200] A 3D representation module, configured to determine a score distillation loss according to the second image, the noise image, and the added noise, and adjust the neural radiance field according to the score distillation loss;

[0201] A model generation module, configured to determine a 3D model of the target object according to the adjusted neural radiance field.

[0202] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principle. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solution formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0203] Moreover, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented separately or in any suitable subcombination in multiple embodiments.

[0204] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A method for generating three-dimensional data, characterized in that, comprising: obtaining a first image of a target object and a description text corresponding to the first image; obtaining a noise image; wherein, the noise image includes a rendered image with added noise; the rendered image includes an image rendered from a neural radiance field at a preset viewing angle; through a diffusion network, cross-attention fusion is performed on the features of the first image, the features of the description text, and the features of the noise image to generate a second image; wherein, the second image includes an image of the target object at the preset viewing angle; determining a score distillation loss according to the second image, the noise image, and the added noise, and adjusting the neural radiance field according to the score distillation loss; determining a three-dimensional model of the target object according to the adjusted neural radiance field.

2. The method according to claim 1, characterized in that, the cross-attention fusion of the features of the first image, the features of the description text, and the features of the noise image through the diffusion network includes: performing feature processing on the noise image through each denoising sub-network in the diffusion network; performing feature processing on the first image through a mirror network having the same structure as the diffusion network; performing feature extraction on the description text through a text feature extraction network; through the cross-attention sub-network in the diffusion network, cross-attention fusion is performed on the features of the noise image processed by the denoising sub-network in the diffusion network, the features of the first image processed by the corresponding denoising sub-network in the mirror network, and the features of the description text.

3. The method according to claim 2, characterized in that, the cross-attention fusion of the features of the noise image processed by the denoising sub-network in the diffusion network, the features of the first image processed by the corresponding denoising sub-network in the mirror network, and the features of the description text includes: performing cross-attention fusion on the features of the noise image processed by at least part of the denoising sub-networks in the diffusion network, the features of the first image processed by the corresponding denoising sub-networks in the mirror network, and the features of the description text.

4. The method according to claim 2, characterized in that, the cross-attention fusion of the features of the noise image processed by the denoising sub-network in the diffusion network, the features of the first image processed by the corresponding denoising sub-network in the mirror network, and the features of the description text includes: performing cross-attention fusion on the features of the noise image processed by at least part of the network layers in the denoising sub-network, the features of the first image processed by the corresponding network layers in the corresponding denoising sub-network, and the features of the description text.

5. The method according to claim 1, characterized in that, when the target object includes a target human object, the method further includes: obtaining a third image of a preset three-dimensional model at the preset viewing angle; wherein, the preset three-dimensional model includes a preset human model in a preset pose; Correspondingly, the cross-attention fusion of the features of the first image, the features of the description text, and the features of the noise image through the diffusion network to generate a second image includes: Generating a second image through a diffusion network based on the features of the first image, the features of the description text, the features of the noise image, and the features of the third image; wherein, the second image includes an image of the target human object in the preset pose from the preset perspective.

6. The method according to claim 5, wherein, after determining the three-dimensional model of the target human object, further comprising: Binding the three-dimensional model of the target human object to the bone model in the preset pose; wherein, the three-dimensional model of the bound bone model is driven based on the pose control parameters.

7. The method according to claim 1, wherein, the obtaining of the first image of the target object includes: Detecting the object in the target video to determine the video frame containing the target object; Determining the target video frame from the video frames containing the target object; Segmenting the target video frame to obtain the first image of the target object.

8. The method according to claim 7, wherein, obtaining the description text corresponding to the first image includes: Generating the description text corresponding to the first image through an image description network based on the first image of the target object.

9. The method according to claim 7, wherein, when the target object includes a target human object, the method further comprises: Performing motion estimation on the target human object in the target video to determine a sequence of pose control parameters; Driving the three-dimensional model of the target human object according to the sequence of pose control parameters.

10. The method according to claim 7, wherein, after determining the three-dimensional model of the target object, further comprising: Generating a three-dimensional scene corresponding to the target video according to the three-dimensional model of the target object.

11. A three-dimensional data generation device, wherein, comprising: A first acquisition module for acquiring the first image of the target object and the description text corresponding to the first image; A second acquisition module for acquiring a noise image; wherein, the noise image includes a rendered image with added noise; the rendered image includes an image rendered from a neural radiance field from a preset perspective; A diffusion module for performing cross-attention fusion on the features of the first image, the features of the description text, and the features of the noise image through a diffusion network to generate a second image; wherein, the second image includes an image of the target object from the preset perspective; A three-dimensional representation module for determining a score distillation loss according to the second image, the noise image, and the added noise, and adjusting the neural radiance field according to the score distillation loss; A model generation module for determining the three-dimensional model of the target object according to the adjusted neural radiance field.

12. An electronic device, wherein, the electronic device includes: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method for generating three-dimensional data as described in any one of claims 1-10.

13. A storage medium containing computer-executable instructions that, when executed by a computer processor, are used to execute the method for generating three-dimensional data as described in any one of claims 1-10.

Citation Information

Patent Citations

  • Three-dimensional model generation method and device, computer equipment and storage medium

    CN116824092A

  • Electronic device and method for generating an image

    KR1020220017109A

  • Three-dimensional scene rendering method, device, and storage medium

    WO2023138471A1

Cited By

  • Method and device for generating three-dimensional model, electronic equipment and storage medium

    CN120953504A