Three-dimensional data generation method and apparatus, electronic device, and storage medium

By integrating image features, description text and noise image features in three-dimensional content creation through diffusion networks, a high-quality three-dimensional model is generated, which solves the cumbersome and complex problems of the traditional three-dimensional content creation process and achieves rapid and efficient three-dimensional data generation.

WO2025113388A1PCT designated stage expired Publication Date: 2025-06-05BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/134286
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-29
Filing Date
2024-11-25
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

The traditional three-dimensional content creation process is complicated and requires a lot of manpower and material costs, making it difficult to quickly and efficiently generate high-quality three-dimensional data.

Method used

The image features, description text features and noise image features of the target object are fused through the diffusion network to generate a second image, and the fractional distillation loss is determined based on the second image, the noise image and the added noise, and the neural radiation field is adjusted to determine the three-dimensional model of the target object.

Benefits of technology

It realizes the rapid and efficient generation of high-quality three-dimensional models based on the first image and description text, simplifies the three-dimensional content creation process and reduces the human and material costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024134286_05062025_PF_FP_ABST
    Figure CN2024134286_05062025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose three-dimensional data generation method and apparatus, an electronic device, and a storage medium. The method comprises: acquiring a first image of a target object, and a description text corresponding to the first image; acquiring a noise image, the noise image comprising a rendered image to which noise is added, and the rendered image comprising an image obtained by rendering a neural radiance field at a preset angle of view; by means of a diffusion network, performing cross-attention fusion on features of the first image, features of the description text and features of the noise image, to generate a second image, the second image comprising an image of the target object at the preset viewing angle; on the basis of the second image, the noise image and the added noise, determining a score distillation loss, and on the basis of the score distillation loss, adjusting the neural radiance field; and on the basis of the adjusted neural radiance field, determining a three-dimensional model of the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, electronic device and storage medium for generating three-dimensional data

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to Chinese patent application No. 2023116192922, filed on November 29, 2023, entitled “A method, device, electronic device and storage medium for generating three-dimensional data”. The entire contents of that application are incorporated herein by reference. Technical Field

[0003] The embodiments of the present disclosure relate to the field of computer technology, and in particular to a method, device, electronic device, and storage medium for generating three-dimensional data. Background Art

[0004] The traditional three-dimensional (3D) content creation process is tedious and complicated, requiring a lot of manpower and material resources. Summary of the Invention

[0005] Embodiments of the present disclosure provide a method, device, electronic device, and storage medium for generating three-dimensional data.

[0006] In a first aspect, an embodiment of the present disclosure provides a method for generating three-dimensional data, comprising:

[0007] Acquire a first image of a target object and a description text corresponding to the first image;

[0008] Acquire a noise image; wherein the noise image includes a rendered image with added noise; the rendered image includes an image rendered from a neural radiation field at a preset viewing angle;

[0009] Cross-attentionally fusing features of the first image, features of the description text, and features of the noise image through a diffusion network to generate a second image; wherein the second image includes an image of the target object at the preset viewing angle;

[0010] determining a fractional distillation loss based on the second image, the noisy image, and the added noise, and adjusting the neural radiation field based on the fractional distillation loss;

[0011] A three-dimensional model of the target object is determined according to the adjusted neural radiation field.

[0012] In a second aspect, an embodiment of the present disclosure further provides a device for generating three-dimensional data, including:

[0013] A first acquisition module is used to acquire a first image of a target object and a description text corresponding to the first image;

[0014] A second acquisition module is configured to acquire a noise image; wherein the noise image includes a rendered image with added noise; and the rendered image includes an image rendered from a neural radiation field at a preset viewing angle;

[0015] a diffusion module, configured to perform cross-attention fusion on features of the first image, features of the description text, and features of the noise image through a diffusion network to generate a second image; wherein the second image includes an image of the target object at the preset viewing angle;

[0016] a three-dimensional representation module, configured to determine a fractional distillation loss based on the second image, the noisy image, and the added noise, and adjust the neural radiation field based on the fractional distillation loss;

[0017] A model generation module is used to determine a three-dimensional model of the target object according to the adjusted neural radiation field.

[0018] In a third aspect, an embodiment of the present disclosure further provides an electronic device, the electronic device comprising:

[0019] one or more processors;

[0020] a storage device for storing one or more programs,

[0021] When the one or more programs are executed by the one or more processors, the one or more processors implement the method for generating three-dimensional data as described in any one of the embodiments of the present disclosure.

[0022] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to execute the method for generating three-dimensional data as described in any one of the embodiments of the present disclosure.

[0023] The technical solution of the embodiment of the present disclosure is to obtain a first image of a target object and a descriptive text corresponding to the first image; obtain a noise image; wherein the noise image includes a rendered image with added noise; the rendered image includes an image rendered by a neural radiation field at a preset perspective; through a diffusion network, cross-attentionally fuse the features of the first image, the features of the descriptive text, and the features of the noise image to generate a second image; wherein the second image includes an image of the target object at a preset perspective; determine a fractional distillation loss based on the second image, the noise image, and the added noise, and adjust the neural radiation field based on the fractional distillation loss; and determine a three-dimensional model of the target object based on the adjusted neural radiation field.

[0024] By cross-attentionally fusing the features of the first image, the descriptive text, and the noise image using a diffusion network, a second image can be generated from the noise image, using the first image and descriptive text as control conditions. By determining a rendered image at a preset perspective based on the neural radiance field, adding noise to the rendered image to generate a noise image for input into the diffusion network, and constructing the neural radiance field using a fractional distillation loss, the 3D model expressed based on the neural radiance field can have the same feature distribution as the second image generated from the first image and descriptive text. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.

[0026] FIG1 is a schematic flow chart of a method for generating three-dimensional data according to an embodiment of the present disclosure;

[0027] FIG2 is a schematic diagram of a framework of a method for generating three-dimensional data provided by an embodiment of the present disclosure;

[0028] FIG3 is a schematic diagram of feature cross-attention fusion in a method for generating three-dimensional data provided by an embodiment of the present disclosure;

[0029] FIG4 is a schematic flow chart of a method for generating three-dimensional data provided by an embodiment of the present disclosure;

[0030] FIG5 is a schematic diagram of a third image in a method for generating three-dimensional data provided by an embodiment of the present disclosure;

[0031] FIG6 is a schematic flow chart of a method for generating three-dimensional data provided by an embodiment of the present disclosure;

[0032] FIG7 is a schematic structural diagram of a device for generating three-dimensional data provided by an embodiment of the present disclosure;

[0033] FIG8 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0034] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0035] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0036] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.

[0037] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0038] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0039] How to quickly and efficiently generate high-quality three-dimensional data is a technical problem that needs to be solved. The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not intended to limit the scope of these messages or information.

[0040] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0041] Figure 1 is a flow chart of a method for generating three-dimensional data according to an embodiment of the present disclosure. This embodiment of the present disclosure is applicable to generating similar three-dimensional models based on images and text. The method can be performed by a three-dimensional data generation device, which can be implemented in software and / or hardware and can be configured in an electronic device, such as a computer.

[0042] As shown in FIG1 , the method for generating three-dimensional data provided in this embodiment may include:

[0043] S110: Acquire a first image of a target object and a description text corresponding to the first image.

[0044] In the disclosed embodiments, the target object can be a physical object, a biological object, or the like. The first image includes an image containing the target object, and the descriptive text can be considered to be text describing the target object in the first image, with the first image and the descriptive text having consistent content. The first image and the descriptive text can serve as control conditions for generating a three-dimensional model of the target object, ensuring that the generated three-dimensional model is similar to the first image and the descriptive text.

[0045] The first image and the description text may be obtained through user input, or by reading a preset storage space, etc., which are not exhaustively listed here.

[0046] S120. Acquire a noise image; wherein the noise image includes a rendered image with added noise; and the rendered image includes an image rendered from a neural radiation field at a preset viewing angle.

[0047] In the disclosed embodiments, a Neural Radiance Field (NeRF) can be randomly initialized, and camera parameters at a preset viewing angle and point light sources around the camera can be sampled in spherical coordinates. The NeRF field can then be rendered at the preset viewing angle to produce a rendered image.

[0048] After obtaining a NeRF rendered image at a preset viewing angle, noise (e.g., Gaussian noise) may be added to the rendered image to obtain a noisy image. The preset viewing angle may include at least one viewing angle; accordingly, the rendered image may include at least one image, and the noisy image may also include at least one image.

[0049] For example, Figure 2 is a schematic diagram of a framework for a method for generating three-dimensional data provided by an embodiment of the present disclosure. Referring to Figure 2 , rendering the NeRF field at a preset viewing angle yields a rendered image a; adding noise b to the rendered image yields a noisy image c.

[0050] S130. Cross-attentionally fuse the features of the first image, the features of the descriptive text, and the features of the noise image through a diffusion network to generate a second image; wherein the second image includes an image of the target object at a preset viewing angle.

[0051] In the disclosed embodiment, the diffusion network may include a two-dimensional diffusion network, and the diffusion network may be a pre-built network model (i.e., the parameters of the diffusion network in this embodiment are fixed). Referring again to FIG. 2 , the diffusion network may perform inverse diffusion on the input noise image c to generate a second image e. Because the noise image c is generated based on the rendered image a at a preset perspective, the second image e also belongs to an image at the preset perspective.

[0052] In Figure 2, when the diffusion network generates the second image, the features of the extracted noise image c can be cross-attention fused with the features of the first image d and the features of the description text t, so that the generated second image has similarity with the first image and the description text.

[0053] Exemplarily, FIG3 is a schematic diagram of feature cross-attention fusion in a method for generating three-dimensional data provided by an embodiment of the present disclosure. Referring to FIG3, in some optional implementations, cross-attention fusion of features of the first image, features of the descriptive text, and features of the noise image is performed through a diffusion network, which may include: feature processing of the noise image through each denoising sub-network in the diffusion network; feature processing of the first image through a mirror network with the same structure as the diffusion network; feature extraction of the descriptive text through a text feature extraction network; and cross-attention fusion of features of the noise image processed by the denoising sub-network in the diffusion network, features of the first image processed by the corresponding denoising sub-network in the mirror network, and features of the descriptive text through a cross-attention sub-network in the diffusion network.

[0054] Referring to Figure 3, the diffusion network may include multiple denoising sub-networks (such as the U-Net network in Figure 3), and feature processing may be performed on the input noisy image based on the denoising sub-network. In Figure 3, feature processing may be performed on the first image by a mirror network having the same structure as the diffusion network to obtain the processed features of the first image; feature extraction may be performed on the description text based on an existing text feature extraction network to obtain the features of the description text. In some optional implementations, feature processing of the noisy image by each denoising sub-network in the diffusion network may include: extracting features from the noisy image by an image encoder to obtain a noise feature map, and then performing feature processing on the noise feature map by each denoising sub-network in the diffusion network; feature processing of the first image by a mirror network having the same structure as the diffusion network may include: extracting features from the first image by an image encoder to obtain a first feature map, and then performing feature processing on the first feature map by a mirror network having the same structure as the diffusion network. It should be noted that the same structure of the mirror network and the diffusion network may mean that both include the same number of denoising sub-networks.

[0055] In Figure 3, the diffusion network can also include a cross-attention sub-network. Through the cross-attention sub-network, the features of the noisy image processed by each denoising sub-network can be cross-attended and fused with the features of the first image processed by the corresponding denoising sub-network in the mirror network, as well as the features of the descriptive text. That is, the features of the noisy image processed by the i-th denoising sub-network in the diffusion network can be cross-attended and fused with the features of the first image processed by the i-th denoising sub-network in the mirror network, as well as the features of the descriptive text.

[0056] In some optional implementations, cross-attention fusion of features of the noisy image processed by a denoising sub-network in the diffusion network, features of the first image processed by a corresponding denoising sub-network in the mirror network, and features of the descriptive text may include: cross-attention fusion of features of the noisy image processed by at least a portion of the denoising sub-network in the diffusion network, features of the first image processed by a corresponding denoising sub-network in the mirror network, and features of the descriptive text. Furthermore, in the diffusion network, a corresponding cross-attention sub-network may be deployed for each denoising sub-network requiring cross-attention fusion.

[0057] In some optional implementations, cross-attention fusing the features of the noisy image processed by the denoising sub-network in the diffusion network, the features of the first image processed by the corresponding denoising sub-network in the mirror network, and the features of the descriptive text may include: cross-attention fusing the features of the noisy image processed by at least part of the network layers in the denoising sub-network of the diffusion network, the features of the first image processed by the corresponding network layers in the corresponding denoising sub-network in the mirror network, and the features of the descriptive text.

[0058] In the embodiment of the present disclosure, a mirror network consistent with the diffusion network is used to process the first image, and the intermediate features of the noisy image processed by at least part of the denoising sub-network in the mirror network are fused into the corresponding denoising sub-networks in the diffusion network through the cross-attention sub-network. At the same time, the cross-attention sub-network can also fuse the features of the descriptive text into the corresponding denoising sub-networks in the diffusion network, so that the generated second image has a high similarity with the first image and the descriptive text.

[0059] Moreover, at least part of the denoising sub-network that needs to perform cross-attention fusion, as well as at least one network layer in the denoising sub-network that needs to perform cross-attention fusion, can be set according to the actual application scenario, thereby saving network computational power and improving the efficiency of generating the second image to a certain extent while ensuring that the second image has similarity with the first image and the description text.

[0060] S140. Determine a fractional distillation loss based on the second image, the noisy image, and the added noise, and adjust the neural radiation field based on the fractional distillation loss.

[0061] Referring again to Figure 2, the second image e and the noisy image c can be compared to determine the added noise f predicted by the diffusion network. Based on the predicted added noise f and the actual added noise b, the Score Distillation Sampling (SDS) loss can be determined. The closer the NeRF-rendered image is to the feature distribution of the second image, the closer the noise predicted by the diffusion network is to the actual added noise. Therefore, the parameters in NeRF can be adjusted based on the Score Distillation loss to make the NeRF rendering result similar to the feature distribution of the second image. This allows for an implicit representation of the three-dimensional model based on the loss of the two-dimensional image.

[0062] S150: Determine a three-dimensional model of the target object according to the adjusted neural radiation field.

[0063] In the embodiment of the present disclosure, the three-dimensional model may include a three-dimensional mesh model. The initial three-dimensional model can be extracted based on the constructed NeRF based on an existing algorithm. For example, the initial three-dimensional model can be extracted based on the constructed NeRF based on the Deep Marching Tetrahedra (DMTet) algorithm. In addition, a high-resolution three-dimensional model can be generated based on the initial three-dimensional model based on an existing algorithm. The generated three-dimensional model of the target object can have a high similarity with the first image and the description text. Compared with the generation of traditional three-dimensional models, the technical solution of the embodiment of the present disclosure can realize the rapid and efficient generation of high-quality three-dimensional models based on the first image and the description text.

[0064] The technical solution of the embodiment of the present disclosure is to obtain a first image of a target object and a descriptive text corresponding to the first image; obtain a noise image; wherein the noise image includes a rendered image with added noise; the rendered image includes an image rendered by a neural radiation field at a preset perspective; through a diffusion network, cross-attentionally fuse the features of the first image, the features of the descriptive text, and the features of the noise image to generate a second image; wherein the second image includes an image of the target object at a preset perspective; determine a fractional distillation loss based on the second image, the noise image, and the added noise, and adjust the neural radiation field based on the fractional distillation loss; and determine a three-dimensional model of the target object based on the adjusted neural radiation field.

[0065] By cross-attentionally fusing the features of the first image, the descriptive text, and the noise image based on a diffusion network, it is possible to generate a second image based on the noise image, using the first image and descriptive text as control conditions. By determining a rendered image at a preset perspective based on the neural radiation field, generating a noise image for input into the diffusion network by adding noise to the rendered image, and constructing the neural radiation field using fractional distillation loss, the 3D model expressed based on the neural radiation field can have the same feature distribution as the second image generated based on the first image and descriptive text. This allows for the rapid and efficient generation of high-quality 3D data based on the first image and descriptive text.

[0066] The various optional solutions in the three-dimensional data generation method provided in the embodiments of this disclosure can be combined with those provided in the above embodiments. The three-dimensional data generation method provided in this embodiment provides a detailed description of the control conditions for generating the second image when the target object includes a target human object. By adding a third image in a preset pose as a control condition for generating the second image, the generated three-dimensional model can also be in the preset pose, thereby facilitating subsequent processing of the three-dimensional model, such as skeletal binding.

[0067] FIG4 is a flow chart of a method for generating three-dimensional data provided by an embodiment of the present disclosure. As shown in FIG4 , the method for generating three-dimensional data provided by this embodiment, when the target object includes a target human object, may include:

[0068] S410: Acquire a first image of a target human object and a description text corresponding to the first image.

[0069] S420. Acquire a noise image; wherein the noise image includes a rendered image with added noise; and the rendered image includes an image rendered from a neural radiation field at a preset viewing angle.

[0070] S430: Acquire a third image of a preset three-dimensional model at a preset viewing angle; wherein the preset three-dimensional model includes a preset human body model in a preset posture.

[0071] In this embodiment, the preset 3D model may include a 3D mesh model. The preset 3D model may be generated based on an existing neural network conditioned on posture control parameters. For example, a Skinned Multi-Person Linear eXpressive (SMPL-X) model may be used to deform a preconfigured standard model based on posture control parameters for each preset region to generate a preset human body model in a preset posture.

[0072] The preset 3D model can be rendered from various preset perspectives to obtain respective third images; each third image can present the full surface of the preset human body. For example, FIG5 is a schematic diagram of a third image in a method for generating 3D data provided by an embodiment of the present disclosure. Referring to FIG5 , the preset 3D model is in a natural standing posture, and third images of the preset 3D model can be obtained from front and back perspectives.

[0073] S440. Generate a second image through a diffusion network based on features of the first image, features of the descriptive text, features of the noise image, and features of the third image; wherein the second image includes an image of the target human object in a preset posture at a preset perspective.

[0074] In this embodiment, while cross-attentionally fusing the features of the first image, the description text, and the noise image through the diffusion network, features of the third image are also extracted and introduced into the denoising sub-network of the diffusion network. This allows the generated second image to maintain similarity to the first image and the description text while also assuming the same pose as the pre-set human model in the third image, thereby adjusting the pose of the 3D model.

[0075] S450. Determine a fractional distillation loss based on the second image, the noisy image, and the added noise, and adjust the neural radiation field based on the fractional distillation loss.

[0076] S460: Determine a three-dimensional model of the target human subject based on the adjusted neural radiation field.

[0077] In some optional implementations, after determining the three-dimensional model of the target human object, it may also include: binding the three-dimensional model of the target human object to a skeleton model in a preset posture; wherein the three-dimensional model bound to the skeleton model is driven based on posture control parameters.

[0078] In these optional implementations, the preset pose can be, for example, a pose presented by an existing skeletal model. Because the 3D model and the skeletal model present the same pose, standard key points of the 3D model and the skeletal model can be easily aligned. Existing binding tools can be used to bind the 3D model to the skeletal model, and the skeletal model can then be driven by pose control parameters to achieve the generation of a dynamic 3D model.

[0079] The technical solution of the embodiment of the present disclosure describes in detail the control conditions for generating the second image when the target object includes a target human object. By adding a third image in a preset posture as a control condition for generating the second image, the generated three-dimensional model can also be in a preset posture, which is beneficial for subsequent further processing of the three-dimensional model, such as skeletal binding and other processing. The method for generating three-dimensional data provided by the embodiment of the present disclosure and the method for generating three-dimensional data provided by the above-mentioned embodiment belong to the same disclosed concept. The technical details not fully described in this embodiment can be referred to the above-mentioned embodiment, and the same technical features have the same beneficial effects in this embodiment and the above-mentioned embodiment.

[0080] The various optional schemes in the method for generating three-dimensional data provided in the embodiments of the present disclosure can be combined with the above embodiments. The method for generating three-dimensional data provided in this embodiment provides a detailed description of the application scenario. By detecting and segmenting human objects and scene objects in the target video, a segmented image is obtained. The segmented image can be used as a control condition for generating a three-dimensional model, so that the generated three-dimensional model has similarity with the human objects and scene objects in the target video. Based on this, it is possible to build and render a three-dimensional scene similar to the target video.

[0081] FIG6 is a flow chart of a method for generating three-dimensional data provided by an embodiment of the present disclosure. As shown in FIG6 , the method for generating three-dimensional data provided by this embodiment may include:

[0082] S610: Detect an object in a target video and determine a video frame containing the target object.

[0083] In this embodiment, based on existing image processing algorithms, objects (e.g., scene objects and human subjects) can be detected and tracked for each frame in the target video. When a new object is detected, the detected object instance can be saved, and the video frame containing the object instance can be recorded. At least one of the detected objects can be identified as the target object. Thus, the video frame containing the target object can be determined in the target video.

[0084] S620: Determine a target video frame from video frames containing the target object.

[0085] The target video frame can be determined from the videos containing the target object based on a preset selection strategy. For example, the video frame with the highest target object area ratio among all video frames can be determined as the target video frame; another example is the video frame with the highest target object clarity among all video frames. Other selection strategies can also be applied, and these are not exhaustive here.

[0086] S630: Segment the target video frame to obtain a first image of the target object.

[0087] In this embodiment, the target object region in the target video frame may be segmented based on an existing image segmentation algorithm to obtain a first image of the target object.

[0088] S640: Generate a description text corresponding to the first image based on the first image of the target object through an image description network.

[0089] In this embodiment, the image caption network may include an existing image annotation generation network, such as a Transformer-based model, such as Bidirectional Encoder Representations from Transformers (BERT). A first image may be input into the image caption network, causing it to output corresponding descriptive text. This allows the acquisition of a first image and descriptive text of a target object in a target video frame.

[0090] S650. Acquire a noise image; wherein the noise image includes a rendered image with added noise; and the rendered image includes an image rendered from a neural radiation field at a preset viewing angle.

[0091] S660. Cross-attentionally fuse the features of the first image, the features of the descriptive text, and the features of the noise image through a diffusion network to generate a second image; wherein the second image includes an image of the target object at a preset viewing angle.

[0092] Wherein, in the case where the target object includes a target human object, the method may further include: acquiring a third image of a preset three-dimensional model at a preset viewing angle; wherein the preset three-dimensional model includes a preset human body model in a preset posture;

[0093] Correspondingly, cross-attention fusion of the features of the first image, the features of the descriptive text, and the features of the noise image is performed through a diffusion network to generate a second image, which may include: generating a second image based on the features of the first image, the features of the descriptive text, the features of the noise image, and the features of the third image through a diffusion network; wherein the second image includes an image of a target human object in a preset posture at a preset perspective.

[0094] S670. Determine a fractional distillation loss based on the second image, the noisy image, and the added noise, and adjust the neural radiation field based on the fractional distillation loss.

[0095] S680: Determine a three-dimensional model of the target object based on the adjusted neural radiation field.

[0096] The corresponding neural radiation field can be adjusted for each object to determine the three-dimensional model corresponding to each target object.

[0097] S690: Generate a three-dimensional scene corresponding to the target video according to the three-dimensional model of the target object.

[0098] The 3D model corresponding to the target object in the target video can be used to build a 3D scene based on the spatial position and viewing angle of the target object in the target video, thereby generating a 3D scene with a high degree of similarity to the target video. This allows for the rapid construction of a corresponding 3D scene based on a 2D video, has broad application prospects, and can meet users' needs for 3D scene construction.

[0099] In addition, you can also place preset 3D model materials in the constructed 3D scene to enrich the 3D scene. You can also perform 3D scene rendering with different viewing angles, lighting, etc. from the target video to further enrich the presentation of the 3D scene.

[0100] In some optional implementations, when the target object includes a target human object, the method for generating three-dimensional data may also include: performing motion estimation on the target human object in the target video and determining a posture control parameter sequence; and driving the three-dimensional model of the target human object according to the posture control parameter sequence.

[0101] Among them, based on the existing human posture estimation model (such as a dense posture model, etc.), the motion estimation of the target human object in the video frame corresponding to the target human object can be performed to determine the posture parameters of each joint point of the target human object in each frame (such as position, angle and other parameters), thereby obtaining the posture control parameter sequence of each joint point.

[0102] After determining the posture control parameter sequence, the posture control sequence for each joint can be loaded into the corresponding 3D model to drive the 3D model of the target human subject. This allows the 3D model of the target human subject to move in a 3D scene with high similarity to the target video. This allows the creation of dynamic scenes and improves the user experience.

[0103] The technical solution of the embodiment of the present disclosure describes the application scenario in detail. By detecting and segmenting human objects and scene objects in the target video, a segmented image is obtained, and the segmented image can be used as a control condition for generating a three-dimensional model, so that the generated three-dimensional model has similarities with the human objects and scene objects in the target video. Based on this, it is possible to build and render a three-dimensional scene similar to the target video. The method for generating three-dimensional data provided by the embodiment of the present disclosure and the method for generating three-dimensional data provided by the above embodiment belong to the same public concept. The technical details not fully described in this embodiment can be referred to the above embodiment, and the same technical features have the same beneficial effects in this embodiment and the above embodiment.

[0104] Figure 7 is a schematic diagram of the structure of a three-dimensional data generation device provided by an embodiment of the present disclosure. The three-dimensional data generation device provided by this embodiment is suitable for generating similar three-dimensional models based on images and texts.

[0105] As shown in FIG7 , the apparatus for generating three-dimensional data provided by an embodiment of the present disclosure may include:

[0106] A first acquisition module 710 is configured to acquire a first image of a target object and a description text corresponding to the first image;

[0107] The second acquisition module 720 is configured to acquire a noise image; wherein the noise image includes a rendered image with added noise; and the rendered image includes an image rendered from a neural radiation field at a preset viewing angle.

[0108] Diffusion module 730, configured to perform cross-attention fusion of features of the first image, features of the descriptive text, and features of the noise image through a diffusion network to generate a second image; wherein the second image includes an image of the target object at a preset viewing angle;

[0109] a three-dimensional representation module 740 for determining a fractional distillation loss based on the second image, the noisy image, and the added noise, and adjusting the neural radiance field based on the fractional distillation loss;

[0110] The model generation module 750 is used to determine a three-dimensional model of the target object according to the adjusted neural radiation field.

[0111] In some optional implementations, the diffusion module can be used to:

[0112] The noise image is processed by each denoising sub-network in the diffusion network;

[0113] Perform feature processing on the first image through a mirror network with the same structure as the diffusion network;

[0114] Through the text feature extraction network, the description text is extracted;

[0115] Through the cross-attention sub-network in the diffusion network, the features of the noisy image processed by the denoising sub-network in the diffusion network are cross-attended and fused with the features of the first image processed by the corresponding denoising sub-network in the mirror network and the features of the description text.

[0116] In some optional implementations, the diffusion module can be used to:

[0117] The features of the noisy image processed by at least part of the denoising sub-network in the diffusion network are cross-attended with the features of the first image processed by the corresponding denoising sub-network in the mirror network and the features of the descriptive text.

[0118] In some optional implementations, the diffusion module can be used to:

[0119] Cross-attention fusion is performed on the features of the noisy image processed by at least part of the network layers in the denoising subnetwork, the features of the first image processed by the corresponding network layers in the corresponding denoising subnetwork, and the features of the descriptive text.

[0120] In some optional implementations, when the target object includes a target human object, the three-dimensional data generating device may further include:

[0121] A third acquisition module is configured to acquire a third image of a preset three-dimensional model at a preset viewing angle, wherein the preset three-dimensional model includes a preset human body model in a preset posture;

[0122] Accordingly, the diffusion module can be used to:

[0123] A second image is generated through a diffusion network based on features of the first image, features of the descriptive text, features of the noise image, and features of the third image; wherein the second image includes an image of a target human object in a preset posture at a preset perspective.

[0124] In some optional implementations, the three-dimensional data generating device may further include:

[0125] A model driving module is used to bind the three-dimensional model of the target human object to the skeleton model in a preset posture after determining the three-dimensional model of the target human object;

[0126] Among them, the three-dimensional model of the bound skeleton model is driven based on the posture control parameters.

[0127] In some optional implementations, the first acquisition module may be configured to:

[0128] Detecting the object in the target video and determining the video frame containing the target object;

[0129] Determine a target video frame from video frames containing the target object;

[0130] The target video frame is segmented to obtain a first image of the target object.

[0131] In some optional implementations, the first acquisition module may also be configured to:

[0132] Through the image description network, based on the first image of the target object, a description text corresponding to the first image is generated.

[0133] In some optional implementations, the model-driven module can be used to:

[0134] Perform motion estimation on the target human object in the target video and determine the posture control parameter sequence;

[0135] The three-dimensional model of the target human object is driven according to the posture control parameter sequence.

[0136] In some optional implementations, the model generation module may also be used to:

[0137] Based on the 3D model of the target object, a 3D scene corresponding to the target video is generated.

[0138] The three-dimensional data generation device provided in the embodiments of the present disclosure can execute the three-dimensional data generation method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0139] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present disclosure.

[0140] Reference is now made to FIG8 , which illustrates a schematic diagram of the structure of an electronic device (e.g., a terminal device or server in FIG8 ) 800 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device illustrated in FIG8 is merely an example and should not limit the functionality and scope of use of the embodiments of the present disclosure.

[0141] As shown in Figure 8, the electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the electronic device 800 are also stored in the RAM 803. The processing device 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0142] Typically, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or by wire to exchange data. Although FIG8 shows the electronic device 800 with various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.

[0143] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above-mentioned functions defined in the method for generating three-dimensional data of the embodiment of the present disclosure are performed.

[0144] The electronic device provided by the embodiment of the present disclosure and the method for generating three-dimensional data provided by the above embodiment belong to the same disclosed concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0145] An embodiment of the present disclosure provides a computer storage medium having a computer program stored thereon. When the program is executed by a processor, the method for generating three-dimensional data provided in the above embodiment is implemented.

[0146] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM) or flash memory (FLASH), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0147] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.

[0148] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0149] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device:

[0150] Obtain a first image of a target object and a description text corresponding to the first image; obtain a noise image; wherein the noise image includes a rendered image with added noise; the rendered image includes an image rendered by a neural radiation field at a preset perspective; through a diffusion network, cross-attentionally fuse the features of the first image, the features of the description text, and the features of the noise image to generate a second image; wherein the second image includes an image of the target object at a preset perspective; determine a fractional distillation loss based on the second image, the noise image, and the added noise, and adjust the neural radiation field based on the fractional distillation loss; determine a three-dimensional model of the target object based on the adjusted neural radiation field.

[0151] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0152] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0153] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the names of the units and modules do not, in certain circumstances, limit the units and modules themselves.

[0154] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and the like.

[0155] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0156] According to one or more embodiments of the present disclosure, a method for generating three-dimensional data is provided, the method comprising:

[0157] Acquire a first image of a target object and a description text corresponding to the first image;

[0158] Acquire a noise image; wherein the noise image includes a rendered image with added noise; the rendered image includes an image rendered from a neural radiation field at a preset viewing angle;

[0159] Cross-attentionally fusing features of the first image, features of the description text, and features of the noise image through a diffusion network to generate a second image; wherein the second image includes an image of the target object at the preset viewing angle;

[0160] determining a fractional distillation loss based on the second image, the noisy image, and the added noise, and adjusting the neural radiation field based on the fractional distillation loss;

[0161] A three-dimensional model of the target object is determined according to the adjusted neural radiation field.

[0162] According to one or more embodiments of the present disclosure, a method for generating three-dimensional data is provided, further comprising:

[0163] In some optional implementations, performing cross-attention fusion on the features of the first image, the features of the description text, and the features of the noise image through a diffusion network includes:

[0164] Performing feature processing on the noisy image through each denoising sub-network in the diffusion network;

[0165] performing feature processing on the first image through a mirror network having the same structure as the diffusion network;

[0166] Performing feature extraction on the description text through a text feature extraction network;

[0167] Through the cross-attention sub-network in the diffusion network, the features of the noisy image processed by the denoising sub-network in the diffusion network, the features of the first image processed by the corresponding denoising sub-network in the mirror network, and the features of the description text are cross-attended and fused.

[0168] According to one or more embodiments of the present disclosure, a method for generating three-dimensional data is provided, further comprising:

[0169] In some optional implementations, cross-attention fusion of features of the noisy image processed by the denoising sub-network in the diffusion network, features of the first image processed by the corresponding denoising sub-network in the mirror network, and features of the description text includes:

[0170] Cross-attention fusion is performed on the features of the noisy image processed by at least part of the denoising sub-network in the diffusion network, the features of the first image processed by the corresponding denoising sub-network in the mirror network, and the features of the description text.

[0171] According to one or more embodiments of the present disclosure, a method for generating three-dimensional data is provided, further comprising:

[0172] In some optional implementations, cross-attention fusion of features of the noisy image processed by the denoising sub-network in the diffusion network, features of the first image processed by the corresponding denoising sub-network in the mirror network, and features of the description text includes:

[0173] Cross-attention fusion is performed on the features of the noisy image processed by at least part of the network layers in the denoising subnetwork, the features of the first image processed by the corresponding network layers in the corresponding denoising subnetwork, and the features of the description text.

[0174] According to one or more embodiments of the present disclosure, a method for generating three-dimensional data is provided, further comprising:

[0175] In some optional implementations, when the target object includes a target human object, the method further includes:

[0176] Acquire a third image of a preset three-dimensional model at the preset viewing angle; wherein the preset three-dimensional model includes a preset human body model in a preset posture;

[0177] Accordingly, the cross-attention fusion of the features of the first image, the features of the description text, and the features of the noise image through the diffusion network to generate the second image includes:

[0178] A second image is generated through a diffusion network based on features of the first image, features of the descriptive text, features of the noise image, and features of the third image; wherein the second image includes an image of the target human object in the preset posture at the preset perspective.

[0179] According to one or more embodiments of the present disclosure, a method for generating three-dimensional data is provided, further comprising:

[0180] In some optional implementations, after determining the three-dimensional model of the target human object, the method further includes:

[0181] Binding the three-dimensional model of the target human object to the skeleton model in the preset posture;

[0182] The three-dimensional model bound to the skeleton model is driven based on posture control parameters.

[0183] According to one or more embodiments of the present disclosure, a method for generating three-dimensional data is provided, further comprising:

[0184] In some optional implementations, acquiring a first image of the target object includes:

[0185] Detecting the object in the target video and determining the video frame containing the target object;

[0186] Determining a target video frame from the video frames containing the target object;

[0187] The target video frame is segmented to obtain a first image of the target object.

[0188] According to one or more embodiments of the present disclosure, a method for generating three-dimensional data is provided, further comprising:

[0189] In some optional implementations, obtaining a description text corresponding to the first image includes:

[0190] Generate a description text corresponding to the first image based on the first image of the target object through an image description network.

[0191] According to one or more embodiments of the present disclosure, a method for generating three-dimensional data is provided, further comprising:

[0192] In some optional implementations, when the target object includes a target human object, the method further includes:

[0193] Performing motion estimation on the target human object in the target video to determine a posture control parameter sequence;

[0194] The three-dimensional model of the target human object is driven according to the posture control parameter sequence.

[0195] According to one or more embodiments of the present disclosure, a method for generating three-dimensional data is provided, further comprising:

[0196] In some optional implementations, after determining the three-dimensional model of the target object, the method further includes:

[0197] A three-dimensional scene corresponding to the target video is generated according to the three-dimensional model of the target object.

[0198] According to one or more embodiments of the present disclosure, a device for generating three-dimensional data is provided, the device comprising:

[0199] A first acquisition module is used to acquire a first image of a target object and a description text corresponding to the first image;

[0200] A second acquisition module is configured to acquire a noise image; wherein the noise image includes a rendered image with added noise; and the rendered image includes an image rendered from a neural radiation field at a preset viewing angle;

[0201] a diffusion module, configured to perform cross-attention fusion on features of the first image, features of the description text, and features of the noise image through a diffusion network to generate a second image; wherein the second image includes an image of the target object at the preset viewing angle;

[0202] a three-dimensional representation module, configured to determine a fractional distillation loss based on the second image, the noisy image, and the added noise, and adjust the neural radiation field based on the fractional distillation loss;

[0203] A model generation module is used to determine a three-dimensional model of the target object according to the adjusted neural radiation field.

[0204] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0205] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0206] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A method for generating three-dimensional data, comprising: Acquire a first image of a target object and a description text corresponding to the first image; Acquire a noise image; wherein the noise image includes a rendered image with added noise; the rendered image includes an image rendered by a neural radiation field at a preset viewing angle; Through a diffusion network, cross-attention fusion is performed on the features of the first image, the features of the description text, and the features of the noise image to generate a second image; wherein the second image includes an image of the target object at the preset viewing angle; determining a fractional distillation loss based on the second image, the noisy image, and the added noise, and adjusting the neural radiation field based on the fractional distillation loss; A three-dimensional model of the target object is determined according to the adjusted neural radiation field.

2. The method according to claim 1, wherein the cross-attention fusion of the features of the first image, the features of the description text, and the features of the noise image through the diffusion network comprises: Performing feature processing on the noisy image through each denoising sub-network in the diffusion network; Performing feature processing on the first image through a mirror network having the same structure as the diffusion network; Extracting features from the description text through a text feature extraction network; Through the cross-attention sub-network in the diffusion network, the features of the noisy image processed by the denoising sub-network in the diffusion network, the features of the first image processed by the corresponding denoising sub-network in the mirror network, and the features of the description text are cross-attended and fused.

3. The method according to claim 2, wherein the cross-attention fusion of the features of the noisy image processed by the denoising sub-network in the diffusion network, the features of the first image processed by the corresponding denoising sub-network in the mirror network, and the features of the description text comprises: The features of the noisy image processed by at least part of the denoising sub-network in the diffusion network, the features of the first image processed by the corresponding denoising sub-network in the mirror network, and the features of the description text are cross-attended and fused.

4. The method according to claim 2, wherein the cross-attention fusion of the features of the noisy image processed by the denoising sub-network in the diffusion network, the features of the first image processed by the corresponding denoising sub-network in the mirror network, and the features of the description text comprises: The features of the noisy image processed by at least part of the network layers in the denoising subnetwork, the features of the first image processed by the corresponding network layers in the corresponding denoising subnetwork, and the features of the description text are cross-attention fused.

5. The method according to claim 1, wherein in the case where the target object comprises a target human subject, the method further comprises: Acquire a third image of a preset three-dimensional model at the preset viewing angle; wherein the preset three-dimensional model includes a preset human body model in a preset posture; Correspondingly, the feature of the first image, the feature of the description text and the feature of the noise image are cross-attended and fused through the diffusion network to generate the second image, including: A second image is generated through a diffusion network based on features of the first image, features of the description text, features of the noise image, and features of the third image; wherein the second image includes an image of a target human object in the preset posture at the preset viewing angle.

6. The method according to claim 5, wherein after determining the three-dimensional model of the target human object, the method further comprises: Binding the three-dimensional model of the target human object with the skeleton model in the preset posture; Wherein, the three-dimensional model bound to the skeleton model is driven based on posture control parameters.

7. The method according to claim 1, wherein acquiring a first image of the target object comprises: Detecting the object in the target video and determining the video frame containing the target object; Determining a target video frame from the video frames containing the target object; The target video frame is segmented to obtain a first image of the target object.

8. The method according to claim 7, wherein obtaining the description text corresponding to the first image comprises: Based on the first image of the target object, a description text corresponding to the first image is generated through an image description network.

9. The method according to claim 7, wherein in the case where the target object comprises a target human subject, the method further comprises: Performing motion estimation on the target human object in the target video to determine a posture control parameter sequence; The three-dimensional model of the target human object is driven according to the posture control parameter sequence.

10. The method according to claim 7, wherein after determining the three-dimensional model of the target object, the method further comprises: A three-dimensional scene corresponding to the target video is generated according to the three-dimensional model of the target object.

11. A device for generating three-dimensional data, comprising: A first acquisition module, used to acquire a first image of a target object and a description text corresponding to the first image; A second acquisition module is used to acquire a noise image; wherein the noise image includes a rendered image with added noise; and the rendered image includes an image rendered by a neural radiation field at a preset viewing angle; A diffusion module, configured to perform cross-attention fusion on the features of the first image, the features of the description text, and the features of the noise image through a diffusion network to generate a second image; wherein the second image includes an image of the target object at the preset viewing angle; a three-dimensional representation module, configured to determine a fractional distillation loss based on the second image, the noisy image, and the added noise, and adjust the neural radiation field based on the fractional distillation loss; The model generation module is used to determine the three-dimensional model of the target object according to the adjusted neural radiation field.

12. An electronic device, comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method for generating three-dimensional data as described in any one of claims 1 to 10.

13. A storage medium comprising computer executable instructions, wherein the computer executable instructions are used to execute the method for generating three-dimensional data according to any one of claims 1 to 10 when executed by a computer processor.

Citation Information

Patent Citations

  • Three-dimensional object generation method based on non-photorealistic picture

    CN116778061A

  • Three-dimensional model generation method and device, computer equipment and storage medium

    CN116824092A

  • Training method, 3D object generation method and device, equipment and medium

    CN116883587A

  • Synthetic depth image generation from CAD data using generative adversarial neural networks for enhancement

    WO2019032481A1

  • Method and apparatus for generating computer-generated holographic field on basis of neural radiance field

    WO2023201771A1