Object generation method, 4D object generation method, and video generation method

By determining the object generation parameters and multiple object image sets corresponding to multiple time points in the neural network model, dynamic objects are generated, and the problem of being unable to accurately generate dynamic objects in the prior art is solved, and efficient and accurate dynamic object generation is achieved.

CN118972669BActive Publication Date: 2025-05-27ALIBABA DAMO (HANGZHOU) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410892921.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-03
Publication Date
2025-05-27
Estimated Expiration
2044-07-03

AI Technical Summary

Technical Problem

In the prior art, neural network models cannot accurately generate dynamic objects and cannot meet the actual needs of users to generate dynamic objects.

Method used

By determining the object generation parameters and multiple object image sets corresponding to multiple time points, input the object generation model to generate dynamic objects. The object generation model processes multiple object image sets based on object generation parameters, determines multiple object dynamic data corresponding to multiple perspectives, and generates dynamic objects based on these data.

Benefits of technology

It realizes the use of neural network models to accurately generate dynamic objects, meets the actual needs of users to generate dynamic objects, and reduces the time and labor cost of users to generate dynamic objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118972669B_ABST
    Figure CN118972669B_ABST
Patent Text Reader

Abstract

The embodiments of this specification provide an object generation method, a 4D object generation method, and a video generation method. The object generation method includes: determining object generation parameters and multiple sets of object images corresponding to multiple time points, where the multiple time points are arranged in chronological order, and each set of object images includes object images of a target object from multiple perspectives at the corresponding time point; inputting the object generation parameters and the multiple sets of object images into an object generation model to obtain a dynamic object corresponding to the target object, where the object dimension of the target object is smaller than the object dimension of the dynamic object, and the object generation model processes the multiple sets of object images according to the object generation parameters, determines multiple object dynamic data corresponding to the multiple perspectives, and generates the dynamic object according to the multiple object dynamic data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of computer technology, and in particular, to a method for generating an object; one or more embodiments of this specification also relate to a method for generating a 4D object, another method for generating an object, a method for generating a video, a computing device, a computer-readable storage medium, and a computer program product. Background Art

[0002] With the continuous development of artificial intelligence technology, by providing users with trained neural network models, it can help users solve various types of data processing problems, thereby reducing the time and labor costs of users' data processing and meeting their actual needs.

[0003] In the prior art, neural network models can generate static objects based on data provided by users; however, neural network models cannot accurately generate dynamic objects, resulting in the inability to meet the actual needs of users to generate dynamic objects. Therefore, how to use neural network models to accurately generate dynamic objects has become an urgent problem to be solved. Summary of the invention

[0004] In view of this, an embodiment of this specification provides an object generation method. One or more embodiments of this specification also involve a 4D object generation method, two other object generation methods, an object generation model training method, a video generation method, an object generation device, a 4D object generation device, two other object generation devices, an object generation model training device, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defect that the neural network model in the prior art cannot accurately generate dynamic objects.

[0005] According to a first aspect of an embodiment of this specification, there is provided a method for generating an object, including:

[0006] Determining object generation parameters and a plurality of object image sets corresponding to a plurality of time points, wherein the plurality of time points are arranged in chronological order, and each object image set includes object images of a target object at a corresponding time point and from a plurality of perspectives;

[0007] The object generation parameters and the multiple object image sets are input into an object generation model to obtain a dynamic object corresponding to the target object, wherein the object dimension of the target object is smaller than the object dimension of the dynamic object. The object generation model processes the multiple object image sets according to the object generation parameters, determines the multiple object dynamic data corresponding to the multiple perspectives, and generates the dynamic object according to the multiple object dynamic data.

[0008] According to a second aspect of an embodiment of this specification, there is provided an object generation device, including:

[0009] A data determination module is configured to determine object generation parameters and a plurality of object image sets corresponding to a plurality of time points, wherein the plurality of time points are arranged in chronological order, and each object image set includes object images of a target object at a corresponding time point and from a plurality of perspectives;

[0010] The object generation module is configured to input the object generation parameters and the multiple object image sets into an object generation model to obtain a dynamic object corresponding to the target object, wherein the object dimension of the target object is smaller than the object dimension of the dynamic object, the object generation model processes the multiple object image sets according to the object generation parameters, determines the multiple object dynamic data corresponding to the multiple perspectives, and generates the dynamic object according to the multiple object dynamic data.

[0011] According to a third aspect of an embodiment of this specification, a 4D object generation method is provided, including:

[0012] Determining object generation parameters and a plurality of 3D object image sets corresponding to a plurality of time points, wherein the plurality of time points are arranged in chronological order, and each 3D object image set includes 3D object images of the 3D object at a corresponding time point and from a plurality of perspectives;

[0013] The object generation parameters and the multiple 3D object image sets are input into an object generation model to obtain a dynamic 4D object corresponding to the 3D object, wherein the object dimension of the 3D object is smaller than the object dimension of the dynamic 4D object, the object generation model processes the multiple 3D object image sets according to the object generation parameters, determines multiple object dynamic data corresponding to the multiple perspectives, and generates the dynamic 4D object according to the multiple object dynamic data.

[0014] According to a fourth aspect of an embodiment of this specification, a 4D object generation device is provided, including:

[0015] a data determination module configured to determine object generation parameters and a plurality of 3D object image sets corresponding to a plurality of time points, wherein the plurality of time points are arranged in chronological order, and each 3D object image set includes 3D object images of the 3D object at a corresponding time point and from a plurality of perspectives;

[0016] The object generation module is configured to input the object generation parameters and the multiple 3D object image sets into an object generation model to obtain a dynamic 4D object corresponding to the 3D object, wherein the object dimension of the 3D object is smaller than the object dimension of the dynamic 4D object, the object generation model processes the multiple 3D object image sets according to the object generation parameters, determines multiple object dynamic data corresponding to the multiple perspectives, and generates the dynamic 4D object according to the multiple object dynamic data.

[0017] According to a fifth aspect of an embodiment of this specification, there is provided an object generation method, including:

[0018] Determining object generation parameters and a plurality of object image sets corresponding to a plurality of time points, wherein the plurality of time points are arranged in chronological order, and each object image set includes object images of a target object at a corresponding time point and from a plurality of perspectives;

[0019] Inputting the object generation parameters and the plurality of object image sets into an object generation model, wherein the object generation model comprises a dynamic data determination unit and an object generation unit;

[0020] Using the dynamic data determination unit, the plurality of object image sets are processed according to the object generation parameters to determine a plurality of object dynamic data corresponding to the plurality of viewing angles;

[0021] The object generation unit is used to generate the dynamic object according to the plurality of object dynamic data.

[0022] According to a sixth aspect of an embodiment of this specification, there is provided an object generation device, including:

[0023] A first data determination module is configured to determine object generation parameters and a plurality of object image sets corresponding to a plurality of time points, wherein the plurality of time points are arranged in chronological order, and each object image set includes object images of a target object at a corresponding time point and from a plurality of perspectives;

[0024] a data input module configured to input the object generation parameters and the plurality of object image sets into an object generation model, wherein the object generation model includes a dynamic data determination unit and an object generation unit;

[0025] A second data determination module is configured to use the dynamic data determination unit to process the plurality of object image sets according to the object generation parameters to determine a plurality of object dynamic data corresponding to the plurality of viewing angles;

[0026] The object generation module is configured to use the object generation unit to generate the dynamic object according to the multiple object dynamic data.

[0027] According to a seventh aspect of an embodiment of this specification, there is provided an object generation method, which is applied to a server, and includes:

[0028] Receiving object generation parameters sent by a client, and a plurality of object image sets corresponding to a plurality of time points, wherein the plurality of time points are arranged in chronological order, and each object image set includes object images of a target object at a corresponding time point and from a plurality of perspectives;

[0029] Inputting the object generation parameters and the plurality of object image sets into an object generation model to obtain a dynamic object corresponding to the target object, wherein the object dimension of the target object is smaller than the object dimension of the dynamic object, the object generation model processes the plurality of object image sets according to the object generation parameters, determines a plurality of object dynamic data corresponding to the plurality of perspectives, and generates the dynamic object according to the plurality of object dynamic data;

[0030] The dynamic object is sent to the client.

[0031] According to an eighth aspect of an embodiment of this specification, there is provided an object generation device, which is applied to a server, and includes:

[0032] A data receiving module is configured to receive object generation parameters sent by a client, and a plurality of object image sets corresponding to a plurality of time points, wherein the plurality of time points are arranged in chronological order, and each object image set includes object images of a target object at a corresponding time point and from a plurality of perspectives;

[0033] an object generation module, configured to input the object generation parameters and the plurality of object image sets into an object generation model, obtain a dynamic object corresponding to the target object, wherein the object dimension of the target object is smaller than the object dimension of the dynamic object, the object generation model processes the plurality of object image sets according to the object generation parameters, determines a plurality of object dynamic data corresponding to the plurality of perspectives, and generates the dynamic object according to the plurality of object dynamic data;

[0034] The object sending module is configured to send the dynamic object to the client.

[0035] According to a ninth aspect of the embodiments of this specification, there is provided an object generation model training method, comprising:

[0036] Determine training samples and sample labels for the object generation model to be trained, wherein the training samples include sample object generation parameters and a plurality of sample object image sets corresponding to a plurality of time points, the plurality of time points are arranged in chronological order, and each sample object image set includes sample object images of the sample object at a plurality of viewing angles at a corresponding time point;

[0037] Processing the plurality of sample object image sets according to the sample object generation parameters using the to-be-trained object generation model to determine a plurality of sample object dynamic data corresponding to the plurality of viewing angles, and generating sample dynamic objects corresponding to the sample objects according to the plurality of sample object dynamic data;

[0038] Based on the sample dynamic object and the sample label, the parameters of the object generation model to be trained are adjusted to obtain a trained object generation model.

[0039] According to a tenth aspect of an embodiment of this specification, there is provided an object generation model training device, comprising:

[0040] A training data determination module is configured to determine training samples and sample labels for a generation model of an object to be trained, wherein the training samples include sample object generation parameters and a plurality of sample object image sets corresponding to a plurality of time points, the plurality of time points are arranged in chronological order, and each sample object image set includes sample object images of the sample object at a plurality of viewing angles at a corresponding time point;

[0041] an object generation module, configured to process the plurality of sample object image sets according to the sample object generation parameters using the object generation model to be trained, determine a plurality of sample object dynamic data corresponding to the plurality of viewing angles, and generate sample dynamic objects corresponding to the sample objects according to the plurality of sample object dynamic data;

[0042] The model training module is configured to adjust the parameters of the object generation model to be trained based on the sample dynamic object and the sample label to obtain a trained object generation model.

[0043] According to an eleventh aspect of the embodiments of this specification, a video generation method is provided, which is applied to a video generation platform, including:

[0044] Receive the object video generated text and multiple object image sets corresponding to multiple time points sent by the client, wherein the multiple time points are arranged in chronological order, and each object image set includes object images of the target object at the corresponding time point and multiple perspectives;

[0045] Inputting the object video generation text and the plurality of object image sets into the object generation model to obtain the object video corresponding to the target object, wherein the object video is a video observed from any perspective, the object generation model processes the plurality of object image sets according to the object video generation text, determines a plurality of object dynamic data corresponding to the plurality of perspectives, and generates the arbitrary perspective object video according to the plurality of object dynamic data;

[0046] The object video is sent to the client.

[0047] According to a twelfth aspect of the embodiments of this specification, a computing device is provided, including:

[0048] Memory and processor;

[0049] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above-mentioned various object generation methods, 4D object generation methods or object generation model training methods are implemented.

[0050] According to the thirteenth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores a computer program / instruction, which, when executed by a processor, implements the above-mentioned multiple object generation methods, 4D object generation methods, or object generation model training steps.

[0051] According to the fourteenth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instruction, which, when executed by a processor, implements the above-mentioned multiple object generation methods, 4D object generation methods or object generation model training steps.

[0052] The object generation method provided by one or more embodiments of the present specification, in the process of generating dynamic objects by using an object generation model, first, by determining object generation parameters, multiple time points, and multiple object image sets, which are large in number and rich in types, it is convenient to subsequently generate accurate and vivid dynamic objects; then, the object generation model is used to process the multiple object image sets according to the object generation parameters, so as to determine multiple object dynamic data corresponding to multiple perspectives, and based on the content-rich data of the multiple object dynamic data, accurate and vivid dynamic objects can be generated, thereby realizing the accurate generation of dynamic objects by using a neural network model, meeting the actual needs of users for generating dynamic objects, and reducing the time and labor cost of users for generating dynamic objects. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1It is an application schematic diagram of an object generation method provided by an embodiment of this specification;

[0054] Figure 2 is a flow chart of an object generation method provided by an embodiment of this specification;

[0055] Figure 3 is a process flow chart of a method for generating an object provided by an embodiment of the present specification;

[0056] Figure 4 is a flow chart of a 4D object generation method provided by an embodiment of this specification;

[0057] Figure 5 is a flow chart of another object generation method provided by an embodiment of this specification;

[0058] Figure 6 is a flowchart of another object generation method provided by an embodiment of this specification;

[0059] Figure 7 is a flow chart of an object generation model training method provided by an embodiment of this specification;

[0060] Figure 8 is a flow chart of a video generation method provided by an embodiment of this specification;

[0061] Fig. 9 It is a structural block diagram of a computing device provided by an embodiment of this specification. DETAILED DESCRIPTION

[0062] Many specific details are described in the following description to facilitate a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of this specification, so this specification is not limited to the specific implementation disclosed below.

[0063] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0064] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0065] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0066] In one or more embodiments of this specification, a large model refers to a deep learning model with large-scale model parameters, which usually contains hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than 10 trillion model parameters. A large model can also be called a foundation model / foundation model. The large model is pre-trained with large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks, and the model has good generalization ability, such as a large-scale language model (LLM), a multi-modal pre-training model, etc.

[0067] When the big model is used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. The big model can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, it can be applied to computer vision tasks such as visual question answering (VQA), image caption (IC), image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of the big model include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.

[0068] First, the terms involved in one or more embodiments of this specification are explained.

[0069] MV-VDM: Multi-View Video Diffusion Model, a 4D base model that can generate multi-view model action videos.

[0070] MV-Video: Multi-View Video dataset, a 4D dataset.

[0071] text-to-3D: Text to 3D generation.

[0072] image-to-3D: Image to 3D generation.

[0073] video-to-4D: Video to 4D generation.

[0074] NeRF: Neural Radiance Field, a 3D representation model.

[0075] 3DGS: 3D Gaussian splattering, a more explicit and efficient 3D representation model.

[0076] 4DGS: 4D Gaussian splatter, learn motion modeling based on 3DGS.

[0077] SDS: Score-Distillation-Sampling, a strategy for distilling large pre-trained models.

[0078] Pipeline: refers to a workflow, which means connecting multiple operations or processing steps in series to form an orderly workflow. Each step is responsible for a part of the processing task, and its output is directly used as the input of the next step.

[0079] With the continuous development of artificial intelligence technology, by providing users with trained neural network models, it can help users solve various types of data processing problems, thereby reducing the time and labor costs of users' data processing and meeting their actual needs.

[0080] In a 4D content generation solution provided in this specification, a user can generate corresponding 4D content using a neural network model by providing corresponding data; and in terms of 4D generation technology, the solution mainly focuses on generating 4D content by distilling a diffusion model conditioned by pre-trained text or single-view images; and such operation causes the solution to have major defects, making it difficult for the solution to utilize various ready-made 3D resources with complex multi-view attributes, and due to the inherent ambiguity of these supervisory signals, the generated 4D results often have spatiotemporal inconsistencies.

[0081] In another 4D content generation scheme provided in this specification, a pre-trained text-to-image model can be used to maintain the model's appearance, and a pre-trained text-to-video model can be used to supervise the learning of motion; or, a pre-trained single-view condition model can be used to supervise the model's new perspective morphology, and a pre-trained image-to-video model can be used to reconstruct the motion. However, the disadvantages of this scheme are: first, the separate morphology and motion supervision easily lead to the inconsistency of the generated 4D in time and space; second, it mainly relies on text-conditioned or single-view image conditioned models to optimize appearance, and cannot use multi-view as a condition, resulting in the inability to better maintain the multi-view details of the original 3D model while modeling motion.

[0082] Based on this, in this specification, an object generation method is provided. One or more embodiments of this specification also involve a 4D object generation method, two other object generation methods, an object generation model training method, a video generation method, an object generation device, a 4D object generation device, two other object generation devices, an object generation model training device, a computing device, a computer-readable storage medium and a computer program product, which are described in detail one by one in the following embodiments.

[0083] See also Figure 1 , Figure 1 A schematic diagram of an application of an object generation method provided according to an embodiment of the present specification is shown. Figure 1 It can be seen that the user can send the 3D model image and prompt text obtained by collecting images of the 3D model from multiple perspectives to the server 104 through the terminal 102. The server 104 inputs the 3D model image and prompt text into the object generation model to obtain a 4D model corresponding to the 3D model output by the object generation model. By sending the 4D model to the terminal 102, the user can observe the 4D model from different angles and time points.

[0084] See also Figure 2 , Figure 2 A flowchart of an object generation method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.

[0085] Step 202: Determine object generation parameters and multiple object image sets corresponding to multiple time points, wherein the multiple time points are arranged in chronological order, and each object image set includes object images of the target object at the corresponding time point from multiple perspectives.

[0086] Among them, the object generation parameter can be understood as a parameter for generating a dynamic object, for example, the object generation parameter can be an object generation text, an object generation time, etc.; the object generation text can be understood as prompt information for generating the dynamic object, and the object generation text can be a prompt text describing the action posture of the dynamic object to be generated. For example, in the case where the target object is a static 3D eagle model, the object generation text can be "The eagle is flying"; the object generation time can be understood as a time parameter for generating the dynamic object, and the object generation time can be used to guide the object generation model to generate a dynamic object corresponding to the object generation. For example, in the case where the target object is a static 3D eagle model, the object generation time can be 1 minute, which is used to guide the object generation model to generate a dynamic object with a duration of 1 minute.

[0087] An object image set can be understood as a set of object images; wherein an object image set includes object images obtained by capturing images of the target object through multiple perspectives at the corresponding time point. For example, the corresponding time point may be the first time point among multiple time points arranged in chronological order, the multiple perspectives may be orthogonal 4 perspectives, and the target object is a 3D eagle model. Based on this, the object image set corresponding to the first time point includes object images obtained by capturing images of the 3D eagle model through orthogonal 4 perspectives at the first time point. It should be noted that the above is an example of an object image set, and other object image sets can be determined in the same way.

[0088] The target object may be understood as a 3D object, for example, the target object may be a 3D model, 3D point cloud data, etc.; the time point may be understood as a timestamp.

[0089] Multiple viewing angles can be set according to actual application scenarios. For example, the multiple viewing angles can be orthogonal 4 viewing angles, horizontal viewing angles, upward viewing angles, downward viewing angles, etc.; it should be noted that the orthogonal 4 viewing angles can be understood as four orthogonal directions, and the object image captured using the orthogonal 4 viewing angles can be an orthogonal 4 view, that is, a picture taken from four orthogonal directions. For example, the pitch angles of the four cameras are all 0 degrees, and the azimuth angles are 0, 90, 180, and 270 degrees respectively. The pictures are aimed at the center of the object (i.e., the target object) and the pictures taken in this way are orthogonal four views;

[0090] In one or more embodiments provided in this specification, in order to provide object generation services to users, the object generation server used by the object generation method provided in this specification can receive object generation parameters sent by the user based on the client, and the multiple object image sets corresponding to multiple time points. The specific implementation method is as follows.

[0091] The determining of the object generation parameters and a plurality of object image sets corresponding to a plurality of time points comprises:

[0092] Receiving the object generation parameters and the multiple object image sets corresponding to multiple time points sent by the client, wherein the object generation parameters and the multiple object image sets are generated by the client based on a data sending operation performed by a user on a data sending page;

[0093] Among them, the data sending page can be understood as a page displayed in the client that enables the user to send data. The data sent by the user refers to the object generation parameters and the multiple object image sets corresponding to multiple time points. The data sending page can be a human-computer interaction page such as a web page and an application page.

[0094] Among them, the data sending operation can be understood as the user providing data to the client by triggering the data sending control corresponding to the data sending page, so that the client forwards it to the object generation server; the data sending control can be a page component in the data sending page (such as a page button, input box, etc.), or the data sending control can be a hardware device corresponding to the data sending page, such as a keyboard device, a touch screen, a button device, a scanning device, an image acquisition device, etc.

[0095] The client may be a hardware device such as a terminal, a smart device, or a server, or the client may be a software such as an application, a script, or a mini-program.

[0096] Among them, the object generation server can be understood as the server to which the object generation method is applied. The server can be a hardware device such as a server or a server cluster, or a software device such as a virtual machine, a cloud server, or a container.

[0097] Specifically, when the user needs to generate an object, the user can send the object generation parameters and the plurality of object image sets corresponding to the plurality of time points to the client through the data sending page and the data sending control provided by the client for the user, so that the client sends them to the object generation server;

[0098] The object generation server can receive the above data and perform object generation operations based on the data, thereby meeting the user's dynamic object generation needs.

[0099] In one or more embodiments provided in this specification, the object generation method provided in this specification can also receive object generation parameters and a target object; by performing a multi-perspective image acquisition operation on the target object, a plurality of object image sets corresponding to a plurality of time points are obtained; wherein the multi-perspective image acquisition operation can be implemented through a neural network model or other application programs; that is, a neural network model or application program can be used to perform a multi-perspective image acquisition operation on the target object to obtain a plurality of object image sets corresponding to a plurality of time points.

[0100] In one or more embodiments provided in the present specification, the target object may be obtained by reconstructing an initial object (e.g., an incomplete object, a rough object, etc.), for example, reconstructing an incomplete 3D model or a rough 3D model to obtain the complete or detailed 3D model; it should be noted that the object reconstruction operation may be implemented using a neural network model, for example, using a trained 3D reconstruction model to reconstruct the initial object to obtain the target object.

[0101] In one or more embodiments provided in the present specification, the target object may be an object generated by generating text from the target object and / or generating an image from the target object, wherein the text generated by the target object may be a prompt text for generating the target object, for example, the text generated by the target object may be "generate a 3D model of an eagle"; wherein the image generated by the target object may be understood as an image used to generate the target object, for example, the image generated by the target object may be an image of an eagle; it should be noted that the operation of generating the target object by generating text from the target object and / or generating an image from the target object may be implemented by a neural network model, for example, the text generated by the target object and / or the image generated by the target object may be input into a 3D generation module for object generation, thereby obtaining the target object.

[0102] Based on the above embodiments, it can be known that the target object can be obtained through an object reconstruction operation or an object generation operation, for example, a 3D model is generated through a model reconstruction operation and a model generation operation.

[0103] Step 204: Input the object generation parameters and the multiple object image sets into the object generation model to obtain a dynamic object corresponding to the target object, wherein the object dimension of the target object is smaller than the object dimension of the dynamic object, and the object generation model processes the multiple object image sets according to the object generation parameters, determines the multiple object dynamic data corresponding to the multiple perspectives, and generates the dynamic object according to the multiple object dynamic data.

[0104] Among them, the object generation model can be understood as a neural network model for generating dynamic objects, the object generation model can be a neural network model for generating 4D models; the object generation model can be a neural network model for generating 4D point cloud data; or, the object generation model can be a neural network model for generating 4D videos.

[0105] Among them, the dynamic object can be understood as a model that can display actions or postures in a dynamic manner. The dynamic object can be a 4D model, a 4D video or 4D point cloud data; the dynamic object can be an object that can be observed from any perspective or angle and can display actions in a dynamic manner; for example, the dynamic object can be a 4D eagle model that can be observed from any angle and displays flying actions; or the dynamic object can be a 4D eagle video that can be observed from any angle and displays flying actions.

[0106] The object dimension of the target object may be 3D, and the object dimension of the dynamic object may be 4D.

[0107] The object dynamic data may be understood as data used to characterize the dynamic features of a dynamic object. For example, the object dynamic data may be a dynamic video of a dynamic object at a certain viewing angle; or the object dynamic data may be dynamic point cloud data of a dynamic object at a certain viewing angle.

[0108] In one or more embodiments provided in this specification, when a user sends data through a client to generate an object, a dynamic object generated using an object generation model can be sent to the client for display, as described below.

[0109] After inputting the object generation parameters and the plurality of object image sets into the object generation model to obtain the dynamic object corresponding to the target object, the method further includes:

[0110] The dynamic object is sent to the client and displayed to the user using the data sending page in the client.

[0111] Specifically, after a dynamic object is generated using an object generation model, the dynamic object is sent to a client and displayed to a user through a data sending page in the client, thereby satisfying the user's need to generate a dynamic object and reducing the user's time and labor cost for generating a dynamic object.

[0112] In one or more embodiments provided in the present specification, the object generation model includes a dynamic data determination unit and an object generation unit, wherein the dynamic data determination unit is used to process the multiple object image sets according to the object generation parameters to determine the multiple object dynamic data corresponding to the multiple perspectives; the object generation unit is used to generate the dynamic object according to the multiple object dynamic data, and through the cooperation of the two units, accurate and vivid dynamic objects are generated.

[0113] Among them, the dynamic data determination unit can be understood as one or more network layers in the object generation model, or the dynamic data determination unit can be understood as a neural network model in the object generation model. For example, the dynamic data determination unit can be a 4D base model MV-VDM, and the 4D base model MV-VDM (multi-view video diffusion model) can generate multi-view multi-frame temporally and spatially consistent videos (i.e., dynamic objects).

[0114] Among them, the object generation unit can be understood as one or more network layers in the object generation model, or the object generation unit can be understood as a neural network model in the object generation model. For example, the object generation unit can be other parts in the 4D optimization pipeline (workflow) except the dynamic data determination unit, or the object generation unit can be other network layers in the 4D generation framework except the 4D base model MV-VDM; the object generation unit can process the object dynamic data output by the dynamic data determination unit; it should be noted that in one or more embodiments provided in this specification, based on MV-VDM, an efficient 4D optimization pipeline is proposed, which can better maintain the details of the original 3D model by combining reconstruction and 4D-SDS distillation strategies, and can drive the existing 3D model with high quality according to the text. In addition, the driving of the existing 3D model provided in this specification can be understood as converting the 3D model into a dynamic 4D object or 4D video.

[0115] In one or more embodiments provided in the present specification, in order to meet the needs of users to generate dynamic objects, in the process of inputting the object generation parameters and the multiple object image sets into the object generation model to obtain the dynamic object corresponding to the target object, the object generation model can be used to determine the object generation parameter characteristics of the object generation parameters, as well as determine the multiple object image set characteristics of the multiple object image sets, and then generate the dynamic object based on the object generation parameter characteristics and the multiple object image set characteristics, thereby meeting the needs of users to generate dynamic objects. The specific implementation method is as follows.

[0116] The step of inputting the object generation parameters and the plurality of object image sets into an object generation model to obtain a dynamic object corresponding to the target object comprises steps 1 to 3:

[0117] Step 1: Input the object generation parameters and the multiple object image sets into an object generation model, and use the object generation model to determine the object generation parameter features of the object generation parameters and determine the multiple object image set features of the multiple object image sets.

[0118] The object generation parameter feature may be understood as a feature determined based on the object generation parameter, and the object generation parameter feature may be a feature matrix or a feature vector;

[0119] The object image set features can be understood as features determined based on object generation parameters, and the object image set features can be feature matrices or feature vectors. It should be noted that in the process of determining the object image set features, feature extraction is performed on the multiple object images contained in the object image set to obtain image features corresponding to each object image, and the image features corresponding to the multiple object images contained in the object image set are used as the object image set features of the object image set. In other words, the object image set features can be composed of image features corresponding to multiple object images contained in the object image set.

[0120] Specifically, the object generation parameters and the multiple object image sets are input into the object generation model, and the object generation model is used to extract features of the object generation parameters to obtain object generation parameter features, and to extract features of the multiple object image sets to obtain multiple object image set features.

[0121] In one or more embodiments provided in the present specification, since object generation parameters, multiple time points, and multiple object image sets are data of large quantity and rich types, in order to better mine the information contained in the object generation parameters and the multiple object image sets, the object generation model can be used to determine the multi-view features corresponding to the multiple object image sets, as well as the video features corresponding to the multiple object image sets, and the object generation parameter encoding of the object generation parameters and the multiple object image set encoding of the multiple object image sets can be determined by the object generation model; the multi-view features, video features, object generation parameter encoding, and multiple object image set encoding are used to generate multiple object dynamic data; the specific implementation method is as follows.

[0122] The determining the object generation parameter features of the object generation parameters by using the object generation model, and determining a plurality of object image set features of the plurality of object image sets, comprises:

[0123] Using the object generation model, performing multi-view feature conversion on the multiple object image sets to obtain multi-view features corresponding to the multiple object image sets, and performing video feature conversion on the multiple object image sets to obtain video features corresponding to the multiple object image sets;

[0124] Determining, using the object generation model, object generation parameter encoding of the object generation parameter and a plurality of object image set encodings of the plurality of object image sets;

[0125] The object generation parameter feature is generated by using the multi-view feature, the video feature and the object generation parameter encoding, and the multiple object image set features are generated by using the multi-view feature, the video feature and the multiple object image set encoding.

[0126] Among them, multi-view feature conversion can be understood as performing multi-view feature enhancement on the features corresponding to multiple object image sets to obtain multi-view features after feature enhancement; the multi-view feature can be understood as the feature after multi-view feature enhancement; the multi-view feature enhancement can be understood as weighting the feature through weight data determined by training, and the features corresponding to the object image set can be obtained through feature extraction.

[0127] Among them, the video feature conversion can be understood as converting the object image set features corresponding to the object image set into features adapted to the video; the video feature can be understood as features adapted to the video.

[0128] The object generation parameter encoding can be understood as the encoded data obtained by encoding the object generation parameter; the object image set encoding can be understood as the encoded data obtained by encoding the object image set.

[0129] Specifically, using the object generation model, feature extraction is performed on the multiple object image sets to obtain features corresponding to each object image set, multi-view feature conversion is performed on the features corresponding to each object image set to obtain multi-view features corresponding to the multiple object image sets, and video feature conversion is performed on the features corresponding to each object image set to obtain video features corresponding to the multiple object image sets;

[0130] Using the object generation model, encoding the object generation parameters to obtain object generation parameter encoding, and encoding the multiple object image sets to obtain multiple object image set encodings;

[0131] Finally, the multi-view features, the video features and the object generation parameter encoding are fused to generate the object generation parameter features; and the multi-view features, the video features and the multiple object image set encodings are fused to generate the multiple object image set features.

[0132] The object generation method provided in this specification is used to illustrate the scene of driving any static 3D model to generate 4D content, wherein the object generation model can be a 4D generation framework (hereinafter referred to as the 4D generation framework) for driving any static 3D model, which is used to drive any static 3D model to generate 4D content; the 4D generation framework includes a 4D base model; the target object is a static 3D model (i.e., a 3D eagle model); multiple object image sets can be multi-view renderings (Multi-view Renderings) of the 3D model, and the multi-view renderings include clean renderings and noise renderings; the multi-view refers to orthogonal 4 perspectives, and the noise rendering refers to a rendering with noise added; multiple time points can be understood as timestamps (Timestep), which are the timestamps corresponding to each rendering in the multi-view rendering; when processing a multi-view rendering sequence, the 4D generation framework will process the data step by step according to timesteps, that is, process the renderings in the sequence one by one or batch by batch. The object generation parameter is the object generation prompt text ("Eagle is flying"); the multi-view rendering has corresponding camera parameters (Camera), and the camera parameters may include the focal length, position and orientation information of the camera during the rendering acquisition process.

[0133] Based on this, after determining the above data, the multi-view rendering, timestamp, camera parameters and the object generation prompt text (hereinafter referred to as prompt text) can be input into the 4D base model in the 4D generation framework; wherein, the multi-view rendering, timestamp, and camera parameters can be input into the 4D base model as a whole, for example, the multi-view rendering, timestamp, and camera parameters are spliced ​​or fused, so as to add the timestamp and camera parameters to the multi-view rendering and input into the 4D base model.

[0134] The 4D base model first processes the input multi-view rendering image based on the multi-view self-attention mechanism to obtain multi-view enhanced image features; and converts the multi-view rendering image into video features adapted to the video;

[0135] Secondly, the multi-view rendering and the prompt text are encoded to obtain the rendering encoding (i.e., the encoding of the set of multiple object images) and the prompt text encoding (i.e., the encoding of the object generation parameters);

[0136] Finally, the multi-view enhanced image features, video features and rendering image encoding are fused to obtain fusion features determined based on the rendering image (i.e., multiple object image set features); the multi-view enhanced image features, video features and prompt text encoding are fused to obtain fusion features determined based on the prompt text (i.e., object generation parameter features).

[0137] In the above embodiment, in the process of using the object generation model to determine the object generation parameter features of the object generation parameters, and determining the multiple object image set features of the multiple object image sets, the object generation model can be used to determine multi-view features, video features, and multiple object image set encodings using the multiple object image sets, and the object generation parameters can be used to determine the object generation parameter encodings. Then, based on the multi-view features, video features, object generation parameter encodings, and multiple object image set encodings, object generation parameter features and multiple object image set features are generated, thereby achieving accurate mining of the information contained in the object generation parameters and the multiple object image sets.

[0138] In one or more embodiments provided in the present specification, in the process of determining multiple object dynamic data, the multi-view attention unit in the object generation model can be used to process multiple object image sets to obtain multi-view features; and the video feature extraction unit in the object generation model can be used to process multiple object image sets to obtain video features. Subsequently, multiple object dynamic data can be generated based on the multi-view features and video features. The specific implementation method is as follows.

[0139] The performing multi-view feature conversion on the multiple object image sets to obtain multi-view features corresponding to the multiple object image sets, and performing video feature conversion on the multiple object image sets to obtain video features corresponding to the multiple object image sets, include:

[0140] Using the multi-view attention unit in the object generation model, extracting features from each of the multiple object image sets to obtain the multi-view features corresponding to the multiple object image sets;

[0141] The video feature extraction unit in the object generation model is used to perform video feature conversion on each of the object image sets to obtain video features corresponding to the multiple object image sets.

[0142] Among them, the multi-view attention unit can be understood as one or more network layers in the object generation model, or a sub-model in the object generation model; the multi-view attention unit is used to implement the multi-view self-attention mechanism.

[0143] The video feature extraction unit can be understood as one or more network layers in the object generation model, or a sub-model in the object generation model; the video feature extraction unit is used to convert the multi-view rendering image into one that is compatible with the video.

[0144] Continuing with the above example, the multi-view attention unit is a multi-view self-attention module that implements a multi-view self-attention mechanism (Multi-view 3DAttention), and the video feature extraction unit can be a multi-view image to video adaptation module (MV2V-Adatper); based on this, after the multi-view rendering image is input into the MV-VDM model, it is processed using the multi-view self-attention module and the multi-view image to video adaptation module in the MV-VDM model to obtain a feature vector.

[0145] In the above embodiment, in the process of performing multi-view feature conversion on the multiple object image sets to obtain multi-view features corresponding to the multiple object image sets, and performing video feature conversion on the multiple object image sets to obtain video features corresponding to the multiple object image sets, the multi-view attention unit and the video feature extraction unit in the object generation model can be implemented, so as to facilitate the subsequent generation of multiple object dynamic data based on the multi-view features and video features.

[0146] In one or more embodiments provided in the present specification, in the process of object generation, the object generation model can be used to perform feature extraction on each object image sequence, image acquisition parameters and time series to obtain an image sequence feature matrix of each object image sequence, a parameter feature matrix of the image acquisition parameters and a time series feature matrix of the time series, and feature fusion is performed on the image sequence feature matrix, the parameter feature matrix and the time series feature matrix to obtain video features corresponding to the multiple object image sets, and then multiple object dynamic videos are generated based on the video features, wherein each object image sequence, image acquisition parameters and time series are determined based on the multiple object image sets, and the specific implementation method is as follows.

[0147] The step of using the video feature extraction unit in the object generation model to perform video feature conversion on each of the object image sets to obtain video features corresponding to the multiple object image sets includes:

[0148] Using the video feature extraction unit in the object generation model, the object images included in each of the plurality of object image sets are sorted to obtain a plurality of object image sequences;

[0149] Determine an image acquisition parameter corresponding to each object image set, and determine the multiple time points arranged in time sequence as a time series, wherein the image acquisition parameter is a parameter used to acquire the object image;

[0150] Using the self-attention layer in the video feature extraction unit, feature extraction is performed on each object image sequence, image acquisition parameter and time series to obtain an image sequence feature matrix of each object image sequence, a parameter feature matrix of the image acquisition parameter and a time series feature matrix of the time series;

[0151] The multi-head attention layer in the video feature extraction unit is used to perform feature fusion on the image sequence feature matrix, the parameter feature matrix and the time series feature matrix to obtain video features corresponding to the multiple object image sets.

[0152] Using the above example, the multi-view image to video adaptation module processes the multi-view rendering as follows:

[0153] 1. Along the spatial dimension, the noise frames (i.e., multi-view renderings) are connected in series to obtain series frames (i.e., a sequence of images of each object); wherein the spatial dimension refers to: a rendering of the same action shot from four perspectives, i.e., a column of renderings of the above-mentioned "multi-view renderings".

[0154] 2. Use the self-attention layer of the 3D diffusion model (i.e., video feature extraction unit) to extract features from the serial frames, corresponding timestamps, and camera parameters to obtain the corresponding projection matrix W Q ′、W K , W V (i.e., image sequence feature matrix, parameter feature matrix, time series feature matrix);

[0155] 3. Use the multi-head attention mechanism to calculate the projection matrix W Q ′、W K , W V Process and obtain the projection matrix W O ′ (i.e. video features).

[0156] In the above embodiment, in the process of using the video feature extraction unit in the object generation model to perform video feature conversion on the object image sets and obtain the video features corresponding to the multiple object image sets, the self-attention layer in the video feature extraction unit can be used to extract features from the object image sequences, image acquisition parameters and time series to obtain the image sequence feature matrix of the object image sequences, the parameter feature matrix of the image acquisition parameters and the time series feature matrix of the time series, and then the multi-head attention layer in the video feature extraction unit is used to determine the video features corresponding to the multiple object image sets based on the image sequence feature matrix, the parameter feature matrix and the time series feature matrix; by using the video feature extraction unit in the object generation process, the object generation model has the ability of multi-view condition (considering multiple perspectives), the realism of dynamic objects is improved, and a vivid 4D model can be observed from any perspective.

[0157] In one or more embodiments provided in this specification, in order to achieve high-quality driving of any existing static 3D model through a sentence of text, thereby obtaining a 4D model, object generation text is used to generate dynamic objects; in the process of object generation, after the object generation text is determined, the object generation text can be text-encoded by a text encoder in the object generation model to obtain an object generation parameter code; and, the multiple object image sets are image-encoded by an image encoder in the object generation model to obtain the multiple object image set codes, and subsequently multiple object dynamic data can be generated based on the object generation parameter code and the multiple object image set codes. The specific implementation method is as follows.

[0158] The object generation parameter is an object generation text, and the object generation parameter code is an object generation text code;

[0159] The object generation parameter encoding of the object generation parameter is determined by using the object generation model, and the multiple object image set encodings of the multiple object image sets include:

[0160] Using the text encoder in the object generation model, the object generation text is encoded to obtain the object generation parameter encoding;

[0161] Using an image encoder in the object generation model, performing image encoding on the plurality of object image sets to obtain encodings of the plurality of object image sets;

[0162] Continuing with the above example, the prompt text is encoded using the text encoder in the MV-VDM model to obtain the text encoding; the multi-view rendering image is encoded using the image encoder in the MV-VDM model to obtain the image encoding, thereby realizing the mining of object generation parameters and the information contained in multiple object image sets corresponding to multiple time points in multiple ways, thereby improving the accuracy and diversity of the information.

[0163] In one or more embodiments provided in the present specification, after obtaining a variety of information, in order to integrate the various information, multi-view features, video features, object generation parameter encodings, and multiple object image set encodings can be fused through a cross-attention mechanism to obtain object generation parameter features and multiple object image set features, so as to facilitate the subsequent generation of multiple object dynamic data based on the object generation parameter features and multiple object image set features; the specific implementation method is as follows.

[0164] The step of using the multi-view feature, the video feature, and the object generation parameter encoding to generate the object generation parameter feature, and using the multi-view feature, the video feature, and the plurality of object image set encodings to generate the plurality of object image set features comprises:

[0165] Using a first cross attention module in the object generation model, the multi-view feature, the video feature and the object generation parameter encoding are subjected to feature fusion to generate the object generation parameter feature;

[0166] The second cross-attention module in the object generation model is used to perform feature fusion on the multi-view features, the video features and the multiple object image set encodings to generate the multiple object image set features.

[0167] Continuing with the above example, after using the MV2V-Adapter layer and Multi-view 3D Attention to obtain the corresponding feature output, add the output of the MV2V-Adapter layer to the output of "Multi-view 3D Attention" to obtain the combined feature;

[0168] Then, the combined feature and the text code are input into the first cross-attention layer for processing, and the first feature vector is output; and the combined feature and the image code are input into the second cross-attention layer for processing, and the second feature vector is output; subsequently, the first feature vector and the second feature vector can be input into the spatiotemporal attention module "Spatiotemporal Attention" in the MV-VDM model for processing.

[0169] Step 2: Based on the object generation parameter features and the multiple object image set features, generate the multiple object dynamic data corresponding to the multiple viewing angles.

[0170] In one or more embodiments provided in this specification, in the process of generating an object, the spatiotemporal attention unit in the object generation model can be used to process the object generation parameter features and the features of multiple object image sets to generate the multiple object dynamic data corresponding to the multiple perspectives. The specific implementation method is as follows.

[0171] The generating the plurality of object dynamic data corresponding to the plurality of viewing angles based on the object generation parameter feature and the plurality of object image set features comprises:

[0172] Using the spatiotemporal attention unit in the object generation model, performing temporal feature processing on the object generation parameter features and the plurality of object image set features to obtain temporal features;

[0173] Performing spatial feature processing on the object generation parameter features and the plurality of object image set features to obtain spatial features;

[0174] Performing feature fusion on the spatial features and the temporal features to obtain fusion features corresponding to the multiple perspectives;

[0175] The output unit in the object generation model is utilized to generate the multiple object dynamic data corresponding to the multiple perspectives based on the fusion features corresponding to the multiple perspectives.

[0176] Among them, the spatiotemporal attention unit can be understood as a unit for enhancing the spatiotemporal consistency of dynamic objects. The spatiotemporal attention unit can be one or more network layers in the object generation model, or the spatiotemporal attention unit can be a sub-model in the object generation model.

[0177] The temporal feature can be understood as the feature that characterizes the temporal information of a dynamic object; and the spatial feature can be understood as the feature that characterizes the spatial information of a dynamic object.

[0178] Continuing with the above example, the spatiotemporal attention unit may be a spatiotemporal attention module (Spatiotemporal Attention) which may include two parallel processing branches: one branch for spatial attention processing and the other branch for temporal attention processing.

[0179] Among them, Z∈R(b×f)×(n×h×w)×c,Z∈R(b×n×h×w)×f×c are the inputs of the spatial attention branch and the temporal attention branch respectively. The b, n, f, h, w, c are the batch size, view, number of frames, height, width, and number of channels of the features respectively.

[0180] For the spatial attention branch, the execution steps are:

[0181] (1) determining the input corresponding to the spatial attention branch from the first eigenvector and the second eigenvector, and encoding the input data to obtain encoded data;

[0182] (2) Perform multi-view self-attention enhancement on the encoded data to obtain the corresponding feature vector (i.e., spatial feature);

[0183] For the temporal attention branch, the execution steps are:

[0184] (1) extracting the input corresponding to the temporal attention branch from the first eigenvector and the second eigenvector, and encoding the input data to obtain encoded data;

[0185] (2) Perform temporal enhancement on the encoded data and output the corresponding feature vector (i.e., temporal feature).

[0186] After obtaining the features of the above two branches, they are fused to obtain fused features containing spatial information and temporal information, and the fused features are rendered into multiple 4D model videos corresponding to multiple perspectives through a rendering unit. Each 4D model video is a video in which the 3D model displays actions in a dynamic manner at a corresponding perspective.

[0187] The output unit may be a rendering unit that renders the fused features as video or point cloud data. For example, the output unit may be one or more network layers in an object generation model, or a sub-model in the object generation model.

[0188] In one or more embodiments provided in this specification, the using of the spatiotemporal attention unit in the object generation model to perform temporal feature processing on the object generation parameter features and the plurality of object image set features to obtain the temporal features includes:

[0189] Determine a spatiotemporal attention unit in the object generation model, and use a temporal coding layer in the spatiotemporal attention unit to encode the object generation parameter features and the plurality of object image set features to obtain a temporal coding;

[0190] Using the temporal motion layer in the spatiotemporal attention unit, feature mapping and weighting processing are performed on the time code to obtain the time feature;

[0191] Continuing with the above example, for the temporal attention branch, the execution steps are:

[0192] (1) The input corresponding to the temporal attention branch from the first eigenvector and the second eigenvector is input to the branch;

[0193] (2) Using the temporal encoding module “Temporal Encoding” (i.e., temporal encoding layer) to encode the input data to obtain encoded data;

[0194] (3) Use the temporal motion module (i.e., temporal motion layer) to perform feature mapping and weighting processing on the encoded data and output the corresponding feature vector.

[0195] It should be noted that the temporal motion module is a pre-trained video diffusion model.

[0196] The performing spatial feature processing on the object generation parameter feature and the plurality of object image set features to obtain the spatial feature comprises:

[0197] Determine a spatiotemporal attention unit in the object generation model, and use a spatial coding layer in the spatiotemporal attention unit to encode the object generation parameter features and the plurality of object image set features to obtain a spatial code;

[0198] Using the multi-view self-attention layer in the spatiotemporal attention unit, extracting features from the spatial encoding to obtain the spatial encoding;

[0199] Continuing with the above example, for the spatial attention branch, the execution steps are:

[0200] (1) The input corresponding to the spatial attention branch from the first eigenvector and the second eigenvector is input to the branch;

[0201] (2) Using the spatial encoding module “Spatial Encoding” (i.e., spatial encoding layer) to encode the input data to obtain encoded data;

[0202] (3) Input the encoded data into the multi-view self-attention module “Multi-view 3D Attention” (i.e., multi-view self-attention layer) for processing and output the corresponding feature vector.

[0203] The step of fusing the spatial features and the temporal features to obtain fusion features corresponding to the multiple perspectives includes:

[0204] The feature fusion layer in the spatiotemporal attention unit is used to fuse the spatial features and the temporal features to obtain fusion features corresponding to the multiple perspectives.

[0205] Continuing with the above example, based on the feature vectors output by the two branches, an alpha blending layer is used to fuse the two feature vectors to obtain a feature vector with enhanced spatiotemporal consistency (i.e., the output of the spatiotemporal attention module).

[0206] Step three: fusing the dynamic data of the multiple objects to obtain the dynamic object.

[0207] When the object dynamic data is an object dynamic video, multiple object dynamic videos can be fused to obtain a dynamic object; or, when the object dynamic data is an object dynamic point cloud data, multiple object dynamic point cloud data can be fused to obtain a dynamic object.

[0208] In one or more embodiments provided in this specification, in the process of determining a dynamic object, initial dynamic data can be determined based on multiple object dynamic data, and the initial dynamic data can be optimized to obtain optimized dynamic data, thereby ensuring the accuracy and vividness of the dynamic data.

[0209] In one or more embodiments provided in this specification, the object dynamic data is an object dynamic video;

[0210] The step of fusing the dynamic data of the multiple objects to obtain the dynamic object comprises:

[0211] Performing video fusion on the multiple object dynamic videos to obtain an initial dynamic object;

[0212] The initial dynamic object is optimized by using the object optimization unit in the object generation model to obtain the dynamic object.

[0213] The initial dynamic object may be understood as a rough dynamic object or an incomplete dynamic object; the dynamic object is obtained by optimizing the initial object, and the dynamic object may be a fine dynamic object or a complete dynamic object.

[0214] The object optimization unit may be a unit for optimizing the initial dynamic object, and the object optimization unit may be one or more network layers in the object generation model, or the object optimization unit may be a sub-model in the object generation model; for example, the object optimization unit may be a network layer that performs a distillation model operation (i.e., SDS, 4D-SDS) based on a rough motion video to obtain a 4DGS; wherein 4D-SDS may be understood as a strategy for obtaining a 4D object based on SDS.

[0215] Continuing with the above example, in the process of 4D object generation, the object generation method provided in this specification constructs a 4D optimization pipeline; the pipeline is a 4D optimization pipeline that combines reconstruction and 4D-SDS distillation; the specific execution method of the pipeline is: based on the multi-view video of orthogonal perspectives generated by the trained MV-VDM model, coarse motion is directly reconstructed based on the multi-view video; and then the 4D-SDS mechanism is introduced to distill the MV-VDM model from any perspective, so as to model fine-level motion (i.e., 4D video).

[0216] That is to say, based on multi-view video, coarse motion is reconstructed; 4D-SDS is introduced to distill the MV-VDM model from coarse motion video of any viewpoint, thereby modeling fine-level motion and obtaining 4D video (4DGS).

[0217] In one or more embodiments provided in this specification, before using the object generation model to generate a dynamic object, the object generation model needs to be trained. The specific training is as follows.

[0218] Before inputting the object generation parameters and the plurality of object image sets into the object generation model to obtain the dynamic object corresponding to the target object, the method further includes:

[0219] Determine training samples and sample labels of a to-be-trained object generation model, wherein the training samples include object generation parameters of a sample object and a plurality of sample object image sets corresponding to a plurality of time points, the plurality of time points are arranged in chronological order, and each sample object image set includes sample object images of the sample object at a plurality of viewing angles at a corresponding time point;

[0220] Processing the plurality of sample object image sets according to the object generation parameters using the object generation model to be trained, determining a plurality of sample object dynamic data corresponding to the plurality of viewing angles, and generating sample dynamic objects corresponding to the sample objects according to the plurality of sample object dynamic data;

[0221] Based on the sample dynamic object and the sample label, the parameters of the object generation model to be trained are adjusted to obtain the trained object generation model.

[0222] Among them, the sample object can be understood as the target object serving as a training sample, the sample object image can be understood as the image of the sample object, and the sample object image set can be understood as a set composed of sample object images. For the interpretation of the sample object image set, reference can be made to the object image set in the above embodiments.

[0223] The sample dynamic object can be understood as a dynamic object determined based on the training sample.

[0224] The sample label can be understood as a dynamic object serving as a sample label, for example, a 4DGS model, where the sample label is used to calculate a loss function; the loss function is calculated based on the sample label and the sample dynamic object, and the loss function is used to adjust model parameters of the model generated by the training object.

[0225] Using the above example, in the training phase, the training target of the MV-VDM model is: a multi-viewmulti-frame video with camera parameters rendered based on a 4D model. After this model is trained, it is used as the base model for the 4D optimization part.

[0226] The training input for the 4D optimization part: 4 views of an existing static 3D model (3DGS) and a text prompt;

[0227] Input in the testing phase: 4 views of an existing static 3D model (3DGS) and a text prompt;

[0228] The output of the 4D generation framework: a 4D model, in this case a 4DGS model;

[0229] Overall training process:

[0230] 1. Train MV-VDM first; pre-train MV-VDM on 4D dataset; after training, use it as the base model for subsequent 4D driver optimization;

[0231] 2. Regardless of whether it is generated or reconstructed, a static 3D model (represented as 3DGS) is obtained, and four orthogonal views are rendered based on the static 3D model and sent to MV-VDM to generate a multi-frame video of four orthogonal views;

[0232] 3. Use the generated multi-view video to perform rough motion reconstruction, i.e., coarse motion reconstruction, to obtain rough quality 4D motion generation (4DGS);

[0233] 4. Introduce 4D-SDS loss and combine it with text prompt to optimize the 4DGS in the previous step at a fine-level.

[0234] 5. The optimized 4DGS model is obtained. This model is the action modeling of the initial static 3DGS and has good spatiotemporal consistency;

[0235] It should be noted that the above-mentioned model training steps are the same as the steps of object generation performed by the object generation method in one or more of the above-mentioned embodiments. For the explanation of the model training steps, please refer to the corresponding or corresponding contents in one or more of the above-mentioned embodiments.

[0236] The object generation method in one or more embodiments of the present specification provides a 4D generation framework for driving any static 3D model. To address the problem of inconsistency in time and space of 4D objects during the modeling process, a large 4D model is pre-trained on large-scale 4D data to jointly model temporal and spatial consistency. Compared with the above-mentioned 4D content generation solution, which uses 3D and video models for separate supervision of appearance and motion, the present specification provides a supervision signal for driving a 4D generation framework for any static 3D model that is more consistent in time and space.

[0237] At the same time, in the above-mentioned 4D content generation solution, the appearance is optimized by relying on a text-conditioned or single-view image conditioned model, which makes it difficult to apply to various existing 3D objects with multi-view details; this specification provides a 4D generation framework for driving any static 3D model, using the multi-view of the static 3D model as a condition, which can better maintain the appearance details of the original model, and thus has a wider range of applications.

[0238] Therefore, this specification provides a 4D generation framework that drives any static 3D model, and achieves better 4D generation results than the comparative methods by designing a more efficient algorithm framework and training on 4D big data;

[0239] The object generation method provided by one or more embodiments of the present specification, in the process of generating dynamic objects by using an object generation model, first, by determining object generation parameters, multiple time points, and multiple object image sets, which are large in number and rich in types, it is convenient to subsequently generate accurate and vivid dynamic objects; then, the object generation model is used to process the multiple object image sets according to the object generation parameters, so as to determine multiple object dynamic data corresponding to multiple perspectives, and based on the content-rich data of the multiple object dynamic data, accurate and vivid dynamic objects can be generated, thereby realizing the accurate generation of dynamic objects by using a neural network model, meeting the actual needs of users for generating dynamic objects, and reducing the time and labor cost of users for generating dynamic objects.

[0240] The following combination Figure 3 Taking the application of the object generation method provided in this specification in driving any static 3D model to generate 4D content as an example, the object generation method is further described. Figure 3 A flowchart of a method for generating an object according to an embodiment of the present specification is shown. Figure 3 It can be seen that the object generation model can be a 4D generation framework that drives any static 3D model. The framework includes a 4D generation base model MV-VDM and a 4D optimization pipeline that combines reconstruction and 4D-SDS distillation. Correspondingly, the model training is also divided into two parts, namely: 1. Multi-view video generation model training; 2. Static object drive training.

[0241] The method for "multi-view video generation model training" is as follows:

[0242] Step 1: Determine the training data.

[0243] The training samples include:

[0244] Multi-perspective renderings of 3D models: The multi-perspective renderings include clean renderings and noise renderings; the multi-perspective refers to: orthogonal 4 perspectives.

[0245] Timestamp: This timestamp is the timestamp corresponding to each rendering in the multi-view rendering; when processing a sequence of multi-view renderings, the model will process the data step by step according to timesteps, that is, process the renderings in the sequence one by one or batch by batch.

[0246] Camera parameters: mainly include the camera's focal length, camera position, and orientation information.

[0247] Prompt text: That is, "An eagle is flying."

[0248] Among them, the sample label is: multi-view video corresponding to the 3D model (ground truth).

[0249] Step 2: Input the training sample into the MV-VDM model, and process it using the "multi-view self-attention module" and "multi-view image to video adaptation module" in the model to obtain a feature vector.

[0250] Specifically, first, the multi-view self-attention module is used to process training samples such as multi-view renderings, timestamps, and camera parameters, and the processed feature vectors are output;

[0251] Secondly, the multi-view image to video adaptation module is executed as follows:

[0252] (1) The noise frames are concatenated along the spatial dimension to obtain concatenated frames; wherein the spatial dimension refers to: a rendering of the same action shot from four perspectives, that is, a column of renderings of the above-mentioned “multi-perspective renderings”.

[0253] (2) Using the self-attention layer of the 3D diffusion model, we extract features from the serial frames, corresponding timestamps, and camera parameters to obtain the corresponding projection matrix W Q ′、W K , W V ;

[0254] (3) Using the multi-head attention mechanism, the projection matrix W Q ′、W K , W V Process and obtain the projection matrix W o ′.

[0255] (4) Add the output of the MV2V-Adapter layer to the output of “Multi-view 3D Attention” and record it as “combined feature”

[0256] Step 3: Use two cross-attention layers to process the “combined features” to align textual cues and maintain the identity features of the 3D model.

[0257] The specific implementation method is:

[0258] (1) Using the text encoder in the MV-VDM model, the prompt text is encoded to obtain the text encoding;

[0259] (2) Using the image encoder in the MV-VDM model, the multi-view rendering image is encoded to obtain image encoding;

[0260] (3) Input the “comprehensive feature vector” and the text encoding into the first cross-attention layer for processing and output the first feature vector;

[0261] (4) The “comprehensive feature vector” and the image code are input into the second cross-attention layer for processing, and the second feature vector is output;

[0262] Step 4: Input the first eigenvector and the second eigenvector into the spatiotemporal attention module “Spatiotemporal Attention” in the MV-VDM model for processing.

[0263] Based on the above figures, it can be seen that the spatiotemporal attention module in this article includes two parallel branches: the left branch is for spatial attention and the right branch is for temporal attention.

[0264] Among them, Z∈R(b×f)×(n×h×w)×c,Z∈R(b×n×h×w)×f×c are the inputs of the spatial attention branch and the temporal attention branch respectively. The b, n, f, h, w, c are the batch size, view, number of frames, height, width, and number of channels of the image features respectively.

[0265] For the spatial attention branch, the execution steps are:

[0266] (1) The input corresponding to the spatial attention branch from the first eigenvector and the second eigenvector is input to the branch;

[0267] (2) using a spatial encoding module to encode the input data to obtain encoded data;

[0268] (3) Input the encoded data into the multi-view self-attention module for processing and output the corresponding feature vector;

[0269] For the temporal attention branch, the execution steps are:

[0270] (1) The input corresponding to the temporal attention branch from the first eigenvector and the second eigenvector is input to the branch;

[0271] (2) using a time coding module to encode the input data to obtain coded data;

[0272] (3) Use the temporal motion module to process the encoded data and output the corresponding feature vector.

[0273] The specific processing method of the timing motion module is:

[0274] First, the encoded data is mapped to the target dimension by mapping the entry;

[0275] Secondly, through the self-attention mechanism, the encoded data mapped to the target dimension is weighted averaged to obtain the weighted encoded data;

[0276] Finally, the weighted encoded data is mapped to the original dimension by mapping the output, thereby outputting the feature vector.

[0277] It should be noted that the temporal motion module is a pre-trained video diffusion model.

[0278] Finally, based on the feature vectors output by the two branches, an alpha blending layer is used to fuse the two feature vectors to obtain a feature vector with enhanced spatiotemporal consistency (i.e., the output of the spatiotemporal attention module).

[0279] Step 5: Utilize the output layer of the MV-VDM model to render the feature vectors that enhance spatiotemporal consistency and generate a multi-view video.

[0280] Step 6: Calculate the loss function based on the multi-view video and sample labels, and adjust the model parameters of the MV-VDM model based on the loss function until the model training stop condition is reached.

[0281] Among them, the method for "static object driving training" is as follows:

[0282] Step 1: Determine the training data.

[0283] The training sample includes: four orthogonal views of an existing static 3D model (3DGS) and corresponding prompt text (An eagle is flying); the 3D model is a three-dimensional Gaussian splash generated / reconstructed.

[0284] Among them, the sample label is: real 4D video (4DGS).

[0285] It should be noted that the 3D model (3DGS) in this solution is generated or reconstructed based on the pre-trained 3D diffusion model (MVDream).

[0286] Step 2: Output the four orthogonal views of the 3D model (3DGS) to the trained MV-VDM model to obtain a multi-view video of the 3D object.

[0287] Step 3: The 4D generation framework reconstructs coarse motion based on multi-view video.

[0288] Step 4: The 4D generation framework introduces 4D-SDS (four-dimensional fractional distillation sampling) to distill the MV-VDM model from the rough action video of any perspective, allowing the MV-VDM model to learn the details of the original 3D model, thereby modeling fine motion movements and obtaining fine 4D video (4DGS, four-dimensional Gaussian splatter).

[0289] It should be noted that the 4D-SDS is implemented through one or more network layers in the 4D generation framework, and the network layer (i.e., the object optimization unit) is used to perform the distillation MV-VDM model to obtain the 4DGS operation.

[0290] Step 5: Calculate the loss function using the predicted video (4DGS), prompt text, and sample label (ground truth), and optimize the motion field of the static object based on the loss function until the stopping condition is reached.

[0291] It should be noted that the motion field is a complete space for simulating natural motion in a 4D generative framework, which is used to describe every possible way for the model to move from a given state, thereby generating natural movements.

[0292] After completing the training of the 4D generation framework through the above training steps, in the process of applying the 4D generation framework, a 3D model (3DGS) and a prompt text can be input into the 4D generation framework to obtain a 4D video (4DGS) corresponding to the 3D model, so that the user can observe the 4D model from different angles and time points.

[0293] Based on the above embodiments, it can be seen that the present specification provides a 4D generation framework for driving any static 3D model. The 4D generation framework provides a 4D generation foundation model (i.e., MV-VDM). Based on this foundation model, an efficient pipeline (i.e., 4D optimization pipeline) is proposed to achieve high-quality driving of any existing static 3D model through a sentence of text, thereby obtaining a 4D object; specifically, based on the trained MV-VDM, a multi-view multi-frame temporally and spatially consistent video can be generated; and, based on the MV-VDM, an efficient 4D optimization pipeline is proposed, which combines reconstruction and 4D-SDS distillation strategies to better maintain the details of the original 3D model and can drive the existing 3D model with high quality according to the text.

[0294] See also Figure 4 , Figure 4 A flowchart of a 4D object generation method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.

[0295] Step 402: determining object generation parameters and a plurality of 3D object image sets corresponding to a plurality of time points, wherein the plurality of time points are arranged in chronological order, and each 3D object image set includes 3D object images of the 3D object at a corresponding time point and from a plurality of perspectives;

[0296] Step 404: Input the object generation parameters and the multiple 3D object image sets into the object generation model to obtain a dynamic 4D object corresponding to the 3D object, wherein the object dimension of the 3D object is smaller than the object dimension of the dynamic 4D object, the object generation model processes the multiple 3D object image sets according to the object generation parameters, determines multiple object dynamic data corresponding to the multiple perspectives, and generates the dynamic 4D object according to the multiple object dynamic data.

[0297] The 4D object generation method provided by one or more embodiments of the present specification, in the process of generating a dynamic 4D object by using an object generation model, firstly, by determining a large number of data with rich types, such as object generation parameters, multiple time points, and multiple 3D object image sets, it is convenient to subsequently generate accurate and vivid dynamic 4D objects; then, the object generation model is used to process the multiple 3D object image sets according to the object generation parameters, so as to determine multiple object dynamic data corresponding to multiple perspectives, and based on the rich content of the multiple object dynamic data, it is possible to generate accurate and vivid dynamic 4D objects, thereby realizing accurate generation of dynamic 4D objects by using a neural network model, meeting the actual needs of users to generate dynamic 4D objects, and reducing the time and labor cost of users to generate dynamic 4D objects.

[0298] The above is a schematic scheme of a 4D object generation method of this embodiment. It should be noted that the technical scheme of the 4D object generation method and the technical scheme of the above-mentioned object generation method belong to the same concept, and the details not described in detail in the technical scheme of the 4D object generation method can be referred to the description of the technical scheme of the above-mentioned object generation method.

[0299] See also Figure 5 , Figure 5 A flowchart of another object generation method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.

[0300] Step 502: determining object generation parameters and a plurality of object image sets corresponding to a plurality of time points, wherein the plurality of time points are arranged in chronological order, and each object image set includes object images of a target object at a corresponding time point and from a plurality of perspectives;

[0301] Step 504: inputting the object generation parameters and the plurality of object image sets into an object generation model, wherein the object generation model includes a dynamic data determination unit and an object generation unit;

[0302] Step 506: using the dynamic data determination unit to process the plurality of object image sets according to the object generation parameters to determine a plurality of object dynamic data corresponding to the plurality of viewing angles;

[0303] Step 508: Utilize the object generation unit to generate the dynamic object according to the plurality of object dynamic data.

[0304] The object generation method provided in one or more embodiments of the present specification can utilize an object generation model including a dynamic data determination unit and an object generation unit for processing during the process of generating dynamic objects; first, by determining object generation parameters, multiple time points, and multiple object image sets, which are data of a large number and rich types, it is convenient to subsequently generate accurate and vivid dynamic objects; then, the dynamic data determination unit in the object generation model is used to process multiple object image sets according to the object generation parameters, so as to determine multiple object dynamic data corresponding to multiple perspectives; and, the object generation unit can generate accurate and vivid dynamic objects based on the content-rich data of multiple object dynamic data, thereby realizing the accurate generation of dynamic objects using a neural network model, meeting the actual needs of users for generating dynamic objects, and reducing the time and labor cost of users for generating dynamic objects.

[0305] The above is a schematic scheme of another object generation method of this embodiment. It should be noted that the technical scheme of the other object generation method and the technical scheme of the above object generation method belong to the same concept, and the details of the technical scheme of the other object generation method that are not described in detail can all be referred to the description of the technical scheme of the above object generation method.

[0306] See also Figure 6 , Figure 6 A flowchart of another object generation method provided according to an embodiment of the present specification is shown. The object generation method is applied to a server and specifically includes the following steps.

[0307] Step 602: receiving object generation parameters sent by a client, and a plurality of object image sets corresponding to a plurality of time points, wherein the plurality of time points are arranged in chronological order, and each object image set includes object images of a target object at a corresponding time point and from a plurality of perspectives;

[0308] Step 604: inputting the object generation parameters and the plurality of object image sets into an object generation model to obtain a dynamic object corresponding to the target object, wherein the object dimension of the target object is smaller than the object dimension of the dynamic object, the object generation model processes the plurality of object image sets according to the object generation parameters, determines a plurality of object dynamic data corresponding to the plurality of perspectives, and generates the dynamic object according to the plurality of object dynamic data;

[0309] Step 606: Send the dynamic object to the client.

[0310] One or more embodiments of the present specification provide an object generation method applied to a server. After receiving the object generation parameters sent by the client and the multiple object image sets corresponding to multiple time points, the object generation model can be used to process the data to generate a dynamic object. In the process of generating the dynamic object using the object generation model, first, the object generation parameters, multiple time points, and multiple object image sets, which are large in number and rich in types, are determined to facilitate the subsequent generation of accurate and vivid dynamic objects. Then, the object generation model is used to process the multiple object image sets according to the object generation parameters to determine the multiple object dynamic data corresponding to multiple perspectives, and based on the rich content of the multiple object dynamic data, accurate and vivid dynamic objects can be generated, thereby realizing the accurate generation of dynamic objects using a neural network model. The dynamic object is then sent to the client, thereby meeting the actual needs of users to generate dynamic objects and reducing the time and labor cost of users to generate dynamic objects.

[0311] The above is a schematic scheme of another object generation method of this embodiment. It should be noted that the technical scheme of the another object generation method and the technical scheme of the above object generation method belong to the same concept, and the details not described in detail in the technical scheme of the another object generation method can be referred to the description of the technical scheme of the above object generation method.

[0312] See also Figure 7 , Figure 7 A flowchart of an object generation model training method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.

[0313] Step 702: Determine training samples and sample labels for the object generation model to be trained, wherein the training samples include sample object generation parameters and a plurality of sample object image sets corresponding to a plurality of time points, the plurality of time points are arranged in chronological order, and each sample object image set includes sample object images of the sample object at a plurality of viewing angles at a corresponding time point;

[0314] Step 704: using the to-be-trained object generation model to process the plurality of sample object image sets according to the sample object generation parameters, determining a plurality of sample object dynamic data corresponding to the plurality of viewing angles, and generating sample dynamic objects corresponding to the sample objects according to the plurality of sample object dynamic data;

[0315] Step 706: Based on the sample dynamic object and the sample label, adjust the parameters of the object generation model to be trained to obtain a trained object generation model.

[0316] The object generation model training method provided in one or more embodiments of the present specification requires the use of training samples and sample labels to train the object generation model to be trained before using the object generation model to generate dynamic objects; in the process of model training, first, sample data such as sample object generation parameters, multiple time points, and multiple sample object image sets, which are large in number and rich in types, are determined to facilitate subsequent training to obtain an object generation model that can generate accurate and vivid dynamic objects; secondly, the object generation model to be trained is used to process multiple sample object image sets according to the sample object generation parameters, so as to determine multiple sample object dynamic data corresponding to multiple perspectives, and based on the content-rich data of the multiple sample object dynamic data, accurate and vivid sample dynamic objects can be generated; finally, the sample dynamic objects and sample labels are used to adjust the parameters of the object generation model to be trained, so as to obtain an object generation model that can generate accurate and vivid dynamic objects; the accurate generation of dynamic objects using a neural network model is achieved, the actual needs of users to generate dynamic objects are met, and the time and labor cost of users to generate dynamic objects are reduced.

[0317] The above is a schematic scheme of the object generation model training method of this embodiment. It should be noted that the technical scheme of the object generation model training method and the technical scheme of the above-mentioned object generation method belong to the same concept, and the details not described in detail in the technical scheme of the object generation model training method can be referred to the description of the technical scheme of the above-mentioned object generation method.

[0318] See also Figure 8 , Figure 8 A flow chart of a video generation method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.

[0319] Step 802: receiving the object video generated text and a plurality of object image sets corresponding to a plurality of time points sent by the client, wherein the plurality of time points are arranged in chronological order, and each object image set includes object images of a target object at a corresponding time point and from a plurality of perspectives;

[0320] Step 804: inputting the object video generation text and the plurality of object image sets into the object generation model to obtain the object video corresponding to the target object, wherein the object video is a video observed from any perspective, the object generation model processes the plurality of object image sets according to the object video generation text, determines a plurality of object dynamic data corresponding to the plurality of perspectives, and generates the arbitrary perspective object video according to the plurality of object dynamic data;

[0321] Step 806: Send the object video to the client.

[0322] The object video generation text can refer to the explanation of the object generation parameters in the above embodiment, which will not be repeated here; the object video can refer to the explanation of the dynamic object in the above embodiment, which will not be repeated here;

[0323] One or more embodiments of the present specification provide a video generation method for a video generation platform. In the process of generating an object video, by determining a large number of data with rich types, such as an object video generation text, multiple time points, and multiple object image sets, it is convenient to subsequently generate an accurate and vivid object video; then, the object generation model is used to process the multiple object image sets according to the object video generation text, so as to determine multiple object dynamic data corresponding to multiple perspectives, and based on the content-rich data of the multiple object dynamic data, an accurate and vivid object video can be generated, thereby realizing the accurate generation of the object video using a neural network model; by sending the object video to the client, the actual needs of users to generate object videos are met, and the time and labor cost of users to generate object videos are reduced.

[0324] The above is a schematic scheme of the object generation platform of this embodiment. It should be noted that the technical scheme of the object generation platform and the technical scheme of the above-mentioned object generation method belong to the same concept, and the details not described in detail in the technical scheme of the object generation platform can be referred to the description of the technical scheme of the above-mentioned object generation method.

[0325] Corresponding to the above method embodiment, this specification also provides an object generation device embodiment, the device comprising:

[0326] A data determination module is configured to determine object generation parameters and a plurality of object image sets corresponding to a plurality of time points, wherein the plurality of time points are arranged in chronological order, and each object image set includes object images of a target object at a corresponding time point and from a plurality of perspectives;

[0327] The object generation module is configured to input the object generation parameters and the multiple object image sets into an object generation model to obtain a dynamic object corresponding to the target object, wherein the object dimension of the target object is smaller than the object dimension of the dynamic object, the object generation model processes the multiple object image sets according to the object generation parameters, determines the multiple object dynamic data corresponding to the multiple perspectives, and generates the dynamic object according to the multiple object dynamic data.

[0328] Optionally, the object generation module is further configured to:

[0329] Inputting the object generation parameters and the plurality of object image sets into an object generation model, and determining object generation parameter features of the object generation parameters and a plurality of object image set features of the plurality of object image sets using the object generation model;

[0330] Based on the object generation parameter features and the plurality of object image set features, generating the plurality of object dynamic data corresponding to the plurality of viewing angles;

[0331] The dynamic data of the multiple objects are fused to obtain the dynamic object.

[0332] Optionally, the object generation module is further configured to:

[0333] Using the object generation model, performing multi-view feature conversion on the multiple object image sets to obtain multi-view features corresponding to the multiple object image sets, and performing video feature conversion on the multiple object image sets to obtain video features corresponding to the multiple object image sets;

[0334] Determining, using the object generation model, object generation parameter encoding of the object generation parameter and a plurality of object image set encodings of the plurality of object image sets;

[0335] The object generation parameter feature is generated by using the multi-view feature, the video feature and the object generation parameter encoding, and the multiple object image set features are generated by using the multi-view feature, the video feature and the multiple object image set encoding.

[0336] Optionally, the object generation module is further configured to:

[0337] Using the multi-view attention unit in the object generation model, extracting features from each of the multiple object image sets to obtain the multi-view features corresponding to the multiple object image sets;

[0338] The video feature extraction unit in the object generation model is used to perform video feature conversion on each of the object image sets to obtain video features corresponding to the multiple object image sets.

[0339] Optionally, the object generation module is further configured to:

[0340] Using the video feature extraction unit in the object generation model, the object images included in each of the plurality of object image sets are sorted to obtain a plurality of object image sequences;

[0341] Determine an image acquisition parameter corresponding to each object image set, and determine the multiple time points arranged in time sequence as a time series, wherein the image acquisition parameter is a parameter used to acquire the object image;

[0342] Using the self-attention layer in the video feature extraction unit, feature extraction is performed on each object image sequence, image acquisition parameter and time series to obtain an image sequence feature matrix of each object image sequence, a parameter feature matrix of the image acquisition parameter and a time series feature matrix of the time series;

[0343] The multi-head attention layer in the video feature extraction unit is used to perform feature fusion on the image sequence feature matrix, the parameter feature matrix and the time series feature matrix to obtain video features corresponding to the multiple object image sets.

[0344] Optionally, the object generation parameter is an object generation text, and the object generation parameter code is an object generation text code;

[0345] The object generation module is further configured to:

[0346] Using the text encoder in the object generation model, the object generation text is encoded to obtain the object generation parameter encoding;

[0347] Using an image encoder in the object generation model, performing image encoding on the plurality of object image sets to obtain encodings of the plurality of object image sets;

[0348] The object generation module is further configured to:

[0349] Using a first cross attention module in the object generation model, the multi-view feature, the video feature and the object generation parameter encoding are subjected to feature fusion to generate the object generation parameter feature;

[0350] The second cross-attention module in the object generation model is used to perform feature fusion on the multi-view features, the video features and the multiple object image set encodings to generate the multiple object image set features.

[0351] Optionally, the object generation module is further configured to:

[0352] Using the spatiotemporal attention unit in the object generation model, performing temporal feature processing on the object generation parameter features and the plurality of object image set features to obtain temporal features;

[0353] Performing spatial feature processing on the object generation parameter features and the plurality of object image set features to obtain spatial features;

[0354] Performing feature fusion on the spatial features and the temporal features to obtain fusion features corresponding to the multiple perspectives;

[0355] The output unit in the object generation model is utilized to generate the multiple object dynamic data corresponding to the multiple perspectives based on the fusion features corresponding to the multiple perspectives.

[0356] Optionally, the object generation module is further configured to:

[0357] Determine a spatiotemporal attention unit in the object generation model, and use a temporal coding layer in the spatiotemporal attention unit to encode the object generation parameter features and the plurality of object image set features to obtain a temporal coding;

[0358] Using the temporal motion layer in the spatiotemporal attention unit, feature mapping and weighting processing are performed on the time code to obtain the time feature;

[0359] The object generation module is further configured to:

[0360] Determine a spatiotemporal attention unit in the object generation model, and use a spatial coding layer in the spatiotemporal attention unit to encode the object generation parameter features and the plurality of object image set features to obtain a spatial code;

[0361] Using the multi-view self-attention layer in the spatiotemporal attention unit, extracting features from the spatial encoding to obtain the spatial encoding;

[0362] The object generation module is further configured to:

[0363] The feature fusion layer in the spatiotemporal attention unit is used to fuse the spatial features and the temporal features to obtain fusion features corresponding to the multiple perspectives.

[0364] Optionally, the object dynamic data is an object dynamic video;

[0365] The object generation module is further configured to:

[0366] Performing video fusion on the multiple object dynamic videos to obtain an initial dynamic object;

[0367] The initial dynamic object is optimized by using the object optimization unit in the object generation model to obtain the dynamic object.

[0368] Optionally, the data determination module is further configured to:

[0369] Receiving the object generation parameters and the multiple object image sets corresponding to multiple time points sent by the client, wherein the object generation parameters and the multiple object image sets are generated by the client based on a data sending operation performed by a user on a data sending page;

[0370] The object generation device further includes an object sending module, which is configured to:

[0371] The dynamic object is sent to the client and displayed to the user using the data sending page in the client.

[0372] Optionally, the object generation device further includes a model training module configured to:

[0373] Determine training samples and sample labels of a to-be-trained object generation model, wherein the training samples include object generation parameters of a sample object and a plurality of sample object image sets corresponding to a plurality of time points, the plurality of time points are arranged in chronological order, and each sample object image set includes sample object images of the sample object at a plurality of viewing angles at a corresponding time point;

[0374] Processing the plurality of sample object image sets according to the object generation parameters using the object generation model to be trained, determining a plurality of sample object dynamic data corresponding to the plurality of viewing angles, and generating sample dynamic objects corresponding to the sample objects according to the plurality of sample object dynamic data;

[0375] Based on the sample dynamic object and the sample label, the parameters of the object generation model to be trained are adjusted to obtain the trained object generation model.

[0376] The object generation device provided by one or more embodiments of the present specification, in the process of generating dynamic objects using an object generation model, first, by determining object generation parameters, multiple time points, and multiple object image sets, which are large in number and rich in types, to facilitate the subsequent generation of accurate and vivid dynamic objects; then, the object generation model is used to process the multiple object image sets according to the object generation parameters, so as to determine multiple object dynamic data corresponding to multiple perspectives, and based on the content-rich data of the multiple object dynamic data, accurate and vivid dynamic objects can be generated, thereby realizing the accurate generation of dynamic objects using a neural network model, meeting the actual needs of users for generating dynamic objects, and reducing the time and labor cost of users for generating dynamic objects.

[0377] The above is a schematic scheme of an object generation device of this embodiment. It should be noted that the technical scheme of the object generation device and the technical scheme of the object generation method described above are of the same concept, and the details not described in detail in the technical scheme of the object generation device can be found in the description of the technical scheme of the object generation method described above.

[0378] Corresponding to the above method embodiment, this specification also provides a 4D object generation device embodiment, the device comprising:

[0379] a data determination module configured to determine object generation parameters and a plurality of 3D object image sets corresponding to a plurality of time points, wherein the plurality of time points are arranged in chronological order, and each 3D object image set includes 3D object images of the 3D object at a corresponding time point and from a plurality of perspectives;

[0380] The object generation module is configured to input the object generation parameters and the multiple 3D object image sets into an object generation model to obtain a dynamic 4D object corresponding to the 3D object, wherein the object dimension of the 3D object is smaller than the object dimension of the dynamic 4D object, the object generation model processes the multiple 3D object image sets according to the object generation parameters, determines multiple object dynamic data corresponding to the multiple perspectives, and generates the dynamic 4D object according to the multiple object dynamic data.

[0381] The 4D object generation device provided by one or more embodiments of the present specification, in the process of generating a dynamic 4D object by using an object generation model, firstly, determines a large number of data with rich types, such as object generation parameters, multiple time points, and multiple 3D object image sets, so as to facilitate the subsequent generation of accurate and vivid dynamic 4D objects; then, uses the object generation model to process the multiple 3D object image sets according to the object generation parameters, so as to determine multiple object dynamic data corresponding to multiple perspectives, and can generate accurate and vivid dynamic 4D objects based on the rich content of the multiple object dynamic data, thereby realizing the accurate generation of dynamic 4D objects by using a neural network model, meeting the actual needs of users to generate dynamic 4D objects, and reducing the time and labor cost of users to generate dynamic 4D objects.

[0382] The above is a schematic scheme of a 4D object generation device of this embodiment. It should be noted that the technical scheme of the 4D object generation device and the technical scheme of the above-mentioned 4D object generation method belong to the same concept, and the details not described in detail in the technical scheme of the 4D object generation device can be referred to the description of the technical scheme of the above-mentioned 4D object generation method.

[0383] Corresponding to the above method embodiment, this specification also provides another object generation device embodiment, the device comprising:

[0384] A first data determination module is configured to determine object generation parameters and a plurality of object image sets corresponding to a plurality of time points, wherein the plurality of time points are arranged in chronological order, and each object image set includes object images of a target object at a corresponding time point and from a plurality of perspectives;

[0385] a data input module configured to input the object generation parameters and the plurality of object image sets into an object generation model, wherein the object generation model includes a dynamic data determination unit and an object generation unit;

[0386] A second data determination module is configured to use the dynamic data determination unit to process the plurality of object image sets according to the object generation parameters to determine a plurality of object dynamic data corresponding to the plurality of viewing angles;

[0387] The object generation module is configured to use the object generation unit to generate the dynamic object according to the multiple object dynamic data.

[0388] The object generation device provided by one or more embodiments of the present specification can utilize an object generation model including a dynamic data determination unit and an object generation unit for processing during the process of generating dynamic objects; first, by determining object generation parameters, multiple time points, and multiple object image sets, which are large in number and rich in types, it is convenient to subsequently generate accurate and vivid dynamic objects; then, the dynamic data determination unit in the object generation model is utilized to process multiple object image sets according to the object generation parameters, thereby determining multiple object dynamic data corresponding to multiple perspectives; and, the object generation unit can generate accurate and vivid dynamic objects based on the content-rich data of multiple object dynamic data, thereby realizing the accurate generation of dynamic objects using a neural network model, meeting the actual needs of users for generating dynamic objects, and reducing the time and labor cost of users for generating dynamic objects.

[0389] The above is a schematic scheme of another object generation device of this embodiment. It should be noted that the technical scheme of the another object generation device and the technical scheme of the another object generation method mentioned above belong to the same concept, and the details of the technical scheme of the another object generation device that are not described in detail can all be referred to the description of the technical scheme of the another object generation method mentioned above.

[0390] Corresponding to the above method embodiment, this specification also provides another object generation device embodiment, which is applied to a server and includes:

[0391] A data receiving module is configured to receive object generation parameters sent by a client, and a plurality of object image sets corresponding to a plurality of time points, wherein the plurality of time points are arranged in chronological order, and each object image set includes object images of a target object at a corresponding time point and from a plurality of perspectives;

[0392] an object generation module, configured to input the object generation parameters and the plurality of object image sets into an object generation model, obtain a dynamic object corresponding to the target object, wherein the object dimension of the target object is smaller than the object dimension of the dynamic object, the object generation model processes the plurality of object image sets according to the object generation parameters, determines a plurality of object dynamic data corresponding to the plurality of perspectives, and generates the dynamic object according to the plurality of object dynamic data;

[0393] The object sending module is configured to send the dynamic object to the client.

[0394] The object generation device applied to the server provided by one or more embodiments of the present specification can use the object generation model to process the data after receiving the object generation parameters sent by the client and the multiple object image sets corresponding to multiple time points, so as to generate a dynamic object; in the process of using the object generation model to generate the dynamic object, first, determine the object generation parameters, multiple time points and multiple object image sets, which are large in number and rich in types, to facilitate the subsequent generation of accurate and vivid dynamic objects; then, use the object generation model to process the multiple object image sets according to the object generation parameters, so as to determine the multiple object dynamic data corresponding to multiple perspectives, and based on the rich content of the multiple object dynamic data, it is possible to generate accurate and vivid dynamic objects, thereby realizing the accurate generation of dynamic objects using a neural network model; then, the dynamic object is sent to the client, thereby meeting the actual needs of users to generate dynamic objects and reducing the time and labor cost of users to generate dynamic objects.

[0395] The above is a schematic scheme of another object generation device of this embodiment. It should be noted that the technical scheme of the yet another object generation device and the technical scheme of the yet another object generation method described above belong to the same concept, and the details not described in detail in the technical scheme of the yet another object generation device can all be referred to the description of the technical scheme of the yet another object generation method described above.

[0396] Corresponding to the above method embodiment, this specification also provides an object generation model training device embodiment, the device comprising:

[0397] A training data determination module is configured to determine training samples and sample labels for a generation model of an object to be trained, wherein the training samples include sample object generation parameters and a plurality of sample object image sets corresponding to a plurality of time points, the plurality of time points are arranged in chronological order, and each sample object image set includes sample object images of the sample object at a plurality of viewing angles at a corresponding time point;

[0398] an object generation module, configured to process the plurality of sample object image sets according to the sample object generation parameters using the object generation model to be trained, determine a plurality of sample object dynamic data corresponding to the plurality of viewing angles, and generate sample dynamic objects corresponding to the sample objects according to the plurality of sample object dynamic data;

[0399] The model training module is configured to adjust the parameters of the object generation model to be trained based on the sample dynamic object and the sample label to obtain a trained object generation model.

[0400] The object generation model training device provided by one or more embodiments of the present specification needs to use training samples and sample labels to train the object generation model to be trained before using the object generation model to generate dynamic objects; in the process of model training, first, sample data such as sample object generation parameters, multiple time points and multiple sample object image sets, which are large in number and rich in types, are determined to facilitate subsequent training to obtain an object generation model that can generate accurate and vivid dynamic objects; secondly, the object generation model to be trained is used to process multiple sample object image sets according to the sample object generation parameters, so as to determine multiple sample object dynamic data corresponding to multiple perspectives, and based on the content-rich data of the multiple sample object dynamic data, accurate and vivid sample dynamic objects can be generated; finally, the sample dynamic objects and sample labels are used to adjust the parameters of the object generation model to be trained, so as to obtain an object generation model that can generate accurate and vivid dynamic objects; the accurate generation of dynamic objects using a neural network model is realized, the actual needs of users to generate dynamic objects are met, and the time and labor cost of users to generate dynamic objects are reduced.

[0401] The above is a schematic scheme of an object generation model training device of this embodiment. It should be noted that the technical scheme of the object generation model training device and the technical scheme of the object generation model training method described above belong to the same concept, and the details not described in detail in the technical scheme of the object generation model training device can be found in the description of the technical scheme of the object generation model training method described above.

[0402] Fig. 9 The structure block diagram of a computing device 900 provided according to an embodiment of the present specification is shown. The components of the computing device 900 include but are not limited to a memory 910 and a processor 920. The processor 920 is connected to the memory 910 via a bus 930, and a database 950 is used to store data.

[0403] The computing device 900 also includes an access device 940 that enables the computing device 900 to communicate via one or more networks 960. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 940 may include one or more of any type of network interface (e.g., a network interface card (NIC)) that is wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a world-wide interoperability for microwave access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, and a near field communication (NFC).

[0404] In one embodiment of the present specification, the above components of the computing device 900 and Fig. 9 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Fig. 9 The computing device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0405] The computing device 900 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 900 may also be a mobile or stationary server.

[0406] Among them, the processor 920 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned various object generation methods, 4D object generation methods or object generation model training methods.

[0407] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the computing device embodiment, since it is basically similar to the multiple object generation method, 4D object generation method or object generation model training method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the multiple object generation method, 4D object generation method or object generation model training method embodiment.

[0408] An embodiment of the present specification also provides a computer-readable storage medium storing a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned various object generation methods, 4D object generation methods, or object generation model training methods.

[0409] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the computer-readable storage medium embodiment, since it is basically similar to the multiple object generation method, 4D object generation method or object generation model training method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the multiple object generation method, 4D object generation method or object generation model training method embodiment.

[0410] An embodiment of the present specification also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned multiple object generation methods, 4D object generation methods, or object generation model training methods.

[0411] The above is a schematic scheme of a computer program product of this embodiment. It should be noted that the technical scheme of the computer program product and the technical scheme of the above-mentioned multiple object generation method, 4D object generation method or object generation model training method belong to the same concept, and the details not described in detail in the technical scheme of the computer program product can be referred to the description of the technical scheme of the above-mentioned multiple object generation method, 4D object generation method or object generation model training method.

[0412] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0413] The computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0414] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0415] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0416] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that technicians in the relevant technical field can well understand and use this specification. This specification is only limited by the claims and their full scope and equivalents.

Claims

1. A method for generating an object, comprising: Determining object generation parameters and a plurality of object image sets corresponding to a plurality of time points, wherein the plurality of time points are arranged in chronological order, and each object image set includes object images of a target object at a corresponding time point and from a plurality of perspectives; Inputting the object generation parameters and the plurality of object image sets into an object generation model, and determining object generation parameter features of the object generation parameters and a plurality of object image set features of the plurality of object image sets using the object generation model; Using the spatiotemporal attention unit in the object generation model, performing temporal feature processing on the object generation parameter features and the plurality of object image set features to obtain temporal features; Performing spatial feature processing on the object generation parameter features and the plurality of object image set features to obtain spatial features; Performing feature fusion on the spatial features and the temporal features to obtain fusion features corresponding to the multiple perspectives; Using an output unit in the object generation model, based on the fusion features corresponding to the multiple perspectives, a plurality of object dynamic data corresponding to the multiple perspectives are generated; The dynamic data of the multiple objects are fused to obtain a dynamic object, wherein the object dimension of the target object is smaller than the object dimension of the dynamic object.

2. The object generation method according to claim 1, wherein the determining the object generation parameter features of the object generation parameters using the object generation model and determining the plurality of object image set features of the plurality of object image sets comprises: Using the object generation model, performing multi-view feature conversion on the multiple object image sets to obtain multi-view features corresponding to the multiple object image sets, and performing video feature conversion on the multiple object image sets to obtain video features corresponding to the multiple object image sets; Determining, using the object generation model, object generation parameter encoding of the object generation parameter and a plurality of object image set encodings of the plurality of object image sets; The object generation parameter feature is generated by using the multi-view feature, the video feature and the object generation parameter encoding, and the multiple object image set features are generated by using the multi-view feature, the video feature and the multiple object image set encoding.

3. The object generation method according to claim 2, wherein the performing multi-view feature conversion on the plurality of object image sets to obtain multi-view features corresponding to the plurality of object image sets, and performing video feature conversion on the plurality of object image sets to obtain video features corresponding to the plurality of object image sets, comprises: Using the multi-view attention unit in the object generation model, extracting features from each of the multiple object image sets to obtain the multi-view features corresponding to the multiple object image sets; The video feature extraction unit in the object generation model is used to perform video feature conversion on each of the object image sets to obtain video features corresponding to the multiple object image sets.

4. The object generation method according to claim 3, wherein the step of using the video feature extraction unit in the object generation model to perform video feature conversion on each of the object image sets to obtain video features corresponding to the plurality of object image sets comprises: Using the video feature extraction unit in the object generation model, the object images included in each of the plurality of object image sets are sorted to obtain a plurality of object image sequences; Determine an image acquisition parameter corresponding to each object image set, and determine the multiple time points arranged in time sequence as a time series, wherein the image acquisition parameter is a parameter used to acquire the object image; Using the self-attention layer in the video feature extraction unit, feature extraction is performed on each object image sequence, image acquisition parameter and time series to obtain an image sequence feature matrix of each object image sequence, a parameter feature matrix of the image acquisition parameter and a time series feature matrix of the time series; The multi-head attention layer in the video feature extraction unit is used to perform feature fusion on the image sequence feature matrix, the parameter feature matrix and the time series feature matrix to obtain video features corresponding to the multiple object image sets.

5. The object generation method according to claim 2, wherein the object generation parameter is an object generation text, and the object generation parameter code is an object generation text code; The object generation parameter encoding of the object generation parameter is determined by using the object generation model, and the multiple object image set encodings of the multiple object image sets include: Using the text encoder in the object generation model, the object generation text is encoded to obtain the object generation parameter encoding; Using an image encoder in the object generation model, performing image encoding on the plurality of object image sets to obtain encodings of the plurality of object image sets; The step of using the multi-view feature, the video feature, and the object generation parameter encoding to generate the object generation parameter feature, and using the multi-view feature, the video feature, and the plurality of object image set encodings to generate the plurality of object image set features comprises: Using a first cross attention module in the object generation model, the multi-view feature, the video feature and the object generation parameter encoding are subjected to feature fusion to generate the object generation parameter feature; The second cross-attention module in the object generation model is used to perform feature fusion on the multi-view features, the video features and the multiple object image set encodings to generate the multiple object image set features.

6. The object generation method according to claim 1, wherein the step of performing temporal feature processing on the object generation parameter features and the plurality of object image set features using the spatiotemporal attention unit in the object generation model to obtain the temporal features comprises: Determine a spatiotemporal attention unit in the object generation model, and use a temporal coding layer in the spatiotemporal attention unit to encode the object generation parameter features and the plurality of object image set features to obtain a temporal coding; Using the temporal motion layer in the spatiotemporal attention unit, feature mapping and weighting processing are performed on the time code to obtain the time feature; The performing spatial feature processing on the object generation parameter feature and the plurality of object image set features to obtain the spatial feature comprises: Determine a spatiotemporal attention unit in the object generation model, and use a spatial coding layer in the spatiotemporal attention unit to encode the object generation parameter features and the plurality of object image set features to obtain a spatial code; Using the multi-view self-attention layer in the spatiotemporal attention unit, extracting features from the spatial encoding to obtain the spatial encoding; The step of fusing the spatial features and the temporal features to obtain fusion features corresponding to the multiple perspectives includes: The feature fusion layer in the spatiotemporal attention unit is used to fuse the spatial features and the temporal features to obtain fusion features corresponding to the multiple perspectives.

7. The object generation method according to claim 1, wherein the object dynamic data is an object dynamic video; The step of fusing the dynamic data of the multiple objects to obtain the dynamic object comprises: Performing video fusion on the multiple object dynamic videos to obtain an initial dynamic object; The initial dynamic object is optimized by using the object optimization unit in the object generation model to obtain the dynamic object.

8. The object generation method according to claim 1, wherein determining the object generation parameters and a plurality of object image sets corresponding to a plurality of time points comprises: Receiving the object generation parameters and the multiple object image sets corresponding to multiple time points sent by the client, wherein the object generation parameters and the multiple object image sets are generated by the client based on a data sending operation performed by a user on a data sending page; After fusing the dynamic data of the multiple objects to obtain the dynamic object, the method further includes: The dynamic object is sent to the client and displayed to the user using the data sending page in the client.

9. A method for generating a 4D object, comprising: Determining object generation parameters and a plurality of 3D object image sets corresponding to a plurality of time points, wherein the plurality of time points are arranged in chronological order, and each 3D object image set includes 3D object images of the 3D object at a corresponding time point and from a plurality of perspectives; inputting the object generation parameters and the plurality of 3D object image sets into an object generation model, and determining object generation parameter features of the object generation parameters and determining a plurality of 3D object image set features of the plurality of 3D object image sets using the object generation model; Using the spatiotemporal attention unit in the object generation model, performing temporal feature processing on the object generation parameter features and the plurality of 3D object image set features to obtain temporal features; Performing spatial feature processing on the object generation parameter feature and the plurality of 3D object image set features to obtain spatial features; Performing feature fusion on the spatial features and the temporal features to obtain fusion features corresponding to the multiple perspectives; Using an output unit in the object generation model, based on the fusion features corresponding to the multiple perspectives, a plurality of object dynamic data corresponding to the multiple perspectives are generated; The dynamic data of the multiple objects are fused to obtain a dynamic 4D object, wherein the object dimension of the 3D object is smaller than the object dimension of the dynamic 4D object.

10. An object generation method, applied to a server, comprising: Receiving object generation parameters sent by a client, and a plurality of object image sets corresponding to a plurality of time points, wherein the plurality of time points are arranged in chronological order, and each object image set includes object images of a target object at a corresponding time point and from a plurality of perspectives; Inputting the object generation parameters and the plurality of object image sets into an object generation model, and determining object generation parameter features of the object generation parameters and a plurality of object image set features of the plurality of object image sets using the object generation model; Using the spatiotemporal attention unit in the object generation model, performing temporal feature processing on the object generation parameter features and the plurality of object image set features to obtain temporal features; Performing spatial feature processing on the object generation parameter features and the plurality of object image set features to obtain spatial features; Performing feature fusion on the spatial features and the temporal features to obtain fusion features corresponding to the multiple perspectives; Using an output unit in the object generation model, based on the fusion features corresponding to the multiple perspectives, a plurality of object dynamic data corresponding to the multiple perspectives are generated; Performing data fusion on the multiple object dynamic data to obtain a dynamic object, wherein the object dimension of the target object is smaller than the object dimension of the dynamic object; The dynamic object is sent to the client.

11. A video generation method, applied to a video generation platform, comprising: Receive the object video generated text and multiple object image sets corresponding to multiple time points sent by the client, wherein the multiple time points are arranged in chronological order, and each object image set includes object images of the target object at the corresponding time point and multiple perspectives; Inputting the object video generated text and the plurality of object image sets into an object generation model, and determining the object video generated text features of the object video generated text and the plurality of object image set features of the plurality of object image sets by using the object generation model; Using the spatiotemporal attention unit in the object generation model, performing temporal feature processing on the object video generation text features and the plurality of object image set features to obtain temporal features; Performing spatial feature processing on the text features generated by the object video and the set features of the plurality of object images to obtain spatial features; Performing feature fusion on the spatial features and the temporal features to obtain fusion features corresponding to the multiple perspectives; Using an output unit in the object generation model, based on the fusion features corresponding to the multiple perspectives, a plurality of object dynamic data corresponding to the multiple perspectives are generated; fusing the multiple object dynamic data to obtain an object video corresponding to the target object, wherein the object video is a video observed from any viewing angle; The object video is sent to the client.

12. A computing device comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the object generation method according to any one of claims 1 to 8, the 4D object generation method according to claim 9, the object generation method according to claim 10, or the video generation method according to claim 11 are implemented.

13. A computer-readable storage medium storing a computer program / instruction, which, when executed by a processor, implements the steps of the object generation method described in any one of claims 1 to 8, the 4D object generation method described in claim 9, the object generation method described in claim 10, or the video generation method described in claim 11.

14. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the object generation method described in any one of claims 1 to 8, the 4D object generation method described in claim 9, the object generation method described in claim 10, or the video generation method described in claim 11.

Citation Information

Patent Citations

  • Dynamic video generation method and device, equipment, medium and program product

    CN115690637A

  • Image processing model training method and three-dimensional object model construction method

    CN115731344A