Driving method, device and electronic equipment of three-dimensional model
By acquiring motion video and initial point cloud, and using point cloud encoder and diffusion model to generate motion point cloud sequence, the problem of high cost and low efficiency in generating motion 3D models in existing technologies is solved, and efficient and accurate 3D model driving is achieved.
Patent Information
- Application Number
- CN202411783331.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2044-12-05
AI Technical Summary
Existing technologies are costly and inefficient in generating 3D models with motion, especially when combined with motion description text or images, making it difficult to efficiently construct 3D models of motion.
By acquiring the initial point cloud of the motion video and the 3D model, motion point cloud corresponding to the motion image is generated, and the point cloud encoder, diffusion model and point cloud decoder are used for training to generate motion point cloud sequences to drive the 3D model.
It reduces the driving cost of 3D models, improves the driving efficiency of 3D models, ensures the accuracy of motion point clouds, and achieves efficient 3D animation generation.
Smart Images

Figure CN119722880B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of artificial intelligence, in particular to the technical field of computer vision, deep learning, large model, etc., which can be applied to the scene of three-dimensional animation, and in particular to a three-dimensional model driving method and device and electronic equipment. BACKGROUND
[0002] At present, relying on artificial intelligence (AI) generative technology, a three-dimensional model can be constructed based on text or image, etc. When a three-dimensional model with action needs to be generated, a three-dimensional model with action also needs to be constructed in combination with action description text or action image, etc., which is high in cost and poor in efficiency. SUMMARY
[0003] The present disclosure provides a three-dimensional model driving method, device and electronic equipment.
[0004] According to an aspect of the present disclosure, a three-dimensional model driving method is provided, which includes: acquiring an action video and an initial point cloud of a three-dimensional model; the action video includes action images; generating an action point cloud corresponding to the action images according to the action images and the initial point cloud; the action described by the action point cloud is consistent with the action described by the action images; determining an action point cloud sequence of the three-dimensional model according to the action point cloud corresponding to each action image; and the action point cloud sequence is used for driving processing of the three-dimensional model.
[0005] According to another aspect of the present disclosure, a training method of an action driving model is provided, which includes: acquiring an initial action driving model; the action driving model includes a point cloud encoder, a diffusion model and a point cloud decoder connected in sequence; the diffusion model is used to generate an action point cloud encoding vector in combination with an initial point cloud encoding vector output by the point cloud encoder and an action image encoding vector; acquiring training data; the training data includes an initial point cloud of a three-dimensional model and a sample action point cloud, and a sample action image; and training the point cloud encoder, the point cloud decoder and the diffusion model using the training data to obtain a trained point cloud encoder, a trained point cloud decoder and a trained diffusion model.
[0006] According to another aspect of the present disclosure, there is provided a driving device of a three-dimensional model, the device comprising: an acquisition module configured to acquire an action video and an initial point cloud of a three-dimensional model; the action video comprising action images; a generation module configured to generate, according to the action images and the initial point cloud, action point clouds corresponding to the action images; the action described by the action point cloud being consistent with the action described by the action image; a determination module configured to determine, according to the action point clouds corresponding to each of the action images, a sequence of action point clouds of the three-dimensional model; the sequence of action point clouds being used for driving processing of the three-dimensional model.
[0007] According to another aspect of the present disclosure, there is provided a training device of an action driving model, the device comprising: a first acquisition module configured to acquire an initial action driving model; the action driving model comprising a point cloud encoder, a diffusion model and a point cloud decoder connected in sequence; the diffusion model being configured to generate an action point cloud encoding vector in combination with an action image encoding vector and an initial point cloud encoding vector output by the point cloud encoder; a second acquisition module configured to acquire training data; the training data comprising an initial point cloud of a three-dimensional model and a sample action point cloud, and a sample action image; and a training processing module configured to perform training processing on the point cloud encoder, the point cloud decoder and the diffusion model using the training data, to obtain a trained point cloud encoder, a trained point cloud decoder and a trained diffusion model.
[0008] According to another aspect of the present disclosure, there is provided an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the driving method of a three-dimensional model as described above; or perform the training method of an action driving model as described above.
[0009] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the driving method of a three-dimensional model as described above; or perform the training method of an action driving model as described above.
[0010] According to another aspect of the present disclosure, there is provided a computer program product comprising a computer program which, when executed by a processor, implements the steps of the driving method of a three-dimensional model as described above; or implements the steps of the training method of an action driving model as described above.
[0011] It should be appreciated that the content described in this section is not intended to identify key or important features of the embodiments of the disclosure, nor is it intended to limit the scope of the disclosure. Other features of the disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0012] The accompanying drawings are used to better understand the present scheme and do not constitute a limitation on the disclosure. Among them:
[0013] Figure 1 is a schematic diagram according to a first embodiment of the disclosure;
[0014] Figure 2 is a schematic diagram according to a second embodiment of the disclosure;
[0015] Figure 3 is a schematic diagram according to a third embodiment of the disclosure;
[0016] Figure 4 is a flowchart of the training of the action driving model;
[0017] Figure 5 is a schematic diagram according to a fourth embodiment of the disclosure;
[0018] Figure 6 is a schematic diagram according to a fifth embodiment of the disclosure;
[0019] Figure 7 is a block diagram of an electronic device for implementing the driving method of the three-dimensional model or the training method of the action driving model according to the embodiments of the disclosure. DETAILED DESCRIPTION
[0020] Exemplary embodiments of the disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the disclosure to help understanding, and should be considered as merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the disclosure. Also, for the sake of clarity and conciseness, the description below omits the description of well-known functions and structures.
[0021] Currently, relying on artificial intelligence (AI) generative technology, three-dimensional models can be constructed based on text or images, etc. When a three-dimensional model with action needs to be generated, a three-dimensional model with action also needs to be constructed in combination with action description text or action image, etc., which is high in cost and poor in efficiency.
[0022] To solve the above problems, the disclosure provides a driving method, device and electronic equipment for a three-dimensional model.
[0023] Figure 1is a schematic diagram according to the first embodiment of the present disclosure. It should be noted that the driving method of the three-dimensional model according to the embodiments of the present disclosure can be applied to a driving device of the three-dimensional model, which can be configured in an electronic device to enable the electronic device to perform the driving function of the three-dimensional model.
[0024] The electronic device can be any device with computing capability, such as a personal computer (PC), a mobile terminal, a server, etc. The mobile terminal can be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, a smart speaker, a server, a server cluster, etc. hardware devices with various operating systems, touch screens and / or display screens.
[0025] The driving device of the three-dimensional model can also be software in the electronic device, such as driving software of the three-dimensional model, etc. In the following embodiments, the execution subject is taken as an example to illustrate the electronic device.
[0026] As shown in Figure 1 The driving method of the three-dimensional model can include the following steps:
[0027] Step 101, obtaining an action video and an initial point cloud of a three-dimensional model; the action video includes action images.
[0028] In the embodiments of the present disclosure, the three-dimensional model can be a three-dimensional model of an object that needs to generate a sequence of action point clouds. The object can be, for example, an animal, a person, a virtual character, etc., which can be set according to actual needs.
[0029] In the embodiments of the present disclosure, the action video can be a video determined based on the three-dimensional model and a sequence of action descriptions; or, it can be an action video of another object similar in geometric structure to the object to which the three-dimensional model belongs.
[0030] In one example, the action video can be a video determined based on the three-dimensional model and a sequence of action descriptions. Correspondingly, the process of the electronic device performing step 101 can be, for example, obtaining a perspective image of the three-dimensional model and a sequence of action descriptions; the sequence of action descriptions includes action descriptions; generating action images corresponding to the action descriptions according to the action descriptions and the perspective image; and determining the action video according to the action images corresponding to each action description.
[0031] The perspective image of the three-dimensional model can be an image of the three-dimensional model from one perspective; or, it can be an image of the three-dimensional model from multiple perspectives. The electronic device can obtain the perspective image by rendering the three-dimensional model. The rendering strategy can be, for example, rendering the three-dimensional model in combination with a rendering software such as blender software.
[0032] The electronic device can input the action description and the perspective image into the image generation model for each action description in the action description sequence, obtain an action image corresponding to the action description output by the image generation model, and then combine to obtain the action video.
[0033] As an alternative, the electronic device can input the perspective image of the three-dimensional model and the action description sequence into the video generation model to obtain the action video output by the video generation model.
[0034] The action video generation process in combination with the perspective image of the three-dimensional model and the action description sequence can flexibly adjust the action descriptions in the action description sequence and generate the required action video, thereby improving the efficiency of obtaining the action video and improving the flexibility of the action described by the action video.
[0035] In another example, the action video can be an action video of another object similar in geometric structure to the object to which the three-dimensional model belongs. Correspondingly, the process of step 101 performed by the electronic device can be, for example, obtaining an action description sequence and a candidate action video corresponding to the action description sequence; obtaining first geometric structure information of a candidate object to which the candidate action video belongs and second geometric structure information of the three-dimensional model; and determining the candidate action video as the action video in a case where the first geometric structure information matches the second geometric structure information.
[0036] The determination of whether the first geometric structure information matches the second geometric structure information can be made by shape comparison of the first geometric structure information and the second geometric structure information, which is not limited here.
[0037] As an alternative, the electronic device can obtain any one of the candidate action videos of the candidate object; and determine the candidate action video as the action video in a case where the first geometric structure information of the candidate object matches the second geometric structure information of the three-dimensional model.
[0038] In a case where the first geometric structure information of the candidate object matches the second geometric structure information of the three-dimensional model, the candidate action video of the candidate object is determined as the action video, which can avoid video generation processing and use the existing candidate action video, thereby reducing the cost of obtaining the action video and improving the efficiency of obtaining the action video.
[0039] In step 102, an action point cloud corresponding to the action image is generated according to the action image and the initial point cloud; the action described by the action point cloud is consistent with the action described by the action image.
[0040] In the embodiments of the present disclosure, the process of step 102 performed by the electronic device may, for example, be that the action image and the initial point cloud are input into a point cloud generation model, and an action point cloud output by the point cloud generation model is obtained. The point cloud generation model may be trained in combination with a sample action image, a sample initial point cloud, and a sample action point cloud.
[0041] Step 103: determining an action point cloud sequence of the three-dimensional model according to the action point cloud corresponding to each action image; the action point cloud sequence is used for driving processing of the three-dimensional model.
[0042] The electronic device may, for example, combine each action point cloud according to the order of each action image in the action video to obtain the action point cloud sequence of the three-dimensional model.
[0043] In the embodiments of the present disclosure, after step 103, the electronic device may further perform the following process: performing three-dimensional modeling processing according to the action point cloud sequence to realize driving processing of the three-dimensional model.
[0044] Specifically, after performing three-dimensional modeling processing according to the action point cloud sequence, the data after the modeling processing may be provided to the object, so that the object can view a high-quality three-dimensional animation, and the generation efficiency of the three-dimensional animation is improved.
[0045] The driving method of the three-dimensional model in the embodiments of the present disclosure comprises the following steps: obtaining an action video and an initial point cloud of a three-dimensional model; the action video comprises action images; generating an action point cloud corresponding to each action image according to the action image and the initial point cloud; the action described by the action point cloud is consistent with the action described by the action image; determining an action point cloud sequence of the three-dimensional model according to the action point cloud corresponding to each action image; the action point cloud sequence is used for driving processing of the three-dimensional model; and the action described by the action point cloud in the action point cloud sequence is consistent with the action described by the corresponding action image, thereby realizing driving processing of the three-dimensional model based on the action video, avoiding multiple constructions of the three-dimensional model, reducing the driving cost of the three-dimensional model, and improving the driving efficiency of the three-dimensional model.
[0046] To further improve the determination accuracy of the action point cloud, the action image and the initial point cloud may be encoded, and an encoded vector of the action point cloud is generated in combination with the encoded vector obtained by the encoding to determine the action point cloud. Figure 2 As shown in Figure 2 is a schematic diagram according to a second embodiment of the present disclosure, Figure 2 may comprise the following steps:
[0047] Step 201: obtaining an action video and an initial point cloud of a three-dimensional model; the action video comprises action images.
[0048] In step 202, a motion image encoding vector of the motion image is determined.
[0049] In the embodiments of the present disclosure, the motion image can be an image in one view. In one example, the process of step 202 performed by the electronic device may, for example, be that the motion image is input into an image encoder, and a motion image encoding vector output by the image encoder is obtained. The image encoder may, for example, be combined with an image decoder and trained by an image set.
[0050] In another example, the process of step 202 performed by the electronic device may, for example, be that, based on the motion image, images in multiple views are determined; the images in the multiple views are input into an image encoder, and a motion image encoding vector output by the image encoder is obtained.
[0051] The electronic device may, for example, input the motion image into an image-to-multiview model to obtain images in multiple views output by the image-to-multiview model. The images in the multiple views may, for example, be images obtained by adjusting the view of the motion in the motion image.
[0052] The image-to-multiview model may, for example, be a Zero-1-to-3 model.
[0053] In the embodiments of the present disclosure, the image encoder may, for each image in a view, determine an image encoding vector of the image in the view; and the image encoding vectors are spliced to obtain the motion image encoding vector.
[0054] The motion image encoding vector is determined based on the image encoder. The large number of parameters of the image encoder enables the motion image encoding vector to accurately describe the motion features in the motion image, ensures the accuracy of the motion image encoding vector, and further ensures the accuracy of the motion point cloud.
[0055] In step 203, an initial point cloud encoding vector of an initial point cloud is determined.
[0056] In the embodiments of the present disclosure, the process of step 203 performed by the electronic device may, for example, be that the initial point cloud is input into a point cloud encoder, and an initial point cloud encoding vector output by the point cloud encoder is obtained.
[0057] The point cloud encoder and the point cloud decoder may, for example, be trained in combination with a sample view image, an initial point cloud, and a sample motion point cloud. The trained point cloud encoder may, for example, be used to encode the point cloud to obtain a point cloud encoding vector. The trained point cloud decoder may, for example, be used to decode the point cloud encoding vector to obtain a point cloud.
[0058] The initial point cloud encoding vector is determined based on a point cloud encoder. The point cloud encoder has a large number of parameters, so that the initial point cloud encoding vector can accurately describe the features of the initial point cloud, and the accuracy of the initial point cloud encoding vector is ensured, and the accuracy of the action point cloud is ensured.
[0059] In step 204, the action point cloud encoding vector is generated according to the action image encoding vector and the initial point cloud encoding vector.
[0060] In the embodiments of the present disclosure, the process of step 204 performed by the electronic device may, for example, be that the action image encoding vector is determined as the condition information of the diffusion model; and the initial point cloud encoding vector is input into the diffusion model to obtain the action point cloud encoding vector output by the diffusion model.
[0061] The diffusion model may, for example, be trained in combination with the initial point cloud encoding vector, the image encoding vector of the sample action image, and the sample action point cloud encoding vector of the sample action point cloud.
[0062] The action point cloud encoding vector is determined based on the diffusion model. The diffusion model has a large number of parameters, so that the accuracy of the determined action point cloud encoding vector is improved, and the accuracy of the determined action point cloud is further improved.
[0063] In step 205, the action point cloud corresponding to the action image is determined according to the action point cloud encoding vector.
[0064] In the embodiments of the present disclosure, the process of step 205 performed by the electronic device may, for example, be that the action point cloud encoding vector is input into the point cloud decoder to obtain the action point cloud output by the point cloud decoder.
[0065] In step 206, the action point cloud sequence of the three-dimensional model is determined according to the action point cloud corresponding to each action image. The action point cloud sequence is used for driving processing of the three-dimensional model.
[0066] It should be noted that the details of steps 201 and 206 can refer to steps 101 and 103 in the embodiments shown in Figure 1 The details of steps 101 and 103 in the embodiments shown in
[0067] The driving method of the three-dimensional model according to the embodiments of the present disclosure comprises the following steps: obtaining a motion video and an initial point cloud of a three-dimensional model; the motion video comprises motion images; determining a motion image encoding vector of the motion images; determining an initial point cloud encoding vector of the initial point cloud; generating a motion point cloud encoding vector according to the motion image encoding vector and the initial point cloud encoding vector; determining a motion point cloud corresponding to the motion images according to the motion point cloud encoding vector; determining a motion point cloud sequence of the three-dimensional model according to the motion point clouds corresponding to the motion images; and using the motion point cloud sequence for driving processing of the three-dimensional model. According to the embodiments of the present disclosure, the motion point cloud encoding vector is determined according to the motion image encoding vector and the initial point cloud encoding vector, and then the motion point cloud is determined, so that more features can be considered in the process of determining the motion point cloud, thereby further improving the determination accuracy of the motion point cloud.
[0068] Figure 3 is a schematic diagram according to the third embodiment of the present disclosure. It should be noted that the training method of the motion driving model according to the embodiments of the present disclosure can be applied to a training device of the motion driving model. The device can be configured in an electronic device, so that the electronic device can perform the training function of the motion driving model.
[0069] The electronic device can be any device with computing capability, such as a personal computer (PC), a mobile terminal, a server, etc. The mobile terminal can be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, a smart speaker, a server, a server cluster, etc. The hardware device has various operating systems, touch screens and / or display screens.
[0070] The training device of the motion driving model can also be software in the electronic device, such as a training software of the motion driving model. In the following embodiments, the execution subject is taken as an example to illustrate the electronic device.
[0071] As shown in Figure 3 , the training method of the motion driving model can comprise the following steps:
[0072] Step 301: obtaining an initial motion driving model; the motion driving model comprises a point cloud encoder, a diffusion model and a point cloud decoder connected in sequence; the diffusion model is used to generate a motion point cloud encoding vector by combining a motion image encoding vector and an initial point cloud encoding vector output by the point cloud encoder.
[0073] In the embodiments of the present disclosure, the point cloud encoder can be used to encode the initial point cloud to obtain the initial point cloud encoding vector. The point cloud decoder can be used to decode the motion point cloud encoding vector to obtain the motion point cloud.
[0074] Among them, the point cloud encoder and point cloud decoder can be the encoder and decoder in a variational autoencoder (VAE).
[0075] Step 302: Obtain training data; the training data includes the initial point cloud of the 3D model, the sample action point cloud, and the sample action images.
[0076] In this embodiment of the disclosure, the electronic device may perform step 302 as follows: acquire sample motion point clouds and initial point clouds of a 3D model; perform 3D modeling processing based on the sample motion point clouds to obtain a modeled 3D model; acquire sample view images of the modeled 3D model; and determine training data based on the sample motion point clouds, initial point clouds, and sample view images.
[0077] The sample viewpoint images can be images from one viewpoint of the modeled 3D model, or images from multiple viewpoints of the modeled 3D model. Using multiple viewpoint images can expand the amount of training data and further improve the accuracy of the trained action-driven model.
[0078] It should be noted that the 3D model involved in the training data can be a model of any 3D object, without any specific restrictions here.
[0079] Among them, combining the sample action point cloud of the 3D model to determine the sample viewpoint image, and then determining the training data, can reduce the cost of obtaining training data and improve the efficiency of obtaining training data.
[0080] Step 303: Use training data to train the point cloud encoder, point cloud decoder and diffusion model to obtain the trained point cloud encoder, trained point cloud decoder and trained diffusion model.
[0081] In this embodiment of the disclosure, the electronic device may perform step 303 as follows: training the point cloud encoder and the point cloud decoder with training data to obtain the trained point cloud encoder and the trained point cloud decoder; and training the diffusion model with the trained point cloud encoder and the training data to obtain the trained diffusion model.
[0082] The process involves first training a point cloud encoder and a point cloud decoder, and then using the trained point cloud encoder and training data to train the diffusion model, which ensures the accuracy of the trained diffusion model.
[0083] In the embodiments of the present disclosure, the training process of the electronic device on the point cloud encoder and the point cloud decoder may, for example, be determining a sample image encoding vector of a sample action image; inputting an initial point cloud and the sample image encoding vector into the point cloud encoder and the point cloud decoder connected in sequence to obtain a predicted action point cloud output by the point cloud decoder; determining a first loss function value according to a sample action point cloud, the predicted action point cloud, and a loss function of the point cloud encoder and the point cloud decoder; and performing parameter adjustment processing on the point cloud encoder and the point cloud decoder according to the first loss function value to obtain a trained point cloud encoder and a trained point cloud decoder.
[0084] The loss function of the point cloud encoder and the point cloud decoder may be, for example, a KL divergence between the sample action point cloud and the predicted action point cloud.
[0085] Determining the first loss function value according to the sample action point cloud, the predicted action point cloud, and the loss function of the point cloud encoder and the point cloud decoder, and then performing parameter adjustment processing, can improve the accuracy of the trained point cloud encoder and the point cloud decoder.
[0086] In the embodiments of the present disclosure, the training process of the electronic device on the diffusion model may, for example, be inputting an initial point cloud and a sample action point cloud into the trained point cloud encoder respectively to obtain an initial point cloud encoding vector and a sample action point cloud encoding vector output by the point cloud encoder; determining a sample image encoding vector of a sample action image; determining the sample image encoding vector as condition information of the diffusion model, inputting the initial point cloud encoding vector into the diffusion model, and obtaining a predicted action point cloud encoding vector output by the diffusion model; determining a second loss function value according to the sample action point cloud encoding vector, the predicted action point cloud encoding vector, and a loss function of the diffusion model; and performing parameter adjustment processing on the diffusion model according to the second loss function value to obtain a trained diffusion model.
[0087] Determining the second loss function value according to the sample action point cloud encoding vector, the predicted action point cloud encoding vector, and the loss function of the diffusion model, and then performing parameter adjustment processing, can improve the accuracy of the trained diffusion model.
[0088] The method for training the action driving model according to the embodiments of the present disclosure comprises the following steps: obtaining an initial action driving model; the action driving model comprises a point cloud encoder, a diffusion model and a point cloud decoder connected in sequence; the diffusion model is used to generate an action point cloud encoding vector by combining an action image encoding vector and an initial point cloud encoding vector output by the point cloud encoder; obtaining training data; the training data comprises an initial point cloud of a three-dimensional model, a sample action point cloud and a sample action image; the point cloud encoder, the point cloud decoder and the diffusion model are trained by using the training data, so as to obtain a trained point cloud encoder, a trained point cloud decoder and a trained diffusion model; wherein the step-by-step training of the point cloud encoder, the diffusion model and the point cloud decoder can ensure the accuracy of the trained point cloud encoder and the point cloud decoder, and further improve the accuracy of the trained diffusion model.
[0089] The following examples are used for illustration. As shown in Figure 4 , it is a flowchart of driving of a three-dimensional model. In Figure 4 , the method can specifically comprise the following steps: step 401, obtaining a three-dimensional model; step 402, generating a motion video by combining the three-dimensional model and a video generation algorithm (i.e., determining an action video by combining a perspective image of the three-dimensional model and an action description sequence); step 403, performing point sampling on the three-dimensional model to obtain a three-dimensional point cloud (an initial point cloud) of the three-dimensional model; step 404, inputting the three-dimensional point cloud and the motion video into a point cloud encoder to obtain point cloud encoding (an initial point cloud encoding vector); step 405, inputting the point cloud encoding and the motion video into a diffusion model to obtain a point cloud encoding sequence output by the diffusion model (i.e., a plurality of action point cloud encoding vectors); and step 406, inputting the point cloud encoding sequence into a point cloud decoder to obtain a model animation sequence (i.e., a plurality of action point clouds).
[0090] To achieve the above-mentioned embodiments, the present disclosure further provides a driving device of a three-dimensional model. As shown in Figure 5 , Figure 5 is a schematic diagram according to the fourth embodiment of the present disclosure. The driving device 50 of the three-dimensional model can comprise an obtaining module 501, a generating module 502 and a determining module 503.
[0091] The obtaining module 501 is configured to obtain an action video and an initial point cloud of a three-dimensional model; the action video comprises an action image; the generating module 502 is configured to generate an action point cloud corresponding to the action image according to the action image and the initial point cloud; the action described by the action point cloud is consistent with the action described by the action image; and the determining module 503 is configured to determine an action point cloud sequence of the three-dimensional model according to the action point cloud corresponding to each action image; the action point cloud sequence is used for driving processing of the three-dimensional model.
[0092] As a possible implementation manner of the embodiment of the present disclosure, the acquisition module 501 is specifically configured to acquire a perspective image of the three-dimensional model and an action description sequence; the action description sequence comprises an action description; an action image corresponding to the action description is generated according to the action description and the perspective image; and the action video is determined according to the action image corresponding to each action description.
[0093] As a possible implementation manner of the embodiment of the present disclosure, the acquisition module 501 is specifically configured to acquire an action description sequence and a candidate action video corresponding to the action description sequence; acquire first geometric structure information of a candidate object to which the candidate action video belongs and second geometric structure information of the three-dimensional model; and in a case where the first geometric structure information matches the second geometric structure information, the candidate action video is determined as the action video.
[0094] As a possible implementation manner of the embodiment of the present disclosure, the generation module 502 comprises a first determination unit, a second determination unit, a generation unit and a third determination unit; the first determination unit is configured to determine an action image encoding vector of the action image; the second determination unit is configured to determine an initial point cloud encoding vector of the initial point cloud; the generation unit is configured to generate an action point cloud encoding vector according to the action image encoding vector and the initial point cloud encoding vector; and the third determination unit is configured to determine an action point cloud corresponding to the action image according to the action point cloud encoding vector.
[0095] As a possible implementation manner of the embodiment of the present disclosure, the first determination unit is specifically configured to input the action image into an image encoder to acquire the action image encoding vector output by the image encoder; or, determine a plurality of perspective images according to the action image; input the plurality of perspective images into an image encoder to acquire the action image encoding vector output by the image encoder.
[0096] As a possible implementation manner of the embodiment of the present disclosure, the second determination unit is specifically configured to input the initial point cloud into a point cloud encoder to acquire the initial point cloud encoding vector output by the point cloud encoder; and the third determination unit is specifically configured to input the action point cloud encoding vector into a point cloud decoder to acquire the action point cloud output by the point cloud decoder.
[0097] As a possible implementation manner of the embodiment of the present disclosure, the generation unit is specifically configured to determine the action image encoding vector as condition information of a diffusion model; and input the initial point cloud encoding vector into the diffusion model to acquire the action point cloud encoding vector output by the diffusion model.
[0098] The driving device of the three-dimensional model of the embodiment of the present disclosure, by acquiring an action video and an initial point cloud of a three-dimensional model; the action video includes action images; according to the action images and the initial point cloud, an action point cloud corresponding to the action images is generated; the action described by the action point cloud is consistent with the action described by the action images; according to the action point cloud corresponding to each action image, an action point cloud sequence of the three-dimensional model is determined; the action point cloud sequence is used for driving processing of the three-dimensional model; wherein in the action point cloud sequence, the action described by the action point cloud is consistent with the action described by the corresponding action image, thereby realizing the driving processing of the three-dimensional model based on the action video, avoiding multiple constructions of the three-dimensional model, thereby reducing the driving cost of the three-dimensional model and improving the driving efficiency of the three-dimensional model.
[0099] In order to realize the above-mentioned embodiment, the present disclosure also provides a training device of an action driving model. As shown in Figure 6 Figure 6 is a schematic diagram according to the fifth embodiment of the present disclosure. The training device of the action driving model 60 can include a first acquisition module 601, a second acquisition module 602 and a training processing module 603.
[0100] The first acquisition module 601 is configured to acquire an initial action driving model; the action driving model includes a point cloud encoder, a diffusion model and a point cloud decoder connected in sequence; the diffusion model is configured to generate an action point cloud encoding vector by combining an action image encoding vector and an initial point cloud encoding vector output by the point cloud encoder; the second acquisition module 602 is configured to acquire training data; the training data includes an initial point cloud of a three-dimensional model and a sample action point cloud, and a sample action image; the training processing module 603 is configured to train the point cloud encoder, the point cloud decoder and the diffusion model using the training data to obtain a trained point cloud encoder, a trained point cloud decoder and a trained diffusion model.
[0101] As a possible implementation manner of the embodiment of the present disclosure, the second acquisition module 602 is specifically configured to acquire a sample action point cloud and an initial point cloud of the three-dimensional model; perform three-dimensional model modeling processing according to the sample action point cloud to obtain a modeled three-dimensional model; acquire a sample perspective image of the modeled three-dimensional model; and determine the training data according to the sample action point cloud, the initial point cloud and the sample perspective image.
[0102] As a possible implementation manner of the embodiment of the present disclosure, the training processing module 603 comprises a first training processing unit and a second training processing unit; the first training processing unit is configured to train the point cloud encoder and the point cloud decoder by using the training data, to obtain a trained point cloud encoder and a trained point cloud decoder; and the second training processing unit is configured to train the diffusion model by using the trained point cloud encoder and the training data, to obtain a trained diffusion model.
[0103] As a possible implementation manner of the embodiment of the present disclosure, the first training processing unit is specifically configured to determine a sample image encoding vector of the sample action image; input the initial point cloud and the sample image encoding vector into the point cloud encoder and the point cloud decoder connected in sequence, to obtain a predicted action point cloud output by the point cloud decoder; determine a first loss function value according to the sample action point cloud, the predicted action point cloud, and a loss function of the point cloud encoder and the point cloud decoder; and perform parameter adjustment processing on the point cloud encoder and the point cloud decoder according to the first loss function value, to obtain the trained point cloud encoder and the trained point cloud decoder.
[0104] As a possible implementation manner of the embodiment of the present disclosure, the second training processing unit is specifically configured to input the initial point cloud and the sample action point cloud into the trained point cloud encoder respectively, to obtain an initial point cloud encoding vector and a sample action point cloud encoding vector output by the point cloud encoder; determine a sample image encoding vector of the sample action image; determine the sample image encoding vector as condition information of the diffusion model, and input the initial point cloud encoding vector into the diffusion model, to obtain a predicted action point cloud encoding vector output by the diffusion model; determine a second loss function value according to the sample action point cloud encoding vector, the predicted action point cloud encoding vector, and a loss function of the diffusion model; and perform parameter adjustment processing on the diffusion model according to the second loss function value, to obtain the trained diffusion model.
[0105] The training device of the action driving model in the embodiment of the present disclosure comprises the following steps: obtaining an initial action driving model; the action driving model comprises a point cloud encoder, a diffusion model and a point cloud decoder connected in sequence; the diffusion model is used for combining an action image coding vector and an initial point cloud coding vector output by the point cloud encoder to generate an action point cloud coding vector; obtaining training data; the training data comprises an initial point cloud of a three-dimensional model and a sample action point cloud, and a sample action image; the point cloud encoder, the point cloud decoder and the diffusion model are trained by using the training data to obtain a trained point cloud encoder, a trained point cloud decoder and a trained diffusion model; wherein, the step-by-step training of the point cloud encoder, the diffusion model and the point cloud decoder can ensure the accuracy of the trained point cloud encoder and the point cloud decoder, and further improve the accuracy of the trained diffusion model.
[0106] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information are all carried out on the premise of obtaining the consent of the user and in accordance with relevant laws and regulations, and do not violate public order and good customs.
[0107] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.
[0108] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.
[0109] As shown in Figure 7 The device 700 includes a computing unit 701 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 702 or a computer program loaded into a random access memory (RAM) 703 from a storage unit 708. Various programs and data required for the operation of the device 700 can also be stored in the RAM 703. The computing unit 701, the ROM 702 and the RAM 703 are connected to each other through a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0110] A plurality of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0111] The computing unit 701 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 701 performs various methods and processes described above, such as the driving method of a three-dimensional model or the training method of an action-driven model. For example, in some embodiments, the driving method of a three-dimensional model or the training method of an action-driven model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the driving method of a three-dimensional model or the training method of an action-driven model described above can be performed. Alternatively, in other embodiments, the computing unit 701 can be configured to perform the driving method of a three-dimensional model or the training method of an action-driven model by any other appropriate means, such as by means of firmware.
[0112] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0113] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a standalone software package, or entirely on a remote machine or server.
[0114] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0115] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0116] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0117] The computer system can include clients and servers. This relationship can be. The servers are typically remote from the clients with the interactions between them occurring over a communication network. The relationship between a client and a server is one of client-server. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0118] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be executed in parallel, executed in series, or executed in different orders, as long as the desired results of the technical solutions of the present disclosure are achieved, and the present disclosure is not limited herein.
[0119] The specific implementation described above does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements within the spirit and principles of the present disclosure should be included in the protection scope of the present disclosure.
Claims
1. A driving method of a three-dimensional model, the method comprising: obtaining an action video and an initial point cloud of a three-dimensional model; the action video comprising action images; determining an action image encoding vector of the action image; determining an initial point cloud encoding vector of the initial point cloud; determining the action image encoding vector as conditional information of a diffusion model; inputting the initial point cloud encoding vector into the diffusion model to obtain an action point cloud encoding vector output by the diffusion model; determining an action point cloud corresponding to the action image according to the action point cloud encoding vector; the action described by the action point cloud being consistent with the action described by the action image; determining an action point cloud sequence of the three-dimensional model according to the action point cloud corresponding to each of the action images; the action point cloud sequence being used for driving processing of the three-dimensional model.
2. The method of claim 1, wherein, obtaining an action video, comprising: obtaining a perspective image of the three-dimensional model and an action description sequence; the action description sequence comprising action descriptions; generating an action image corresponding to the action description according to the action description and the perspective image; determining the action video according to the action image corresponding to each of the action descriptions.
3. The method of claim 1, wherein, obtaining an action video, comprising: obtaining an action description sequence and a candidate action video corresponding to the action description sequence; obtaining first geometric structure information of a candidate object to which the candidate action video belongs and second geometric structure information of the three-dimensional model; in a case where the first geometric structure information matches the second geometric structure information, determining the candidate action video as the action video.
4. The method of claim 1, wherein, the determining of the action image encoding vector of the action image comprises: inputting the action image into an image encoder to obtain the action image encoding vector output by the image encoder; or determining images of multiple perspectives according to the action image; inputting the images of the multiple perspectives into an image encoder to obtain the action image encoding vector output by the image encoder.
5. The method of claim 1, wherein, the determining of the initial point cloud encoding vector of the initial point cloud comprises: inputting the initial point cloud into a point cloud encoder to obtain the initial point cloud encoding vector output by the point cloud encoder; the determining of the action point cloud corresponding to the action image according to the action point cloud encoding vector comprises: inputting the action point cloud encoding vector into a point cloud decoder to obtain the action point cloud output by the point cloud decoder.
6. A training method of an action driving model, the method comprising: obtaining an initial action driving model; the action driving model comprising a point cloud encoder, a diffusion model and a point cloud decoder connected in sequence; the diffusion model is configured to generate an action point cloud encoding vector according to an initial point cloud encoding vector output by the point cloud encoder, with an action image encoding vector as conditional information; obtaining training data; the training data comprising an initial point cloud of a three-dimensional model and a sample action point cloud, and a sample action image; training the point cloud encoder, the point cloud decoder and the diffusion model using the training data to obtain a trained point cloud encoder, a trained point cloud decoder and a trained diffusion model.
7. The method of claim 6, wherein, The acquisition training data comprises: Acquiring a sample action point cloud and an initial point cloud of the three-dimensional model; According to the sample action point cloud, a three-dimensional model modeling process is performed to obtain a modeled three-dimensional model; Acquiring a sample view image of the modeled three-dimensional model; According to the sample action point cloud, the initial point cloud and the sample view image, the training data is determined.
8. The method of claim 6, wherein, The training data is used to train the point cloud encoder, the point cloud decoder and the diffusion model, and the trained point cloud encoder, the trained point cloud decoder and the trained diffusion model are obtained, comprising: The training data is used to train the point cloud encoder and the point cloud decoder, and the trained point cloud encoder and the trained point cloud decoder are obtained; The trained point cloud encoder and the training data are used to train the diffusion model, and the trained diffusion model is obtained.
9. The method of claim 8, wherein, The training data is used to train the point cloud encoder and the point cloud decoder, and the trained point cloud encoder and the trained point cloud decoder are obtained, comprising: Determine the sample image encoding vector of the sample action image; The initial point cloud and the sample image encoding vector are input into the point cloud encoder and the point cloud decoder connected in turn, and a predicted action point cloud output by the point cloud decoder is acquired; According to the sample action point cloud, the predicted action point cloud and the loss function of the point cloud encoder and the point cloud decoder, a first loss function value is determined; According to the first loss function value, the parameters of the point cloud encoder and the point cloud decoder are adjusted to obtain the trained point cloud encoder and the trained point cloud decoder.
10. The method of claim 8, wherein, The trained point cloud encoder and the training data are used to train the diffusion model, and the trained diffusion model is obtained, comprising: The initial point cloud and the sample action point cloud are input into the trained point cloud encoder respectively, and the initial point cloud encoding vector and the sample action point cloud encoding vector output by the trained point cloud encoder are acquired; Determine the sample image encoding vector of the sample action image; The sample image encoding vector is determined as the condition information of the diffusion model, and the initial point cloud encoding vector is input into the diffusion model to obtain a predicted action point cloud encoding vector output by the diffusion model; According to the sample action point cloud encoding vector, the predicted action point cloud encoding vector and the loss function of the diffusion model, a second loss function value is determined; According to the second loss function value, the parameters of the diffusion model are adjusted to obtain the trained diffusion model.
11. A driving device of a three-dimensional model, the device comprising: An acquisition module for acquiring an action video and an initial point cloud of a three-dimensional model; The action video comprises an action image; The generating module comprises a first determining unit, a second determining unit, a generating unit and a third determining unit; the first determining unit is configured to determine a motion image encoding vector of the motion image; the second determining unit is configured to determine an initial point cloud encoding vector of the initial point cloud; the generating unit is configured to determine the motion image encoding vector as condition information of a diffusion model, input the initial point cloud encoding vector into the diffusion model, and obtain a motion point cloud encoding vector output by the diffusion model; and the third determining unit is configured to determine a motion point cloud corresponding to the motion image according to the motion point cloud encoding vector; the motion described by the motion point cloud is consistent with the motion described by the motion image; The determining module is configured to determine a motion point cloud sequence of the three-dimensional model according to the motion point clouds corresponding to the respective motion images; and the motion point cloud sequence is used for driving processing of the three-dimensional model.
12. The apparatus of claim 11, wherein, The obtaining module is specifically configured to, obtain a perspective image of the three-dimensional model and a motion description sequence; and the motion description sequence comprises a motion description; generate a motion image corresponding to the motion description according to the motion description and the perspective image; determine a motion video according to the motion image corresponding to each motion description.
13. The apparatus of claim 11, wherein, The obtaining module is specifically configured to, obtain a motion description sequence and a candidate motion video corresponding to the motion description sequence; obtain first geometric structure information of a candidate object to which the candidate motion video belongs and second geometric structure information of the three-dimensional model; in a case where the first geometric structure information matches the second geometric structure information, determine the candidate motion video as the motion video.
14. The apparatus of claim 11, wherein, The first determining unit is specifically configured to, input the motion image into an image encoder to obtain the motion image encoding vector output by the image encoder; or determine a plurality of perspective images according to the motion image; input the plurality of perspective images into an image encoder to obtain the motion image encoding vector output by the image encoder.
15. The apparatus of claim 11, wherein, The second determining unit is specifically configured to input the initial point cloud into a point cloud encoder to obtain the initial point cloud encoding vector output by the point cloud encoder. The third determining unit is specifically configured to input the motion point cloud encoding vector into a point cloud decoder to obtain the motion point cloud output by the point cloud decoder.
16. A training device of a motion driving model, the device comprising: a first obtaining module configured to obtain an initial motion driving model; the motion driving model comprises a point cloud encoder, a diffusion model and a point cloud decoder connected in sequence; the diffusion model is configured to generate a motion point cloud encoding vector according to an initial point cloud encoding vector output by the point cloud encoder, with a motion image encoding vector as condition information; a second obtaining module configured to obtain training data; the training data comprises initial point clouds and sample motion point clouds of a three-dimensional model, and a sample motion image; and a third obtaining module configured to train the initial motion driving model according to the training data. The training processing module is configured to perform training processing on the point cloud encoder, the point cloud decoder, and the diffusion model by using the training data, to obtain a trained point cloud encoder, a trained point cloud decoder, and a trained diffusion model.
17. The apparatus of claim 16, wherein, The second obtaining module is specifically configured to, obtain a sample action point cloud and an initial point cloud of the three-dimensional model; perform three-dimensional model modeling processing according to the sample action point cloud, to obtain a modeled three-dimensional model; obtain a sample perspective image of the modeled three-dimensional model; determine the training data according to the sample action point cloud, the initial point cloud, and the sample perspective image.
18. The apparatus of claim 16, wherein, The training processing module includes a first training processing unit and a second training processing unit. The first training processing unit is configured to perform training processing on the point cloud encoder and the point cloud decoder by using the training data, to obtain a trained point cloud encoder and a trained point cloud decoder. The second training processing unit is configured to perform training processing on the diffusion model by using the trained point cloud encoder and the training data, to obtain a trained diffusion model.
19. The apparatus of claim 18, wherein, The first training processing unit is specifically configured to, determine a sample image encoding vector of the sample action image; input the initial point cloud and the sample image encoding vector into the point cloud encoder and the point cloud decoder connected in sequence, to obtain a predicted action point cloud output by the point cloud decoder; determine a first loss function value according to the sample action point cloud, the predicted action point cloud, and a loss function of the point cloud encoder and the point cloud decoder; perform parameter adjustment processing on the point cloud encoder and the point cloud decoder according to the first loss function value, to obtain the trained point cloud encoder and the trained point cloud decoder.
20. The apparatus of claim 18, wherein, The second training processing unit is specifically configured to, input the initial point cloud and the sample action point cloud into the trained point cloud encoder respectively, to obtain an initial point cloud encoding vector and a sample action point cloud encoding vector output by the trained point cloud encoder; determine a sample image encoding vector of the sample action image; determine the sample image encoding vector as condition information of the diffusion model, and input the initial point cloud encoding vector into the diffusion model, to obtain a predicted action point cloud encoding vector output by the diffusion model; determine a second loss function value according to the sample action point cloud encoding vector, the predicted action point cloud encoding vector, and a loss function of the diffusion model; perform parameter adjustment processing on the diffusion model according to the second loss function value, to obtain the trained diffusion model.
21. An electronic device, comprising: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 5; or, perform the method of any one of claims 6 to 10.
22. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are for causing the computer to perform the method according to any one of claims 1 to 5; or, to perform the method according to any one of claims 6 to 10.
23. A computer program product comprising computer instructions which, when executed by a processor, implement the method according to any one of claims 1 to 5; or, implement the method according to any one of claims 6 to 10.
Citation Information
Patent Citations
3D animation making method
CN105825539A
Three-dimensional digital human generation and interaction method and system
CN117496072A