Video generation model acquisition method and video generation method

By preprocessing and training the video generation model with posture signal sequences and reference image sequences, and using random masks and loss value correction, the problems of unnatural movements and color difference in video generation are solved, achieving higher quality video generation.

CN120602743APending Publication Date: 2025-09-05BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510821019.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing video generation models are prone to unnatural motion and excessive color difference when generating long videos, resulting in a decline in video quality.

Method used

By preprocessing the sample video, obtaining the posture signal sequence and reference image sequence, using the initial generation model for training, and through random mask and loss value correction, optimizing the video generation model to improve the inter-frame transition capability and video quality.

Benefits of technology

The training effect of the video generation model is improved, making the generated videos more natural and smooth, reducing unnatural movements and color differences, and improving the quality of video generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602743A_ABST
    Figure CN120602743A_ABST
Patent Text Reader

Abstract

The invention provides an acquisition method of a video generation model and a video generation method, relates to the technical field of artificial intelligence, in particular to the technical fields of computer vision, deep learning, large models and the like, and can be applied to scenes of digital human, artificial intelligence-based content generation and the like. The video generation model obtaining method comprises the steps that a sample video is preprocessed, a first attitude signal sequence and a reference image sequence are obtained, and the matching degree between at least one attitude signal in the first attitude signal sequence and a corresponding attitude signal in the sample video is smaller than a matching degree threshold value; a first reference image in the reference image sequence is a first frame image in the sample video; respectively inputting the first attitude signal sequence and the reference image sequence into an initial generation model to obtain a prediction result output by the initial generation model; determining a loss value based on the prediction result and the sample video; and based on the loss value, correcting the initial generation model until a target video generation model is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to technical fields such as computer vision, deep learning, and large models, and specifically to a method for acquiring a video generation model and a method for generating a video, which can be applied to scenarios such as digital humans and content generation based on artificial intelligence. Background Art

[0002] Digital humans exist in the non-physical world, are created and used by computers, and possess multiple human characteristics, including physical appearance, performance, and interaction abilities. Digital humans can perform tasks that would be impossible for real people, and have significant development implications in fields such as education, advertising, and film and television. Summary of the Invention

[0003] The present disclosure aims to solve one of the technical problems in the related art at least to a certain extent.

[0004] To this end, the purpose of the present invention is to propose a method for acquiring a video generation model and a video generation method. By training a video generation model that can read random masked posture signals, unnatural posture signals in the process of generating long videos are masked, so that the video generated by the model is more natural, and problems such as unnatural movements and excessive color difference in the video are reduced, thereby improving the generation quality of virtual videos.

[0005] According to a first aspect of the present disclosure, a method for obtaining a video generation model is provided, comprising:

[0006] Preprocessing the sample video to obtain a first gesture signal sequence and a reference image sequence, wherein a matching degree between at least one gesture signal in the first gesture signal sequence and a corresponding gesture signal in the sample video is less than a matching degree threshold, and a first reference image in the reference image sequence is a first frame image in the sample video;

[0007] inputting the first posture signal sequence and the reference image sequence into an initial generation model respectively to obtain a prediction result output by the initial generation model;

[0008] Determining a loss value based on the prediction result and the sample video;

[0009] Based on the loss value, the initial generation model is modified until a target video generation model is obtained.

[0010] According to a second aspect of the present disclosure, a video generation method is provided, comprising:

[0011] determining a target image and a first posture signal sequence;

[0012] parsing the first posture signal sequence to determine abnormal feature points in the first posture signal sequence;

[0013] Masking the abnormal feature points to obtain a second posture signal sequence;

[0014] The second posture signal sequence and the target image are input into the trained video generation model to obtain the target video.

[0015] According to a third aspect of the present disclosure, a device for acquiring a video generation model is provided, comprising:

[0016] A first processing module is configured to pre-process the sample video to obtain a first gesture signal sequence and a reference image sequence, wherein a matching degree between at least one gesture signal in the first gesture signal sequence and a corresponding gesture signal in the sample video is less than a matching degree threshold, and a first reference image in the reference image sequence is a first frame image in the sample video;

[0017] a first generation module, configured to input the first posture signal sequence and the reference image sequence into an initial generation model respectively, and obtain a prediction result output by the initial generation model;

[0018] A first determining module, configured to determine a loss value based on the prediction result and the sample video;

[0019] A correction module is used to correct the initial generation model based on the loss value until a target video generation model is obtained.

[0020] According to a fourth aspect of the present disclosure, a video generating apparatus is provided, comprising:

[0021] A second determination module is used to determine the target image and the first posture signal sequence;

[0022] a third determining module, configured to parse the first posture signal sequence and determine abnormal feature points in the first posture signal sequence;

[0023] A second processing module is used to perform mask processing on the abnormal feature points to obtain a second posture signal sequence;

[0024] The second generation module is used to input the second posture signal sequence and the target image into the trained video generation model to obtain the target video.

[0025] According to a fifth aspect of the present disclosure, there is provided an electronic device, including:

[0026] at least one processor; and

[0027] a memory communicatively connected to the at least one processor; wherein,

[0028] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for obtaining the video generation model as described in the first aspect, or the video generation method as described in the second aspect.

[0029] According to a sixth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method for obtaining a video generation model as described in the first aspect, or the video generation method as described in the second aspect.

[0030] According to a seventh aspect of the present disclosure, a computer program product is provided, comprising computer instructions, which, when executed by a processor, implement the method for obtaining a video generation model as described in the first aspect, or the steps of the video generation method as described in the second aspect.

[0031] The video generation model acquisition method and video generation method provided by the present disclosure have the following beneficial effects:

[0032] First, a gesture signal sequence corresponding to a sample video, containing at least one masked gesture signal, and a reference image sequence containing the first frame, are obtained. The initial generative model's predictions for the gesture signal sequence and the reference image sequence are then obtained. The initial generative model is then corrected based on the loss between the predictions and the sample video. Using masked data to train the video generative model ensures that the trained model has better inter-frame transition capabilities. This allows the model to generate high-quality videos even after masking problematic gesture signals, improving the training effectiveness of the video generative model.

[0033] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The above and / or additional aspects and advantages of the present disclosure will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, which are provided for a better understanding of the present solution and do not constitute a limitation of the present disclosure.

[0035] Figure 1 is a flowchart of a method for obtaining a video generation model proposed according to an embodiment of the present disclosure;

[0036] Figure 2is a flowchart of a method for obtaining a video generation model proposed according to another embodiment of the present disclosure;

[0037] Figure 3 is a flowchart of a method for obtaining a video generation model proposed according to another embodiment of the present disclosure;

[0038] Figure 4 is a flowchart of a method for obtaining a video generation model proposed according to another embodiment of the present disclosure;

[0039] Figure 5 is a flowchart of a method for obtaining a video generation model proposed according to another embodiment of the present disclosure;

[0040] Figure 6 is a flowchart of a method for obtaining a video generation model proposed according to another embodiment of the present disclosure;

[0041] Figure 7 is a schematic diagram of a model training process proposed in an embodiment of the present disclosure;

[0042] Figure 8 is a flowchart of a video generation method proposed according to an embodiment of the present disclosure;

[0043] Figure 9 is a schematic diagram of a video generation process proposed in an embodiment of the present disclosure;

[0044] Figure 10 1 is a schematic structural diagram of a device for acquiring a video generation model according to an embodiment of the present disclosure;

[0045] Figure 11 is a structural diagram of a video generating device proposed according to an embodiment of the present disclosure;

[0046] Figure 12 A block diagram of an exemplary electronic device suitable for implementing embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0047] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0048] The embodiments of the present disclosure relate to the field of artificial intelligence technology, and in particular to technical fields such as computer vision, deep learning, and large models.

[0049] Among them, Artificial Intelligence (AI) is a new technical science that studies and develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence.

[0050] Computer vision refers to the use of cameras and computers to replace the human eye to identify, track, and measure targets, and further perform graphic processing so that the computer processing becomes an image that is more suitable for human eye observation or transmission to instrument detection.

[0051] Deep learning involves learning the inherent patterns and representational hierarchies of sample data. The information gained from this learning process is highly helpful in interpreting data such as text, images, and sounds. The ultimate goal of deep learning is to enable machines to have the same analytical and learning capabilities as humans, enabling them to recognize data such as text, images, and sounds.

[0052] Large AI models (abbreviated as "large models") are a class of AI models with a large number of parameters, constructed using artificial neural networks. They are typically pre-trained on massive amounts of data through self-supervised or semi-supervised learning, and their performance and capabilities are further optimized through methods such as instruction fine-tuning and human alignment. Large models are characterized by a large number of parameters, extensive training data, and extensive computing resources, enabling them to solve common tasks, follow human instructions, and perform complex reasoning.

[0053] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0054] The following describes a method for acquiring a video generation model and a method for generating a video according to an embodiment of the present disclosure with reference to the accompanying drawings.

[0055] It should be noted that the execution subject of the method for obtaining the video generation model of this embodiment is the acquisition device of the video generation model, and the execution subject of the video generation method is the video generation device. Both devices can be implemented by software and / or hardware and can be configured in electronic devices. The electronic devices may include but are not limited to terminals, server terminals, etc.

[0056] Figure 1 It is a flowchart of a method for obtaining a video generation model proposed according to an embodiment of the present disclosure.

[0057] like Figure 1 As shown, the method for obtaining the video generation model includes:

[0058] S101: Preprocessing a sample video to obtain a first posture signal sequence and a reference image sequence.

[0059] The first posture signal sequence is a series of human or object posture data that changes over time. These data may include information such as joint angles, positions, and speeds, and are used to describe the dynamic changes of posture.

[0060] In the disclosed embodiment, a sample video can be obtained by real-person recording or live broadcast. When pre-processing the sample video, the sample video can be cropped with the person centered, and then the cropped video can be divided into short videos of a fixed duration (such as 10 seconds) in the time dimension. After that, the existing posture estimation model can be used to obtain the posture data of any frame for each short video, so that the first posture signal sequence corresponding to the sample data can be obtained.

[0061] In the present disclosure, at least one posture signal in the first posture signal sequence may have a matching degree with the corresponding posture signal in the sample video that is less than a matching degree threshold, so that when the video generation model is trained using the first posture signal sequence, the model's inter-frame transition capability can be retained, and the trained model can generate smooth videos even if problematic posture signals are masked during inference, thereby obtaining more natural results and optimizing the video generation effect.

[0062] It should be noted that the matching threshold can be set according to the accuracy requirements of the model training in actual application. The higher the matching threshold, the stronger the ability of the trained model to generate smooth video after masking the erroneous posture signal.

[0063] In an embodiment of the present disclosure, after obtaining a gesture signal sequence corresponding to a sample video using a gesture estimation model, at least one gesture signal may be randomly selected from the sequence and subjected to random masking to obtain a first gesture signal sequence. The random masking may be performed by fully masking the at least one randomly selected gesture signal or by partially masking the at least one randomly selected gesture signal.

[0064] The first reference image in the reference image sequence is the first frame image in the sample video.

[0065] In this embodiment of the present application, after dividing the sample data into multiple short videos of fixed length, one frame can be extracted from each short video as a reference image. A reference image sequence of all reference images can then be obtained in chronological order. Because the first frame of a video typically contains the initial position and appearance of the target object, the first frame of the sample video can be determined as the first reference image in the reference image sequence.

[0066] S102: Inputting the first posture signal sequence and the reference image sequence into the initial generation model respectively to obtain a prediction result output by the initial generation model.

[0067] The initial generation model can be any open source model with video generation capabilities, for example, a generation model DIT (Diffusion Transformer) that combines diffusion models and transformer architecture.

[0068] In the disclosed embodiments, for initial generative models of different architectures, the loss calculation methods used to train the models may be different, such as image scale loss or feature level loss, so the prediction results output by the initial generative model may be predicted images generated for other frames in the sample video other than the reference image, or may be noise that should be reduced predicted at the feature level.

[0069] For example, the initial generative model is a generative model DIT that combines the diffusion model and the Transformer architecture, and the initial generative model includes a pre-trained variational autoencoder (VAE). After the first posture signal sequence and the reference image sequence are respectively input into the initial generative model, the first posture signal sequence and the reference image sequence can be respectively encoded into corresponding feature vectors through VAE, and then the feature vectors can be spliced ​​with the noise, and the noise that should be reduced can be predicted through DIT to obtain the prediction result output by the initial generative model.

[0070] S103: Determine a loss value based on the prediction result and the sample video.

[0071] In the disclosed embodiments, when the prediction result is a predicted image generated for a frame other than the reference image in the sample video, the difference between the predicted image and the corresponding frame in the sample video can be calculated to obtain a loss value. Alternatively, when the prediction result is noise that should be reduced, the difference between the predicted noise and the true noise value in the sample video can be calculated to obtain a loss value.

[0072] S104: Based on the loss value, the initial generation model is modified until the target video generation model is obtained.

[0073] In the disclosed embodiment, after determining the loss value, the initial generative model can be modified by adjusting the learning rate, for example. After each modification, the modified generative model can be used to repeat steps 102 and 103 to calculate a new loss value. When the number of iterations reaches a maximum limit or the loss value falls below a preset threshold, the generative model at that point is determined as the target video generative model.

[0074] In this embodiment, a gesture signal sequence corresponding to a sample video, including at least one masked gesture signal, and a reference image sequence containing the first frame image are first obtained. The initial generative model's predictions for the gesture signal sequence and the reference image sequence are then obtained. The initial generative model is then corrected based on the loss between the predictions and the sample video. By using masked data to train the video generative model, the trained video generative model can be ensured to have better inter-frame transition capabilities. This allows the model to generate high-quality videos even after masking problematic gesture signals, thereby improving the training effectiveness of the video generative model.

[0075] Figure 2 It is a flowchart of a method for obtaining a video generation model proposed in another embodiment of the present disclosure.

[0076] like Figure 2 As shown, the method for obtaining the video generation model includes:

[0077] S201: Perform posture estimation on a sample video to determine a second posture signal sequence corresponding to the sample video.

[0078] The second posture signal sequence is a plurality of posture signals directly extracted from the sample video. The second posture sequence may not contain a posture signal whose matching degree with the corresponding posture in the sample video is less than a matching threshold.

[0079] In an embodiment of the present disclosure, the sample video can be divided into short videos of a fixed length (such as 10 seconds) in the time dimension, and a posture estimation model is used for each short video to generate a posture signal of at least one video frame. All posture signals are then arranged in time sequence of the video frames to obtain a second posture signal sequence corresponding to the sample data.

[0080] S202: Perform random masking on the second gesture signal sequence to obtain a first gesture signal sequence.

[0081] In the disclosed embodiment, random masking is beneficial to improving the robustness of the video generation model to partial occlusion by randomly occluding posture key points or joint angle signals. When the second posture signal sequence is subjected to random masking, at least one frame of posture signal images in the second posture signal sequence may be masked in its entirety, or partial signals in partial frames of posture signal images may be masked. The posture signal sequence that has completed the masking process is the first posture signal sequence, which contains at least one posture signal whose matching degree with the corresponding posture in the sample video is less than the matching threshold.

[0082] S203: Inputting the first posture signal sequence and the reference image sequence of the sample data into the initial generation model respectively to obtain a prediction result output by the initial generation model.

[0083] S204: Determine a loss value based on the prediction result and the sample video.

[0084] S205: Based on the loss value, the initial generation model is modified until the target video generation model is obtained.

[0085] The description of S203 to S205 can be found in the above embodiment and will not be repeated here.

[0086] In this embodiment, the posture signal is randomly masked and the masked posture signal is used as training data to train the model, which provides data conditions for improving the inter-frame transition capability of the trained video generation model, and is conducive to improving the quality of the model-generated video.

[0087] It should be noted that in addition to first estimating the posture of the sample video and then randomly masking part of the posture signals in the posture signal sequence to obtain the first posture signal sequence, it is also possible to obtain the first posture signal sequence by directly and randomly not extracting posture signals from certain frames in the video sample, or only extracting posture signals from partial areas in certain frames, so that at least one posture signal has a matching degree with the posture signal of the original video frame that is less than a matching degree threshold.

[0088] Optionally, posture estimation may be performed on a second image other than the first image in the sample video to determine a first posture signal sequence.

[0089] The first image is determined based on a random number and is an image of a frame in the sample video for which no posture estimation is performed, which is equivalent to masking the posture signal of the image.

[0090] It should be noted that there can be multiple random numbers, each corresponding to a corresponding sequence of image frames in a sample video. The maximum number of random numbers is the total number of frames in the sample video. The more random numbers there are, the more frames in the sample video are not subjected to pose estimation, and the higher the training intensity of the obtained first pose signal sequence for the model. The number of random numbers can be set according to actual needs.

[0091] In the disclosed embodiment, multiple numbers can be randomly generated first, and the frame number of each number in the sample video corresponding to the image is determined as the first image. For example, when the sample video has 10 frames and the random numbers are 2, 4, 6, and 8, the images of the 2nd, 4th, 6th, and 8th frames in the sample video are the first images. Then, when extracting the posture signal, the posture of the second image in the sample video other than the first image can be estimated, and the estimation results can be sorted according to the order of the second image to obtain the first posture signal sequence. At this time, in the first posture signal sequence, the position corresponding to the first image is empty.

[0092] Alternatively, posture estimation may be performed on a partial area of ​​the first image and the entire area of ​​the second image in the sample video to determine the first posture signal sequence.

[0093] In the disclosed embodiment, posture estimation can be performed on a partial area in the first image, and the posture signal obtained in this way can also meet the condition that the matching degree between the posture signal and the original sample video is lower than the matching degree threshold, thereby obtaining a first posture signal sequence.

[0094] In the disclosed embodiment, a posture signal sequence is obtained by performing no posture estimation or only performing posture estimation on a partial area of ​​an image of at least one random frame in a sample video, thereby enriching the method of determining a posture signal sequence with a masking effect, and also enabling the trained video generation model to have better inter-frame transition capabilities, which is beneficial to improving the quality of the video generated by the model.

[0095] Figure 3 It is a flowchart of a method for obtaining a video generation model proposed in another embodiment of the present disclosure.

[0096] like Figure 3 As shown, the method for obtaining the video generation model includes:

[0097] S301: Preprocessing the sample video to obtain a first posture signal sequence and a reference image sequence.

[0098] The description of S301 can be found in the above embodiment and will not be repeated here.

[0099] S302: Determine a first random probability.

[0100] The first random probability is a number randomly generated for any sample video when the model is supervised for training. It represents the probability that the last frame of the sample data will be obtained during training. The first random probability can be uniformly distributed between 0 and 1.

[0101] In the model training process disclosed herein, whether to add the tail frame is determined by random probability, which allows the model to learn the process from the first frame to the last frame while ensuring the ability to infer the tail frame without giving the tail frame.

[0102] S303: When the first random probability is greater than a first threshold, obtain a last frame image in the sample video.

[0103] The first threshold may be any value set according to actual needs, such as 0.2, 0.5, etc. The first threshold may determine whether the last frame of the video needs to be extracted together with the first frame as reference images when training a model using a sample video.

[0104] In the embodiment of the present disclosure, when the first random probability is greater than the first threshold, the last frame of the sample video can be obtained as a reference image for model training.

[0105] S304: Using the last frame image, replace the last reference image frame in the reference image sequence to obtain an updated reference image sequence.

[0106] In the embodiment of the present disclosure, the last reference image frame in the reference image sequence may be any frame image of the last short video among the multiple short videos split from the sample video, and the last frame image of the sample video is also the last frame image of the last short video. In this case, the last frame image can be used to replace the last reference image frame in the reference image sequence to obtain an updated reference image sequence.

[0107] In the embodiment of the present disclosure, during the model training process, only the first frame or the first frame and the last frame are selected as reference images according to a certain probability, so as to increase the frame insertion capability of the trained model. In the early stage of long video inference, the last frame information can be added without adding the last frame information, and the last frame information can be added to the last video segment of the long video inference to connect back to the first frame, so as to better meet the needs of actual business.

[0108] S305: Inputting the first posture signal sequence and the updated reference image sequence into the initial generation model respectively to obtain a prediction result output by the initial generation model.

[0109] S306: Determine a loss value based on the prediction result and the sample video.

[0110] S307: Based on the loss value, the initial generation model is modified until the target video generation model is obtained.

[0111] The description of S305 to S307 can be found in the above embodiment and will not be repeated here.

[0112] In this embodiment, by adding the last frame as a reference with a certain probability during the training stage, the trained model can have better connection capabilities between the first and last frames when generating videos. When the model generates long videos in segments, it can ensure that the last generation result can transition back to the first frame, thereby improving the quality of the generated video.

[0113] Figure 4 It is a flowchart of a method for obtaining a video generation model proposed in another embodiment of the present disclosure.

[0114] like Figure 4 As shown, the method for obtaining the video generation model includes:

[0115] S401: Preprocessing the sample video to obtain a first posture signal sequence and a reference image sequence.

[0116] The description of S401 can be found in the above embodiment and will not be repeated here.

[0117] S402: Determine a background sequence corresponding to the sample video.

[0118] The background sequence refers to a series of continuous frames formed over time by the background part of the video that does not change or changes slowly over time (such as static scenes, fixed objects, etc.).

[0119] In the disclosed embodiments, when training a model, if only a few reference images in a reference image sequence are available (e.g., only the first frame, or the first and last two frames), the model cannot obtain background information for the remaining frames, potentially resulting in unstable backgrounds in the generated video. Therefore, it is possible to obtain background sequences corresponding to sample videos to train the model, resulting in a more stable background when the trained model generates videos.

[0120] In an embodiment of the present disclosure, in a reference image sequence, for frames other than the first and last frames of the sample, a posture signal can be used to obtain the bounding rectangle of the posture graph, and the bounding rectangle can be used to mask out the portrait information in the image and retain the background information, thereby obtaining a background sequence corresponding to the sample data.

[0121] Optionally, a mask sequence may be determined based on the third posture signal sequence, and then mask processing may be performed on the sample video based on the mask sequence to obtain a background sequence.

[0122] The third posture signal sequence may be a series of human or object posture data that changes with time, extracted from frames other than the first frame or the first and last two frames in the sample video.

[0123] In the disclosed embodiment, a bounding rectangle corresponding to each third gesture signal in the third gesture signal sequence can be determined. A mask sequence is then determined based on the size of the bounding rectangle, its position within each video frame, and the corresponding temporal sequence. In other words, each bounding rectangle in the mask sequence is represented by a gray square. By masking the sample video with the person in the corresponding video frame, masking is performed on the sample video to obtain a background sequence. This ensures that the background sequence does not contain the person's gesture information, thereby improving the accuracy and efficiency of background sequence acquisition.

[0124] In the embodiment of the present disclosure, in addition to obtaining the background sequence through mask processing, the background can also be directly extracted from the sample video to obtain the background sequence.

[0125] S403: Using the background sequence, replace the other reference images in the reference image sequence except the first reference image to obtain an updated reference image sequence.

[0126] In the disclosed embodiment, in order to improve the intensity of model training so that the postures of the characters in the video are smoother and more stable when the trained model generates a video, the reference image sequence may only contain the first frame or the first and last two frames of the sample video. However, when there are only reference images of the first frame or the first and last two frames, the model cannot obtain the background information of the remaining frames, which may cause the background of the generated video to be unstable. Therefore, the background sequence can be used to replace the other reference images in the reference image sequence except the first reference image to obtain an updated reference image sequence. If the reference image sequence includes the last frame image, the background sequence replaces the other images except the first reference image and the last reference image in the reference image sequence to obtain an updated reference image sequence.

[0127] It should be noted that, in the embodiment of the present disclosure, all reference images except the first reference image in the reference image sequence may not be replaced, but may be selectively replaced according to a certain probability.

[0128] Optionally, the second probabilities currently corresponding to the other reference images can be determined first. In the disclosed embodiment, a second probability can be randomly generated for each reference image in the reference image sequence except the first reference image. The second probability is used to select the reference image in the reference image sequence to be replaced with the background image.

[0129] Then, the corresponding background images in the background sequence may be used to replace the corresponding reference images whose second probabilities are greater than the second threshold, so as to obtain an updated reference image sequence.

[0130] The second threshold may be a minimum probability set as needed. When the probability is greater than the second threshold, the background image is used to replace the reference image in the reference image sequence to achieve selective replacement of the reference image.

[0131] In the disclosed embodiment, a second probability is randomly assigned to each reference image in the reference image sequence except the first reference image, and based on the relationship between the second probability and the threshold, some reference images in the reference image sequence are selectively replaced with background images, thereby further improving the diversity and quality of model training.

[0132] S404: Inputting the first posture signal sequence and the updated reference image sequence into the initial generation model respectively to obtain a prediction result output by the initial generation model.

[0133] S405: Determine a loss value based on the prediction result and the sample video.

[0134] S406: Based on the loss value, the initial generation model is modified until the target video generation model is obtained.

[0135] The description of S404 to S406 can be found in the above embodiments and will not be repeated here.

[0136] In this embodiment, the video generation model is trained by injecting background information into the reference image sequence, so that the trained model can reduce the color difference between the generated result and the first and last frames of the input when generating the video, further reducing the accumulated color difference during long video inference, which is beneficial to improving the quality of the generated video.

[0137] Figure 5 It is a flowchart of a method for obtaining a video generation model proposed in another embodiment of the present disclosure.

[0138] like Figure 5 As shown, the method for obtaining the video generation model includes:

[0139] S501: Preprocessing the sample video to obtain a first posture signal sequence and a reference image sequence.

[0140] The description of S501 can be found in the above embodiment and will not be repeated here.

[0141] S502: Determine description information corresponding to the sample video.

[0142] The description information is a text description of some details in the sample video (such as background, character hairstyle and clothing, etc.).

[0143] In the disclosed embodiment, a large video understanding model can be used to detect and generate video descriptions to obtain description information corresponding to the sample video.

[0144] S503: Inputting the description information, the first posture signal sequence and the reference image sequence into the initial generation model respectively to obtain the prediction result output by the initial generation model.

[0145] In this embodiment, by adding description information during the model training process, the model can learn more information about sample videos, thereby improving the accuracy of the model-generated videos.

[0146] S504: Determine a loss value based on the prediction result and the sample video.

[0147] S505: Based on the loss value, the initial generation model is modified until the target video generation model is obtained.

[0148] The description of S504 and S505 can be found in the above embodiments, which will not be repeated here.

[0149] It should be noted that, in the embodiment of the present disclosure, the specific training process for the initial generation model of different architectures is not exactly the same. Figure 6This article describes in detail the training process for a generative model using a variational autoencoder (VAE) plus a diffusion transformer (DIT) architecture. DIT is an open-source model that combines diffusion models with the Transformer architecture.

[0150] Figure 6 It is a flowchart of a method for obtaining a video generation model proposed in another embodiment of the present disclosure.

[0151] like Figure 6 As shown, the method for obtaining the video generation model includes:

[0152] S601: Preprocess the sample video to obtain a first posture signal sequence and a reference image sequence.

[0153] The description of S601 can be found in the above embodiment and will not be repeated here.

[0154] S602: Input the first posture signal sequence and the reference image sequence into the encoding network in the initial generation model to obtain a feature vector.

[0155] Here, the encoding network refers to the pre-trained variational autoencoder VAE.

[0156] In the disclosed embodiment, the first posture signal sequence and the reference image sequence are respectively encoded into latent features through the encoding network in the initial generation model to obtain their corresponding feature vectors.

[0157] S603: After fusing the feature vector with the true value of the noise corresponding to the sample video, the feature vector is input into the diffusion network in the initial generation model to obtain the predicted noise output by the diffusion network.

[0158] The true value of the noise is generated by gradually adding Gaussian noise to the sample video, and the noise addition can satisfy the normal distribution.

[0159] Among them, the diffusion network is a DIT that combines the diffusion model (Diffusion Models) and the Transformer architecture.

[0160] In the disclosed embodiment, the feature vector of the posture signal, the feature vector of the reference image, and the true value of the noise can be spliced ​​together, and the spliced ​​result can be input into the diffusion network and converted into the smallest semantic unit (token) read by the diffusion network to predict the noise that should be reduced and obtain the output predicted noise.

[0161] S604: Determine a loss value based on the predicted noise and the true value of the noise.

[0162] In the disclosed embodiment, the diffusion network gradually adds Gaussian noise to the image, converting the image into pure noise (forward process), and then learns to recover the original data from the noise (reverse process). The training goal is to predict the added noise through the neural network, thereby optimizing the denoising ability of the reverse process. Therefore, any loss function can be used to calculate the difference between the predicted noise and the true noise value to obtain a loss value, which can be used to train the diffusion network to achieve the purpose of training the video generation model.

[0163] S605: Based on the loss value, the diffusion network is modified until the target video generation model is obtained.

[0164] In the disclosed embodiment, after determining the loss value, the gradient of the loss with respect to the diffusion network parameters is calculated, and the parameters are then updated based on the gradient to complete the correction of the diffusion network. After each correction, the corrected generative model is used to repeat steps 602 to 604 to calculate a new loss value. When the number of iterations reaches a maximum limit or the loss value falls below a preset threshold, the generative model at that point is determined as the target video generative model.

[0165] In this embodiment, during the training process of the generative model of the VAE plus DIT structure, the noise that should be reduced is predicted from the feature level, and the loss value is calculated to correct the diffusion network in the video generation model, thereby further improving the quality of model training.

[0166] Below is Figure 7 Take the example to describe how the video generation model is obtained. Figure 7 It is a schematic diagram of the generative model training process.

[0167] Depend on Figure 7 It can be seen that the sample video 701 used for training can be split into multiple videos first, and images can be extracted from each video to obtain a reference image sequence. Then, the character's action posture (Pose) can be identified for each image in the reference image sequence, and the posture signal sequence 702 can be extracted, and the first frame and the last frame of the sample video can be obtained. After that, the posture signal sequence 702 and the two reference images of the first frame and the last frame can be input into the VAE encoder of the initial generation model respectively to obtain the feature vector 703 of the posture signal and the feature vector 704 of the reference image. Then, after splicing these two feature vectors 703 and 704 with noise, they are input into the DIT diffusion network, that is, the diffusion network can be corrected according to the output predicted noise to obtain the target video generation model.

[0168] Figure 8 It is a flowchart of a video generation method proposed in another embodiment of the present disclosure.

[0169] like Figure 8 As shown, the video generation method includes:

[0170] S801: Determine a target image and a first posture signal sequence.

[0171] In the disclosed embodiment, an insertion point may be selected in a video of a digital human or a real person, the insertion point frame may be selected as a target image, and a gesture signal may be generated using a speech-driven gesture model to obtain a first gesture signal sequence.

[0172] Optionally, in an embodiment of the present disclosure, the target image may be determined in any of the following ways: when the target video is the first sub-video, the previous image frame corresponding to the insertion position in the video insertion instruction may be determined as the target image.

[0173] Alternatively, when the target video is not the first sub-video, the last image frame in the previous sub-video adjacent to the target video may be determined as the target image.

[0174] Alternatively, when the target video is the last sub-video, the last image frame in the previous sub-video adjacent to the target video and the next image frame corresponding to the insertion position may be determined as the target image.

[0175] In the embodiment of the present disclosure, by using the previous image frame of the insertion position or the last image frame of the previously generated video as the reference image for the target video to be generated, it can be ensured that when multiple segments of the video are generated, the transition between adjacent video segments is more natural, thereby improving the quality of long videos generated by the model.

[0176] S802: Analyze the first gesture signal sequence to determine abnormal feature points in the first gesture signal sequence.

[0177] In the disclosed embodiments, when a speech-driven gesture model is used to generate a gesture signal, anomalies may exist in the generated gesture signal due to factors such as noise, sensor failure, abnormal human movement, or environmental interference. Abnormal feature points in the first gesture signal sequence may be abnormal compared to the upper and lower frames, such as sudden jitter or violent movement. Abnormal feature points in the first gesture signal sequence can be identified through statistical, time series analysis, and optical flow analysis.

[0178] S803: Mask the abnormal feature points to obtain a second posture signal sequence.

[0179] In the disclosed embodiment, abnormal feature points can be masked so that the posture signal in the second posture signal sequence is more consistent with the actual action trajectory, and then a video is generated based on the second posture signal sequence, which can effectively reduce unnatural postures in the video and improve video quality.

[0180] S804: Input the second posture signal sequence and the target image into the trained video generation model to obtain the target video.

[0181] In the disclosed embodiment, since the gesture signal sequence is randomly masked during the training of the video generation model, the trained video generation model can also generate gesture-coherent video content based on the second gesture signal sequence containing the mask.

[0182] In the disclosed embodiment, the reference image of the first video segment is derived from the insertion point. After the video is generated through the posture signal, the reference image of the second video segment is derived from the last frame of the first video segment. Multiple videos are generated in sequence. The last video segment uses the last frame of the previous video segment and the insertion point image as the first and last frame reference images to generate a video that can connect back to the insertion point. Finally, multiple videos are combined to generate a long video.

[0183] In this embodiment, after masking the abnormal feature points in the posture signal, the target image of the previous frame adjacent to the insertion position is input into the trained video generation model to obtain the target video. This allows the character's posture to be more natural, stable, and free of color difference when the model generates a long video, meeting the requirement that the generated video can connect back to the first frame.

[0184] Below is Figure 9 Take the video generation process as an example. Figure 9 A schematic diagram of a video generation process provided in an embodiment of the present disclosure.

[0185] like Figure 9 As shown, firstly, the posture signal sequence 901 is obtained by recognizing the voice, and an insertion point is selected in a digital human video as needed, and the frame before the insertion point is the first frame 902 of the target generated video, and the frame after the insertion point is the last frame 903 of the target generated video. The first frame 902 and the last frame 903 are the target images in the above steps. The first frame and the last frame are obtained as references for generating the video, so that when the generated video is inserted into the original video, there will be no color difference or large pixel changes at the same position, making the video more natural and smooth. The posture signal sequence 901, the first frame 902 and the last frame 903 are input into the trained video generation model. The video generation model can use the VAE encoder to encode the posture signal sequence 901, the first frame 902 and the last frame 903 into corresponding feature vectors, that is, Figure 9The pose signal features 904 and image features 905 in the image are concatenated with noise and then input into the diffusion network to obtain the output of the diffusion network. The output of the diffusion network is then passed through the VAE decoder to generate the target video. This process can be repeated to generate multiple video segments to obtain a long video.

[0186] Figure 10 It is a structural diagram of a device for acquiring a video generation model proposed in an embodiment of the present disclosure.

[0187] like Figure 10 As shown, the video generation model acquisition device 100 includes:

[0188] A first processing module 1001 is configured to pre-process the sample video to obtain a first gesture signal sequence and a reference image sequence, wherein a matching degree between at least one gesture signal in the first gesture signal sequence and a corresponding gesture signal in the sample video is less than a matching degree threshold, and a first reference image in the reference image sequence is the first frame image in the sample video;

[0189] The first generation module 1002 is used to input the first posture signal sequence and the reference image sequence into the initial generation model respectively to obtain the prediction result output by the initial generation model;

[0190] A first determining module 1003 is configured to determine a loss value based on the prediction result and the sample video;

[0191] The correction module 1004 is used to correct the initial generation model based on the loss value until the target video generation model is obtained.

[0192] In some embodiments, the apparatus 100 for obtaining a video generation model further includes:

[0193] A posture estimation module, configured to perform posture estimation on the sample video and determine a second posture signal sequence corresponding to the sample video;

[0194] The mask processing module is used to perform random mask processing on the second posture signal sequence to obtain the first posture signal sequence.

[0195] In some embodiments, the first processing module 1001 may be used for any of the following:

[0196] Performing posture estimation on a second image other than the first image in the sample video to determine a first posture signal sequence, wherein the first image is determined based on a random number;

[0197] Performing posture estimation on a partial area of ​​a first image and an entire area of ​​a second image in a sample video to determine a first posture signal sequence, wherein the first image is determined based on a random number.

[0198] In some embodiments, the first processing module 1001 may also be used to:

[0199] determining a first random probability;

[0200] When the first random probability is greater than a first threshold, obtaining a last frame image in the sample video;

[0201] The last reference image frame in the reference image sequence is replaced by the tail frame image to obtain an updated reference image sequence.

[0202] In some embodiments, the first processing module 1001 may also be used to:

[0203] Determine the background sequence corresponding to the sample video;

[0204] The background sequence is used to replace the other reference images in the reference image sequence except the first reference image to obtain an updated reference image sequence.

[0205] In some embodiments, the first processing module 1001 may be specifically configured to:

[0206] determining a mask sequence based on the third posture signal sequence;

[0207] Based on the mask sequence, the sample video is masked to obtain the background sequence.

[0208] In some embodiments, the first processing module 1001 may be specifically configured to:

[0209] Determine second probabilities currently corresponding to other reference images respectively;

[0210] The corresponding background images in the background sequence are used to replace the corresponding reference images whose second probability is greater than the second threshold, so as to obtain an updated reference image sequence.

[0211] In some embodiments, the first generating module 1002 may be specifically configured to:

[0212] Determine description information corresponding to the sample video;

[0213] The description information, the first posture signal sequence and the reference image sequence are respectively input into the initial generation model to obtain the prediction result output by the initial generation model.

[0214] In some embodiments, the first generating module 1002 may be specifically configured to:

[0215] Input the first posture signal sequence and the reference image sequence into the encoding network in the initial generation model to obtain a feature vector;

[0216] After fusing the feature vector with the true noise value corresponding to the sample video, the vector is input into the diffusion network in the initial generation model to obtain the predicted noise output by the diffusion network.

[0217] In some embodiments, the first determining module 1003 may be specifically configured to:

[0218] Determine the loss value based on the predicted noise and the true value of the noise;

[0219] Based on the loss value, the initial generation model is modified until the target video generation model is obtained, including:

[0220] Based on the loss value, the diffusion network is modified until the target video generation model is obtained.

[0221] It should be noted that the aforementioned explanation of the method for obtaining the video generation model is also applicable to the device for obtaining the video generation model of this embodiment, and will not be repeated here.

[0222] In this embodiment, a gesture signal sequence corresponding to a sample video, including at least one masked gesture signal, and a reference image sequence containing the first frame image are first obtained. The initial generative model's predictions for the gesture signal sequence and the reference image sequence are then obtained. The initial generative model is then corrected based on the loss between the predictions and the sample video. By using masked data to train the video generative model, the trained video generative model can be ensured to have better inter-frame transition capabilities. This allows the model to generate high-quality videos even after masking problematic gesture signals, thereby improving the training effectiveness of the video generative model.

[0223] Figure 11 It is a structural diagram of a video generating device proposed in one embodiment of the present disclosure.

[0224] like Figure 11 As shown, the video generating device 110 includes:

[0225] The second determining module 1101 is used to determine the target image and the first posture signal sequence;

[0226] A third determining module 1102 is configured to analyze the first gesture signal sequence and determine abnormal feature points in the first gesture signal sequence;

[0227] The second processing module 1103 is used to perform mask processing on the abnormal feature points to obtain a second posture signal sequence;

[0228] The second generation module 1104 is configured to input the second posture signal sequence and the target image into the trained video generation model to obtain the target video.

[0229] In some embodiments, determining the target image includes any of the following:

[0230] When the target video is the first sub-video, the previous image frame corresponding to the insertion position in the video insertion instruction is determined as the target image;

[0231] In the case where the target video is not the first sub-video, the last image frame in the previous sub-video adjacent to the target video is determined as the target image;

[0232] When the target video is the last sub-video, the last image frame in the previous sub-video adjacent to the target video and the next image frame corresponding to the insertion position are determined as the target image.

[0233] It should be noted that the above explanation of the video generating method is also applicable to the video generating device of this embodiment and will not be repeated here.

[0234] In this embodiment, after masking the abnormal feature points in the posture signal, the target image of the previous frame adjacent to the insertion position is input into the trained video generation model to obtain the target video. This allows the character's posture to be more natural, stable, and free of color difference when the model generates a long video, meeting the requirement that the generated video can connect back to the first frame.

[0235] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0236] Figure 12 A schematic block diagram of an example electronic device 1200 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0237] like Figure 12As shown, the device 1200 includes a computing unit 1201, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1202 or a computer program loaded from a storage unit 1208 into a random access memory (RAM) 1203. Various programs and data required for the operation of the device 1200 can also be stored in the RAM 1203. The computing unit 1201, the ROM 1202, and the RAM 1203 are connected to each other via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.

[0238] Various components in device 1200 are connected to I / O interface 1205, including an input unit 1206, such as a keyboard and mouse; an output unit 1207, such as various types of displays and speakers; a storage unit 1208, such as a magnetic disk and optical disk; and a communication unit 1209, such as a network card, a modem, a wireless communication transceiver, etc. Communication unit 1209 allows device 1200 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0239] The computing unit 1201 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1201 performs the various methods and processes described above, such as the method for obtaining a video generation model. For example, in some embodiments, the method for obtaining a video generation model can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 1208. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 1200 via the ROM 1202 and / or the communication unit 1209. When the computer program is loaded into the RAM 1203 and executed by the computing unit 1201, one or more steps of the method for obtaining a video generation model described above can be performed. Alternatively, in other embodiments, the computing unit 1201 may be configured to execute the method for acquiring the video generation model in any other appropriate manner (for example, by means of firmware).

[0240] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0241] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0242] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0243] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0244] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.

[0245] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship is established by computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or simply "VPS"). The server may also be a server in a distributed system or a server integrated with blockchain.

[0246] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0247] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present disclosure, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined. In the description of the present disclosure, the words "if" and "if" used can be interpreted as "at the time of" or "when" or "in response to a determination" or "under the circumstances of".

[0248] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A method for obtaining a video generation model, comprising: Preprocessing the sample video to obtain a first gesture signal sequence and a reference image sequence, wherein a matching degree between at least one gesture signal in the first gesture signal sequence and a corresponding gesture signal in the sample video is less than a matching degree threshold, and a first reference image in the reference image sequence is a first frame image in the sample video; inputting the first posture signal sequence and the reference image sequence into an initial generation model respectively to obtain a prediction result output by the initial generation model; Determining a loss value based on the prediction result and the sample video; Based on the loss value, the initial generation model is modified until a target video generation model is obtained.

2. The method according to claim 1, wherein The method further comprises: Performing posture estimation on the sample video to determine a second posture signal sequence corresponding to the sample video; The second posture signal sequence is subjected to random masking processing to obtain the first posture signal sequence.

3. The method according to claim 1, wherein The preprocessing of the sample video to obtain the first gesture signal sequence includes any one of the following: Performing posture estimation on a second image other than the first image in the sample video to determine the first posture signal sequence, wherein the first image is determined based on a random number; Performing posture estimation on a partial area of ​​a first image and an entire area of ​​a second image in the sample video to determine the first posture signal sequence, wherein the first image is determined based on a random number.

4. The method according to claim 1, wherein The method further comprises: determining a first random probability; When the first random probability is greater than a first threshold, obtaining a last frame image in the sample video; The last reference image frame in the reference image sequence is replaced by the last frame image to obtain an updated reference image sequence.

5. The method according to claim 1, wherein The method further comprises: Determining a background sequence corresponding to the sample video; The background sequence is used to replace other reference images in the reference image sequence except the first reference image to obtain an updated reference image sequence.

6. The method according to claim 5, wherein: The determining of the background sequence corresponding to the sample video includes: determining a mask sequence based on the third posture signal sequence; Based on the mask sequence, mask processing is performed on the sample video to obtain the background sequence.

7. The method according to claim 5, wherein: The step of replacing the reference images except the first reference image in the reference image sequence with the background sequence to obtain an updated reference image sequence includes: Determine second probabilities currently corresponding to the other reference images respectively; The corresponding background images in the background sequence are used to replace the corresponding reference images whose second probability is greater than the second threshold, so as to obtain the updated reference image sequence.

8. The method of claim 1, wherein: The step of inputting the first posture signal sequence and the reference image sequence into an initial generation model to obtain a prediction result output by the initial generation model includes: Determining description information corresponding to the sample video; The description information, the first posture signal sequence and the reference image sequence are respectively input into an initial generation model to obtain a prediction result output by the initial generation model.

9. The method according to any one of claims 1 to 8, wherein: The step of inputting the first posture signal sequence and the reference image sequence into an initial generation model to obtain a prediction result output by the initial generation model includes: Inputting the first posture signal sequence and the reference image sequence into the encoding network in the initial generation model to obtain a feature vector; After the feature vector is fused with the true value of the noise corresponding to the sample video, the feature vector is input into the diffusion network in the initial generation model to obtain the predicted noise output by the diffusion network.

10. The method of claim 9, wherein: The determining of the loss value based on the prediction result and the sample video includes: Determining a loss value based on the predicted noise and the true value of the noise; The step of modifying the initial generation model based on the loss value until a target video generation model is obtained includes: Based on the loss value, the diffusion network is modified until the target video generation model is obtained.

11. A video generation method, comprising: determining a target image and a first posture signal sequence; parsing the first posture signal sequence to determine abnormal feature points in the first posture signal sequence; Masking the abnormal feature points to obtain a second posture signal sequence; The second posture signal sequence and the target image are input into the trained video generation model to obtain the target video.

12. The method of claim 11, wherein: Determining the target image includes any one of the following: When the target video is the first sub-video, the previous image frame corresponding to the insertion position in the video insertion instruction is determined as the target image; In a case where the target video is not the first sub-video, determining the last image frame in the previous sub-video adjacent to the target video as the target image; In the case that the target video is the last sub-video, the last image frame in the previous sub-video adjacent to the target video and the next image frame corresponding to the insertion position are determined as the target image.

13. A device for acquiring a video generation model, comprising: A first processing module is configured to pre-process the sample video to obtain a first gesture signal sequence and a reference image sequence, wherein a matching degree between at least one gesture signal in the first gesture signal sequence and a corresponding gesture signal in the sample video is less than a matching degree threshold, and a first reference image in the reference image sequence is a first frame image in the sample video; a first generation module, configured to input the first posture signal sequence and the reference image sequence into an initial generation model respectively, and obtain a prediction result output by the initial generation model; A first determining module, configured to determine a loss value based on the prediction result and the sample video; A correction module is used to correct the initial generation model based on the loss value until a target video generation model is obtained.

14. A video generating device, comprising: A second determination module is used to determine the target image and the first posture signal sequence; a third determining module, configured to parse the first posture signal sequence and determine abnormal feature points in the first posture signal sequence; A second processing module is used to perform mask processing on the abnormal feature points to obtain a second posture signal sequence; The second generation module is used to input the second posture signal sequence and the target image into the trained video generation model to obtain the target video.

15. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for obtaining a video generation model described in any one of claims 1-10, or the video generation method described in any one of claims 11-12.

16. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: in, The computer instructions are used to enable the computer to execute the method for obtaining a video generation model according to any one of claims 1 to 10, or the video generation method according to any one of claims 11 to 12.

17. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the method for obtaining a video generation model according to any one of claims 1 to 10, or the steps of the video generation method according to any one of claims 11 to 12.

Citation Information

Cited By

  • Action migration method and device, electronic equipment, computer readable storage medium and program product

    CN121354212A