A video generation method, system and model

By dynamically adjusting the sampling frame rate and adaptively adjusting the motion generation density according to the motion amplitude, the problems of motion distortion, blurred details, stuttering, or screen flickering in video generation are solved, improving the smoothness and realism of the generated video.

CN122073636BActive Publication Date: 2026-07-24HITHINK ROYALFLUSH INFORMATION NETWORK CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HITHINK ROYALFLUSH INFORMATION NETWORK CO LTD
Filing Date
2026-04-22
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing technologies struggle to handle inconsistent frame rates for different ranges of motion when generating videos of character movements, resulting in issues such as motion distortion, blurred details, stuttering, or screen flickering.

Method used

By acquiring reference images and action sequences, motion evaluation metrics representing the actions are determined, and the sampling frame rate is dynamically adjusted to adapt to the amplitude of the actions, thereby generating the target video.

Benefits of technology

It achieves adaptive adjustment of motion generation density based on motion amplitude, improving the smoothness and realism of video motion while also taking into account generation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122073636B_ABST
    Figure CN122073636B_ABST
Patent Text Reader

Abstract

The present application relates to the field of video generation, and particularly relates to a video generation method, system and model. The method comprises: obtaining a reference image and a reference action sequence, the reference action sequence comprising action representations corresponding to different time points; determining motion evaluation indexes of the action representations corresponding to different time points based on the reference action sequence; determining a sampling frame rate of different time points based on the motion evaluation indexes corresponding to different time points; and generating a target video based on the sampling frame rate, the reference image and the reference video. According to the motion evaluation indexes of the reference action sequence, the sampling frame rate is dynamically determined, and the target video is generated based on the sampling frame rate, the reference image and the reference video, so that the density of action generation can be adaptively adjusted according to the action amplitude, the action constraint in the violent motion area is more fine, and the reasonable calculation amount in the gentle motion area is maintained, so that the action fluency and authenticity of the generated video are improved, and the generation efficiency is considered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of video generation, and in particular to a video generation method, system, and model. Background Technology

[0002] Character motion video generation is an important research direction in the field of computer vision and artificial intelligence content generation. Its goal is to generate videos that retain the appearance features of the reference image and accurately reproduce the actions in the reference video, based on a given reference image (such as a person, cartoon character, or pet) and a reference video containing motion information. For example, it could make a character (such as a human) in the reference image perform actions in the reference video. However, in practical applications, the original frame rates of reference videos vary, and the amplitude of motion within the video is not uniformly distributed. Traditional methods use a fixed sampling rate to uniformly sample all motion segments, which easily leads to computational redundancy in slow-motion segments and loss of information in fast-motion segments. Furthermore, in high-speed motion scenes, the reference signal is relatively sparse (e.g., the difference in motion between adjacent frames is too large), making it difficult for video generation models to achieve accurate interpolation, easily resulting in problems such as motion distortion, blurred details, stuttering, or screen flickering.

[0003] Therefore, there is an urgent need for a video generation method, system, and model that uses different frame rates for different action segments to solve problems such as motion distortion, blurred details, stuttering, or screen flickering in the generated videos. Summary of the Invention

[0004] One embodiment of this specification provides a video generation method, the method comprising: acquiring a reference image and a reference action sequence, the reference action sequence including action representations corresponding to different time points; determining a motion evaluation index for the action representations corresponding to the different time points based on the reference action sequence; determining a sampling frame rate for the different time points based on the motion evaluation index; and generating a target video based on the sampling frame rate, the reference image, and the reference action sequence.

[0005] One embodiment of this specification provides a video generation system, the system comprising: an acquisition module configured to acquire a reference image and a reference action sequence, the reference action sequence including action representations corresponding to different time points; a first determination module configured to determine a motion evaluation index of the action representations corresponding to the different time points based on the reference action sequence; a second determination module configured to determine a sampling frame rate at the different time points based on the motion evaluation index; and a generation module configured to generate a target video based on the sampling frame rate, the reference image, and the reference action sequence.

[0006] One embodiment of this specification provides a video generation model, which includes an encoding module, a video generation module, a decoding module, and a post-processing module. The encoding module is configured to encode an intermediate action sequence to generate an encoded action sequence. The video generation module is configured to generate a video latent variable sequence based on the encoded action sequence and a reference image. The decoding module is configured to process the video latent variable sequence to generate an intermediate video. The post-processing module is configured to resample the intermediate video based on the sampling frame rate to generate a target video.

[0007] This invention dynamically determines the sampling frame rate based on the motion evaluation index of the reference action sequence, and generates a target video based on the sampling frame rate, reference image, and reference video. It can adaptively adjust the density of motion generation according to the motion amplitude, so that the region of violent motion can obtain more refined motion constraints, and the region of smooth motion can maintain a reasonable amount of computation. It solves problems such as motion distortion, blurred details, stuttering, or screen flickering in the generated video, thereby improving the smoothness and realism of the generated video motion while taking into account the generation efficiency. Attached Figure Description

[0008] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein:

[0009] Figure 1 These are schematic diagrams illustrating application scenarios of the video generation system according to some embodiments of this specification; Figure 2 This is an exemplary flowchart of a video generation method according to some embodiments of this specification; Figure 3 This is an exemplary flowchart illustrating the generation of a reference action sequence according to some embodiments of this specification; Figure 4 This is an exemplary flowchart illustrating the determination of motion evaluation metrics representing actions according to some embodiments of this specification; Figure 5 This is an exemplary flowchart illustrating the generation of a target video according to some embodiments of this specification; Figure 6 This is an exemplary flowchart of generating a target video according to other embodiments of this specification; Figure 7 This is an exemplary flowchart illustrating the acquisition of a first input according to some embodiments of this specification; Figure 8 This is an exemplary flowchart illustrating the acquisition of a second input according to some embodiments of this specification; Figure 9This is an exemplary block diagram of a video generation system according to some embodiments of this specification; Figure 10 This is a schematic diagram of a video generation model according to some embodiments of this specification. Detailed Implementation

[0010] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.

[0011] It should be understood that the terms “system,” “device,” “unit,” and / or “module” used herein are one way to distinguish different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.

[0012] As indicated in this specification and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of expressly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.

[0013] Flowcharts are used in this specification to illustrate the operations performed by the system according to embodiments of this specification. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the steps can be processed in reverse order or simultaneously. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.

[0014] Figure 1 This is a schematic diagram illustrating an application scenario of a video generation system according to some embodiments of this specification.

[0015] In some embodiments, in application scenario 100 of the video generation system, the video generation method can be implemented by carrying out the methods and / or processes disclosed in this specification.

[0016] like Figure 1 As shown, the application scenario 100 of the video generation system involved in the embodiments of this specification includes reference images and reference videos 110, processor 120, terminal 130, storage device 140 and network 150.

[0017] Reference images are static image assets used to provide baseline visual information such as the appearance, identity, style, and texture of the main character performing actions in the target video. A character refers to the object performing actions in the generated target video. Reference images include images of characters capable of performing actions, such as people, cartoon characters, and pets. The target video is the final output video generated by the video generation system, in which the character from the reference images fully performs the corresponding actions in the reference video.

[0018] Reference video is raw video footage containing complete and continuous motion information, used to extract reference motion sequences. The reference video encapsulates the actions the character in the target video needs to perform.

[0019] For more information on the reference images and videos, please refer to [link / reference]. Figure 2 .

[0020] Processor 120 can process data and / or information acquired from terminal 130, storage device 140, network 150, and / or other components of application scenario 100, such as video generation system. For example, processor 120 can acquire reference images and reference videos 110 and analyze and process them.

[0021] In some embodiments, processor 120 may be a single server or a group of servers. The server group may be centralized or distributed. In some embodiments, processor 120 may be local or remote. For example, processor 120 may access information and / or data from terminal 130 and / or storage device 140 via network 150. As another example, processor 120 may be directly connected to terminal 130 and / or storage device 140 to access information and / or data (e.g., processor 120 may acquire reference images and reference videos 110 from terminal 130 and / or storage device 140). In some embodiments, processor 120 may be implemented on a cloud platform. For example, the cloud platform may include private cloud, public cloud, hybrid cloud, community cloud, distributed cloud, inter-cloud cloud, multi-cloud, etc., or any combination thereof.

[0022] Terminal 130 can provide functional components related to user interaction and enable user interaction functions (such as providing or displaying information and data to the user). By way of example only, terminal 130 can be one or any combination of other devices with input and / or output functions, such as mobile devices, tablet computers, laptop computers, and desktop computers. By way of example only, output functions include, but are not limited to, one or more combinations of sound output such as voice, display screen, haptic transmission such as vibration, and electromagnetic wave signals such as light. By way of example only, input functions can include, but are not limited to, one or more combinations of keyboard input, touchscreen input, voice input, motion event input such as device tilt / shake / rotation / swing, and electromagnetic wave signal input such as light.

[0023] In some embodiments, a user can input information and / or data through terminal 130, and can also acquire information and / or data through terminal 130. For example, a user can input a reference image and / or reference video through terminal 130. As another example, a user can acquire a target video through terminal 130.

[0024] Storage device 140 may store data, instructions, and / or any other information. In some embodiments, storage device 140 may store data (e.g., reference images and / or reference videos, etc.) acquired from processor 120, terminal 130, and / or other sources. In some embodiments, storage device 140 may store data and / or instructions used by processor 120 to perform or use in order to complete the exemplary methods described herein.

[0025] In some embodiments, storage device 140 may include one or more storage components, each of which may be a separate device or part of another device. In some embodiments, storage device 140 may include random access memory (RAM), read-only memory (ROM), mass storage, removable memory, volatile read-write memory, and any combination thereof. In some embodiments, storage device 140 may be implemented on a cloud platform.

[0026] Network 150 may include any suitable network capable of facilitating information and / or data exchange. In some embodiments, at least one component of application scenario 100 of the video generation system (e.g., processor 120, terminal 130, storage device 140) may exchange information and / or data with at least one other component of application scenario 100 of the video generation system via network 150. For example, processor 120 may send target video to terminal 130 via network 150.

[0027] It should be noted that the application scenario 100 of the video generation system is provided for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can make various modifications or variations based on the description in this specification. For example, the application scenario 100 of the video generation system can achieve similar or different functions on other devices. However, these changes and modifications will not depart from the scope of this specification.

[0028] Figure 2 This is an exemplary flowchart of a video generation method according to some embodiments of this specification. In some embodiments, process 200 may be executed by a processing device (e.g., processor 120). Figure 2 As shown, process 200 may include the following steps.

[0029] Step 210: Obtain the reference image and reference action sequence.

[0030] A reference action sequence is a sequence formed by arranging the actions of a character at different points in time in a reference video in chronological order.

[0031] In some embodiments, the reference action sequence includes action representations at different time points.

[0032] A time point refers to the instantaneous moment corresponding to a single action representation in a reference action sequence. One time point corresponds to one frame of the sampled reference video. Specifically, a reference action sequence is a sequence formed by arranging multiple action representations in chronological order of the actual actions. Each independent action representation corresponds to a specific instantaneous moment in the execution of the actual action, which is a time point.

[0033] Motion representation is data used to characterize a character's posture at a given point in time. It is the fundamental action unit for motion amplitude evaluation and video generation. Motion representation includes the two-dimensional coordinate information of the character's key skeletal points. In some embodiments, motion representation can be obtained by recognizing a single frame of reference video image using a pre-trained human keypoint detection network.

[0034] In some embodiments, the reference image can be obtained through user input. For example, a user can input a reference image through terminal 130, and processor 120 can obtain the reference image from terminal 130. In some embodiments, the reference image can also be stored in storage device 140, and processor 120 can obtain the reference image from storage device 140 through an interface.

[0035] In some embodiments, the processor 120 can directly acquire a set of pre-extracted, time-ordered action representations of characters from a reference video as a reference action sequence. Each action representation in the reference action sequence corresponds to a unique time point.

[0036] In some embodiments, the processor 120 may sample the reference video based on the target output frame rate of the target video to obtain a reference action sequence. Further details on obtaining the reference action sequence can be found in [link to relevant documentation]. Figure 3 .

[0037] Step 220: Based on the reference action sequence, determine the motion evaluation index corresponding to the action representation at different time points.

[0038] Exercise evaluation indicators are numerical metrics used to quantify the intensity of movement at a specific time point within a reference movement sequence. Higher exercise evaluation indicator values ​​indicate more intense movement changes and greater amplitude at that time point. Exercise evaluation indicators for multiple time points are arranged chronologically to form an exercise scoring sequence.

[0039] In some embodiments, the processor 120 can perform frame-by-frame motion amplitude evaluation and quantification on the reference action sequence to generate motion evaluation metrics for each time point.

[0040] In some embodiments, the processor 120 can base its motion score on the motion representations corresponding to reference time points at different time points. In some embodiments, the processor 120 can perform motion amplitude evaluation on the motion representations at any time point to obtain an initial motion score sequence. In some embodiments, the processor 120 can determine a target motion score sequence and a motion evaluation index based on the initial motion score sequence. Further explanation regarding the determination of the motion evaluation index can be found in [link to relevant documentation]. Figure 4 .

[0041] Step 230: Determine the sampling frame rate at different time points based on the motion evaluation index corresponding to different time points.

[0042] The sampling frame rate is a control parameter used to control the sampling density of a reference action sequence in the time dimension. In some embodiments, the sampling frame rate is positively correlated with a motion evaluation metric. For example, the higher the motion evaluation metric value, the higher the sampling frame rate (i.e., the denser the sampling).

[0043] In some embodiments, there is a preset correspondence between motion evaluation metrics and sampling frame rates, and the processor 120 can determine the sampling frame rate at different time points based on the preset correspondence between motion evaluation metrics and sampling frame rates.

[0044] In some embodiments, the processor 120 can determine the target motion amplitude level corresponding to different time points based on motion evaluation indicators; and determine the target repetition factor corresponding to different time points based on the target motion amplitude level and the first mapping relationship.

[0045] The target motion amplitude level refers to the discrete levels into which motion amplitude is divided based on the magnitude of motion evaluation indicators. For example, target motion amplitude levels include low amplitude, medium amplitude, high amplitude, and extremely high amplitude.

[0046] The first mapping relationship represents the correspondence between the preset target motion amplitude level and the repetition factor. For example, low amplitude corresponds to repetition factor 0, medium amplitude corresponds to repetition factor 1, high amplitude corresponds to repetition factor 2, and very high amplitude corresponds to repetition factor 3, etc.

[0047] The target repetition factor is a numerical value used to control the number of times an action representation is copied. In some embodiments, the target repetition factor can reflect the sampling frame rate, for example, the number of times the action representation in a reference action sequence is copied. For example, if the target repetition factor of a frame in the reference action sequence is 2, it means that the action representation corresponding to that frame needs to be copied twice. In some embodiments, the sampling frame rate does not modify the timestamp of the sampled image, but rather makes the action information contained in a unit of time more dense by copying the corresponding frames.

[0048] In some embodiments, the processor 120 may acquire a second mapping relationship and determine the target motion amplitude level based on the motion evaluation index and the second mapping relationship.

[0049] The second mapping relationship represents the correspondence between preset target motion amplitude levels and motion evaluation indicators. For example, processor 120 can preset multiple thresholds, such as a first threshold, a second threshold, and a third threshold. When the motion evaluation indicator is less than the first threshold, the target motion amplitude level is low. When the motion evaluation indicator is greater than the first threshold but less than the second threshold, the target motion amplitude level is medium. When the motion evaluation indicator is greater than the second threshold but less than the third threshold, the target motion amplitude level is high. When the motion evaluation indicator is greater than the third threshold, the target motion amplitude level is extremely high, and so on.

[0050] In some embodiments, the first mapping relationship and the second mapping relationship may be a threshold table, a function, a configuration file, etc.

[0051] In some embodiments, the processor 120 can also directly establish a one-to-one correspondence between motion evaluation indicators and sampling frame rates through preset functional relationships, mapping tables, etc., in order to determine the sampling frame rate.

[0052] In some embodiments, the processor 120 can also analyze motion scores using audio (such as speech or background music). For example, a repetition factor can be determined based on the volume or rhythm of the audio (such as decibels or beats). For instance, the sampling frame rate can be increased during the climax of a song to make lip movements and head movements more nuanced.

[0053] In some embodiments, the processor 120 may also combine a large language model to parse text instructions (such as “make a slow Tai Chi starting posture, then punch quickly”) into expected motion score curves, thereby generating a semantically consistent sampling strategy.

[0054] In some embodiments, the processor 120 may also determine the repetition factor based on a resampling model.

[0055] Resampling models are lightweight, learnable temporal prediction modules. For example, resampling models include lightweight networks such as gated recurrent units (GRUs) or multilayer perceptrons (MLPs).

[0056] In some embodiments, the processor 120 can input the motion scoring sequence into the resampling model. The resampling model outputs a repetition factor for each frame. The resampling model can be jointly trained end-to-end with the video generation model, using video generation quality as the core optimization objective to construct a loss function. This allows the model to automatically learn temporal sampling strategies adapted to different action scenarios, achieving adaptive and precise control of the speed of action sequences. Further details regarding the video generation model can be found in [link to relevant documentation]. Figures 6-9 .

[0057] In some embodiments of this specification, by mapping motion evaluation indicators to corresponding motion amplitude levels and determining corresponding repetition factors based on motion amplitude levels, the sampling frame rate can be adaptively determined in a hierarchical quantization manner. This enables precise matching between sampling density and motion amplitude, and the control logic is simple and reliable. This is beneficial for improving the efficiency and robustness of motion amplitude judgment, while ensuring the accuracy and stability of motion sequence sampling density control.

[0058] Step 240: Generate the target video based on the sampling frame rate, reference image, and reference action sequence.

[0059] In some embodiments, the processor 120 may resample the reference video based on the sampling frame rate to generate a resampled reference action sequence, and generate a target video based on the characters in the reference image and the resampled reference action sequence.

[0060] In some embodiments, the processor 120 may further process the reference action sequence based on the sampling frame rate to obtain an intermediate action sequence; generate an intermediate video using a video generation model based on the intermediate action sequence and the reference image; and resample the intermediate video based on the sampling frame rate to generate a target video. Further details on generating the target video can be found in [link to relevant documentation]. Figure 5 .

[0061] In some embodiments of this specification, by dynamically determining the sampling frame rate based on the motion evaluation index of the reference action sequence, and generating the target video based on the sampling frame rate, the reference image and the reference video, the density of motion generation can be adaptively adjusted according to the motion amplitude, so that the region of violent motion can obtain more refined motion constraints, and the region of smooth motion can maintain a reasonable amount of computation, thereby improving the smoothness and realism of the generated video motion while taking into account the generation efficiency.

[0062] Figure 3 This is an exemplary flowchart illustrating the generation of a reference sequence of actions according to some embodiments of this specification. In some embodiments, process 300 may be executed by a processing device (e.g., processor 120). Figure 3 As shown, process 300 may include the following steps.

[0063] Step 310: Obtain the reference video.

[0064] Reference videos consist of continuous video footage of a character performing actions. For example, reference videos can include complete and continuous motion information, serving as raw data for extracting motion representations and generating reference motion sequences. Reference videos include subjects such as people and animals capable of performing corresponding actions, and the motion flow of the reference videos can serve as the complete motion that a character in a reference image in the target video needs to perform.

[0065] In some embodiments, the reference video can be obtained through user input. For example, a user can input a reference video through terminal 130, and processor 120 can obtain the reference video from terminal 130. In some embodiments, the reference video can also be stored in storage device 140, and processor 120 can obtain the reference video from storage device 140 by calling an interface.

[0066] Step 320: Sample the reference video based on the target output frame rate of the target video to generate a reference image frame sequence.

[0067] The target output frame rate refers to the preset frame rate of the final target video. The target output frame rate can be user-defined or set by the system default, such as 24 frames per second, 30 frames per second, or 60 frames per second.

[0068] A reference image frame sequence is a collection of consecutive static image frames obtained by resampling a reference video based on the target output frame rate and arranging them in chronological order. Each frame in the reference image frame sequence corresponds to a time point of the target output frame rate.

[0069] In some embodiments, resampling the reference video based on the target output frame rate can align the frame rates of the reference video and the target video, eliminating the deviation caused by the difference in the original frame rates of different reference videos.

[0070] In some embodiments, the processor 120 may resample the reference video based on the target output frame rate. If the frame rate of the reference video is higher than the target output frame rate, the processor 120 downsamples the reference video and removes redundant frames. If the frame rate of the original reference video is lower than the target output frame rate, the processor 120 upsamples the reference video to interpolate frames. After the processor 120 resamples the reference video, a final reference image frame sequence is generated that perfectly matches the target output frame rate and is arranged at equal intervals in chronological order.

[0071] In some embodiments, each frame in the resampled reference image frame sequence corresponds to a time point of the target output frame rate. The number of frames in the resampled reference image frame sequence matches the preset duration and target output frame rate of the target video.

[0072] Step 330: Extract action representations from each reference image frame in the reference image frame sequence to generate a reference action sequence.

[0073] In some embodiments, the processor 120 can extract key points frame by frame from the reference image frame sequence to generate a reference action sequence.

[0074] For example, processor 120 can employ a pre-trained human keypoint detection network as a tool for action representation extraction. The human keypoint detection network can identify the skeletal keypoints of a character in the input image and output the corresponding two-dimensional coordinate information. Processor 120 can also input each reference image frame in the reference image frame sequence frame by frame into the human keypoint detection network, extracting the corresponding two-dimensional coordinates of the character's skeletal keypoints for each frame to generate the action representation corresponding to that frame. Processor 120 can also arrange the action representations corresponding to all reference image frames according to the chronological order of the reference image frame sequence to generate a complete reference action sequence. Each action representation in the reference action sequence corresponds to a specific time point, which corresponds one-to-one with the resampled reference image frame.

[0075] In some embodiments of this specification, by resampling the reference video at the target output frame rate of the target video and extracting the action sequence, reference videos from different sources and with different frame rates can be uniformly aligned to the target output frame rate, avoiding action timing disorder caused by frame rate mismatch, and ensuring the timing consistency and accuracy of subsequent motion evaluation and video generation.

[0076] Figure 4 This is an exemplary flowchart illustrating the determination of motion evaluation metrics representing actions according to some embodiments of this specification. In some embodiments, process 400 may be executed by a processing device (e.g., processor 120). Figure 4 As shown, process 400 may include the following steps.

[0077] Step 410: Based on the action representations corresponding to reference time points at different time points, perform motion amplitude evaluation on the action representations at any time point to obtain an initial motion score sequence.

[0078] A reference time point is a baseline time point compared to the time point where motion amplitude evaluation needs to be performed. The reference time point can be any time point (such as the time point corresponding to the first frame) or the previous adjacent time point of the time point where motion amplitude evaluation needs to be performed. For example, if the time point for motion amplitude evaluation is frame 5, the corresponding reference time point could be frame 1 or frame 4 (the previous adjacent time point).

[0079] Any point in time refers to the point in the reference action sequence where the amplitude of motion needs to be evaluated and the corresponding motion evaluation index needs to be calculated.

[0080] Motion amplitude assessment refers to the process of quantifying the intensity of motion changes at any given time point by comparing the differences in motion representation between any given time point and a reference time point.

[0081] The initial motion score sequence refers to the sequence formed by arranging the initial motion scores corresponding to all time points in the reference motion sequence in chronological order. For example, calculating the initial motion score frame by frame for a 10-frame motion sequence yields the initial motion score sequence [0.1, 0.3, 0.5, 0.9, 1.2, 0.8, 0.4, 0.2, 0.1, 0.1].

[0082] In some embodiments, the initial motion score sequence includes initial motion scores at different time points.

[0083] Initial motion score refers to the raw, unprocessed motion amplitude value calculated by directly comparing the motion representation at any time point with that at a reference time point.

[0084] In some embodiments, the action representation includes key points and the location information of the key points.

[0085] Key points refer to the core feature points used to characterize a character's movement and posture. In some embodiments, key points include the character's skeletal joints. For example, key points include skeletal joints such as the tip of the nose, left shoulder, right shoulder, left elbow, right elbow, left hip, right hip, left knee, and right knee.

[0086] The location information of a keypoint refers to its spatial coordinates within the corresponding image frame. The location information of a keypoint can be represented by two-dimensional pixel coordinates (x, y).

[0087] In some embodiments, the processor 120 can spatially align the motion representation corresponding to a reference time point and the motion representation at any time point. In some embodiments, the processor 120 can calculate the distance between keypoints at the reference time point and keypoints at any time point, as well as the angular change of the bone vector formed by the keypoints, based on the aligned motion representation. In some embodiments, the processor 120 can determine an initial motion score for the motion representation corresponding to any time point based on the distance and angular change, to form an initial motion score sequence.

[0088] Spatial alignment refers to the process of eliminating global spatial changes such as overall character displacement, scaling, and rotation between the action representation at any point in time and the reference point in time through spatial transformation, while retaining only the character's own action posture changes.

[0089] The distance between a key point at a reference time point and a key point at any other time point refers to the difference between the position of the same key point at any other time point and its position at the reference time point (such as Euclidean distance).

[0090] In some embodiments, the processor 120 can extract key points and position information from the action representations at a reference time point and at any time point, respectively, and select torso reference key points (such as left shoulder, right shoulder, left hip, right hip, etc.) as alignment points. In some embodiments, the processor 120 can calculate a spatial transformation matrix (such as an affine transformation matrix) between the two sets of action representations based on the reference key points. In some embodiments, the processor 120 can use the transformation matrix to transform the position information of all key points at any time point to complete the spatial alignment of the two sets of action representations. In some embodiments of this specification, by spatially aligning the action representations corresponding to the reference time point and the action representations at any time point, global spatial interference caused by the overall displacement, rotation, and scaling of the character can be eliminated, retaining only the character's own action posture changes.

[0091] A skeletal vector is a vector formed by two adjacent keypoints that are connected. Skeletal vectors are used to represent the spatial orientation and length of the corresponding bone. For example, the vector formed by the left shoulder keypoint and the left elbow keypoint is the skeletal vector of the left arm.

[0092] Angle change refers to the angle difference between the same bone vector at any given time point and a reference time point. Angle change is used to quantify the rotation amplitude of a bone. For example, the angle of the left arm bone vector is 25° in frame 4 and 30° in frame 5, corresponding to an angle change of 5°.

[0093] In some embodiments, the processor 120 can calculate the Euclidean distance between the position information of the same keypoint at two time points for each of the two aligned action representations, thereby obtaining displacement distance data for all keypoints. In some embodiments, the processor 120 can construct the bone vector corresponding to the arbitrary time point and the bone vector corresponding to the reference time point for each aligned action representation at any time point and the reference time point, based on adjacent keypoints with physiological connections. In some embodiments, the processor 120 can calculate the angular change between the bone vector corresponding to the reference time point and the bone vector corresponding to the reference time point for each of the same bone vectors, thereby obtaining rotation angle data for all bone vectors.

[0094] In some embodiments, the processor 120 can fuse the distances of all keypoints with the angular changes of the full skeleton vector to calculate the initial motion score for each time point at any given time. For example, the initial motion score for each time point can be determined based on the following formula (1): (1) in, This represents the initial motion score at a certain point in time. This represents the mean Euclidean distance of all key points; This represents the average angle change of all skeletal vectors. , These are the preset weighting coefficients.

[0095] In some embodiments, after completing the initial motion score calculation for all time points, the processor 120 can arrange all scores in chronological order to generate an initial motion score sequence.

[0096] In some embodiments, during motion amplitude evaluation, in addition to calculating motion evaluation indicators based on keypoint displacement and changes in bone vector angles, the processor 120 can also use the acceleration and jerk of keypoints as supplementary evaluation parameters. Jerk is an indicator that quantifies the acceleration changes of keypoints. For example, the processor 120 can also determine the motion acceleration and jerk at adjacent time points based on the aligned keypoint position information. When acceleration or jerk exceeds a preset physical reasonable threshold, the corresponding frame is determined to be a motion abrupt change frame. By adjusting the repetition factor at the corresponding time point of that frame to increase the sampling frame rate, interpolation smoothing is performed on the motion sequence to make the generated motion conform to the physical laws of human movement, effectively reducing screen flicker and motion jumps.

[0097] In some embodiments, the processor 120 can also sample different frame rates for different regions of the image using a region-adaptive sampling control scheme. For example, the processor 120 can perform semantic segmentation on image frames of a reference image and reference video using a pre-trained image segmentation model, dividing them into multiple regions such as hands, torso, and background. For each region, action representations are extracted and corresponding region motion evaluation metrics are calculated. Based on the region motion evaluation metrics of each region, corresponding region repetition factors and region sampling frame rates are matched. Furthermore, a region mask corresponding to each semantic region is introduced during the model input stage to achieve region-specific sampling and generation control. For instance, when the hand region is in a fast-moving state, the sampling frame rate or generation resolution of that region is increased to ensure the accuracy of fine motion generation. When the background region is stationary, the sampling frame rate of that region is decreased to reduce computational overhead and improve generation efficiency.

[0098] Step 420: Determine the target motion score sequence based on the initial motion score sequence. The motion evaluation index is the target motion score at different time points in the target motion score sequence.

[0099] The target motion score refers to the final motion amplitude value at the corresponding time point obtained after optimizing the initial motion score sequence.

[0100] A target motion score sequence is a sequence formed by arranging the target motion scores corresponding to all time points in a reference motion sequence in chronological order.

[0101] In some embodiments, the processor 120 may perform smoothing optimization on the initial motion scoring sequence to eliminate abrupt noise and outliers in single-frame scores. In some embodiments, the processor 120 may use the optimized initial motion scoring sequence as the target motion scoring sequence. In some embodiments, the processor 120 may use the target motion score corresponding to each time point in the target motion scoring sequence as the corresponding motion evaluation index.

[0102] In some embodiments, the processor 120 may determine the mean of the initial motion scores at any time point in the initial motion score sequence and at one or more adjacent time points. In some embodiments, the processor 120 may specify the mean as the target motion score at any time point.

[0103] For example, processor 120 can iterate through each time point in the initial motion score sequence and use the currently iterated time point as the time point to be processed. Processor 120 can preset a smoothing window size and select one or more time points adjacent to the time point to be processed (e.g., one frame before and after, two frames before and after, etc.). Processor 120 can determine the arithmetic mean of the initial motion scores of the time point to be processed and all selected adjacent time points to obtain the mean of the initial motion scores corresponding to the time point to be processed. Processor 120 can use the mean of the initial motion scores corresponding to the time point to be processed as the target motion score corresponding to the time point to be processed.

[0104] In some embodiments, the target motion scoring sequence can also be determined by a motion scoring assessment model.

[0105] Motion score evaluation models are machine learning models. For example, a motion score evaluation model may include a temporal encoder, a scoring regressor, and an optional smoothing post-processing layer. The temporal encoder is used to capture the temporal dependencies and motion features of the reference motion sequence, the scoring regressor is used to regress the initial motion score corresponding to each time point based on the encoded motion features, and the smoothing post-processing layer is used to perform temporal smoothing optimization on the initial motion score, outputting the final target motion score sequence.

[0106] In some embodiments, the motion score evaluation model can be obtained through training. The training samples for the motion score evaluation model include sample videos, and the labels include sample motion score sequences.

[0107] In some embodiments of this specification, an initial motion score sequence is obtained by performing spatial alignment, displacement, and rotation quantization calculations frame by frame on the action sequence. The final target motion score sequence is then obtained through mean smoothing of adjacent frames. This approach accurately and stably quantifies the magnitude of motion changes at each moment, effectively suppressing single-frame noise and abrupt changes, and eliminating interference from global character displacement on motion evaluation. It also ensures the accuracy and robustness of motion evaluation metrics, providing a reliable quantitative basis for adaptive adjustment of the subsequent sampling frame rate, ultimately improving the motion coherence and realism of the generated video.

[0108] Figure 5 This is an exemplary flowchart illustrating the generation of a target video according to some embodiments of this specification. In some embodiments, process 500 may be executed by a processing device (e.g., processor 120). Figure 5 As shown, process 500 may include the following steps.

[0109] Step 510: Process the reference action sequence based on the sampling frame rate to obtain the intermediate action sequence.

[0110] Intermediate motion sequences refer to motion sequences that are adapted to the input requirements of the video generation model after resampling and copying the reference motion sequence based on the sampling frame rate.

[0111] In some embodiments, the processor 120 can iterate through each time point in the reference action sequence and copy the action representation at each time point a corresponding number of times based on the sampling frame rate (i.e., repetition factor). In some embodiments, the processor 120 can arrange all the copied action representations in the original chronological order to generate an intermediate action sequence. For example, the original reference action sequence is 10 frames long, the repetition factor of the 3rd frame is 1, the repetition factor of the 5th frame is 2, the repetition factor of the 7th frame is 1, and the repetition factor of the remaining frames is 0. After copying, an intermediate action sequence with a length of 10+1+2+1=14 frames is obtained.

[0112] Step 520: Based on the intermediate action sequence and reference image, use a video generation model to generate an intermediate video.

[0113] A video generation model is a machine learning model used to generate videos of character actions based on reference images. In some embodiments, the video generation model includes one or more of the following: diffusion model, generative adversarial network, Transformer model, U-Net network, etc.

[0114] In some embodiments, the input to the video generation model includes an intermediate action sequence and a reference image, and the output includes an intermediate video.

[0115] In some embodiments, the video generation model can be obtained through training. The training samples for the video generation model include sample reference images, sample reference action sequences, and sample sampling frame rates. Labels include intermediate video clips corresponding to the training samples. The sample reference action sequences can be determined by extracting reference video sequences from the sample videos.

[0116] Intermediate video refers to the original generated video produced by the video generation model based on intermediate action sequences and reference images, without final resampling and deduplication. Intermediate video may include duplicate frames due to copying of action representations.

[0117] In some embodiments, the processor 120 can input intermediate action sequences and reference images into a video generation model, and the video generation model outputs an intermediate video. Further explanation of the video generation model and the generation of intermediate videos can be found in the relevant sections below, such as... Figures 6-8 .

[0118] Step 530: Resample the intermediate video based on the sampling frame rate to generate the target video.

[0119] In some embodiments, the processor 120 can determine the number of repeating frames at each time point based on the sampling frame rate. In some embodiments, the processor 120 can perform frame-by-frame filtering on the intermediate video to remove redundant repeating frames caused by motion representation duplication, so that the number of intermediate video frames is consistent with the length of the original reference motion sequence. In some embodiments, the processor 120 can align the deduplicated video sequence to a preset target output frame rate to generate a target video.

[0120] In some embodiments, the processor 120 can also directly process the sampling frame rate, reference image, and reference video using a target video generation model to generate a target video.

[0121] The target video generation model is a machine learning model. In some embodiments, the target video generation model includes one or more of the following: diffusion model, generative adversarial network, Transformer model, U-Net network, etc.

[0122] In some embodiments, the input to the target video generation model includes a sampling frame rate, a reference image, and a reference video. The output of the target video generation model includes the target video.

[0123] The target video generation model can be obtained through training. The training samples for the target video generation model include the sample sampling frame rate, sample reference images, and sample reference videos. The labels for the target video generation model include the sample target videos.

[0124] In some embodiments, the target video generation model and the video generation model can be the same machine learning model or different machine learning models.

[0125] In some embodiments, the processor 120 can also improve the quality of the target video by co-optimizing spatial resolution and temporal sampling rate. For example, the processor 120 can coordinately adjust the spatial resolution and temporal sampling rate of video generation, employing a generation strategy of high temporal sampling rate combined with low spatial resolution for regions of intense motion with high motion scores, and a generation strategy of low temporal sampling rate combined with high spatial resolution for regions of static or gentle motion with low motion scores. Finally, the generation results of different regions are fused and output through a super-resolution network and a frame interpolation network, further improving the overall efficiency and effect of video generation while balancing generation quality and computational overhead.

[0126] In some embodiments, the processor 120 can also optimize the video generation model based on a streaming processing architecture. For example, for online processing needs in real-time interactive generation scenarios such as virtual anchors, the processor 120 can replace the Diffusion Transformer (DiT) backbone network in the video generation model with a state-space model (such as the Mamba model) architecture. By leveraging the temporal modeling advantage of the linear complexity of the state-space model, it can achieve efficient processing of infinitely long action sequences. At the same time, it can combine a sliding window mechanism to complete the real-time prediction of sampling factors, thereby achieving low-latency, streaming real-time action video generation.

[0127] In some embodiments of this specification, a generation process is employed that uses sampling frame rate processing to obtain intermediate motion sequences, then generates intermediate videos based on the intermediate motion sequences and reference images, and finally resamples the intermediate videos to obtain the target video. This process can adaptively adjust the density of the motion sequence according to the motion amplitude, improving the generation detail in areas of large motion. Simultaneously, resampling ensures that the timing and duration of the target video meet expectations, improving the smoothness and controllability of video generation.

[0128] Figure 6 This is an exemplary flowchart illustrating the generation of a target video according to other embodiments of this specification. In some embodiments, process 600 may be executed by a processing device (e.g., processor 120) or a video generation model. In some embodiments, the video generation model includes an encoding module, a video generation module, a decoding module, and a post-processing module. Figure 6 As shown, process 600 may include the following steps.

[0129] Step 610: The intermediate action sequence is encoded by the encoding module to generate an encoded action sequence.

[0130] The encoding module is a machine learning module in a video generation model used to extract features and perform dimensionality reduction encoding on the input action sequence and reference image, transforming high-dimensional image / action data into low-dimensional latent features. In some embodiments, the encoding module includes the encoder portion of a pre-trained Variational Autoencoder (VAE).

[0131] Encoded action sequence refers to the action feature sequence obtained by encoding the intermediate action sequence frame by frame through the encoding module.

[0132] In some embodiments, the processor 120 can input the intermediate action sequence frame by frame into the encoding module, and the encoding module can output the encoded action sequence.

[0133] Step 620: Based on the encoded action sequence and the reference image, generate a video latent variable sequence through the video generation module.

[0134] The video generation module generates a sequence of latent variables for the video that meets the action requirements. In some embodiments, the video generation module includes a diffusion Transformer model. The diffusion Transformer (DiT) model is a diffusion model based on the Transformer architecture. The diffusion Transformer model segments image or video data into a sequence of patches and utilizes the global attention mechanism of the Transformer to process the sequence data, making it particularly suitable for large-scale video generation tasks.

[0135] A video latent variable sequence refers to a set of latent video features arranged in chronological order, generated by the video generation module based on coded action sequences and reference image features. The video latent variable sequence can be reconstructed into a complete video frame sequence using the decoding module.

[0136] In some embodiments, the processor 120 can input the encoded action sequence and reference image into the video generation module, and the video generation module can output a video latent variable sequence.

[0137] In some embodiments, the inputs to the video generation module include a first input and a second input, and the video generation module generates a sequence of latent video variables based on the first and second inputs. Further details regarding the first and second inputs can be found in [link to relevant documentation]. Figure 7 and Figure 8 .

[0138] Step 630: Input the video latent variable sequence into the decoding module for processing to generate intermediate video.

[0139] The decoding module is a machine learning module used to reconstruct the sequence of video latent variables output by the video generation module into a sequence of video pixel frames. In some embodiments, the decoding module includes a decoder portion of a pre-trained variational autoencoder.

[0140] In some embodiments, the processor 120 can input the sequence of video latent variables output by the video generation module into the decoding module group by group. In some embodiments, the decoding module restores the video latent variables into a sequence of video pixel frames. In some embodiments, the processor 120 can arrange all the restored video frames in chronological order to generate an intermediate video.

[0141] Step 640: The intermediate video is resampled based on the sampling frame rate by the post-processing module to generate the target video.

[0142] The post-processing module is a functional module used to perform post-processing operations such as frame rate alignment, deduplication, and image quality optimization on intermediate videos.

[0143] In some embodiments, the post-processing module can obtain the sampling frame rate and repetition factor corresponding to each time point.

[0144] In some embodiments, the post-processing module can directly resample the intermediate video based on the sampling frame rate to obtain the target video.

[0145] In some embodiments, the post-processing module can perform deduplication and resampling on the intermediate video based on a repetition factor to remove redundant duplicate frames. In some embodiments, the post-processing module can align the processed video to the target output frame rate to generate the target video.

[0146] In some embodiments of this specification, a video generation model architecture consisting of an encoding module, a video generation module, a decoding module, and a post-processing module is adopted. This architecture can efficiently encode action sequences and reference images into latent features and generate videos. Then, the target video is output through decoding and post-processing. This makes the model structure clear and the processing flow standardized, reducing the amount of computation while improving the stability and quality of video generation.

[0147] Figure 7 This is an exemplary flowchart illustrating the acquisition of a first input according to some embodiments of this specification. In some embodiments, process 700 may be executed by a processing device (e.g., processor 120). Figure 7 As shown, process 700 may include the following steps.

[0148] Step 710: Encode each action representation in the intermediate action sequence to generate multiple initial action latent variables corresponding to the multiple action representations in the intermediate action sequence.

[0149] Initial action latent variables refer to the latent features of a single frame of action obtained after encoding individual action representations in an intermediate action sequence. Initial action latent variables are the smallest unit constituting the encoded action sequence. They can also be called the encoded action sequence, action feature map, etc.

[0150] In some embodiments, the processor 120 can input each action representation in the intermediate action sequence frame by frame into the encoding module. In some embodiments, the encoding module performs dimensionality reduction encoding on each action representation to generate initial action latent variables corresponding to each action representation, ultimately obtaining T initial action latent variables. The size of each initial action latent variable is Cp×H'×W', where Cp is the number of output channels of the encoding module, and H' and W' are the height and width of the encoded feature map, respectively.

[0151] Step 720: Group the multiple initial action latent variables according to time to obtain multiple groups of initial action latent variables.

[0152] In some embodiments, the number of sets of initial action latent variables is determined based on the temporal downsampling rate of the encoding module that encodes each action representation in the intermediate action sequence.

[0153] For example, the processor 120 obtains the time downsampling rate n of the encoding module, and divides the T initial action latent variables into consecutive groups of n variables each, according to the time sequence, resulting in a total of T / n groups of initial action latent variables. The time downsampling rate n of the encoding module can be preset; for example, n can be 4. For example, with an intermediate action sequence length T = 16 and a time downsampling rate n = 4, the processor 120 divides the 16 initial action latent variables into 4 groups, each containing 4 consecutive initial action latent variables.

[0154] Step 730: Arrange the initial action latent variables in each of the multiple initial action latent variables into a preset channel dimension by temporal pixel rearrangement to obtain action latent variables.

[0155] Temporal pixel reordering refers to the process of concatenating initial action latent variables from multiple frames within the same group along the channel dimension, rearranging multi-frame information from the time dimension to the channel dimension, and achieving lossless encoding of action information. In some embodiments, temporal pixel reordering can be an extension of the classic pixel reordering (spatial compression) operation in the time dimension. For example, temporal pixel reordering includes concatenating a set of 4 initial action latent variables, each with Cp channels, along the channel dimension to form a feature map with 4 × Cp channels.

[0156] The preset channel dimension refers to the channel dimension that the action latent variables need to match after temporal pixel rearrangement.

[0157] In some embodiments, the preset channel dimension is the same as the channel dimension of the noise latent variable.

[0158] The preset channel dimension is used to adapt to the channel splicing requirements of the noise latent variable. For example, if the number of channels of the noise latent variable is C, the preset channel dimension is set to be consistent with C, so that the splicing operation can be performed.

[0159] Noise latent variables are random latent variables that conform to a preset distribution (e.g., a standard normal distribution). Noise latent variables are used to determine the visual appearance, texture, detail, and random variations of each frame of the generated video.

[0160] For example, a noise latent variable can be represented as a standard normally distributed random tensor of size T / n×C×H'×W', where T is the length of the intermediate action sequence, n is the time downsampling rate, and C is the number of channels.

[0161] Motion latent variables are feature variables obtained by performing temporal pixel rearrangement on each group of initial motion latent variables. Motion latent variables are used to capture global motion information in a video sequence. For example, performing temporal pixel rearrangement on the initial motion latent variables grouped every four frames results in a feature map of size (n×Cp)×H'×W', which is the motion latent variable. Here, n is the temporal downsampling rate, and Cp is the number of channels for the initial motion latent variable in a single frame.

[0162] In some embodiments, the time dimension of the action latent variable is consistent with the time dimension of the noise latent variable. The time dimension of the action latent variable is the initial number of groups of the action latent variable, with a value of T / n. In some embodiments, before channel splicing, the processor 120 can set the time dimension of the noise latent variable to make it consistent with the time dimension of the action latent variable. For example, if the time dimension of the action latent variable is 4, the time dimension of the noise latent variable should also be set to 4 to ensure that the lengths of the two variables are completely matched in the time dimension, allowing the channel splicing operation to be performed normally.

[0163] Step 740: The motion latent variable and the noise latent variable are concatenated to obtain the first input to the video generation module.

[0164] Channel concatenation refers to the operation of merging two or more features along the channel dimension. Channel concatenation can fuse motion condition information with basic noise information to form the first input of the video generation module. For example, processor 120 can concatenate a noise latent variable with C channels and a motion latent variable with n×Cp channels along the channel dimension to form a feature variable with C+n×Cp channels.

[0165] The first input refers to the action condition input for the video generation module. It is obtained by concatenating the action latent variables and noise latent variables through multiple channels, and is used to provide action generation constraints for video generation.

[0166] For example, the feature tensor with dimensions T / n×(C+n×Cp)×H'×W' after splicing is the first input to the video generation module.

[0167] In some embodiments of this specification, the first input to the video generation module is constructed by encoding action representations, temporal grouping, temporal pixel rearrangement, and concatenating with latent noise variables. This fully preserves the temporal features of the action sequence and enhances the expressive power of action information in the feature space. This allows the video generation module to learn temporal action changes more accurately, thereby improving the matching accuracy and temporal consistency between the generated video and the generated actions.

[0168] Figure 8 This is an exemplary flowchart illustrating the acquisition of a second input according to some embodiments of this specification. In some embodiments, process 800 may be executed by a processing device (e.g., processor 120). Figure 8 As shown, process 800 may include the following steps.

[0169] Step 810: Encode the reference image to obtain the reference image feature map.

[0170] A reference image feature map is a visual feature map obtained after encoding a reference image using an encoding module. The reference image feature map includes visual information such as the appearance, identity, style, and texture of the characters in the reference image. For more information on reference images and characters, please refer to [link to relevant documentation]. Figure 1 .

[0171] For example, the processor 120 can input a reference image into the encoding module. The encoding module performs feature extraction and dimensionality reduction encoding on the reference image, and outputs a reference image feature map that includes the appearance, identity, style, texture information, etc. of the reference image. In some embodiments, the size of the reference image feature map is Cp×H'×W'.

[0172] Step 820: Segment the reference image feature map into multiple non-overlapping image blocks.

[0173] An image patch refers to a local feature region obtained by segmenting the feature map of a reference image.

[0174] In some embodiments, the processor 120 can uniformly divide the reference image feature map into multiple non-overlapping image blocks according to a preset segmentation size. For example, if the reference image feature map is 64×64 in size and the preset segmentation size is 2×2, the processor 120 can divide the reference image feature map into 1024 non-overlapping 2×2 image blocks.

[0175] Step 830: Project multiple image blocks to obtain the feature vector of the reference image.

[0176] Projection refers to the process of converting segmented image patches into feature vectors of fixed dimensions through a linear mapping network. For example, processor 120 can project a 2×2 image patch into a feature vector of dimension 768 through a linear layer.

[0177] A reference image feature vector is a set of continuous vectors obtained by projecting image patches segmented from a reference image feature map. Each reference image feature vector encodes local visual features, location information, semantic information, and style features of the reference image.

[0178] In some embodiments, the processor 120 may input each segmented image patch into a pre-trained linear projection layer. In some embodiments, the linear projection layer maps the local features of each image patch into a feature vector of fixed dimension. In some embodiments, the processor 120 may arrange the feature vectors corresponding to all image patches in spatial order to obtain a reference image feature vector.

[0179] Step 840: Concatenate the feature vector of the reference image with the video feature vector of the reference video corresponding to the reference action sequence to obtain the second input of the video generation model.

[0180] A video feature vector is a feature vector representation of a video sequence obtained by encoding, segmenting, and projecting a reference video frame sequence corresponding to an intermediate action sequence. The video feature vector encodes both temporal and visual information of the video sequence.

[0181] The second input refers to the input used in the video generation module to determine the character's appearance and style. The character performing the action is the character corresponding to the reference image. The second input includes features related to visual information such as the character's appearance, identity, style, and texture in the reference image. The second input is obtained by concatenating the reference image feature vector and the video feature vector, and is used to provide visual identity constraints for the character during video generation, ensuring that the character in the final generated video remains consistent with the reference image.

[0182] In some embodiments, the processor 120 can acquire a video frame sequence of a reference video corresponding to a reference action sequence, and encode, segment, and project the video frame sequence to obtain a video feature vector. In some embodiments, the processor 120 can concatenate a reference image feature vector to the end of the video feature vector sequence to obtain a second input to the video generation model. In some embodiments, the processor 120 can input the second input into the video generation module, so that each video frame in the reference video can be associated with a reference image feature vector, thereby ensuring that the appearance of the generated character is consistent with the reference image.

[0183] In some embodiments, the processor 120 may also expand the input channels of the video generation module based on the temporal downsampling rate of the encoding module and the number of output channels of the encoding module.

[0184] In some embodiments, before inference of the video generation model, the processor 120 can determine the number of expanded channels as n×Cp based on the temporal downsampling rate n and the number of output channels Cp of the encoding module, and expand the number of input channels C of the video generation module to C+n×Cp. The original channels before expansion are initialized with pre-trained model weights, while the newly added channels are initialized with zero to avoid disrupting the original generation capability of the pre-trained model. For example, when the original number of input channels C=4, the temporal downsampling rate n=4, and the number of output channels Cp=4, the number of input channels can be expanded to 4+4×4=20, where the first 4 channels retain pre-trained weights, and the last 16 channels are initialized with zero.

[0185] In some embodiments of this specification, the input channels of the video generation module are expanded based on the temporal downsampling rate and the number of output channels of the encoding module. Differential weight initialization methods are used for the original and newly added channels. This not only ensures that the latent variables of the actions match the input dimensions of the video generation module, guaranteeing that action conditions can be stably and effectively injected into the video generation process, but also maximizes the reuse of the original weights of the pre-trained model, preserving the model's existing general video generation capabilities. Furthermore, by using zero initialization for the newly added channels, interference with the original performance of the pre-trained model is avoided, significantly reducing the difficulty of model adaptation and the cost of fine-tuning, thus improving the compatibility and engineering practicality of the solution.

[0186] Figure 9 This is an exemplary block diagram of a video generation system according to some embodiments of this specification. In some embodiments, such as Figure 9 As shown, the video generation system 900 includes an acquisition module 910, a first determination module 920, a second determination module 930, and a generation module 940.

[0187] The acquisition module 910 is configured to acquire a reference image and a reference action sequence. The reference action sequence includes action representations at different time points.

[0188] In some embodiments, the acquisition module 910 is further configured to acquire a reference video; sample the reference video based on the target output frame rate of the target video to generate a reference image frame sequence; and extract the action representation from each reference image frame in the reference image frame sequence to generate the reference action sequence.

[0189] The first determining module 920 is configured to determine the motion evaluation index of the action representation corresponding to the different time points based on the reference action sequence.

[0190] In some embodiments, the first determining module 920 is further configured to, for the action representation at any time point among the different time points, perform motion amplitude evaluation on the action representation at the arbitrary time point based on the action representation corresponding to the reference time point among the different time points to obtain an initial motion scoring sequence, the initial motion scoring sequence including the initial motion scores corresponding to the different time points; and determine a target motion scoring sequence based on the initial motion scoring sequence, the motion evaluation index being the target motion scores at the different time points in the target motion scoring sequence.

[0191] In some embodiments, the action representation includes key points and the location information of the key points.

[0192] In some embodiments, the first determining module 920 is further configured to spatially align the action representation corresponding to the reference time point and the action representation at the arbitrary time point; calculate the distance between the keypoint at the reference time point and the keypoint at the arbitrary time point, and the angular change of the skeletal vector formed by the keypoints, based on the aligned action representation; and determine an initial motion score for the action representation corresponding to the arbitrary time point based on the distance and the angular change, so as to form the initial motion score sequence.

[0193] In some embodiments, the first determining module 920 is further configured to determine the mean of the initial motion scores corresponding to any time point and one or more adjacent time points in the initial motion score sequence; and designate the mean as the target motion score corresponding to the arbitrary time point.

[0194] The second determining module 930 is configured to determine the sampling frame rate at different time points based on the motion evaluation index corresponding to the different time points.

[0195] In some embodiments, the second determining module 930 is further configured to determine the target motion amplitude level corresponding to the different time points based on the motion evaluation index; and to determine the target repetition factor corresponding to the different time points based on the target motion amplitude level and a first mapping relationship. The first mapping relationship represents the mapping relationship between the target motion amplitude level and the repetition factor; the target repetition factor represents the magnitude of the sampling frame rate and is used to indicate the number of times the motion representation in the reference motion sequence is copied.

[0196] In some embodiments, the second determining module 930 is further configured to obtain a second mapping relationship; the second mapping relationship represents the mapping relationship between the target motion amplitude level and the motion evaluation index; and to determine the target motion amplitude level based on the motion evaluation index and the second mapping relationship.

[0197] The generation module 940 is configured to generate a target video based on the sampling frame rate, the reference image, and the reference motion sequence.

[0198] In some embodiments, the generation module 940 is further configured to process the reference action sequence based on the sampling frame rate to obtain an intermediate action sequence; generate an intermediate video using a video generation model based on the intermediate action sequence and the reference image; and resample the intermediate video based on the sampling frame rate to generate the target video.

[0199] In some embodiments, the video generation model includes an encoding module, a video generation module, a decoding module, and a post-processing module. In some embodiments, the generation module 940 is further configured to: encode the intermediate action sequence using the encoding module to generate an encoded action sequence; generate a video latent variable sequence using the video generation module based on the encoded action sequence and the reference image; input the video latent variable sequence into the decoding module for processing to generate the intermediate video; and resample the intermediate video using the post-processing module based on the sampling frame rate to generate the target video.

[0200] In some embodiments, the input to the video generation module includes a first input. In some embodiments, the generation module 940 is further configured to encode each action representation in the intermediate action sequence to generate multiple initial action latent variables corresponding to the multiple action representations in the intermediate action sequence; group the multiple initial action latent variables according to time to obtain multiple sets of initial action latent variables; perform temporal pixel rearrangement of each set of initial action latent variables to a preset channel dimension to obtain action latent variables; and concatenate the action latent variables with noise latent variables to obtain the first input to the video generation module.

[0201] In some embodiments, the preset channel dimension is the same as the channel dimension of the noise latent variable.

[0202] In some embodiments, the number of the multiple sets of initial action latent variables is determined based on the temporal downsampling rate of the encoding module that encodes each action representation in the intermediate action sequence.

[0203] In some embodiments, the time dimension of the action latent variable is consistent with the time dimension of the noise latent variable.

[0204] In some embodiments, the input to the video generation module includes a second input. In some embodiments, the generation module 940 is further configured to encode the reference image to obtain a reference image feature map; segment the reference image feature map into multiple non-overlapping image blocks; project the multiple image blocks to obtain a reference image feature vector; and concatenate the reference image feature vector with the video feature vector of the reference video corresponding to the reference action sequence to obtain the second input to the video generation model.

[0205] Figure 10 This is a schematic diagram of a video generation model according to some embodiments of this specification. In some embodiments, such as Figure 10 As shown, the video generation model 1000 includes an encoding module 1010, a video generation module 1020, a decoding module 1030, and a post-processing module 1040.

[0206] The encoding module 1010 is configured to encode the intermediate action sequence to generate an encoded action sequence.

[0207] The video generation module 1020 is configured to generate a video latent variable sequence based on the encoded action sequence and the reference image.

[0208] The decoding module 1030 is configured to process the video latent variable sequence to generate an intermediate video.

[0209] The post-processing module 1040 is configured to resample the intermediate video based on the sampling frame rate to generate the target video.

[0210] The basic concepts have been described above. Obviously, for those skilled in the art, the detailed disclosure above is merely illustrative and does not constitute a limitation of this specification. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to this specification. Such modifications, improvements, and corrections are suggested in this specification and therefore remain within the spirit and scope of the exemplary embodiments described herein.

[0211] Furthermore, this specification uses specific terms to describe embodiments thereof. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Moreover, certain features, structures, or characteristics in one or more embodiments of this specification can be appropriately combined.

[0212] Finally, it should be understood that the embodiments described in this specification are merely illustrative of the principles of the embodiments described herein. Other variations may also fall within the scope of this specification. Therefore, alternative configurations of the embodiments described herein are intended to be illustrative rather than limiting, and should be considered consistent with the teachings of this specification. Accordingly, the embodiments described herein are not limited to those explicitly introduced and described herein.

Claims

1. A video generation method, characterized in that, The method includes: Acquire a reference image and a reference action sequence, wherein the reference action sequence includes action representations corresponding to different time points; Based on the reference action sequence, determine the motion evaluation index represented by the action at different time points; The sampling frame rate at each time point is determined based on the motion evaluation index corresponding to that time point; and... Generate a target video based on the sampling frame rate, the reference image, and the reference motion sequence; wherein, determining the sampling frame rate at different time points based on the motion evaluation index corresponding to the different time points includes: The target motion amplitude level corresponding to the different time points is determined based on the motion evaluation index; and... Based on the target motion amplitude level and the first mapping relationship, the target repetition factor corresponding to the different time points is determined; the first mapping relationship represents the mapping relationship between the target motion amplitude level and the repetition factor; the target repetition factor represents the size of the sampling frame rate and is used to indicate the number of times the motion representation in the reference motion sequence is copied.

2. The method as described in claim 1, characterized in that, The acquisition of the reference action sequence includes: Get the reference video; The reference video is sampled based on the target output frame rate of the target video to generate a reference image frame sequence; and... The motion representation is extracted from each reference image frame in the reference image frame sequence to generate the reference motion sequence.

3. The method as described in claim 1, characterized in that, The step of determining the motion evaluation index for the action representation at different time points based on the reference action sequence includes: For the action representation at any point in time among the different time points, Based on the action representations corresponding to reference time points at different time points, motion amplitude is evaluated on the action representations at any given time point to obtain an initial motion score sequence, wherein the initial motion score sequence includes the initial motion scores corresponding to the different time points; and... The target motion score sequence is determined based on the initial motion score sequence, and the motion evaluation index is the target motion score at different time points in the target motion score sequence.

4. The method as described in claim 3, characterized in that, The motion representation includes key points and their location information. The step of evaluating the motion amplitude of the motion representation at any given time point to obtain an initial motion scoring sequence, based on the motion representation corresponding to a reference time point in the different time points, includes: Spatial alignment is performed on the action representation corresponding to the reference time point and the action representation at any time point. Based on the aligned motion representation, calculate the distance between the keypoint at the reference time point and the keypoint at any time point, as well as the angular change of the bone vector formed by the keypoints; and, Based on the distance and the angle change, an initial motion score is determined for the action representation corresponding to the arbitrary time point, so as to form the initial motion score sequence.

5. The method as described in claim 3, characterized in that, Determining the target motion score sequence based on the initial motion score sequence includes: Determine the mean of the initial motion scores for any time point and one or more adjacent time points in the initial motion score sequence; and, The mean is designated as the target motion score corresponding to the arbitrary time point.

6. The method as described in claim 1, characterized in that, Determining the target motion amplitude level corresponding to different time points based on the motion evaluation index includes: Obtain a second mapping relationship; the second mapping relationship represents the mapping relationship between the target motion amplitude level and the motion evaluation index; and, The target motion amplitude level is determined based on the motion evaluation index and the second mapping relationship.

7. The method as described in claim 1, characterized in that, The process of generating the target video based on the sampling frame rate, the reference image, and the reference action sequence includes: The reference action sequence is processed based on the sampling frame rate to obtain an intermediate action sequence; Based on the intermediate action sequence and the reference image, an intermediate video is generated using a video generation model; and, The intermediate video is resampled based on the sampling frame rate to generate the target video.

8. The method as described in claim 7, characterized in that, The video generation model includes an encoding module, a video generation module, a decoding module, and a post-processing module; the generation of the target video based on the sampling frame rate, the reference image, and the reference action sequence includes: The intermediate action sequence is encoded by the encoding module to generate an encoded action sequence; Based on the encoded action sequence and the reference image, a video latent variable sequence is generated by the video generation module; The video latent variable sequence is input into the decoding module for processing to generate the intermediate video; and... The post-processing module resamples the intermediate video based on the sampling frame rate to generate the target video.

9. The method as described in claim 7, characterized in that, The input to the video generation module includes a first input, and obtaining the first input includes: Encode each action representation in the intermediate action sequence to generate multiple initial action latent variables corresponding to the multiple action representations in the intermediate action sequence; The multiple initial action latent variables are grouped according to time to obtain multiple groups of initial action latent variables; Each of the multiple sets of initial action latent variables is time-series pixel rearranged to a preset channel dimension to obtain action latent variables; and... The action latent variable and the noise latent variable are concatenated to obtain the first input of the video generation module.

10. The method as described in claim 9, characterized in that, The preset channel dimension is the same as the channel dimension of the noise latent variable.

11. The method as described in claim 9, characterized in that, The number of sets of initial action latent variables is determined based on the temporal downsampling rate of the encoding module that encodes each action representation in the intermediate action sequence.

12. The method as described in claim 9, characterized in that, The time dimension of the action latent variable is consistent with the time dimension of the noise latent variable.

13. The method as described in claim 7, characterized in that, The input to the video generation module includes a second input, and obtaining the second input includes: The reference image is encoded to obtain a reference image feature map; The reference image feature map is segmented into multiple non-overlapping image blocks; Projecting the plurality of image patches yields a reference image feature vector; and, The reference image feature vector is concatenated with the video feature vector of the reference video corresponding to the reference action sequence to obtain the second input of the video generation model.

14. The method as described in claim 8, characterized in that, The method further includes: The input channels of the video generation module are expanded based on the temporal downsampling rate of the encoding module and the number of output channels of the encoding module.

15. A video generation system, characterized in that, The system includes: The acquisition module is configured to acquire a reference image and a reference action sequence, wherein the reference action sequence includes action representations corresponding to different time points; The first determining module is configured to determine the motion evaluation index of the action representation at different time points based on the reference action sequence; The second determining module is configured to determine the sampling frame rate at different time points based on the motion evaluation index corresponding to the different time points; and The generation module is configured to generate a target video based on the sampling frame rate, the reference image, and the reference action sequence; wherein, the second determining module is further configured to: The target motion amplitude level corresponding to the different time points is determined based on the motion evaluation index; and... Based on the target motion amplitude level and the first mapping relationship, the target repetition factor corresponding to the different time points is determined; the first mapping relationship represents the mapping relationship between the target motion amplitude level and the repetition factor; the target repetition factor represents the size of the sampling frame rate and is used to indicate the number of times the motion representation in the reference motion sequence is copied.

16. A video generation model, characterized in that, The video generation model includes an encoding module, a video generation module, a decoding module, and a post-processing module; wherein, The encoding module is configured to encode the intermediate action sequence to generate an encoded action sequence. The intermediate action sequence is obtained by processing a reference action sequence based on the sampling frame rate. The reference action sequence includes action representations corresponding to different time points. The video generation module is configured to generate a video latent variable sequence based on the encoded action sequence and the reference image; The decoding module is configured to process the video latent variable sequence to generate an intermediate video; and, The post-processing module is configured to resample the intermediate video based on the sampling frame rate to generate the target video; wherein the sampling frame rate is determined based on the following method: Based on the reference action sequence, determine the motion evaluation index represented by the action at different time points; The target motion amplitude level corresponding to the different time points is determined based on the motion evaluation index; and... Based on the target motion amplitude level and the first mapping relationship, the target repetition factor corresponding to the different time points is determined; the first mapping relationship represents the mapping relationship between the target motion amplitude level and the repetition factor; the target repetition factor represents the size of the sampling frame rate and is used to indicate the number of times the motion representation in the reference motion sequence is copied.