Video generation method and device based on action coherence, equipment and medium

By performing key frame analysis and inter-frame residual vector processing on the video action material set, the diffusion model is optimized to solve the problem of insufficient motion continuity in video generation, achieving high-quality, natural and smooth video generation.

CN120655792APending Publication Date: 2025-09-16PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510763034.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing video generation models have poor motion coherence between frames, resulting in a decrease in video viewing quality. This is especially true in advertising content that requires high coherence and narrative, affecting the accuracy and trust of information conveyed.

Method used

By acquiring a set of video action materials, performing key frame analysis, extracting inter-frame residual vectors, adjusting the initial diffusion model, performing cosine scaling and noise video generation, and combining Bayesian expectation denoising, the diffusion model is optimized to enhance motion coherence.

Benefits of technology

It significantly improves the motion continuity and viewing quality of the video, improves the clarity and restoration effect of the video image, enhances the image resolution and detail expression, and the generated video motion is natural and smooth.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655792A_ABST
    Figure CN120655792A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data analysis, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses a video generation method, device, equipment and medium based on action coherence, the method comprises the following steps: obtaining a video action material set, carrying out key frame analysis on the video action material set to obtain a key video frame sequence, extracting an inter-frame residual vector between consecutive frames in the key video frame sequence, adjusting a preset initial diffusion model by using the inter-frame residual vector to obtain an optimized diffusion model, performing cosine scaling on the inter-frame residual vector by using the optimized diffusion model to obtain a scaled residual vector, and obtaining a video adjustment text, and carrying out noise addition and splicing on the key video frame sequence by using the video adjustment text and the zoom residual vector to obtain a noise video, and carrying out noise reduction on the noise video to obtain a target action video. According to the method and the device, the action coherence in the customized generated video can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data analysis technology, and in particular to a method, device, equipment and medium for generating a video based on motion continuity. Background Art

[0002] AI-powered customized video generation technology is rapidly developing, particularly in areas like advertising and marketing where rapid content updates and personalized content are crucial. Diffusion models, as a mainstream generation paradigm, have achieved significant breakthroughs in image generation and are gradually expanding into text-to-video tasks. However, current mainstream video diffusion models often utilize a structure that combines traditional convolutional neural networks (CNNs) with attention mechanisms. While these models possess considerable generation capabilities, they still suffer from significant shortcomings in handling video continuity.

[0003] In the healthcare field, for example, in scenarios such as personalized health education, preoperative education, and rehabilitation guidance, medical institutions can use text-driven video generation capabilities to quickly generate customized video content based on different diseases, treatment plans, or patient populations, helping patients to more intuitively understand the diagnosis and treatment process and precautions. However, the current video diffusion model's lack of inter-frame coherence may lead to jerky movements or scene jumps during the explanation process, affecting the accurate communication of medical information and the viewing experience.

[0004] In the fintech sector, for example, in scenarios like robo-advisory, financial product promotion, and risk education, financial institutions can quickly generate personalized promotional videos tailored to different customer groups by simply inputting simple product information or user profiles. Especially for complex financial products, generated animated videos can more effectively explain the return structure and risk profiles. However, due to the inconsistency between frames in the current diffusion model during video generation, this can affect the authority and trustworthiness of financial content.

[0005] In summary, during the video generation process, the action connection between frames is unnatural, the background information is incoherent, and problems such as frame skipping, jitter, or scene confusion are prone to occur. This lack of action continuity greatly affects the viewing quality of the video, and is especially unsuitable for advertising content that requires high coherence and narrative.

[0006] Therefore, the current technology has the problem of poor motion continuity in customized generated videos. Summary of the Invention

[0007] The present invention provides a method, apparatus, device and medium for generating a video based on motion continuity, the main purpose of which is to solve the problem of poor motion continuity in customized generated videos.

[0008] In a first aspect, to achieve the above-mentioned objectives, the present invention provides a method for generating a video based on motion continuity, comprising:

[0009] Acquire a video action material set, perform key frame analysis on the video action material set, and obtain a key video frame sequence;

[0010] Extracting inter-frame residual vectors between consecutive frames from the key video frame sequence;

[0011] Using the inter-frame residual vector to adjust the preset initial diffusion model to obtain an optimized diffusion model;

[0012] performing cosine scaling on the inter-frame residual vector using the optimized diffusion model to obtain a scaled residual vector;

[0013] Obtaining a video adjustment text, and using the video adjustment text and the scaled residual vector to add noise to the key video frame sequence and then splice the key video frame sequence to obtain a noisy video;

[0014] The noise video is denoised and restored to obtain a target action video.

[0015] In a second aspect, the present invention further provides a video generation device based on motion continuity, comprising:

[0016] A key frame acquisition module is used to acquire a video action material set, perform key frame analysis on the video action material set, and obtain a key video frame sequence;

[0017] A residual vector extraction module, configured to extract inter-frame residual vectors between consecutive frames from the key video frame sequence;

[0018] a diffusion model adjustment module, configured to adjust a preset initial diffusion model using the inter-frame residual vector to obtain an optimized diffusion model;

[0019] a residual vector scaling module, configured to perform cosine scaling on the inter-frame residual vector using the optimized diffusion model to obtain a scaled residual vector;

[0020] A video frame noise adding module is used to obtain a video adjustment text, and use the video adjustment text and the scaled residual vector to add noise to the key video frame sequence and splice it to obtain a noisy video;

[0021] The noise video denoising module is used to denoise and restore the noise video to obtain the target action video.

[0022] In a third aspect, the present invention further provides an electronic device, comprising:

[0023] at least one processor; and,

[0024] a memory communicatively connected to the at least one processor; wherein,

[0025] The memory stores a computer program that can be executed by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the above-mentioned video generation method based on motion continuity.

[0026] In a fourth aspect, the present invention also provides a computer-readable storage medium, in which at least one computer program is stored. The at least one computer program is executed by a processor in an electronic device to implement the above-mentioned method for generating video based on motion continuity.

[0027] The present invention obtains a set of video action materials, performs key frame analysis on the video action material set, obtains a key video frame sequence, and combines resolution standardization, optical flow analysis and motion intensity screening to not only improve the accuracy and robustness of key frame extraction, but also effectively compress the amount of video information, retain the frame content reflecting the main action features, extracts inter-frame residual vectors between consecutive frames from the key video frame sequence, effectively captures the slight changes and motion features between adjacent action frames in the video, and uses grayscale differences to highlight inter-frame dynamic information while maintaining time continuity. The inter-frame residual vectors are used to adjust the preset initial diffusion model to obtain an optimized diffusion model, which significantly enhances the diffusion model's perception and expression capabilities of the video action temporal dynamics. force, uses the optimized diffusion model to perform cosine scaling on the inter-frame residual vector to obtain a scaled residual vector, which can effectively enhance the motion information of the key time point while suppressing irrelevant or noise parts, and obtain video adjustment text. The video adjustment text and the scaled residual vector are used to add noise to the key video frame sequence and splice it to obtain a noisy video, which realizes the deep fusion and synergy of semantic drive and motion features, denoises and restores the noisy video, and uses Bayesian expectation denoising to achieve high-precision denoising processing, which significantly improves the clarity and restoration effect of the video image, and obtains the target action video by upsampling the transition frames to enhance the image resolution and detail expression, which can effectively improve the action continuity in the customized generated video. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0029] Figure 1A schematic diagram of an application environment of a method for generating video based on motion continuity according to an embodiment of the present invention;

[0030] Figure 2 A flowchart of a method for generating video based on motion continuity provided by one embodiment of the present invention;

[0031] Figure 3 A schematic flow chart of a key video frame noise adding process in a video generation method based on motion continuity provided by one embodiment of the present invention;

[0032] Figure 4 A schematic diagram of modules of a video generation device based on motion continuity provided by one embodiment of the present invention;

[0033] Figure 5 A schematic structural diagram of an electronic device for implementing a method for generating a video based on motion continuity provided by an embodiment of the present invention;

[0034] Figure 6 Another structural diagram of an electronic device for implementing a method for generating video based on motion continuity provided by an embodiment of the present invention.

[0035] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0036] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, and to fully understand and implement how the present disclosure applies technical means to solve technical problems and achieve the corresponding technical effects, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. The embodiments of the present disclosure and the various features in the embodiments can be combined with each other without conflict, and the technical solutions formed are all within the scope of protection of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work should fall within the scope of protection of the present disclosure.

[0037] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, apparatus, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0038] The embodiment of the present application provides a method for generating video based on motion continuity, and the execution subject of the method for generating video based on motion continuity includes but is not limited to at least one of the electronic devices such as a server and a terminal that can be configured to execute the device provided by the embodiment of the present application. In other words, the method for generating video based on motion continuity can be executed by software or hardware installed on a terminal device or a server device. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0039] The embodiment of the present invention provides a video generation method based on action continuity, which can be applied in the following situations: Figure 1application environment. Among them, the client communicates with the server through the network. The server can obtain the video action material set through the client, perform key frame analysis on the video action material set, and obtain a key video frame sequence. Combined with resolution standardization, optical flow analysis and motion intensity screening, it not only improves the accuracy and robustness of key frame extraction, but also effectively compresses the amount of video information, retains the frame content reflecting the main action features, and extracts the inter-frame residual vectors between consecutive frames from the key video frame sequence, effectively capturing the slight changes and motion features between adjacent action frames in the video. On the basis of maintaining time continuity, the grayscale difference is used to highlight the inter-frame dynamic information, and the inter-frame residual vector is used to adjust the preset initial diffusion model to obtain an optimized diffusion model, which significantly enhances the diffusion model's perception and expression capabilities of the video action timing dynamics. The diffusion model performs cosine scaling on the inter-frame residual vector to obtain a scaled residual vector, which can effectively enhance the motion information of the key time point while suppressing irrelevant or noise parts, obtain video adjustment text, and use the video adjustment text and the scaled residual vector to add noise to the key video frame sequence and splice it to obtain a noisy video, realizing the deep fusion and synergy of semantic drive and motion features. The noisy video is denoised and restored, and high-precision denoising processing is achieved using Bayesian expectation denoising, significantly improving the clarity and restoration effect of the video image. By upsampling the transition frame, the image resolution and detail expression are enhanced to obtain the target action video, which can effectively improve the action continuity in the customized generated video, and finally the target action video output is fed back to the user client. Among them, the client can be but is not limited to various personal computers, laptops, smart phones, tablets and portable wearable devices. The server can be implemented with an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.

[0040] The following is an explanation of the description of the present invention. The present invention enhances inter-frame continuity by interpolating the intermediate frames of the noisy video, extracts noise residuals frame by frame in combination with the optimized diffusion model, and uses Bayesian expectation denoising to achieve high-precision denoising processing, thereby significantly improving the clarity and restoration effect of the video image. By upsampling the transition frames, the image resolution and detail expression are enhanced, and finally the updated denoised frame images are orderly spliced ​​to generate a target action video with coherent pictures, natural movements, and higher quality.

[0041] Reference Figure 2 FIG. 1 is a flow chart of a method for generating a video based on motion continuity according to an embodiment of the present invention. In this embodiment, the method for generating a video based on motion continuity includes:

[0042] S1. Obtain a video action material set, perform key frame analysis on the video action material set, and obtain a key video frame sequence.

[0043] In an embodiment of the present invention, typical advertising behavior clips are extracted from the acquired video action material set, such as common marketing actions such as product display, pointing operations, and character transitions. Key frame analysis is performed on these videos, and key video frames with representative changes are screened out using methods such as temporal difference, motion intensity detection, or inter-frame similarity calculation. Ultimately, a video key frame sequence that can be used for training is formed, laying the foundation for motion learning and generation of subsequent models.

[0044] In specific healthcare scenarios, this can be applied to generate animated videos about science education or diagnosis and treatment processes. By performing keyframe analysis on video footage of doctors explaining, instrument operation, and patient diagnosis and treatment, important action nodes (such as injections, testing, and signing informed consent forms) are extracted to generate a clearly structured and logically coherent keyframe sequence. This provides data support for the subsequent generation of standardized health education videos, facilitating improved patient understanding and treatment efficiency.

[0045] In specific FinTech scenarios, this can be used to quickly generate personalized product promotion videos. By extracting keyframes from existing videos promoting banking and insurance products, key action sequences such as "customer entry consultation," "product display," and "contract signing" are extracted. This provides data support for the subsequent generation of video ads tailored to the needs of specific target groups, combining user profiles with text prompts, significantly improving marketing efficiency and automation.

[0046] In an embodiment of the present invention, performing key frame analysis on the video action material set to obtain a key video frame sequence includes:

[0047] Standardizing the resolution of the video action material set to obtain a standard action video;

[0048] Dividing the standard action video into a plurality of continuous action video frames;

[0049] Randomly selecting two consecutive action video frames as a video frame group to be detected;

[0050] Performing optical flow analysis on the video frame groups to be detected one by one to obtain optical flow vectors;

[0051] Calculating the motion intensity of each of the to-be-detected video frame groups according to the optical flow vector, and screening out target video frame groups whose motion intensity is greater than a preset threshold;

[0052] The target video frame group is sorted according to a preset time sequence to obtain a key video frame sequence.

[0053] Specifically, a target resolution is determined, usually based on the highest or most commonly used resolution in the material set. Each video segment is processed using video processing software or programming tools (such as OpenCV), and the video resolution is adjusted to the target resolution through scaling operations. During the process, the frame rate of the video must be kept consistent to avoid affecting the smoothness of video playback due to resolution changes. The resulting standard action video will unify the resolution of all materials to facilitate subsequent comparison and analysis.

[0054] Standard action videos are extracted frame by frame. Video processing tools (such as OpenCV) are usually used to save each frame in the video as a separate image file. By traversing the frames in the video, a frame sequence is obtained. Two consecutive adjacent action video frames are randomly selected to form a video frame group to be detected. These selected frame groups can be used for further action recognition or other analysis tasks.

[0055] For each pair of adjacent video frames in the video frame group to be detected, the motion vector of the pixel point is calculated frame by frame using the optical flow method (such as the dense optical flow method or the sparse optical flow method) to obtain the optical flow vector field between frames. By establishing the brightness constant assumption, that is, assuming that the brightness value of a point remains unchanged during the motion, combined with the time and space gradient information, the optical flow vector of each pixel point is solved. The calculation formula is as follows:

[0056] I x *u+I y *v+I t =0

[0057] Among them, I x Represents the spatial gradient of the pixel point (x, y) in the x-axis direction, I y Represents the spatial gradient of the pixel point (x, y) in the y-axis direction, I t represents the gradient in the time direction, and (u,v) represents the optical flow vector.

[0058] After obtaining the optical flow vector of each video frame group to be detected, the corresponding motion intensity can be calculated based on the optical flow vector. For the optical flow vector (u, v) of each pixel point, the amplitude of the optical flow vector is calculated as the instantaneous motion intensity of the point:

[0059]

[0060] Among them, M(x,y) represents the instantaneous motion intensity of the pixel point (x,y), and (u,v) represents the optical flow vector.

[0061] The motion intensity of all pixels in the entire frame group is statistically summarized, and the average or maximum value is usually taken as the overall motion intensity of the frame group:

[0062]

[0063] Where S represents the motion intensity of the video frame group to be detected, N represents the total number of pixels, and M(x,y) represents the instantaneous motion intensity of the pixel (x,y). The motion intensity of all frame groups is compared with a preset threshold. Target video frame groups with a motion intensity greater than the threshold are screened out and are considered to contain significant motion information, which can be used for subsequent analysis or anomaly detection.

[0064] After selecting target video frame groups that meet the motion intensity threshold, these frame groups must be sorted according to a preset chronological order to ensure temporal consistency and scene coherence for subsequent analysis. Each target frame group is arranged from earliest to latest based on its time index or frame number in the original video, forming an ordered sequence of key video frames.

[0065] The keyframe analysis method described above can efficiently extract keyframe sequences with significant motion changes from a collection of video action footage. Combined with resolution normalization, optical flow analysis, and motion intensity screening, this method not only improves the accuracy and robustness of keyframe extraction, but also effectively compresses the amount of video information, preserving the frames that reflect the main motion characteristics. Chronologically sorted keyframe sequences can be used for tasks such as action recognition, video retrieval, and content summarization, significantly improving subsequent processing efficiency and analysis quality.

[0066] S2. Extracting inter-frame residual vectors between consecutive frames from the key video frame sequence.

[0067] In an embodiment of the present invention, after obtaining a key video frame sequence, an image difference method is used to perform pair-by-pair calculations on adjacent frames in the sequence, and the inter-frame residual vectors between each pair of consecutive frames are extracted, reflecting the motion change characteristics of the object in the time dimension, such as position offset, deformation amplitude, etc., providing accurate temporal motion information for subsequent motion modeling and diffusion generation processes.

[0068] In specific healthcare scenarios, extracting inter-frame residual vectors can be used to analyze and model the dynamics of video of diagnostic and treatment procedures. For example, when generating surgical instructional or nursing guidance videos, extracting inter-frame residual vectors between key frames of a doctor's actions (such as injections, suturing, and device operation) can accurately capture the details and rhythm of the movements.

[0069] In specific FinTech scenarios, inter-frame residual vector extraction can be used to model the rhythm of actions in typical financial service scenarios, such as advisor presentations and customer interactions. For example, when simulating banking or insurance explanation videos, by analyzing the residual vectors between keyframes (such as explanation gestures and page turning movements), the AI ​​model can learn the dynamic changes during the real-life introduction process.

[0070] In this embodiment of the present invention, extracting the inter-frame residual vector between consecutive frames from the key video frame sequence includes:

[0071] Converting each frame of video image in the key video frame sequence into a grayscale image;

[0072] Obtaining grayscale pixel values ​​corresponding to the grayscale images of two consecutive frames in the key video frame sequence, and performing a difference between the grayscale pixel values ​​of the two consecutive frames to obtain a residual image;

[0073] The residual map is converted into an inter-frame residual vector between two consecutive frames.

[0074] In detail, in order to simplify the subsequent feature extraction and calculation process, each frame of video image in the key video frame sequence can be converted into a grayscale image. For each frame of color image, its red, green and blue channels are weighted calculated according to the weighted average method to obtain the corresponding grayscale image, which retains the main structure and brightness information of the image, reduces the calculation dimension, and helps to improve the subsequent image feature extraction.

[0075] After converting each frame in the key video frame sequence into a grayscale image, the grayscale values ​​of all pixels in each frame are obtained. The corresponding grayscale pixel values ​​of any two consecutive frames in the sequence are subtracted point by point to generate an inter-frame residual map. This residual map reflects the local brightness changes between adjacent frames and can be used to capture motion details and dynamic areas. This two-dimensional residual map is flattened by rows or columns to convert it into a one-dimensional inter-frame residual vector, which provides input for subsequent feature analysis, clustering, or classification based on vector space.

[0076] By extracting inter-frame residual vectors between consecutive frames in a key video frame sequence, we can effectively capture subtle changes and motion features between adjacent action frames in the video. While maintaining temporal continuity, we leverage grayscale differences to highlight dynamic information between frames, eliminating static background and redundant content, significantly reducing data dimensionality and computational complexity. Furthermore, the inter-frame residual vector format facilitates subsequent similarity calculations, motion change analysis, and pattern recognition, helping to improve the efficiency and accuracy of video content understanding and motion recognition.

[0077] S3. Using the inter-frame residual vector, adjust the preset initial diffusion model to obtain an optimized diffusion model.

[0078] In an embodiment of the present invention, the inter-frame residual vector is injected into a preset initial diffusion model as motion feature guidance information. The temporal network layer of the initial diffusion model (such as the time dimension module in the Unet structure) is mainly fine-tuned to enable the model to more accurately perceive and reproduce the motion trajectory in the real video, thereby obtaining an optimized diffusion model with stronger motion coherence and higher generation quality, which is suitable for generating video sequences with natural motion transitions.

[0079] In specific healthcare scenarios, by injecting inter-frame residual vectors extracted from surgical procedures or nursing processes into the initial diffusion model and adjusting its temporal network structure, the model can more accurately capture and reproduce the rhythm and standardized process of medical actions. The optimized diffusion model can be used to automatically generate standardized medical operation videos, such as injection steps and surgical sterilization procedures, ensuring the coherence and accuracy of each action.

[0080] In specific FinTech scenarios, the inter-frame residual vectors from customer service or product introduction videos are used to optimize the diffusion model, enabling it to better reproduce dynamic processes such as promotional actions and customer interactions. Fine-tuning the model's temporal structure can enhance the naturalness and coherence of videos generated for bank counter services, insurance consultations, and mobile product introductions, thereby achieving more realistic and humanistic digital customer service or personalized financial product promotion content.

[0081] In the embodiment of the present invention, the adjusting the preset initial diffusion model by using the inter-frame residual vector to obtain the optimized diffusion model includes:

[0082] Performing convolution coding on the inter-frame residual vector to obtain residual vector coding;

[0083] Embedding the residual vector encoding into a temporal attention module of a preset initial diffusion model to obtain a motion perception module;

[0084] Performing residual vector analysis on the key video frame sequence using the initial diffusion model to obtain a model residual vector;

[0085] Constructing a residual alignment loss function according to the inter-frame residual vector and the model residual vector;

[0086] The initial diffusion model is adjusted using the motion perception module and the residual alignment loss function to obtain an optimized diffusion model.

[0087] In detail, in order to extract the local temporal features and change patterns in the inter-frame residual vector, a convolution operation can be performed on the vector. A one-dimensional convolutional neural network (1D-CNN) is used to apply multiple sliding convolution kernels to each residual vector. The local correlation features on the time axis are extracted by local weighted summation. After extraction through multiple convolutional layers, the residual vector encoding containing temporal motion information can be obtained, providing a more discriminative feature representation for subsequent classification, recognition or matching.

[0088] To enhance the diffusion model's ability to perceive motion in videos, the extracted residual vector encoding is embedded into the temporal attention module in the initial diffusion model to construct a motion perception module. The residual vector encoding serves as an additional input or guidance vector to the temporal attention module and is fused with the original time-step embeddings or keyframe features of the diffusion model to participate in the calculation of attention weights. This fusion process enables the model to focus on dynamically changing regions between frames during generation or reconstruction, thereby improving its ability to model temporal structure and motion evolution, achieving more accurate motion perception and content generation.

[0089] In order to constrain the model to accurately capture inter-frame motion features during the learning process, the model generates a residual vector between adjacent frames (i.e., a model residual vector) to reflect the model's prediction of motion changes. The residual vector between adjacent frames generated by the model is compared with the true residual vector calculated from the original key frame sequence to calculate the difference between them. The L2 norm is often used to define the residual alignment loss:

[0090]

[0091] Among them, R gen Represents the model residual vector, R true Represents the inter-frame residual vector.

[0092] Based on the temporal motion features extracted by the motion perception module and the constructed residual alignment loss function, a supervised optimization method is used to adjust the initial diffusion model. During training, the motion perception module is embedded in the model structure, and combined with the residual alignment loss function, the model is guided to learn the actual inter-frame dynamic changes in the key frame sequence. Through backpropagation, the model parameters are continuously updated so that the generated residual vectors more accurately reflect the motion information in the video. Ultimately, an optimized diffusion model is obtained that can more effectively capture temporal motion changes, improving the performance and robustness of video generation and motion analysis.

[0093] By convolutionally encoding the inter-frame residual vector and embedding it into the temporal attention module of the initial diffusion model, a motion perception module is constructed, and the model is adjusted in a targeted manner in combination with the residual alignment loss function. This significantly enhances the diffusion model's ability to perceive and express the temporal dynamics of video actions. It not only improves the model's accuracy in capturing subtle motion changes, but also effectively promotes the temporal consistency and motion coherence of the generated video, improves the overall performance and robustness of motion recognition and video generation, and provides more accurate and stable technical support for video processing in complex dynamic scenes.

[0094] S4. Perform cosine scaling on the inter-frame residual vector using the optimized diffusion model to obtain a scaled residual vector.

[0095] In an embodiment of the present invention, a cosine scaling mechanism is introduced in the optimized diffusion model to process the inter-frame residual vectors. Each group of residual vectors is multiplied by a scaling factor based on the cosine function, thereby compressing their amplitude changes, highlighting low-frequency, coherent motion features, suppressing high-frequency noise interference, and effectively smoothing the motion differences between frames, so that the generated video has more natural motion transitions and a more reasonable rhythm.

[0096] In specific healthcare scenarios, cosine scaling of the inter-frame residual vectors in medical procedure videos effectively smooths out motion variations, such as those encountered during injections, blood draws, or rehabilitation instruction. This effectively emphasizes a standard, coherent flow of movements and suppresses unwanted micro-movements or jitter. The optimized residual vectors enable the diffusion model to generate instructional videos of these procedures with more stable movements and a more natural rhythm, improving both teaching accuracy and the patient viewing experience.

[0097] In specific FinTech scenarios, cosine scaling of the inter-frame residual vectors in videos such as customer service and financial product introductions can filter out high-frequency, irrelevant movements (such as unnecessary gestures or meaningless facial movements) while retaining the rhythmic variations of the core explanation movements. This processing enables the model to generate clearer movements and more fluent communication in promotional videos such as bank counter demonstrations and insurance introductions, helping to enhance the professionalism and customer trust of digital promotions.

[0098] In an embodiment of the present invention, performing cosine scaling on the inter-frame residual vector using the optimized diffusion model to obtain a scaled residual vector includes:

[0099] Calculating a temporal weight factor of the inter-frame residual vector;

[0100] Using the optimized diffusion model to adjust the inter-frame residual vector to obtain a modified residual vector;

[0101] The modified residual vector is scaled using the timing weight factor to obtain a scaled residual vector.

[0102] Specifically, the temporal position of each residual vector in the sequence is encoded, and its relative importance on the timeline or motion intensity information is combined to assign different weights through an attention mechanism or weighting function. Using methods such as temporal position encoding, normalized motion amplitude, or self-attention-based weight calculation, each residual vector is assigned a weight factor that reflects its temporal contribution. This highlights the dynamic changes at key time points, helping the model to more accurately focus on and utilize motion information at critical moments.

[0103] The original residual vector is input into the model. Through its motion perception module and temporal attention mechanism, combined with the dynamic feature representations learned during training, the model filters and corrects noise and errors in the residual vector. Through multiple layers of nonlinear mapping and temporal information fusion, the model outputs a more accurate and consistent corrected residual vector that fully reflects the actual motion changes between video frames, improving the accuracy and stability of subsequent action analysis and video generation.

[0104] When scaling the corrected residual vector using the temporal weight factor, the weight is used as an amplification or suppression coefficient according to the temporal weight corresponding to each residual vector, and is multiplied element-by-element by each component of the corrected residual vector. This adjusts the importance of residual information at different time points, highlights the motion characteristics of key time periods, and weakens the influence of non-key frames, resulting in a more discriminative and expressive scaled residual vector, which helps to improve the model's perception of dynamic changes in time series and subsequent analysis effects.

[0105] By calculating the temporal weighting factor of the inter-frame residual vector, adjusting and correcting the residual vector using an optimized diffusion model, and performing cosine scaling on the corrected residual vector, the team can effectively enhance motion information at key time points while suppressing irrelevant or noisy components. This weighted scaling strategy not only improves the temporal expressiveness and discriminability of the residual vector, but also enhances the model's sensitivity and stability to motion changes, helping to improve the accuracy and robustness of video motion analysis and generation, ultimately boosting overall system performance.

[0106] S5. Obtain video adjustment text, and use the video adjustment text and the scaled residual vector to add noise to the key video frame sequence and then splice it to obtain a noisy video.

[0107] In an embodiment of the present invention, video adjustment text related to the target video content (such as product introductions, operating instructions, etc.) is used as content guidance information. The video adjustment text is combined with the scaled residual vector obtained in the previous step and acted together on the key video frame sequence. The noise addition mechanism in the diffusion model is used to inject motion noise and semantic changes of a specific intensity into each frame, guiding the model to simulate natural motion transitions and content evolution. The noisy frames are sequentially spliced ​​to obtain a structurally complete noise video containing semantic and motion features.

[0108] In specific healthcare scenarios, video adjustment text can include semantic instructions such as "Demonstrates electrocardiogram usage steps" or "Introduces the postoperative rehabilitation process." By combining this text with the scaled residual vector and adding noise to the key operation video frame sequence, we can simulate the natural variations and content logic of standard doctor movements. For example, the entire process from donning the device to reading data can be simulated, ultimately splicing together a noisy video that contains the rhythm of medical operations and semantic guidance, laying the foundation for the subsequent generation of coherent and medically compliant popular science or educational videos.

[0109] In specific FinTech scenarios, video-adjusted text, such as "Introducing High-Yield Wealth Management Products" or "Simulating Counter Service Processes," combined with the extracted scaled residual vectors, can be used to add noise to keyframe sequences, simulating the natural transitions and cadence of explanations in different customer service scenarios. The resulting noisy video, generated by splicing these noisy frames in time, captures the action logic and semantic intent of financial business scenarios, facilitating the subsequent generation of digital marketing videos that better align with real-world promotional scenarios, featuring consistent content and natural movements.

[0110] Figure 3 A flowchart of a key video frame noise adding process in a video generation method based on motion continuity provided by one embodiment of the present invention.

[0111] In an embodiment of the present invention, the step of adding noise to the key video frame sequence and splicing the key video frame sequence using the video adjustment text and the scaled residual vector to obtain a noisy video includes:

[0112] Performing text encoding on the video adjustment text to obtain a text semantic vector;

[0113] Fusing the scaled residual vector and the text semantic vector to obtain a noise control vector;

[0114] Extracting key frame image features of the key video frame sequence;

[0115] Adding temporal noise to the key frame image features using the noise control vector to obtain a noisy key frame;

[0116] The noisy key frames are spliced ​​together to obtain a noisy video.

[0117] In detail, the text is segmented and embedded through a pre-trained text encoding model (such as a Transformer-based language model) to extract its semantic feature representation. The model integrates the semantic information of each word in the text through a multi-layer self-attention mechanism and context modeling, and finally outputs a fixed-dimensional text semantic vector, which effectively captures the overall meaning and detailed information of the text, providing key semantic support for video content understanding, cross-modal alignment, and subsequent semantic-driven video processing.

[0118] When fusing the scaled residual vector with the text semantic vector, the two are first dimensionally aligned and normalized. Fusion methods such as concatenation, weighted summation, or cross-modal attention are then used to effectively combine visual motion information with text semantic features. The resulting noisy control vector incorporates both the dynamic characteristics of key actions in the video and the semantic information of the textual adjustment instructions.

[0119] Deep feature representations are extracted from each keyframe image in the key video frame sequence, capturing its spatial and texture information. A noise control vector, which combines visual motion and textual semantics, is used as a conditional input. Noise perturbations are applied to the keyframe image features in a temporal order to simulate the dynamic changes and uncertainty found in real videos. This guided temporal noise addition process yields noisy keyframes, which help the diffusion model better reproduce the continuity and semantic consistency of video motion in subsequent generation or reconstruction stages. The calculation formula for the noisy keyframes is as follows:

[0120]

[0121] in, Indicates the nth frame with noise added to the key frame. represents the noise parameter, represents the key frame image features of the n+cth frame, represents the image features of the key frame of the nth frame, represents the random noise finally obtained by the noise control vector of the n+cth frame, The random noise finally obtained by the noise control vector of the nth frame is represented. The noisy key frames are spliced ​​in chronological order to form a continuous frame sequence to construct the overall noisy video. The splicing process ensures the temporal coherence and sequential consistency of the noisy key frames.

[0122] By fusing the semantic information of the video's adjusted text with the scaled residual vector to form a noise control vector, and combining it with the image features of key video frame sequences for guided temporal noise addition, the model generates noisy keyframes and splices them into a noisy video in chronological order. This achieves a deep fusion and synergy of semantic driving and motion features. This not only enhances the model's precise control over video content and its ability to dynamically express it, but also effectively simulates natural motion and change in the video, improving the realism and coherence of the generated video and significantly enhancing the quality and effectiveness of subsequent denoising, restoration, and motion reproduction.

[0123] S6. De-noise and restore the noise video to obtain a target action video.

[0124] In this embodiment of the present invention, the generated noisy video is subjected to a video frame denoising module within an optimized diffusion model, which gradually removes noise components in reverse order, restoring clear images and continuous motion in each frame. This denoising process, based on Bayesian optimal a posteriori estimation, combines motion features with textual semantic guidance to accurately restore key action details and coherent background changes in the video, ultimately generating a high-quality, natural, and smooth video of the target action that adheres to the preset content.

[0125] In healthcare scenarios, the denoising and restoration process can clean up noisy instructional or surgical demonstration videos frame by frame, restoring the details and coherence of the doctor's movements, making the videos more realistic and easier to understand. By accurately restoring key actions such as injections and testing, the generated target action videos can be used for remote medical training and patient education, improving the standardization and accessibility of medical services.

[0126] In specific FinTech scenarios, the denoising and restoration module can clearly restore noisy customer service or product introduction videos, ensuring the natural flow of advisor explanations and customer interactions. The resulting high-quality, targeted action videos can be used for digital marketing and online service presentations, enhancing user experience and trust.

[0127] In an embodiment of the present invention, denoising and restoring the noise video to obtain the target action video includes:

[0128] Interpolating the intermediate frames of the noise video to obtain a noise interpolation video;

[0129] Randomly selecting a noise video frame of the noise interpolation video;

[0130] Performing noise analysis on the noisy video frames one by one using the optimized diffusion model to obtain noise residuals;

[0131] Performing Bayesian expectation denoising on the noisy video frame using the noise residual to obtain a plurality of denoised video frame images;

[0132] Upsampling the transition frame image between every two denoised video frame images to obtain a transition update frame image;

[0133] The transition update frame image and the denoised video frame image are spliced ​​to obtain a target action video.

[0134] Specifically, by analyzing the features and motion information of the previous and next frames, an interpolation model (such as optical flow interpolation or a deep learning interpolation network) is used to predict the pixel values ​​and dynamic changes of the intermediate frames, filling the gaps and achieving a smooth transition between frames. The resulting noise-interpolated video contains continuous dynamic details, improving the temporal coherence and visual fluidity of the video, and providing a more complete temporal input for subsequent video restoration and denoising.

[0135] A noisy video frame from the noise-interpolated video is randomly selected and fed into the optimized diffusion model as input. Using a multi-layer temporal attention mechanism and motion perception module, the model gradually analyzes and decomposes the noise components in the noisy video frame, extracting and estimating its noise residual. This helps the model identify and quantify the difference between noise interference and true motion signals, providing accurate noise residual information for subsequent denoising and image quality restoration.

[0136] Based on the noise residual, a probabilistic model of noise and true signal is established. Combining prior information with observed data, Bayesian inference is used to calculate the posterior expected value of each pixel in the noisy video frame. This expected value represents the best estimate of the noise-free image pixel given the noise residual. The model then generates several denoised video frames, effectively suppressing noise interference while preserving video detail and motion characteristics, achieving high-quality image restoration. The calculation formula for the estimation function is as follows:

[0137]

[0138] in, represents the noise parameter, represents the noise keyframe, Z represents the noise residual, Represents a denoised video frame image.

[0139] Interpolation upsampling is performed on the transition frames between denoised video frames. Common methods include bilinear interpolation, bicubic interpolation, or deep learning-based super-resolution reconstruction. Taking bilinear interpolation as an example, during upsampling, the color value of the new pixel is calculated based on the weighted average of neighboring pixels. This amplifies the resolution of the transition frame to the target size, maintaining image smoothness and detail. By increasing the number of pixels and refining image details, the resolution and clarity of the frame image are improved, thus generating the transition update frame image.

[0140] These transition update frame images and denoised video frame images are arranged in chronological order, and spliced ​​together through edge alignment and pixel seamless connection to ensure continuity and smooth transition between frames, and finally synthesize a complete and smooth target action video.

[0141] By interpolating the intermediate frames of the noisy video to enhance the continuity between frames, combining the optimized diffusion model to extract the noise residual frame by frame, and using Bayesian expectation denoising to achieve high-precision denoising processing, the clarity and restoration effect of the video image are significantly improved. By upsampling the transition frames, the image resolution and detail expression are enhanced. Finally, the updated denoised frame images are orderly spliced ​​to generate a target action video with coherent pictures, natural movements and higher quality, which overall improves the stability of video reconstruction and the visual experience.

[0142] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0143] like Figure 4 , which is a functional module diagram of a video generation device based on motion continuity provided by an embodiment of the present invention.

[0144] In an embodiment of the present disclosure, a video generation device based on action continuity is provided, and the video generation device based on action continuity corresponds to the video generation method based on action continuity in the above embodiment. Figure 4 As shown, the video generation device 100 based on motion continuity can be installed in an electronic device. According to the functions to be implemented, the video generation device 100 based on motion continuity includes a key frame acquisition module 101, a residual vector extraction module 102, a diffusion model adjustment module 103, a residual vector scaling module 104, a video frame noise addition module 105, and a noise video denoising module 106. The functional modules are described in detail as follows:

[0145] The key frame acquisition module 101 is used to acquire a video action material set, perform key frame analysis on the video action material set, and obtain a key video frame sequence;

[0146] A residual vector extraction module 102 is configured to extract inter-frame residual vectors between consecutive frames from the key video frame sequence;

[0147] The diffusion model adjustment module 103 is configured to adjust the preset initial diffusion model using the inter-frame residual vector to obtain an optimized diffusion model;

[0148] The residual vector scaling module 104 is configured to perform cosine scaling on the inter-frame residual vector using the optimized diffusion model to obtain a scaled residual vector;

[0149] The video frame noise adding module 105 is used to obtain a video adjustment text, and use the video adjustment text and the scaled residual vector to add noise to the key video frame sequence and then splice it to obtain a noisy video;

[0150] The noise video denoising module 106 is used to denoise and restore the noise video to obtain a target action video.

[0151] In one embodiment, when performing key frame analysis on the video action material set to obtain a key video frame sequence, the key frame acquisition module 101 is configured to:

[0152] Standardizing the resolution of the video action material set to obtain a standard action video;

[0153] Dividing the standard action video into a plurality of continuous action video frames;

[0154] Randomly selecting two consecutive action video frames as a video frame group to be detected;

[0155] Performing optical flow analysis on the video frame groups to be detected one by one to obtain optical flow vectors;

[0156] Calculating the motion intensity of each of the to-be-detected video frame groups according to the optical flow vector, and screening out target video frame groups whose motion intensity is greater than a preset threshold;

[0157] The target video frame group is sorted according to a preset time sequence to obtain a key video frame sequence.

[0158] In one embodiment, when extracting inter-frame residual vectors between consecutive frames from the key video frame sequence, the residual vector extraction module 102 is configured to:

[0159] Converting each frame of video image in the key video frame sequence into a grayscale image;

[0160] Obtaining grayscale pixel values ​​corresponding to the grayscale images of two consecutive frames in the key video frame sequence, and performing a difference between the grayscale pixel values ​​of the two consecutive frames to obtain a residual image;

[0161] The residual map is converted into an inter-frame residual vector between two consecutive frames.

[0162] In one embodiment, when the diffusion model adjustment module 103 adjusts the preset initial diffusion model using the inter-frame residual vector to obtain the optimized diffusion model, it is configured to:

[0163] Performing convolution coding on the inter-frame residual vector to obtain residual vector coding;

[0164] Embedding the residual vector encoding into a temporal attention module of a preset initial diffusion model to obtain a motion perception module;

[0165] Performing residual vector analysis on the key video frame sequence using the initial diffusion model to obtain a model residual vector;

[0166] Constructing a residual alignment loss function according to the inter-frame residual vector and the model residual vector;

[0167] The initial diffusion model is adjusted using the motion perception module and the residual alignment loss function to obtain an optimized diffusion model.

[0168] In one embodiment, when performing cosine scaling on the inter-frame residual vector using the optimized diffusion model to obtain the scaled residual vector, the residual vector scaling module 104 is configured to:

[0169] Calculating a temporal weight factor of the inter-frame residual vector;

[0170] Using the optimized diffusion model to adjust the inter-frame residual vector to obtain a modified residual vector;

[0171] The modified residual vector is scaled using the timing weight factor to obtain a scaled residual vector.

[0172] In one embodiment, when the video frame noise adding module 105 adds noise to the key video frame sequence using the video adjustment text and the scaled residual vector and then splices the key video frame sequence to obtain a noisy video, it is configured to:

[0173] Performing text encoding on the video adjustment text to obtain a text semantic vector;

[0174] Fusing the scaled residual vector and the text semantic vector to obtain a noise control vector;

[0175] Extracting key frame image features of the key video frame sequence;

[0176] Adding temporal noise to the key frame image features using the noise control vector to obtain a noisy key frame;

[0177] The noisy key frames are spliced ​​together to obtain a noisy video.

[0178] In one embodiment, when performing denoising and restoring the noise video to obtain the target action video, the noise video denoising module 106 is configured to:

[0179] Interpolating the intermediate frames of the noise video to obtain a noise interpolation video;

[0180] Randomly selecting a noise video frame of the noise interpolation video;

[0181] Performing noise analysis on the noisy video frames one by one using the optimized diffusion model to obtain noise residuals;

[0182] Performing Bayesian expectation denoising on the noisy video frame using the noise residual to obtain a plurality of denoised video frame images;

[0183] Upsampling the transition frame image between every two denoised video frame images to obtain a transition update frame image;

[0184] The transition update frame image and the denoised video frame image are spliced ​​to obtain a target action video.

[0185] In the present invention, a video generation device based on action continuity is provided. First, the present invention obtains a video action material set, performs key frame analysis on the video action material set, obtains a key video frame sequence, combines resolution standardization, optical flow analysis and motion intensity screening, not only improves the accuracy and robustness of key frame extraction, but also effectively compresses the amount of video information, retains the frame content reflecting the main action features, extracts inter-frame residual vectors between consecutive frames from the key video frame sequence, effectively captures the slight changes and motion features between adjacent action frames in the video, and uses grayscale differences to highlight inter-frame dynamic information while maintaining time continuity. The inter-frame residual vector is used to adjust the preset initial diffusion model to obtain an optimized diffusion model, which significantly enhances the diffusion model's effect on video action. The perception and expression ability of temporal dynamics, the optimized diffusion model is used to perform cosine scaling on the inter-frame residual vector to obtain a scaled residual vector, which can effectively enhance the motion information of the key time points while suppressing irrelevant or noise parts. Then, the video adjustment text is obtained, and the key video frame sequence is denoised and spliced ​​using the video adjustment text and the scaled residual vector to obtain a noisy video, which realizes the deep fusion and synergy of semantic drive and motion features. Finally, the noisy video is denoised and restored, and high-precision denoising is achieved using Bayesian expectation denoising, which significantly improves the clarity and restoration effect of the video image. By upsampling the transition frame, the image resolution and detail expression are enhanced to obtain the target action video, which can effectively improve the action coherence in the customized generated video. The specific definition of a video generation device based on action coherence can be found in the definition of a video generation method based on action coherence above, which will not be repeated here. The various modules in the above-mentioned video generation device based on action coherence can be implemented in whole or in part by software, hardware and their combination. The above modules may be embedded in or independent of the processor in the computer device in the form of hardware, or may be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0186] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 5As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a video generation method based on motion continuity.

[0187] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the client side of a video generation method based on motion continuity.

[0188] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0189] Acquire a video action material set, perform key frame analysis on the video action material set, and obtain a key video frame sequence;

[0190] Extracting inter-frame residual vectors between consecutive frames from the key video frame sequence;

[0191] Using the inter-frame residual vector to adjust the preset initial diffusion model to obtain an optimized diffusion model;

[0192] performing cosine scaling on the inter-frame residual vector using the optimized diffusion model to obtain a scaled residual vector;

[0193] Obtaining a video adjustment text, and using the video adjustment text and the scaled residual vector to add noise to the key video frame sequence and then splice the key video frame sequence to obtain a noisy video;

[0194] The noise video is denoised and restored to obtain a target action video.

[0195] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and apparatuses can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the module division is merely a logical function division, and actual implementation may employ other division methods.

[0196] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or hardware plus software functional modules.

[0197] Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the claims are intended to be embraced therein. Any reference to a figure in a claim should not be construed as limiting the claim to which it relates.

[0198] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0199] In some implementations of this embodiment, a computer-readable storage medium is provided, on which a computer program is stored, characterized in that when the computer program is executed by a processor, the steps of the method described in the above embodiment are implemented.

[0200] The readable storage medium of the present invention stores a computer program, which, when executed by a processor of an electronic device, can implement:

[0201] Acquire a video action material set, perform key frame analysis on the video action material set, and obtain a key video frame sequence;

[0202] Extracting inter-frame residual vectors between consecutive frames from the key video frame sequence;

[0203] Using the inter-frame residual vector to adjust the preset initial diffusion model to obtain an optimized diffusion model;

[0204] performing cosine scaling on the inter-frame residual vector using the optimized diffusion model to obtain a scaled residual vector;

[0205] Obtaining a video adjustment text, and using the video adjustment text and the scaled residual vector to add noise to the key video frame sequence and then splice the key video frame sequence to obtain a noisy video;

[0206] The noise video is denoised and restored to obtain a target action video.

[0207] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0208] The computer-readable storage medium may also store at least one computer-executable program / instruction, such as a computer-readable instruction. Computer-readable storage media include, but are not limited to, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Computer-readable storage media may include, for example, read-only memory (ROM), a hard disk, a flash memory, etc. For example, a non-transitory computer-readable storage medium may be connected to a computing device such as a computer, and then, when the computing device executes the computer-readable instructions stored on the computer-readable storage medium, the various methods described above may be performed.

[0209] In addition, the computer device may also include (but is not limited to) a data bus, an input / output (I / O) bus, a display, and input / output devices (eg, keyboard, mouse, speaker, etc.).

[0210] The processor can communicate with external devices via an I / O bus via a wired or wireless network.

[0211] In one embodiment, the at least one computer executable instruction may also be compiled into or constitute a software product / computer program product, wherein one or more computer executable instructions are executed by a processor to perform the various functions and / or method steps in the embodiments described in the present technology.

[0212] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0213] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0214] In the embodiments provided in the present disclosure, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a portion of code, and the above-mentioned module, program segment or a portion of code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.

[0215] It should be noted that, in this disclosure, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element limited by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0216] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.

[0217] It should be noted that if software tools or components other than those of our company appear in the embodiments of this application, they are only used for illustration and do not represent actual use.

Claims

1. A video generation method based on motion continuity, characterized in that: The method comprises: Acquire a video action material set, perform key frame analysis on the video action material set, and obtain a key video frame sequence; Extracting inter-frame residual vectors between consecutive frames from the key video frame sequence; Using the inter-frame residual vector to adjust the preset initial diffusion model to obtain an optimized diffusion model; performing cosine scaling on the inter-frame residual vector using the optimized diffusion model to obtain a scaled residual vector; Obtaining a video adjustment text, and using the video adjustment text and the scaled residual vector to add noise to the key video frame sequence and then splice the key video frame sequence to obtain a noisy video; The noise video is denoised and restored to obtain a target action video.

2. The method for generating a video based on motion continuity according to claim 1, wherein: The step of performing key frame analysis on the video action material set to obtain a key video frame sequence includes: Standardizing the resolution of the video action material set to obtain a standard action video; Dividing the standard action video into a plurality of continuous action video frames; Randomly selecting two consecutive action video frames as a video frame group to be detected; Performing optical flow analysis on the video frame groups to be detected one by one to obtain optical flow vectors; Calculating the motion intensity of each of the to-be-detected video frame groups according to the optical flow vector, and screening out target video frame groups whose motion intensity is greater than a preset threshold; The target video frame group is sorted according to a preset time sequence to obtain a key video frame sequence.

3. The video generation method based on motion continuity according to claim 1, wherein: The extracting the inter-frame residual vector between consecutive frames from the key video frame sequence includes: Converting each frame of video image in the key video frame sequence into a grayscale image; Obtaining grayscale pixel values ​​corresponding to the grayscale images of two consecutive frames in the key video frame sequence; Subtracting the grayscale pixel values ​​of two consecutive frames to obtain a residual image; The residual map is converted into an inter-frame residual vector between two consecutive frames.

4. The method for generating video based on motion continuity according to claim 1, wherein: The adjusting the preset initial diffusion model by using the inter-frame residual vector to obtain an optimized diffusion model includes: Performing convolution coding on the inter-frame residual vector to obtain residual vector coding; Embedding the residual vector encoding into a temporal attention module of a preset initial diffusion model to obtain a motion perception module; Performing residual vector analysis on the key video frame sequence using the initial diffusion model to obtain a model residual vector; Constructing a residual alignment loss function according to the inter-frame residual vector and the model residual vector; The initial diffusion model is adjusted using the motion perception module and the residual alignment loss function to obtain an optimized diffusion model.

5. The video generation method based on motion continuity according to claim 1, wherein: The method of performing cosine scaling on the inter-frame residual vector by using the optimized diffusion model to obtain a scaled residual vector includes: Calculating a temporal weight factor of the inter-frame residual vector; Using the optimized diffusion model to adjust the inter-frame residual vector to obtain a modified residual vector; The modified residual vector is scaled using the timing weight factor to obtain a scaled residual vector.

6. The method for generating video based on motion continuity according to claim 1, wherein: The step of adding noise to the key video frame sequence and splicing the key video frame sequence using the video adjustment text and the scaled residual vector to obtain a noisy video includes: Performing text encoding on the video adjustment text to obtain a text semantic vector; Fusing the scaled residual vector and the text semantic vector to obtain a noise control vector; Extracting key frame image features of the key video frame sequence; Adding temporal noise to the key frame image features using the noise control vector to obtain a noisy key frame; The noisy key frames are spliced ​​together to obtain a noisy video.

7. The video generation method based on motion continuity according to claim 1, wherein: The denoising and restoring of the noise video to obtain the target action video includes: Interpolating the intermediate frames of the noise video to obtain a noise interpolation video; Randomly selecting a noise video frame of the noise interpolation video; Performing noise analysis on the noisy video frames one by one using the optimized diffusion model to obtain noise residuals; Performing Bayesian expectation denoising on the noisy video frame using the noise residual to obtain a plurality of denoised video frame images; Upsampling the transition frame image between every two denoised video frame images to obtain a transition update frame image; The transition update frame image and the denoised video frame image are spliced ​​to obtain a target action video.

8. A video generation device based on motion continuity, characterized in that: The device comprises: A key frame acquisition module is used to acquire a video action material set, perform key frame analysis on the video action material set, and obtain a key video frame sequence; A residual vector extraction module, configured to extract inter-frame residual vectors between consecutive frames from the key video frame sequence; a diffusion model adjustment module, configured to adjust a preset initial diffusion model using the inter-frame residual vector to obtain an optimized diffusion model; a residual vector scaling module, configured to perform cosine scaling on the inter-frame residual vector using the optimized diffusion model to obtain a scaled residual vector; A video frame noise adding module is used to obtain a video adjustment text, and use the video adjustment text and the scaled residual vector to add noise to the key video frame sequence and splice it to obtain a noisy video; The noise video denoising module is used to denoise and restore the noise video to obtain the target action video.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the video generation method based on motion continuity as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for generating a video based on motion continuity according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Video generation method and related equipment

    CN121284363A

  • Mixed linear layer enhanced image generation method and system

    CN121414909A

  • Dynamic fuzzy scene three-dimensional reconstruction method based on 3DGS technology

    CN121482288A

  • A dynamic fuzzy scene three-dimensional reconstruction method based on 3DGS technology

    CN121482288B

  • Interpretable text semantic driving time sequence generation method based on Diffusion Transform model

    CN121981127A