Animation generation method and device, computer equipment and storage medium

By combining an input encoding layer, an adaptive sparse temporal attention module, and a gated output layer, the high computational complexity and low generation efficiency of traditional animation generation methods are solved, achieving efficient and low-redundancy animation generation, which is suitable for dynamic visualization in the fields of fintech and healthcare.

CN120894476APending Publication Date: 2025-11-04PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510957192.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Traditional image animation generation methods have high computational complexity, making it difficult to capture long-distance temporal dependencies, resulting in low generation efficiency and low information transmission efficiency in dynamic visualization applications in the financial and medical fields.

Method used

An animation generation method based on a preset input encoding layer, an adaptive sparse temporal attention module, and a gated output layer is adopted. Through feature extraction, sparse attention computation, and fusion processing, an efficient animation sequence is generated.

Benefits of technology

It effectively reduces computational complexity, improves the efficiency and quality of image and animation generation, achieves temporal consistency from multiple perspectives, and enhances user decision-making efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120894476A_ABST
    Figure CN120894476A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and relates to an animation generation method and device, computer equipment and a storage medium, and the method comprises the steps: obtaining a to-be-processed key frame sequence; performing feature extraction processing on the key frame sequence based on the input coding layer to obtain a multi-scale feature map; performing key frame selection and sparse attention calculation processing on the multi-scale feature map based on an adaptive sparse time sequence attention module to obtain attention features; performing fusion processing on the multi-scale feature map and the attention features based on a gating output layer to obtain an initial animation sequence; evaluating the initial animation sequence based on a discriminator; if the initial animation sequence passes the evaluation, decoding the initial animation sequence based on a decoder to obtain a target animation; and performing output processing on the target animation. In addition, the target animation can be stored in the block chain. The method can be applied to animation generation scenes in the financial science and technology field and the medical field, and the generation efficiency and the generation quality of the image animation are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and can be applied to the fields of financial technology and digital medical treatment, and particularly relates to an animation generation method and device, a computer device and a storage medium. BACKGROUND

[0002] In the field of image animation generation, traditional technologies mainly model the inter-frame time sequence relationship of a static image sequence through dense time sequence convolution or a recurrent neural network (RNN) to achieve dynamic sequence generation. However, such methods have significant defects: on the one hand, dense time sequence convolution needs to perform global feature extraction on the full sequence, resulting in an exponential increase in computational complexity with the sequence length; on the other hand, RNN-type methods have the problem of gradient vanishing or explosion due to the cyclic structure, and are difficult to capture long-distance time sequence dependencies, and have a large amount of redundant calculation in the training and inference stages, which seriously restricts the animation generation efficiency.

[0003] Specifically, traditional methods usually take frame-by-frame independent processing or simple inter-frame interpolation as the core, and lack dynamic modeling capability for global time sequence features. For example, in the production of film and television special effects, traditional technologies need to manually label key frames and adjust motion parameters frame by frame, resulting in a special effect generation period of several weeks, and it is difficult to guarantee the time sequence consistency under multiple perspectives; in the field of virtual character driving, traditional methods rely on high-precision motion capture devices to generate training data, which is costly and cannot adapt to diversified scene requirements.

[0004] The above problems are particularly prominent in dynamic visualization applications in the fields of finance and medical treatment. For example, in the insurance product demonstration scene in the financial field, traditional methods need to demonstrate the claim settlement process step by step through static charts, and users need to piece together the time sequence logic themselves, resulting in low information transmission efficiency; for another example, in the surgical plan rehearsal scene in the medical field, traditional technologies can only provide static anatomical diagrams or discrete surgical step instructions, making it difficult for doctors to intuitively assess the operation risk and patient prognosis effect.

[0005] Therefore, there is an urgent need for an efficient and low-redundancy image animation generation technology to meet the dynamic visualization needs across different fields. SUMMARY

[0006] The purpose of the embodiments of the present application is to provide an animation generation method and device, a computer device and a storage medium to solve the technical problem of low generation efficiency of existing image animation generation methods.

[0007] In a first aspect, an animation generation method is provided, comprising:

[0008] obtaining a key frame sequence to be processed;

[0009] performing feature extraction processing on the key frame sequence based on a preset input encoding layer to obtain corresponding multi-scale feature maps;

[0010] perform key frame selection and sparse attention calculation processing on the multi-scale feature map based on a preset adaptive sparse temporal attention module to obtain corresponding attention features;

[0011] fuse the multi-scale feature map and the attention features based on a preset gating output layer to obtain a corresponding initial animation sequence;

[0012] evaluate the initial animation sequence based on a preset discriminator;

[0013] if the initial animation sequence passes the evaluation, decode the initial animation sequence based on a preset decoder to obtain a corresponding target animation;

[0014] output the target animation.

[0015] In a second aspect, an animation generation apparatus is provided, which includes:

[0016] an acquisition module configured to acquire a key frame sequence to be processed;

[0017] an extraction module configured to perform feature extraction processing on the key frame sequence based on a preset input encoding layer to obtain a corresponding multi-scale feature map;

[0018] a processing module configured to perform key frame selection and sparse attention calculation processing on the multi-scale feature map based on a preset adaptive sparse temporal attention module to obtain corresponding attention features;

[0019] a fusion module configured to fuse the multi-scale feature map and the attention features based on a preset gating output layer to obtain a corresponding initial animation sequence;

[0020] an evaluation module configured to evaluate the initial animation sequence based on a preset discriminator;

[0021] a decoding module configured to, if the initial animation sequence passes the evaluation, decode the initial animation sequence based on a preset decoder to obtain a corresponding target animation;

[0022] an output module configured to output the target animation.

[0023] In a third aspect, a computer device is provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the above animation generation method when executing the computer program.

[0024] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. The computer program is executed by a processor to implement the steps of the animation generation method.

[0025] In the scheme implemented by the animation generation method, the animation generation device, the computer device, and the storage medium, first, the key frame sequence to be processed is obtained. Then, the input coding layer is used to perform feature extraction processing on the key frame sequence to obtain a corresponding multi-scale feature map. Then, the adaptive sparse temporal attention module is used to perform key frame selection and sparse attention calculation processing on the multi-scale feature map to obtain a corresponding attention feature. Subsequently, the multi-scale feature map and the attention feature are fused by using the preset gating output layer to obtain an initial animation sequence. Further, the initial animation sequence is evaluated by using the preset discriminator. If the initial animation sequence passes the evaluation, the initial animation sequence is decoded by using the preset decoder to obtain a corresponding target animation. Finally, the target animation is output. Based on the above automatic processing procedure, after receiving the key frame sequence to be processed, the multi-scale feature map is obtained by using the input coding layer to perform feature extraction processing on the key frame sequence. Then, by using the adaptive sparse temporal attention module, dynamic key frame selection and sparse cross-frame attention are performed to effectively solve the contradiction between the calculation efficiency and the generation quality of the traditional animation generation method. Moreover, the adaptive sparse temporal attention module can automatically identify the salient regions in the input sequence, and only the key frames are deeply processed, thereby greatly reducing the calculation complexity and effectively improving the generation efficiency of the image animation. Meanwhile, the attention feature and the conventional multi-scale feature map are organically combined by using the gating output layer, which not only retains the richness of the spatial details, but also ensures the continuity of the temporal motion, thereby effectively improving the quality of the target animation generated by processing the initial animation sequence by using the combination of the discriminator and the decoder. BRIEF DESCRIPTION OF DRAWINGS

[0026] In order to more clearly illustrate the schemes in the present application, the drawings needed in the embodiments of the present application will be briefly introduced as follows. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0027] Figure 1 is an exemplary system architecture diagram to which the present application can be applied;

[0028] Figure 2 is a flowchart of one embodiment of the animation generation method according to the present application;

[0029] Figure 3is a structural schematic diagram of one embodiment of the animation generation apparatus according to the present application;

[0030] Figure 4 is a structural schematic diagram of one embodiment of the computer device according to the present application. DETAILED DESCRIPTION

[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terminology used in the description herein is for describing particular embodiments only and is not intended to be limiting of the application; the use herein of terms such as "comprise" and "have" and any variations thereof are intended to cover a non-exclusive inclusion; the use herein of terms such as "first", "second" and the like are used for distinguishing between similar objects and not necessarily for describing a specific sequential or chronological order.

[0032] Reference herein to "embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearances of the phrase that the phrase in the specification in various places does not necessarily all refer to the same embodiment, nor are they necessarily mutually exclusive or alternative embodiments. It is explicitly and implicitly understood that the embodiments described herein can be combined.

[0033] In order to make the technical personnel in the art better understand the scheme of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings.

[0034] As shown in Figure 1 The system architecture 100 can include a terminal device 101, a network 102 and a server 103, the terminal device 101 can be a notebook computer 1011, a tablet computer 1012 or a mobile phone 1013. The network 102 is a medium for providing a communication link between the terminal device 101 and the server 103. The network 102 can include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0035] The user can use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0036] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.

[0037] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.

[0038] It should be noted that the animation generation method provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the animation generation device is generally set in the server / terminal device.

[0039] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0040] Continue to refer to Figure 2 The flowchart illustrates an embodiment of the animation generation method according to this application. Depending on different needs, the order of the steps in the flowchart can be changed, and some steps can be omitted. The animation generation method provided by this application embodiment can be applied to any scenario requiring animation generation, and thus can be applied to products in these scenarios, such as animation generation scenarios in the fintech and medical fields. The animation generation method includes the following steps:

[0041] Step S201: Obtain the keyframe sequence to be processed.

[0042] In this embodiment, the animation generation method runs on an electronic device (e.g., Figure 1The server / terminal device shown can acquire the keyframe sequence to be processed via wired or wireless connection. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra-wideband) connections, and other currently known or future-developed wireless connection methods. The executing entity of this application is specifically an animation generation system, which can be simply referred to as the system. The aforementioned keyframe sequence to be processed can be data containing multiple keyframes input by the user according to actual animation generation requirements.

[0043] The overall system architecture of this application is based on an improved StyleGAN2 generator, and mainly includes the following key modules: 1. Input encoding layer: receives latent variable noise. and keyframe sequences Each keyframe k t The system consists of an H×W×C dimensional feature map, where C is the channel, H is the height, and W is the width, totaling d dimensions, with a time sequence of T. 2. The core ASTA module comprises a saliency calculator, a keyframe selector, and sparse temporal attention blocks. 3. A multi-scale feature fusion network implements ASTA processing at different resolution levels. 4. A gated output layer dynamically mixes ASTA features with the original convolutional features. 5. A discriminator verifies the coherence of the generated sequence through 3D convolution and optical flow constraints. The system workflow is as follows: The input keyframe sequence first extracts basic features through a weighted convolutional network. Then, the ASTA module performs dynamic keyframe selection and sparse attention calculation on the feature maps at each level. Finally, the gated fusion layer generates an animation sequence with spatiotemporal consistency.

[0044] This application can be applied to animation generation scenarios in the fintech and medical fields. The animation generation method provided by this application can transform complex financial insurance terms and high-risk medical procedures into intuitive and credible dynamic content, significantly improving user decision-making efficiency. For example, in the business of dynamically demonstrating insurance claims processes in the financial insurance field, application scenarios may include: insurance companies can use animation to show customers the entire claims process (such as car insurance claims), solving the problem of traditional textual explanations being obscure and difficult to understand, and improving user experience and trust. The keyframe sequence input by the user covers the key nodes of the claims process. Each keyframe contains a scene image + data annotation, specifically including: Keyframe 1: Accident Reporting. Image: Accident scene (such as a rear-end collision between two vehicles), the user clicks "One-Click Reporting" through the APP. Data annotation: Timestamp (2024-03-15 14:30), report number (CL20240315001). Keyframe 2: Damage Assessment. Image: An assessor uses AR equipment to scan vehicle damage, and the system automatically generates a 3D damage model. Data annotation: Damaged parts (front bumper, left front fender), estimated repair cost (¥8,500). Keyframe 3: Document submission. Screen: User uploads electronic documents such as driver's license, vehicle registration certificate, and repair invoice. Data annotation: Document status (uploaded / under review), remaining materials list (missing repair list). Keyframe 4: Claim review. Screen: AI review system automatically compares the terms, green checkmarks indicate "passed", red exclamation marks indicate "requires supplementary materials". Data annotation: Review result (passed / rejected), reason for rejection (e.g., not reported within 48 hours). Keyframe 5: Claim payment received. Screen: Mobile banking interface displays claim payment received (¥7,650, minus 10% deductible). Data annotation: Arrival time (2024-03-18 10:15), claim calculation details.

[0045] In the medical field, applications of surgical procedure simulation and patient communication include: hospitals can use animation to demonstrate surgical steps (such as knee replacement surgery) to patients, helping them understand surgical risks, postoperative recovery, and alleviating anxiety. The user-input keyframe sequence includes medical images, 3D models, and surgical instrument animations, highlighting key operational nodes. For example, it might include: Keyframe 1: Preoperative Examination. Image: 3D reconstruction of the patient's knee joint CT scan, with the lesion area marked (highlighted red osteoarthritis lesions). Data annotation: lesion severity (Kellgren-Lawrence classification IV), surgical indications. Keyframe 2: Anesthesia Induction. Image: The anesthesiologist administers medication; the patient's vital signs monitor shows heart rate and blood pressure gradually stabilizing. Data annotation: anesthesia method (combined spinal-epidural anesthesia), drug dosage (propofol 2mg / kg). Keyframe 3: Surgical Approach. Image: 3D animation showing the anteromedial incision of the knee joint (marked with a red dotted line), cutting through the skin and subcutaneous tissue layer by layer. Data annotation: incision length (8cm), tourniquet pressure (300mmHg). Keyframe 4: Prosthesis Implantation. Visuals: Dynamic demonstration of the osteotomy process using surgical instruments (bone chisel, intramedullary positioning rod), precise implantation of the prosthesis (titanium alloy knee joint). Data annotation: Prosthesis type (ZimmerNexGen LPS), osteotomy amount (5mm tibial plateau osteotomy). Keyframe 5: Postoperative Rehabilitation. Visuals: Patient performs straight leg raise exercises under the guidance of a rehabilitation therapist on postoperative day 1, and walks normally on day 30. Data annotation: Rehabilitation period (6 weeks), expected functional recovery (HSS score ≥ 85 points).

[0046] Furthermore, the specific implementation process for obtaining the keyframe sequence to be processed will be described in more detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0047] Step S202: Based on the preset input encoding layer, feature extraction processing is performed on the key frame sequence to obtain the corresponding multi-scale feature map.

[0048] In this embodiment, the specific implementation process of performing feature extraction processing on the keyframe sequence based on the preset input coding layer to obtain the corresponding multi-scale feature map will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0049] Step S203: Based on the preset adaptive sparse temporal attention module, perform keyframe selection and sparse attention calculation processing on the multi-scale feature map to obtain the corresponding attention features.

[0050] In this embodiment, the specific implementation process of performing keyframe selection and sparse attention calculation on the multi-scale feature map based on the preset adaptive sparse temporal attention module to obtain the corresponding attention features will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0051] Step S204: Based on a preset gated output layer, the multi-scale feature map and the attention feature are fused to obtain the corresponding initial animation sequence.

[0052] In this embodiment, the specific implementation process of fusing the multi-scale feature map and the attention feature based on the preset gated output layer to obtain the corresponding initial animation sequence will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0053] Step S205: Evaluate the initial animation sequence based on a preset discriminator.

[0054] In this embodiment, the specific implementation process of evaluating the initial animation sequence based on the preset discriminator will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0055] Step S206: If the initial animation sequence passes the evaluation, the initial animation sequence is decoded based on a preset decoder to obtain the corresponding target animation.

[0056] In this embodiment, the specific implementation process of decoding the initial animation sequence based on the preset decoder to obtain the corresponding target animation will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0057] Step S207: Output the target animation.

[0058] In this embodiment, the output processing of the target animation can be completed by sending the generated target animation to the corresponding user in a specified output format. The selection of the output method is not specifically limited; for example, it can be sent via email, displayed on a user interface, or sent via SMS.

[0059] This application first obtains a keyframe sequence to be processed; then, based on a preset input encoding layer, it performs feature extraction processing on the keyframe sequence to obtain a corresponding multi-scale feature map; subsequently, based on a preset adaptive sparse temporal attention module, it performs keyframe selection and sparse attention calculation processing on the multi-scale feature map to obtain a corresponding attention feature; subsequently, based on a preset gated output layer, it performs fusion processing on the multi-scale feature map and the attention feature to obtain a corresponding initial animation sequence; further, it evaluates the initial animation sequence based on a preset discriminator; if the initial animation sequence passes the evaluation, it decodes the initial animation sequence based on a preset decoder to obtain a corresponding target animation; finally, it outputs the target animation. Based on the above automated processing flow, after receiving the keyframe sequence to be processed, this application performs feature extraction on the keyframe sequence using the input encoding layer to obtain multi-scale feature maps. Then, based on the use of an adaptive sparse temporal attention module, through dynamic keyframe selection and sparse cross-frame attention, it effectively solves the contradiction between computational efficiency and generation quality in traditional animation generation methods. Furthermore, the adaptive sparse temporal attention module can automatically identify salient regions in the input sequence and perform depth processing only on keyframes, thereby significantly reducing computational complexity and effectively improving the generation efficiency of image animation. At the same time, by using a gated output layer to organically combine attention features with conventional multi-scale feature maps, it not only preserves the richness of spatial details but also ensures the coherence of temporal motion, effectively improving the quality of the target animation generated by processing the initial animation sequence using the combination of discriminator and decoder.

[0060] In some optional implementations of this embodiment, step S201 includes the following steps:

[0061] Receive the initial keyframe sequence input by the user, and obtain the preset latent variable noise.

[0062] In this embodiment, the latent variable noise is a low-dimensional random vector that can be sampled from a standard normal distribution N(0,I). This latent variable noise serves as a random input to the coding layer, introducing randomness and diversity to control the subtle changes in animation details. The initial keyframe sequence refers to multiple keyframes provided by the user. If the user input is video, the keyframe sequence is generated through preprocessing (such as frame sampling and resolution adjustment).

[0063] Obtain the preset fusion strategy.

[0064] In this embodiment, the above-mentioned fusion strategy includes: mapping latent variable noise into a tensor with the same dimension as the keyframe features, and fusing it with the initial keyframe sequence through splicing or AdaIN.

[0065] Based on the fusion strategy, the initial keyframe sequence and the latent variable noise are fused to obtain the corresponding processing features.

[0066] In this embodiment, the initial keyframe sequence and latent variable noise can be fused based on the above-mentioned fusion strategy, and the resulting processed features can be used as the corresponding keyframe sequence, i.e., the input data of the input coding layer.

[0067] The processing features are used as the keyframe sequence.

[0068] In this embodiment, the keyframe sequence generation step provides initial conditions for the animation generation process. Latent variable noise controls the randomness of the generated content, while the initial keyframe sequence defines the spatiotemporal constraints of the animation. The combination of the two allows the generated animation to be both diverse and conform to the user-specified structure.

[0069] This application receives an initial keyframe sequence input by the user and obtains preset latent variable noise; then obtains a preset fusion strategy; subsequently, it fuses the initial keyframe sequence and the latent variable noise based on the fusion strategy to obtain corresponding processing features; and finally, it uses the processing features as the keyframe sequence. Based on the above processing flow, this application, by using a fusion strategy to fuse the initial keyframe sequence input by the user and the obtained latent variable noise, can automatically and intelligently generate corresponding keyframe sequences. Furthermore, by using latent variable noise to control the randomness of the generated content, and by defining the spatiotemporal constraints of the animation in the initial keyframe sequence, the intelligence and adaptability of the generated keyframe sequences are effectively improved.

[0070] In some optional implementations, the input encoding layer includes a shared-weight convolutional network; step S202 includes the following steps:

[0071] The keyframe sequence is initially downsampled based on the shared weight convolutional network to obtain the corresponding first feature map.

[0072] In this embodiment, the initial downsampling process based on the shared-weight convolutional network includes: 1. Convolutional kernel configuration: A 3×3 convolutional kernel is used, with a stride of 2 and padding of 1, to maintain a controllable downsampling ratio for spatial resolution. The kernel weights are generated through random initialization (e.g., Xavier initialization), and the bias term is initialized to 0. 2. Channel expansion: After the input keyframe passes through the convolutional layer, the number of output channels is expanded to C′ (e.g., 64). 3. Nonlinear activation: The ReLU activation function is applied after the convolution operation, introducing nonlinearity: ReLU(x) = max(0,x). The initial downsampling reduces the spatial resolution (halving the size) through a convolution with a stride of 2, while simultaneously expanding the number of channels to encode higher-dimensional features. ReLU activation enhances feature representation and avoids the gradient vanishing problem.

[0073] The first feature map is processed by residual block processing to obtain the corresponding second feature map.

[0074] In this embodiment, residual block processing includes: 1. Residual block structure: Each residual block contains two sub-paths: Main path: Two consecutive 3×3 convolutional layers (stride 1, padding 1), each followed by batch normalization (BatchNorm) and ReLU activation. Shortcut connection: If the input and output dimensions are inconsistent (e.g., the number of channels changes), the dimensions are adjusted through 1×1 convolution; otherwise, they are directly added. 2. Feature propagation: Input feature x generates F(x) through the main path, and then adds it to x from the shortcut connection: y = F(x) + x. ReLU activation is applied again to the added result: ReLU(y). 3. Stacking residual blocks: Multiple residual blocks (e.g., 2 to 4) are repeatedly stacked to gradually extract higher-order features. After each stacking, the feature map resolution remains unchanged, but the number of channels may increase (e.g., from 64 to 128). The residual blocks alleviate the gradient vanishing problem in deep networks through shortcut connections, while batch normalization stabilizes the training process. Stacking multiple residual blocks can capture hierarchical features from local textures to global structures.

[0075] The second feature map is then subjected to a final downsampling process to obtain the corresponding third feature map.

[0076] In this embodiment, the final downsampling process includes: 1. Resolution adjustment: After the residual block sequence, a 3×3 convolutional layer (stride 2, padding 1) is used again to reduce the feature map resolution by half, while keeping the number of channels unchanged (e.g., 64). 2. Output generation:

[0077] The aforementioned network is applied independently to all keyframes, but they share the same set of convolutional weights (including residual block parameters), ultimately generating multi-scale feature maps, which serve as the required multi-scale feature maps. The shared weights ensure that all keyframes are encoded in the same feature space, avoiding semantic inconsistencies caused by independent encoding. Finally, downsampling further compresses spatial information, making the features more suitable for subsequent temporal modeling.

[0078] The implementation process of the weight sharing mechanism includes: 1. Parameter initialization: The parameters of the convolutional layers and residual blocks shared by all keyframes are completely identical during initialization (e.g., generated using the same random seed). 2. Training consistency: During backpropagation, the gradients of all keyframes are backpropagated through the shared network, and the parameter update direction is jointly determined by the gradients of all frames. Weight sharing significantly reduces the number of parameters (e.g., from T×independent networks to 1 network), while forcing all keyframes to be represented in a unified feature space, which is beneficial for subsequent temporal processing.

[0079] The third feature map is used as the multi-scale feature map.

[0080] This application performs initial downsampling on the keyframe sequence based on the shared-weight convolutional network to obtain a corresponding first feature map; then, residual block processing is applied to the first feature map to obtain a corresponding second feature map; subsequently, final downsampling is applied to the second feature map to obtain a corresponding third feature map; and finally, the third feature map is used as the multi-scale feature map. Based on the above processing flow, this application uses a shared-weight convolutional network to initially downsample the keyframe sequence to reduce resolution and expand the number of channels, performs residual block processing to progressively extract high-order features while maintaining gradient smoothness, and performs final downsampling to further compress spatial information, generating a multi-scale feature map. This allows for the efficient and accurate conversion of the keyframe sequence into a compact representation in a unified feature space, effectively ensuring the accuracy of the generated multi-scale feature map.

[0081] In some optional implementations, the adaptive sparse temporal attention module includes at least a saliency scorer, a keyframe selector, and a sparse temporal transformer; step S203 includes the following steps:

[0082] The multi-scale feature map is processed based on the saliency scorer to obtain the corresponding saliency score result.

[0083] In this embodiment, the input to the aforementioned Adaptive Sparse Temporal Attention (ASTA) module is the multi-scale feature map. The ASTA module includes: 1) A Saliency Scorer: A lightweight convolutional network generates a spatial importance map. Input: Temporal feature sequence; Output: Saliency score α∈[0,1] at each position. 2) A Keyframe Selector: Dynamically determines the frames requiring in-depth processing. Adaptive threshold selection based on saliency scores; Output: A set of keyframe indices. 3) A Sparse Temporal Transformer: The core attention computation unit. Query Projection: Converts features into query vectors; Key Projection: Generates key vectors; Attention Weights: Calculates only the weights between keyframes; Value Aggregation: Weighted fusion of feature information.

[0084] Specifically, given input keyframe k t The saliency calculator generates a spatial importance graph using the lightweight MobileNetV3 network:

[0085] α t =σ(DSConv3(ReLU(DSConv2(DSConv1(k t )))))

[0086] Where DSConv represents depthwise separable convolution, σ is the sigmoid function, and t is time step. t ∈[0,1] H×W Each element represents the temporal importance of the corresponding spatial location. α t The saliency of each pixel was calculated.

[0087] The implementation process of the significance scorer calculation includes:

[0088] 1. Input Feature Processing. Input features (from the input encoding layer) are extracted using Depthwise Separable Convolution (DSConv1): Depthwise Convolution: A 3×3 convolution kernel (stride 1, padding 1) is applied to each input channel, keeping the number of output channels unchanged. Pointwise Convolution: The number of channels is expanded to C″ (e.g., 32) using 1×1 convolution, reducing computational cost. ReLU activation is applied to introduce non-linearity. 2. Intermediate Feature Extraction. Further processing is performed using DSConv2 (Depthwise Separable Convolution), adjusting the number of channels to C″′ (e.g., 16) while maintaining spatial resolution. ReLU activation is applied again. 3. High-Level Feature and Saliency Map Generation. DSConv3 compresses the number of channels to 1, generating a single-channel feature map. The output is mapped to the [0,1] range using the Sigmoid function to obtain a spatial importance map, representing the temporal importance of each pixel location.

[0089] Among them, Depthwise Separable Convolution (DSConv) splits standard convolution into depthwise convolution and pointwise convolution, significantly reducing the number of parameters and computational cost. In addition, the αt output by the Sigmoid assigns a weight of 0 to 1 to each spatial location, and subsequent keyframe selection is based on this dynamic scoring, avoiding the limitations of fixed rules (such as uniform sampling).

[0090] Based on the saliency score and the preset dynamic threshold selection formula, the target keyframes that meet the conditions are determined by the keyframe selector.

[0091] In this embodiment, the above-mentioned dynamic threshold selection formula specifically includes:

[0092]

[0093] Where, μ x,y and σ x,y Let x and y represent the saliency mean and standard deviation of all frames at position (x, y), respectively, with λ being an adjustable parameter. This design allows for the automatic allocation of more keyframe resources to high dynamic range regions, where x and y are pixel coordinates.

[0094] Specifically, the mean μ of the significance values ​​at position (x,y) for all frames is calculated. x,y and standard deviation σ x,y And the dynamic threshold is set to μ x,y +λσ x,y λ is a hyperparameter (e.g., 1.5). If α t If (x,y) exceeds the threshold, then mark time t as the keyframe at position (x,y) and generate set K(x,y), which is the target keyframe.

[0095] Furthermore, the dynamic thresholding mechanism allows the model to automatically focus on highly dynamic regions (such as moving objects) while ignoring static backgrounds. This data-driven selection approach is more flexible than fixed-interval sampling and adapts to complex scenarios.

[0096] The attention weights of the target keyframes are calculated using the sparse temporal transformer to obtain the corresponding calculation results.

[0097] In this embodiment, the calculation of attention weights based on the sparse temporal transformer includes: in the selected target keyframes, i.e., the keyframe set. Based on this, attention weight calculation is limited to position pairs that share keyframes:

[0098]

[0099] in, Q is an indicator function. q K pLet Q be the query vector, K be the key vector, and V be the value vector. There are three learned weight matrices: Q, K, and V. The pixel features are multiplied by these three matrices to obtain the query vector, key vector, and value vector, respectively. d is the feature dimension, and T is the transpose. This sparse attention model reduces the O(T) time complexity of traditional global attention. 2 The complexity is reduced to A(q,p) represents the attention weights for the features of pixels q and p. Then, the value vector V... p Perform a weighted summation and use the result as an attention feature:

[0100] h ASTA =V p +A(q,p)*V p

[0101] In addition, through Limit the attention span so that the model adaptively focuses on information-dense areas (such as moving objects) and ignores static backgrounds.

[0102] The calculation result is used as the attention feature.

[0103] In this embodiment, the adaptive sparse temporal attention module can be further optimized, including: 1) Enhanced stability of the saliency scorer. Spatial smoothing: in generating α t Then, 3×3 average pooling (step size 1) is applied to suppress noise and avoid misselection of keyframes due to local saliency abrupt changes. Threshold saturation processing: If the dynamic threshold exceeds 1, it is forcibly truncated to 1 to prevent all frames from being misclassified as keyframes. 2. Boundary processing of sparse attention. Minimum number of keyframes: If Force the selection of the two most recent frames as keyframes to ensure at least the existence of temporal context. Location encoding: Add learnable location encoding to the query / key vector of the keyframe to compensate for the loss of temporal continuity caused by sparse sampling.

[0104] This application processes the multi-scale feature map using the saliency scorer to obtain a corresponding saliency score. Then, based on the saliency score and a preset dynamic threshold selection formula, a keyframe selector determines the target keyframes that meet the conditions. Next, an attention weight is calculated on the target keyframes using the sparse temporal transformer to obtain the corresponding calculation result. This calculation result is then used as the attention feature. Based on this processing flow, this application processes the multi-scale feature map using a sparse temporal attention module. It dynamically evaluates frame importance using a saliency scorer, adaptively filters information frames using a keyframe selector, and efficiently aggregates temporal information using a sparse temporal transformer. This enables the efficient and accurate generation of attention features corresponding to the multi-scale feature map, effectively ensuring the accuracy of the obtained attention features.

[0105] In some optional implementations of this embodiment, step S204 includes the following steps:

[0106] Calculate the gating coefficients corresponding to the multi-scale feature map.

[0107] In this embodiment, the calculation process of the gating coefficient includes: calculating the saliency map α for all frames. 1:T Average pooling is performed to generate a global saliency vector, which is then passed through a fully connected layer W. γ The sigmoid function generates the gating coefficient γ. Specifically:

[0108] γ = sigmoid(W γ [avgpool(α 1:T )])

[0109] This design allows highly dynamic regions to utilize ASTA features, while static regions retain the original convolutional features.

[0110] Obtain the feature fusion formula corresponding to the gating coefficient.

[0111] In this embodiment, the above feature fusion formula fuses ASTA features with the original features through adaptive weighting, including:

[0112] h out =γ⊙h conv +(1-γ)⊙h ASTA

[0113] Among them, h conv The original convolutional features, i.e., multi-scale feature maps, h ASTA This refers to the attention features output by the adaptive sparse temporal attention module.

[0114] The multi-scale feature map and the attention feature are fused based on the feature fusion formula to obtain the corresponding fused feature.

[0115] In this embodiment, based on the fusion processing method of the above feature fusion formula, the above multi-scale feature map and attention feature can be substituted into the corresponding positions in the feature fusion formula for calculation processing, and the obtained fused features can be used as the corresponding initial animation sequence.

[0116] The fusion feature is used as the initial animation sequence.

[0117] In this embodiment, the gating mechanism dynamically balances two feature paths: highly dynamic regions (such as moving objects) rely on ASTA features, while static regions retain the original convolutional features, ensuring the unity of spatial details and temporal coherence.

[0118] This application calculates the gating coefficients corresponding to the multi-scale feature map; then obtains the feature fusion formula corresponding to the gating coefficients; subsequently, it fuses the multi-scale feature map and the attention feature based on the feature fusion formula to obtain the corresponding fused feature; and finally, it uses the fused feature as the initial animation sequence. Based on the above processing flow, this application, by calculating the gating coefficients corresponding to the multi-scale feature map and then using the feature fusion formula corresponding to the gating coefficients, can efficiently and accurately complete the fusion processing of the multi-scale feature map and the attention feature, ensuring the uniformity of spatial details and temporal coherence of the generated initial animation sequence, and guaranteeing the quality of the generated initial animation sequence.

[0119] In some optional implementations of this embodiment, the discriminator includes at least a temporal feature extractor, a temporal consistency evaluator, and an output layer; step S205 includes the following steps:

[0120] Based on the temporal feature extractor, the initial animation sequence is used to extract features to obtain the corresponding temporal features.

[0121] In this embodiment, the discriminator is used to evaluate the authenticity and temporal coherence of the generated sequence, including: input layer: receiving real or generated animation sequences; hidden layer: containing temporal feature extractors such as 3D convolution; temporal consistency evaluator: an auxiliary module specifically for detecting temporal artifacts; and output layer: generating authenticity scores.

[0122] Specifically, after receiving the initial animation sequence, the temporal feature extractor described above uses a 4×4×4 convolution kernel (time×space) to extract spatiotemporal features from the initial animation sequence with a stride of 2×2×2. The output feature dimension is progressively compressed (e.g., 64→128→256 channels). Each 3D convolution layer is followed by BatchNorm and LeakyReLU (slope 0.2). Spatiotemporal downsampling is achieved through stride convolution, thereby obtaining the corresponding temporal features.

[0123] The temporal features are evaluated using the temporal consistency evaluator to obtain the corresponding temporal artifact scores.

[0124] In this embodiment, the evaluation process based on the temporal consistency evaluator includes: Optical flow calculation: For adjacent frames of the input sequence, the optical flow field is calculated using a predefined optical flow algorithm (such as a simplified version of FlowNet). The size of the optical flow field is H×W×2 (horizontal and vertical displacement). Optical flow difference feature: The difference ΔFt between continuous optical flow fields is calculated to reflect temporal continuity. ΔFt is compressed to a low dimension (e.g., 16 channels) through 1×1 convolution and concatenated with the backbone features of 3D convolution in the channel dimension. Temporal artifact scoring: The concatenated features are activated by a fully connected layer (hidden layer dimension 256→128) and a Sigmoid function to output a temporal artifact score stemp∈[0,1] (1 indicates severe artifacts).

[0125] The temporal features are scored based on the output layer to obtain the corresponding authenticity score.

[0126] In this embodiment, the scoring process based on the output layer includes: globally averaging the features of the last layer of the 3D convolutional backbone, then passing them through a fully connected layer (hidden layers 512→256→1) and sigmoid activation to output an overall realism score sreal∈[0,1]. The joint loss is: the total discriminant loss is LD=λ1·BCE(sreal,yreal)+λ2·MSE(stemp,ytemp), where ytemp is the artifact label of the real sequence (usually set to 0).

[0127] If the temporal artifact score is greater than a first preset threshold and the realism score is greater than a second preset threshold, then the initial animation sequence is determined to have passed the evaluation; otherwise, the initial animation sequence is determined to have failed the evaluation.

[0128] In this embodiment, the first preset threshold setting includes: Standard: Inter-frame transitions must be natural, without logical or visual artifacts. Detection method: Optical flow analysis: Calculate the optical flow field of adjacent frames. If the abrupt change region exceeds the threshold (e.g., 10% of pixel motion speed > 50 pixels / frame), it is determined to be jitter. Cyclic consistency check: Input the sequence into the discriminator after reversing it. If the score differs from the forward playback by > 0.2, it indicates a timing asymmetry problem (e.g., the turning action cannot be reversed). Threshold: Timing score must be > 0.6.

[0129] The second preset threshold setting mentioned above includes: Standard: The static content of each frame must closely approximate the distribution of real data. Detection method: Whether the position of facial features conforms to anatomical structure (e.g., the eyes do not extend beyond the skull). Whether the texture details are reasonable (e.g., the fabric texture does not suddenly become a solid color). Threshold: The score of a single frame must be >0.7 (assuming 1 is completely realistic).

[0130] Specifically, the initial animation sequence is deemed to have passed the evaluation only if the detected temporal artifact score is greater than the first preset threshold and the realism score is greater than the second preset threshold; otherwise, the initial animation sequence is deemed to have failed the evaluation.

[0131] This application extracts features from the initial animation sequence using the temporal feature extractor to obtain corresponding temporal features; then, it evaluates the temporal features using the temporal consistency evaluator to obtain corresponding temporal artifact scores; subsequently, it scores the temporal features based on the output layer to obtain corresponding realism scores; if the temporal artifact score is greater than a first preset threshold and the realism score is greater than a second preset threshold, the initial animation sequence is deemed to have passed the evaluation; otherwise, the initial animation sequence is deemed to have failed the evaluation. Based on the above processing flow, this application extracts features from the initial animation sequence using the temporal feature extractor to obtain temporal features, evaluates the temporal features using the temporal consistency evaluator, scores the temporal features based on the output layer, and performs multi-dimensional evaluation of the initial animation sequence by combining the obtained temporal artifact scores and realism scores. This effectively improves the accuracy, diversity, and standardization of the evaluation process for the initial animation sequence and helps ensure that the generated initial animation sequence has high quality.

[0132] In some alternative implementations, step S206 includes the following steps:

[0133] Call the preset decoder.

[0134] In this embodiment, the decoder is a pre-built tool for generating animations.

[0135] The decoder is used to upsample the initial animation sequence to obtain the corresponding first animation data.

[0136] In this embodiment, the upsampling process refers to transposed convolutional upsampling, which includes: using transposed convolutions (stride 2, kernel size 4×4) to gradually restore the resolution to the same level as the keyframe sequence, with the number of channels decreasing layer by layer (e.g., 512→256→128→64→3). Furthermore, each transposed convolution layer is followed by BatchNorm and ReLU activation. Additionally, the decoder provides features at the same resolution, fusing low-level details (such as edges and colors) with high-level semantics through channel concatenation (Concat).

[0137] The first animation data is rendered to obtain the corresponding second animation data.

[0138] In this embodiment, the above rendering process includes: compressing the number of channels from 64 to 3 (RGB) through a 1×1 convolution, followed by Tanh activation (output range [-1,1]), and using the obtained animation data as the final target animation to be output.

[0139] The second animation data is used as the target animation.

[0140] In this embodiment, the decoder can be further optimized, including: subpixel convolution replacement: replacing transposed convolution with PixelShuffle (stride 1 / 2) to reduce checkerboard artifacts. Attention mechanism: adding spatial attention to skip connections and dynamically weighting encoder features (such as the CBAM module). Progressive generation: decoding long sequences in blocks to avoid insufficient GPU memory.

[0141] This application calls a preset decoder; then, based on the decoder, it upsamples the initial animation sequence to obtain corresponding first animation data; subsequently, it renders the first animation data to obtain corresponding second animation data; and finally, it uses the second animation data as the target animation. Based on the above processing flow, this application automatically and efficiently completes the decoding process of the initial animation sequence and outputs high-quality animation frames by using a decoder to upsample and render the initial animation sequence, effectively improving the generation quality and computational efficiency of the obtained target animation.

[0142] In some alternative implementations, the user information obtained is subject to user consent and complies with relevant laws and policies.

[0143] Furthermore, any software tools or components not belonging to our company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.

[0144] Furthermore, this application proposes an image animation generation method based on adaptive sparse temporal attention, replacing traditional dense temporal convolution by integrating a novel adaptive sparse temporal attention (ASTA) mechanism into the generator's hidden layer. This method employs a hybrid sparse-dense attention framework to achieve dynamic keyframe selection and motion feature propagation while maintaining computational efficiency. The core technologies include a two-stage hierarchical temporal processing approach: First, a lightweight convolutional network computes a spatial importance map and dynamically selects keyframes. Then, attention weights are computed within the sparse neighborhood of shared keyframes, reducing the quadratic complexity of traditional attention mechanisms to local linear complexity. The generator uses a gated fusion layer to balance spatial detail and temporal coherence, while the discriminator suppresses flicker artifacts through temporal consistency loss. This application can be seamlessly integrated into existing generation architectures, significantly improving animation generation quality and efficiency, and is suitable for scenarios such as video compositing and character animation.

[0145] Furthermore, by innovatively introducing a dynamic keyframe selection mechanism and sparse cross-frame attention computation, computational efficiency is significantly improved while maintaining generation quality. The ASTA module achieves adaptive modeling of complex motion patterns through a saliency-guided local processing strategy, and its hierarchical processing architecture can flexibly adapt to the feature representation requirements of different resolution levels. The design of the gated fusion layer effectively balances spatial detail and temporal coherence, solving the problem of insufficient feature coupling in traditional hybrid architectures. Among them, dynamic keyframe selection is based on visual saliency (such as motion boundaries and texture changes) and learnable weights, replacing fixed rules; sparse attention only acts on the neighborhood of keyframes, reducing computational complexity from O(N^2) to O(N^2). 2 The computational complexity is reduced to O(kN); the gated fusion layer explicitly separates spatial reconstruction and temporal interpolation to avoid feature confusion. Experiments show that this method improves inference speed by more than 3 times compared to the baseline while maintaining generation quality.

[0146] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0147] It should be emphasized that, to further ensure the privacy and security of the aforementioned target animation, the target animation can also be stored in a blockchain node.

[0148] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0149] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware with computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When executed, the program can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).

[0150] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0151] Further reference Figure 3 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of an animation generation apparatus, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0152] like Figure 3 As shown, the animation generation device 300 described in this embodiment includes: an acquisition module 301, an extraction module 302, a processing module 303, a fusion module 304, an evaluation module 305, a decoding module 306, and an output module 307. Wherein:

[0153] The acquisition module 301 is used to acquire the keyframe sequence to be processed;

[0154] The extraction module 302 is used to perform feature extraction processing on the key frame sequence based on a preset input coding layer to obtain the corresponding multi-scale feature map;

[0155] Processing module 303 is used to perform keyframe selection and sparse attention calculation processing on the multi-scale feature map based on a preset adaptive sparse temporal attention module to obtain the corresponding attention features;

[0156] The fusion module 304 is used to fuse the multi-scale feature map and the attention feature based on a preset gated output layer to obtain the corresponding initial animation sequence;

[0157] Evaluation module 305 is used to evaluate the initial animation sequence based on a preset discriminator;

[0158] Decoding module 306 is used to decode the initial animation sequence based on a preset decoder to obtain the corresponding target animation if the initial animation sequence passes the evaluation.

[0159] Output module 307 is used to output the target animation.

[0160] In some optional implementations of this embodiment, the acquisition module 301 includes:

[0161] The receiving submodule is used to receive the initial keyframe sequence input by the user and to obtain the preset latent variable noise;

[0162] The first acquisition submodule is used to acquire the preset fusion strategy;

[0163] The first fusion submodule is used to fuse the initial keyframe sequence and the latent variable noise based on the fusion strategy to obtain the corresponding processing features;

[0164] The first determining submodule is used to use the processing features as the keyframe sequence.

[0165] In some optional implementations of this embodiment, the input encoding layer includes a shared-weight convolutional network; the extraction module 302 includes:

[0166] The first processing submodule is used to perform initial downsampling processing on the keyframe sequence based on the shared weight convolutional network to obtain the corresponding first feature map;

[0167] The second processing submodule is used to perform residual block processing on the first feature map to obtain the corresponding second feature map.

[0168] The third processing submodule is used to perform final downsampling processing on the second feature map to obtain the corresponding third feature map;

[0169] The second determining submodule is used to use the third feature map as the multi-scale feature map.

[0170] In some optional implementations of this embodiment, the adaptive sparse temporal attention module includes at least a saliency scorer, a keyframe selector, and a sparse temporal transformer; the processing module 303 includes:

[0171] The first calculation submodule is used to perform calculation processing on the multi-scale feature map based on the saliency scorer to obtain the corresponding saliency score result;

[0172] The third determination submodule is used to determine the target keyframes that meet the conditions based on the saliency score results and the preset dynamic threshold selection formula, using the keyframe selector.

[0173] The second calculation submodule is used to calculate the attention weights of the target keyframe based on the sparse temporal transformer to obtain the corresponding calculation results.

[0174] The fourth determination submodule is used to use the calculation result as the attention feature.

[0175] In some optional implementations of this embodiment, the fusion module 304 includes:

[0176] The third calculation submodule is used to calculate the gating coefficients corresponding to the multi-scale feature map;

[0177] The second acquisition submodule is used to acquire the feature fusion formula corresponding to the gate coefficient;

[0178] The second fusion submodule is used to fuse the multi-scale feature map and the attention feature based on the feature fusion formula to obtain the corresponding fused feature;

[0179] The fifth determining submodule is used to use the fusion features as the initial animation sequence.

[0180] In some optional implementations of this embodiment, the discriminator includes at least a temporal feature extractor, a temporal consistency evaluator, and an output layer; the evaluation module 305 includes:

[0181] The extraction submodule is used to extract features from the initial animation sequence based on the temporal feature extractor to obtain the corresponding temporal features;

[0182] The evaluation submodule is used to evaluate the time series features based on the time series consistency evaluator to obtain the corresponding time series artifact scores.

[0183] The scoring submodule is used to score the temporal features based on the output layer to obtain the corresponding authenticity score;

[0184] The determination submodule is used to determine that the initial animation sequence passes the evaluation if the temporal artifact score is greater than a first preset threshold and the realism score is greater than a second preset threshold; otherwise, it determines that the initial animation sequence fails the evaluation.

[0185] In some optional implementations of this embodiment, the decoding module 306 includes:

[0186] Call the submodule to invoke the preset decoder;

[0187] The fourth processing submodule is used to upsample the initial animation sequence based on the decoder to obtain the corresponding first animation data;

[0188] The rendering submodule is used to render the first animation data to obtain the corresponding second animation data;

[0189] The sixth determining submodule is used to use the second animation data as the target animation.

[0190] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed] for details. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.

[0191] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected via a system bus. It should be noted that only the computer device 4 with components 41-43 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0192] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.

[0193] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 4. Of course, the memory 41 may include both the internal storage unit and its external storage device of the computer device 4. In this embodiment, the memory 41 is typically used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for animation generation methods. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or will be output.

[0194] In some embodiments, the processor 42 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 42 is typically used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or to process data, for example, to execute computer-readable instructions for the animation generation method.

[0195] The network interface 43 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 4 and other electronic devices.

[0196] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the animation generation method described above.

[0197] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0198] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.

Claims

1. An animation generation method, characterized in that, Includes the following steps: Obtain the keyframe sequence to be processed; Based on a preset input coding layer, feature extraction processing is performed on the key frame sequence to obtain the corresponding multi-scale feature map; Based on the preset adaptive sparse temporal attention module, keyframe selection and sparse attention calculation are performed on the multi-scale feature map to obtain the corresponding attention features; The multi-scale feature map and the attention feature are fused based on a preset gated output layer to obtain the corresponding initial animation sequence; The initial animation sequence is evaluated based on a preset discriminator; If the initial animation sequence passes the evaluation, the corresponding target animation is obtained by decoding the initial animation sequence based on a preset decoder; The target animation is then output.

2. The animation generation method according to claim 1, characterized in that, The step of obtaining the keyframe sequence to be processed specifically includes: Receive the initial keyframe sequence input by the user, and obtain the preset latent variable noise; Obtain the preset fusion strategy; Based on the fusion strategy, the initial keyframe sequence and the latent variable noise are fused to obtain the corresponding processing features; The processing features are used as the keyframe sequence.

3. The animation generation method according to claim 1, characterized in that, The input coding layer includes a shared-weight convolutional network; the step of performing feature extraction processing on the keyframe sequence based on the preset input coding layer to obtain the corresponding multi-scale feature map specifically includes: The keyframe sequence is initially downsampled based on the shared weight convolutional network to obtain the corresponding first feature map. The first feature map is processed by residual block processing to obtain the corresponding second feature map; The second feature map is then subjected to a final downsampling process to obtain the corresponding third feature map; The third feature map is used as the multi-scale feature map.

4. The animation generation method according to claim 1, characterized in that, The adaptive sparse temporal attention module includes at least a saliency scorer, a keyframe selector, and a sparse temporal transformer; the step of performing keyframe selection and sparse attention calculation on the multi-scale feature map based on the preset adaptive sparse temporal attention module to obtain the corresponding attention features specifically includes: The multi-scale feature map is processed based on the saliency scorer to obtain the corresponding saliency score result; Based on the saliency score and the preset dynamic threshold selection formula, the target keyframes that meet the conditions are determined by the keyframe selector. The attention weights of the target keyframes are calculated based on the sparse temporal transformer to obtain the corresponding calculation results. The calculation result is used as the attention feature.

5. The animation generation method according to claim 1, characterized in that, The step of fusing the multi-scale feature map and the attention feature based on a preset gated output layer to obtain the corresponding initial animation sequence specifically includes: Calculate the gating coefficients corresponding to the multi-scale feature map; Obtain the feature fusion formula corresponding to the gate coefficient; The multi-scale feature map and the attention feature are fused based on the feature fusion formula to obtain the corresponding fused feature; The fusion feature is used as the initial animation sequence.

6. The animation generation method according to claim 1, characterized in that, The discriminator includes at least a temporal feature extractor, a temporal consistency evaluator, and an output layer; the step of evaluating the initial animation sequence based on the preset discriminator specifically includes: Based on the temporal feature extractor, feature extraction is performed on the initial animation sequence to obtain the corresponding temporal features; The temporal features are evaluated based on the temporal consistency evaluator to obtain the corresponding temporal artifact scores; The temporal features are scored based on the output layer to obtain the corresponding authenticity score; If the temporal artifact score is greater than a first preset threshold and the realism score is greater than a second preset threshold, then the initial animation sequence is determined to have passed the evaluation; otherwise, the initial animation sequence is determined to have failed the evaluation.

7. The animation generation method according to claim 1, characterized in that, The step of decoding the initial animation sequence based on a preset decoder to obtain the corresponding target animation specifically includes: Call the preset decoder; The initial animation sequence is upsampled based on the decoder to obtain the corresponding first animation data; The first animation data is rendered to obtain the corresponding second animation data; The second animation data is used as the target animation.

8. An animation generation device, characterized in that, include: The acquisition module is used to acquire the keyframe sequence to be processed; The extraction module is used to perform feature extraction processing on the key frame sequence based on a preset input encoding layer to obtain the corresponding multi-scale feature map; The processing module is used to perform keyframe selection and sparse attention calculation on the multi-scale feature map based on a preset adaptive sparse temporal attention module to obtain the corresponding attention features; The fusion module is used to fuse the multi-scale feature map and the attention feature based on a preset gated output layer to obtain the corresponding initial animation sequence; An evaluation module is used to evaluate the initial animation sequence based on a preset discriminator; A decoding module is used to decode the initial animation sequence based on a preset decoder to obtain the corresponding target animation if the initial animation sequence passes the evaluation. The output module is used to process the target animation for output.

9. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the animation generation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the animation generation method as described in any one of claims 1 to 7.