An AI-based game animation frame-by-frame progressive generation method
By using a dynamic imprinting reference length adaptive algorithm and hierarchical map construction, combined with feature deviation detection and compensation mechanisms, the problems of temporal consistency and multimodal collaboration in game animation generation are solved, achieving efficient and coherent long video generation of game animation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU HURRICANE NETWORK CO LTD
- Filing Date
- 2026-03-05
- Publication Date
- 2026-06-05
AI Technical Summary
Existing AI-based game animation generation methods struggle to guarantee temporal consistency and multimodal collaboration in long videos, leading to issues such as character pose breaks, abrupt changes in lighting and shadow styles, and mismatch between audio-visual physical attributes. Furthermore, current technologies are inefficient and cannot ensure the coherence of visual structure and narrative logic.
A dynamic tracing baseline length adaptive algorithm is used to extract steady-state multi-frame sequences and construct a hierarchical baseline tracing map. Feature fusion is performed through a multi-head cross-attention mechanism and cross-modal coding. Combined with feature deviation detection and compensation mechanisms, frame-level deviation repair is achieved. Furthermore, cross-chapter transition segments are generated through style genetic algorithm and causal inference to ensure the efficient continuity of long game animation videos.
It significantly improves the stability and efficiency of generating long game animation videos, effectively avoids style drift and plot breaks, and supports automated and progressive compositing of high-quality, large-scale long game animation videos.
Smart Images

Figure CN122156412A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of game animation technology, specifically an AI-based method for progressively generating game animation frames. Background Technology
[0002] With the rapid development of artificial intelligence technology, generative models (such as diffusion models and generative adversarial networks) are being applied in fields such as digital cultural and creative software, animation and game production engine software, and virtual reality processing software, providing efficient solutions for digital cultural product production, game and animation software development, and digital cultural and creative design. In the content production of home entertainment product software, educational and news-related cultural content industry software, and digital publishing software, AI-based automatic animation generation technology is also demonstrating enormous application potential.
[0003] Current AI-based game animation generation methods mostly use end-to-end models to directly synthesize short video clips, which makes it difficult to guarantee the temporal consistency and multimodal collaboration of long videos.
[0004] Especially in the process of splicing multiple short videos, the lack of effective modeling and constraints on preceding content often leads to problems such as broken character poses, abrupt changes in lighting and shadow styles, and mismatch between audio-visual physical attributes. Existing technologies usually rely on global regeneration or simple interpolation for connection repair, which is not only inefficient but also fails to ensure the coherence of visual structure and narrative logic.
[0005] Therefore, an AI-based method for progressive generation of game animation frames is provided, which can improve the efficiency and continuity of short video splicing during the game short video splicing process. Summary of the Invention
[0006] To address the aforementioned technical problems, the present invention aims to provide an AI-based method for progressively generating frame-by-frame game animations, which can improve the efficiency and continuity of short video splicing during the game short video splicing process.
[0007] To achieve the above objectives, the present invention provides the following technical solution: an AI-based method for progressively generating frame-by-frame game animations, the method comprising: From the end of the short video generated by the preceding AI, a steady-state multi-frame sequence is extracted using a dynamic imprinting reference length adaptive algorithm, which serves as the reference imprinting frame sequence. Extract the four-dimensional core features of the reference imprint frame sequence; Based on the aforementioned four-dimensional core features, a hierarchical benchmark imprint map is constructed through a multi-head cross-attention mechanism, cross-modal coding, and temporal semantic-sentiment alignment. Using the aforementioned baseline imprint as a constraint, the AI generates a target short video by adjusting the imprint loss function; Based on the benchmark tracing map, the connection region between the target short video and the preceding short video is determined, and frame-level deviation detection is performed on the connection region based on the feature deviation detection mechanism. When a deviation is detected, local feature compensation is performed on the target short video through a compensation plan matching mechanism; Repeat the above steps to progressively stitch together multiple AI-generated short videos to obtain a long game animation video.
[0008] Preferably, the four-dimensional core features include dynamic skeleton trajectory features, color features, lighting features, and texture features, which are fused based on semantic action topology templates and causal inference effect features; wherein the extraction of the dynamic skeleton trajectory features includes: Obtain the plot semantic information associated with the reference imprint frame sequence; Based on the aforementioned plot semantic information, identify character-type elements and their semantic action categories in the reference imprint frame sequence; Based on different semantic action categories, a predefined topology template is adaptively selected and loaded; the topology template defines the main degrees of freedom of motion of the core joints, the motion dependencies between joints, and the allowed range of motion in the semantic action category. Under the constraints of the topological template, based on the reference imprint frame sequence, the coordinates of key points of the character skeleton and joint angle parameters are extracted to generate a semantic character skeleton feature vector. Based on the semantic action category and the character skeleton feature vector, the type of special effect and its motion parameters that the action is expected to trigger are inferred in reverse, forming the corresponding special effect motion features; the special effect types include character-derived special effects and environmental feedback special effects; The semantic character skeleton feature vector is fused with the special effects motion features inferred through causal correlation to obtain the corresponding dynamic skeleton trajectory features.
[0009] Preferably, from the end of the preceding AI-generated short video, a steady-state multi-frame sequence is extracted using a dynamic imprinting reference length adaptive algorithm, which serves as the reference imprinting frame sequence, including: Starting from the last frame of the short video generated by the preceding AI, calculate the inter-frame change rate corresponding to the four-dimensional core features frame by frame. Based on the inter-frame change rate of the four-dimensional core features corresponding to each frame of the short video generated by the preceding AI, and the change rate threshold corresponding to each four-dimensional core feature, the relationship between the inter-frame change rate and the corresponding change rate threshold is compared. When the inter-frame change rate of each frame is less than the corresponding change rate threshold, the current frame is marked as a steady-state frame. Slide a detection window of length N from the end of the video towards the beginning. When all frames in a certain detection window are first detected as steady-state frames, the video segment corresponding to the detection window is determined to be a sequence that has entered the multimodal fidelity steady state, and this continuous frame segment is output as the reference imprint frame sequence. If no detection window that meets the conditions is found after traversing the entire video sequence, the preset number of frames at the end of the video is directly extracted as the reference imprint frame sequence.
[0010] Preferably, based on the four-dimensional core features, a hierarchical baseline imprint map is constructed through a multi-head cross-attention mechanism, cross-modal coding, and temporal semantic-sentiment alignment, including: Based on the preceding AI-generated short video, corresponding rubbing carrier category features and rubbing content category features are generated; the rubbing carrier category features include sound effect features and image physical attribute features; the rubbing content features include plot semantic features and character emotional features. Based on the four-dimensional core features, multi-dimensional feature fusion is performed through a multi-head cross attention mechanism. During the fusion process, topological invariant constraints with the skeleton joint point connected graph as the object are introduced to obtain the benchmark visual-dynamic feature set and form the corresponding basic imprint layer. Based on the imprint carrier class features and the basic imprint layer, feature alignment and association fusion are performed through a cross-modal coding mechanism to establish a corresponding mapping relationship. A physical simulation mechanism is then used to verify and optimize the mapping relationship to obtain a carrier-representation joint feature representation, forming a corresponding carrier constraint layer. Based on the imprint content category features and the carrier-representation joint feature representation, a gated recurrent unit network is used to perform temporal semantic-emotion alignment fusion, and the fusion process is constrained by constructing and injecting an emotion-event causal graph model to obtain a plot-character collaborative feature sequence, forming a corresponding content control layer. Based on the basic imprinting layer, carrier constraint layer, and content control layer, they are integrated into a unified multimodal tensor structure through a hierarchical encoder to form a preliminary hierarchical imprinting map. Based on the preliminary hierarchical tracing map, the corresponding cross-level causal explanatory power index is calculated. It is determined whether the cross-level causal explanatory power index is lower than the preset index threshold. If not, the final baseline tracing map is directly output. Otherwise, the level with the weakest causal transmission is located, and the feature reconstruction process for that level is triggered. After iterative optimization until the index meets the standard, the final baseline tracing map is output.
[0011] Preferably, based on the reference imprint map, the connection region between the target short video and the preceding short video is determined, and frame-level deviation detection is performed on the connection region based on the feature deviation detection mechanism, including: Based on the benchmark tracing map, the connection area between the target short video and the preceding AI-generated short video is determined; Based on the connection area between the target short video and the preceding AI-generated short video, the four-dimensional core features corresponding to all frames within the connection area are extracted to form a feature dataset; wherein the four-dimensional core features of each frame correspond to a data point in the dataset. Based on the isolated forest algorithm, data points in the feature dataset are mapped to a high-dimensional space. The feature dimension is randomly selected and the data points are randomly divided. An isolated forest containing multiple isolated trees is constructed iteratively. During the construction of each isolated tree, the path length from the root to the leaf node of each data point is recorded. For each data point, calculate its path length in all isolated trees and take the average as the average path length of that data point. If the average path length of a data point is less than a preset average path length threshold, it is determined that the frame corresponding to the data point has a feature deviation; if there are a preset number or more frames with feature deviations within the connection detection area, it is determined that the connection area has a deviation overall.
[0012] Preferably, when a deviation is detected, local feature compensation is performed on the target short video through a compensation plan matching mechanism, including: When a deviation is detected in a frame of the connecting region or a deviation is detected in the whole, deviation pattern information is extracted and encoded. The pattern information includes the feature dimension combination where the deviation is located, the time distribution of the deviation, and the intensity vector of the deviation. The encoded deviation pattern information is matched and queried against a historical compensation strategy library, which is automatically constructed from each successful compensation. Each strategy records the correspondence between deviation pattern information and effective compensation operation sequence. If a match is successful, the compensation operation sequence in the historical strategy is directly invoked and fine-tuned to generate a compensation plan for the current deviation; otherwise, the generation of a compensation plan based on causal reasoning is initiated. Based on the generated compensation plan, compensation operations are performed on the deviation frames of the target short video.
[0013] Preferably, the method further includes: Based on the final generated long-form game animation video, the story chapters of the long-form game animation video are identified and divided; for each story chapter, based on the baseline tracing map corresponding to the story chapter, it is decoupled into style genes and content genes and stored in a structured manner to form a chapter-level style gene library. When cross-chapter splicing is required, a style genetic algorithm is launched based on the chapter-level style gene library to output a set of offspring style gene sequences that are smoothly and gradually change in style and adapt to the current context in content as a reference for inheritance and transition. Using the inherited transition rubbing benchmark as a temporal constraint, a corresponding transition rubbing frame sequence is generated based on the frame-by-frame animation generation model. A pre-defined narrative coherence discriminator is run in parallel. The narrative coherence discriminator is based on an emotion-event causal graph model and determines in real time whether the generated transition frame sequence constitutes a credible causal bridge from the previous chapter to the current chapter in terms of plot logic. If not, a dynamic condition injection and local iterative regeneration process based on a causal violation heatmap is triggered to accurately correct and optimize the causal logic node with the lowest confidence in the transition frame sequence. Conversely, the transition frame sequence is inserted between chapters to complete the splicing and output.
[0014] Compared with the prior art, the beneficial effects of the present invention are: By employing a dynamic imprinting baseline length adaptive algorithm, multimodal steady-state frame sequences are automatically identified from the end of preceding short videos. Combined with plot semantic information, action topology templates are loaded, and causal inference effects are integrated to generate four-dimensional core features including dynamic skeleton trajectory (structure + semantics), color, lighting, and texture. This effectively avoids prior biases introduced by transitional or noisy frames, significantly improving the stability and semantic consistency of baseline extraction, and providing high-fidelity, context-aligned starting anchors for subsequent progressive generation.
[0015] A three-layer collaborative baseline topology is constructed (the foundation layer ensures topological coherence, the carrier layer ensures audiovisual and physical consistency, and the content layer maintains narrative logic consistency). Topological invariant constraints, physical direction verification, and emotion-event causal graphs are embedded during the fusion process. Cross-layer causal explanatory power indicators are introduced to drive local re-creation at the distortion level. While ensuring the quality of multimodal generation, structural, physical, and semantic consistency constraints are achieved, and the computational overhead of optimization is significantly reduced, improving the stability and efficiency of long-form video generation.
[0016] Frame-level deviation localization is achieved based on unsupervised anomaly detection, and precise local repair is achieved by combining historical compensation strategy matching and causal reasoning generation. Furthermore, chapter-level imprint maps are decoupled into style genes and content genes, and cross-chapter transition segments are generated through style genetic evolution and causal bridging discrimination. This significantly improves the visual and logical coherence of segment connections, effectively suppresses style drift and plot breaks in long-term generation, and supports automated and progressive compositing of high-quality, large-scale game animation videos. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0018] Figure 1 This is a schematic diagram of a frame-by-frame progressive generation method for game animation based on AI. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be described in detail below. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other implementation methods obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0020] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0021] Example 1 like Figure 1 As shown in the figure, this embodiment discloses an AI-based method for progressively generating game animation frames, the method comprising: From the end of the short video generated by the preceding AI, a steady-state multi-frame sequence is extracted using a dynamic imprinting reference length adaptive algorithm, which serves as the reference imprinting frame sequence. It should be noted that, at the end of the short video generated by the preceding AI, a steady-state multi-frame sequence is extracted using a dynamic imprinting reference length adaptive algorithm. This reference imprinting frame sequence includes: Starting from the last frame of the short video generated by the preceding AI, calculate the inter-frame change rate corresponding to the four-dimensional core features frame by frame. Based on the inter-frame change rate of the four-dimensional core features corresponding to each frame of the short video generated by the preceding AI, and the change rate threshold corresponding to each four-dimensional core feature, the relationship between the inter-frame change rate and the corresponding change rate threshold is compared. When the inter-frame change rate of each frame is less than the corresponding change rate threshold, the current frame is marked as a steady-state frame. Slide a detection window of length N from the end of the video towards the beginning. When all frames in a certain detection window are first detected as steady-state frames, the video segment corresponding to the detection window is determined to be a sequence that has entered the multimodal fidelity steady state, and this continuous frame segment is output as the reference imprint frame sequence. If no detection window that meets the conditions is found after traversing the entire video sequence, the preset number of frames at the end of the video is directly extracted as the reference imprint frame sequence.
[0022] In detail, the calculation logic for the inter-frame change rate of the four-dimensional core features is as follows: To address the rate of change of dynamic skeleton trajectory features, for the feature vector of frame t, the cosine similarity between it and the vector of frame (t-1) is calculated, and (1 - cosine similarity) is used as the rate of change. For structured features like skeleton trajectories, similarity is a better measure of the continuity of motion patterns and the stability of topological structures than absolute differences, and it is more sensitive to subtle but crucial joint vibrations.
[0023] For the rate of change of color / lighting / texture features, a weighted Manhattan distance is used. For example, for color features (statistical moments in the LAB space), the mean component representing the dominant hue is given a higher weight, while the secondary skewness components are given a lower weight, and then a weighted sum is performed to obtain the overall rate of change.
[0024] The logic for setting the rate of change threshold for each feature is as follows: Take the entire video sequence (or the first 1 / 3 of the sequence), calculate the rate of change of each four-dimensional feature across all consecutive frames, resulting in four rate of change sequences. For each change sequence, calculate its corresponding median and interquartile range. The rate of change threshold = median + interquartile range To make the sensitivity coefficient adjustable, the median and interquartile range are used instead of the mean and standard deviation, thus avoiding the destructive impact of a few extreme abrupt frames (such as scene transitions) in the video on the threshold calculation.
[0025] Specifically, the detection window length N is adaptively determined. N is not a fixed value, but is inversely proportional to the average inter-frame change rate of the video. The size of N is selected as the maximum value between the minimum window length and the function value of the function formula that is inversely proportional to the average inter-frame change rate. This allows for the use of a longer window to ensure the reliability of steady state for videos with smooth dynamics, and a shorter window to avoid being too strict for videos with drastic dynamics, thus improving the scene adaptability of the algorithm.
[0026] In this embodiment, the sliding of the detection window is not gradual, but rather a jump sliding window strategy is used for searching. Specifically, when at least one frame within the detection window is marked as a non-stationary frame, the frame with the highest proportion of maximum feature change rates exceeding a threshold is found among all non-stationary frames within the window. Based on the feature change rates corresponding to this frame and the corresponding thresholds, the corresponding instability score is calculated. The jump step size is obtained by non-linear mapping of the instability score. Based on the jump step size, the starting position of the new window after the jump is calculated, ensuring that the starting position of the new window is greater than or equal to zero, i.e., not exceeding the beginning of the video. If jumping would cause an out-of-bounds movement, the window is adjusted to N frames starting from the beginning of the video.
[0027] To avoid missing potential steady-state windows due to excessive jumps, the state of the frame preceding the jump start point is checked after the jump is completed. If that frame is a steady-state frame, the frame is backtracked by 1 frame for re-evaluation.
[0028] Extract the four-dimensional core features of the reference imprint frame sequence; the four-dimensional core features include dynamic skeleton trajectory features, color features, light and shadow features and texture features, which are fused based on semantic action topology templates and causal inference effect features; The extraction of the dynamic skeleton trajectory features includes: Obtain the plot semantic information associated with the reference imprint frame sequence; Based on the aforementioned plot semantic information, character-type elements and their performing semantic action categories are identified in the reference imprint frame sequence. Specifically, this involves not only extracting keywords from the text prompts input during video generation but also analyzing preceding and following video segments and even extracting action intentions from potentially related scripts or storyboards. A knowledge base mapping semantic tags to visual actions is established. The action recognition results from text prompts and visual analysis are weighted and fused based on confidence levels to ultimately determine a primary semantic action category and several auxiliary semantic tags. This approach ensures that action recognition combines the accuracy of textual analysis with the robustness of visual analysis.
[0029] Based on different semantic action categories, predefined topology templates are adaptively selected and loaded. These templates define the primary degrees of freedom (DOF) of motion for core joints, the motion dependencies between joints, and the allowed range of motion for each semantic action category. Specifically, a parameterized topology template library is predefined. Each template not only defines joints and connections but, more importantly, defines: core DEF weights, marking joints that play a decisive role in the action, such as the shoulder, elbow, and wrist in a throwing motion; a motion dependency matrix, defining the motion coupling between joints in matrix form, such as how the knee and ankle should coordinate during hip rotation; and an allowed range of motion envelope, defining the reasonable angle and positional variation range for each joint under a specific action. After loading the basic template based on the identified semantic action category, it is not used directly but rather fine-tuned according to the specific body proportions and movement amplitude of the character in the current sequence.
[0030] Under the constraints of the topological template, based on the reference imprint frame sequence, the coordinates of key points and joint angle parameters of the character skeleton are extracted to generate a semantically meaningful character skeleton feature vector. Specifically, firstly, an initial joint coordinate is obtained using a basic pose estimation network. Then, the initial estimate is projected onto the loaded topological template constraint space for optimization. For example, if the initial estimate yields an elbow joint angle of... This exceeds the definition of the punch template. If the range is specified, the algorithm will correct it to... The wrist position is adjusted based on the dependency matrix. This significantly improves the physical plausibility and stability of pose estimation in complex scenarios. The extracted features are not just corrected coordinates, but are transformed into higher-level parameters with clear physical and semantic meanings, such as normalized joint angle sequences, motion trajectory curvature of core joints, action phase markers, and energy change curves. The action phases are automatically divided into preparation, exertion, or following phases based on the motion velocity curves. The final feature vector is a combination of these high-level parameters.
[0031] Based on the semantic action category and the character skeleton feature vector, the type of special effect and its motion parameters that the action is expected to trigger are inferred in reverse, forming the corresponding special effect motion features; the special effect types include character-derived special effects and environmental feedback special effects; The semantically encoded character skeleton feature vector is fused with the special effects motion features inferred through causal correlation to obtain the corresponding dynamic skeleton trajectory features. Specifically, based on the causal inference results, the special effects features are aligned to the triggering action keyframes on the time axis and to the relevant joint coordinate system in space. A lightweight fusion network is designed, whose inputs are the semantically encoded skeleton feature sequence and the special effects feature sequence. The final output feature is a temporally continuous, multimodal feature sequence.
[0032] Based on the aforementioned four-dimensional core features, a hierarchical benchmark imprint map is constructed through a multi-head cross-attention mechanism, cross-modal coding, and temporal semantic-sentiment alignment. It should be noted that, based on the aforementioned four-dimensional core features, a hierarchical baseline imprint map is constructed through a multi-head cross-attention mechanism, cross-modal coding, and temporal semantic-sentiment alignment, including: Based on the preceding AI-generated short video, corresponding rubbing carrier category features and rubbing content category features are generated; the rubbing carrier category features include sound effect features and image physical attribute features; the rubbing content features include plot semantic features and character emotional features. Based on the four-dimensional core features, multi-dimensional feature fusion is performed through a multi-head cross attention mechanism. During the fusion process, topological invariant constraints with the skeleton joint point connected graph as the object are introduced to obtain the benchmark visual-dynamic feature set and form the corresponding basic imprint layer. Based on the imprint carrier class features and the basic imprint layer, feature alignment and association fusion are performed through a cross-modal coding mechanism to establish a corresponding mapping relationship. A physical simulation mechanism is then used to verify and optimize the mapping relationship to obtain a carrier-representation joint feature representation, forming a corresponding carrier constraint layer. Based on the imprint content category features and the carrier-representation joint feature representation, a gated recurrent unit network is used to perform temporal semantic-emotion alignment fusion, and the fusion process is constrained by constructing and injecting an emotion-event causal graph model to obtain a plot-character collaborative feature sequence, forming a corresponding content control layer. Based on the basic imprinting layer, carrier constraint layer, and content control layer, they are integrated into a unified multimodal tensor structure through a hierarchical encoder to form a preliminary hierarchical imprinting map. Based on the preliminary hierarchical tracing map, the corresponding cross-level causal explanatory power index is calculated. It is determined whether the cross-level causal explanatory power index is lower than the preset index threshold. If not, the final baseline tracing map is directly output. Otherwise, the level with the weakest causal transmission is located, and the feature reconstruction process for that level is triggered. After iterative optimization until the index meets the standard, the final baseline tracing map is output.
[0033] Using the aforementioned baseline imprint as a constraint, the AI generates a target short video by adjusting the imprint loss function; Based on the benchmark tracing map, the connection region between the target short video and the preceding short video is determined, and frame-level deviation detection is performed on the connection region based on the feature deviation detection mechanism. It should be noted that, based on the benchmark tracing map, the connection region between the target short video and the preceding short video is determined, and frame-level deviation detection is performed on the connection region based on the feature deviation detection mechanism, including: Based on the benchmark tracing map, the connection area between the target short video and the preceding AI-generated short video is determined; Based on the connection area between the target short video and the preceding AI-generated short video, the four-dimensional core features corresponding to all frames within the connection area are extracted to form a feature dataset; wherein the four-dimensional core features of each frame correspond to a data point in the dataset. Based on the isolated forest algorithm, data points in the feature dataset are mapped to a high-dimensional space. The feature dimension is randomly selected and the data points are randomly divided. An isolated forest containing multiple isolated trees is constructed iteratively. During the construction of each isolated tree, the path length from the root to the leaf node of each data point is recorded. For each data point, calculate its path length in all isolated trees and take the average as the average path length of that data point. If the average path length of a data point is less than a preset average path length threshold, it is determined that the frame corresponding to the data point has a feature deviation; if there are a preset number or more frames with feature deviations within the connection detection area, it is determined that the connection area has a deviation overall.
[0034] The detailed, dynamically adapted high-dimensional mapping breaks through the limitations of fixed dimensions. During implementation, the discriminative power of each dimension in the feature dataset (such as the variance ratio of color and motion dimensions) is first statistically analyzed. Dimensions with high discriminative power are prioritized for random partitioning, and the dimension combination is dynamically adjusted in each iteration to enhance the capture of minor feature deviations in AI-generated video frames (such as subtle color shifts and slight changes in motion trajectory). The intelligent hierarchical judgment mechanism introduces a weight factor when calculating path length (giving higher weight to feature data points of connecting critical frames, i.e., the feature data points of the preceding end frame and the target first frame), improving the sensitivity of key frame deviation recognition. The dynamic threshold is generated by training normal connecting samples of the same type of video in the benchmark tracing map and is adaptively adjusted according to the video type. Finally, through the hierarchical logic of single-frame deviation initial judgment (average path length < dynamic threshold) - multi-frame cumulative verification (number of deviation frames ≥ preset type threshold), it avoids misjudgment caused by random fluctuations in a single frame and ensures accurate recognition of overall deviations in the connecting area (such as feature changes in consecutive multiple frames). At the same time, it can output a single-frame deviation degree score to support subsequent graded correction.
[0035] When a deviation is detected, local feature compensation is performed on the target short video through a compensation plan matching mechanism; When a deviation is detected, local feature compensation is performed on the target short video through a compensation plan matching mechanism, including: When a deviation is detected in a frame of the connecting region or a deviation is detected in the whole, deviation pattern information is extracted and encoded. The pattern information includes the feature dimension combination where the deviation is located, the time distribution of the deviation, and the intensity vector of the deviation. The encoded deviation pattern information is matched and queried against a historical compensation strategy library, which is automatically constructed from each successful compensation. Each strategy records the correspondence between deviation pattern information and effective compensation operation sequence. If a match is successful, the compensation operation sequence in the historical strategy is directly invoked and fine-tuned to generate a compensation plan for the current deviation; otherwise, the generation of a compensation plan based on causal reasoning is initiated. Based on the generated compensation plan, compensation operations are performed on the deviation frames of the target short video.
[0036] In detail, the deviation pattern information needs to be structured and encoded to form a combination of feature dimensions. That is, the detected deviations are modeled at multiple granularities to generate a 4D Boolean mask, corresponding to whether the four core features (skeleton / color / lighting / texture) are distorted. The time distribution is to construct a time heatmap to represent the temporal concentration of deviations within the connecting window. The intensity vector is to calculate the L2 deviation intensity of each distortion feature and normalize it to a unit vector. The three are then fused into a deviation pattern fingerprint.
[0037] The historical compensation strategy database is a dynamically growing compensation knowledge graph. Each node stores the deviation pattern fingerprint of a successful compensation case and the corresponding compensation operation sequence. Represents policy similarity, weights During matching, a graph attention network (GAT) is used to embed the fingerprint of the current bias pattern, and a k-nearest neighbor subgraph search is performed in the graph G to return the top-K most similar strategies. If the maximum similarity is equal to the preset similarity threshold, the match is considered successful, the corresponding compensation operation sequence is called, and the parameters are fine-tuned.
[0038] If there are not enough similar historical strategies, the causal compensation planner is activated, and the specific process is as follows: Construct a lightweight causal graph: nodes include deviation features, potential causes, and actionable actions; perform counterfactual queries: if action a (e.g., reconstraining the skeleton) is performed, does the expected value of deviation feature d regress to the tolerance range? Calculate the expected repair gain for each action using a causal effect estimation model (e.g., a do-calculus approximation based on TARNet); generate a minimum sequence of actions through a greedy search that minimizes the joint causal effect threshold. Output this sequence as the new compensation plan.
[0039] Repeat the above steps to progressively stitch together multiple AI-generated short videos to obtain a long game animation video.
[0040] The method further includes: Based on the final generated long-form game animation video, the story chapters of the long-form game animation video are identified and divided; for each story chapter, based on the baseline tracing map corresponding to the story chapter, it is decoupled into style genes and content genes and stored in a structured manner to form a chapter-level style gene library. When cross-chapter splicing is required, a style genetic algorithm is launched based on the chapter-level style gene library to output a set of offspring style gene sequences that are smoothly and gradually change in style and adapt to the current context in content as a reference for inheritance and transition. Using the inherited transition rubbing benchmark as a temporal constraint, a corresponding transition rubbing frame sequence is generated based on the frame-by-frame animation generation model. A pre-defined narrative coherence discriminator is run in parallel. The narrative coherence discriminator is based on an emotion-event causal graph model and determines in real time whether the generated transition frame sequence constitutes a credible causal bridge from the previous chapter to the current chapter in terms of plot logic. If not, a dynamic condition injection and local iterative regeneration process based on a causal violation heatmap is triggered to accurately correct and optimize the causal logic node with the lowest confidence in the transition frame sequence. Conversely, the transition frame sequence is inserted between chapters to complete the splicing and output.
[0041] In detail, the segmentation of the plot into chapters is based on the detection of language transitions in the plot text, for example, using a pre-trained plot segmentation model (specifically a chapter boundary classifier based on BERT) to divide the long video into N plot chapters. .
[0042] The benchmark rubbing atlas is decoupled into style genes and content genes, for each chapter Decouple the corresponding baseline rubbings: style genes, extract the topological invariants (character skeleton) from its basic rubbing layers. number , The color-light and shadow joint distribution (HSV histogram + main illumination direction vector) is used. For content genes, the causal graph of plot events (nodes = key events, edges = temporal sequence + causal intensity) and emotional state trajectory (discrete labels + continuous intensity) in the content control layer are extracted. The content genes and style genes are then structured and stored in a chapter-level style gene library.
[0043] The style genetic algorithm is started to generate offspring style genes, when splicing chapters and At that time, by chapter Style genes and content genes, acting as parents, perform constrained crossover, allowing mixing only within the topological homeomorphic space, such as... Consistency is required; otherwise, crossover will fail. A small perturbation is applied to the color-light parameters, but the physical rationality of the carrier constraint layer must still be satisfied after mutation. The M group of offspring style genes is output to form a style gradient sequence.
[0044] Transitional imprinting frame sequence generation involves: using the optimal offspring style gene as a temporal style constraint, inputting it into the frame-by-frame animation generation model, and simultaneously injecting chapters... The ending context and chapter Given the initial context, generate a transition frame imprint sequence of length T.
[0045] The narrative coherence discriminator, based on an emotion-event causal graph model, performs the following operations: constructs a counterfactual query; if no chapter is found... The ending event Can the transition frame imprint sequence reasonably lead into the chapter? The start event This query is used to measure chapters. The ending event This supports the subsequent narrative logic. Parallel computation is performed on each frame t of the narrative content to calculate the causal violation score for that frame. The score calculation logic is the difference between the baseline probability and the counterfactual probability, based on the existence of... Derivation of the Temporal Transition Frame Imprint Sequence Based on the probability of non-existence, subtract the probability of non-existence. Derivation of the Temporal Transition Frame Imprint Sequence The probability is used to calculate the causal violation score of the frame. The higher the score, the better the frame's narrative logic is aligned with the causal violation score. The stronger the dependence, the weaker the logical coherence.
[0046] Based on the causal violation scores of all frames, a causal violation heatmap is generated, visually representing the logical coherence of each frame through color intensity. Then, frames are sorted from highest to lowest score, and the top K frames are marked as weak points in the narrative logic. For these K logically weak frames, a dynamic condition injection mechanism is activated: during the generation of these weak frames, a condition is forcibly injected... The causal effect vector, thereby strengthening This provides logical support for the subsequent narrative; simultaneously, it freezes the parameters of other non-weak frames to avoid interfering with the existing stable narrative content. After injection, a local iterative regeneration operation is performed on these weak frames. The causal violation score is recalculated for the regenerated weak frames, and the injection-regeneration-scoring process is repeated until the causal violation scores of all weak frames are lower than the preset threshold. At this point, the narrative logic is deemed to have met the coherence standard, and iterative optimization stops.
[0047] Insert the optimized transition frame imprint sequence into the chapter. and Between. Output the final long video.
[0048] Optionally, in this embodiment, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0049] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0050] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more electronic devices to execute all or part of the steps of the methods described in the various embodiments of this application.
[0051] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0052] In the several embodiments provided in this application, it should be understood that the disclosed application can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of units or modules may be electrical or other forms.
[0053] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0054] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0055] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for progressively generating frame-by-frame game animation based on AI, characterized in that, The method includes: The method includes: From the end of the short video generated by the preceding AI, a steady-state multi-frame sequence is extracted using a dynamic imprinting reference length adaptive algorithm, which serves as the reference imprinting frame sequence. Extract the four-dimensional core features of the reference imprint frame sequence; Based on the aforementioned four-dimensional core features, a hierarchical benchmark imprint map is constructed through a multi-head cross-attention mechanism, cross-modal coding, and temporal semantic-sentiment alignment. Using the aforementioned baseline imprint as a constraint, the AI generates a target short video by adjusting the imprint loss function; Based on the benchmark tracing map, the connection region between the target short video and the preceding short video is determined, and frame-level deviation detection is performed on the connection region based on the feature deviation detection mechanism. When a deviation is detected, local feature compensation is performed on the target short video through a compensation plan matching mechanism; Repeat the above steps to progressively stitch together multiple AI-generated short videos to obtain a long game animation video.
2. The AI-based game animation frame-by-frame progressive generation method according to claim 1, characterized in that, The four-dimensional core features include dynamic skeleton trajectory features, color features, light and shadow features, and texture features, which are formed by fusing semantic action topology templates and causal inference special effects features. The extraction of the dynamic skeleton trajectory features includes: Obtain the plot semantic information associated with the reference imprint frame sequence; Based on the aforementioned plot semantic information, identify character-type elements and their semantic action categories in the reference imprint frame sequence; Based on different semantic action categories, a predefined topology template is adaptively selected and loaded; the topology template defines the main degrees of freedom of motion of the core joints, the motion dependencies between joints, and the allowed range of motion in the semantic action category. Under the constraints of the topological template, based on the reference imprint frame sequence, the coordinates of key points of the character skeleton and joint angle parameters are extracted to generate a semantic character skeleton feature vector. Based on the semantic action category and the character skeleton feature vector, the type of special effect and its motion parameters that the action is expected to trigger are inferred in reverse, forming the corresponding special effect motion features; the special effect types include character-derived special effects and environmental feedback special effects; The semantic character skeleton feature vector is fused with the special effects motion features inferred through causal correlation to obtain the corresponding dynamic skeleton trajectory features.
3. The AI-based game animation frame-by-frame progressive generation method according to claim 2, characterized in that, From the end of the short video generated by preceding AI, a steady-state multi-frame sequence is extracted using a dynamic imprinting reference length adaptive algorithm. This reference imprinting frame sequence includes: Starting from the last frame of the short video generated by the preceding AI, calculate the inter-frame change rate corresponding to the four-dimensional core features frame by frame. Based on the inter-frame change rate of the four-dimensional core features corresponding to each frame of the short video generated by the preceding AI, and the change rate threshold corresponding to each four-dimensional core feature, the relationship between the inter-frame change rate and the corresponding change rate threshold is compared. When the inter-frame change rate of each frame is less than the corresponding change rate threshold, the current frame is marked as a steady-state frame. Slide a detection window of length N from the end of the video towards the beginning. When all frames in a certain detection window are first detected as steady-state frames, the video segment corresponding to the detection window is determined to be a sequence that has entered the multimodal fidelity steady state, and this continuous frame segment is output as the reference imprint frame sequence. If no detection window that meets the conditions is found after traversing the entire video sequence, the preset number of frames at the end of the video is directly extracted as the reference imprint frame sequence.
4. The AI-based game animation frame-by-frame progressive generation method according to claim 3, characterized in that, Based on the aforementioned four-dimensional core features, a hierarchical benchmark imprint map is constructed through a multi-head cross-attention mechanism, cross-modal encoding, and temporal semantic-sentiment alignment, including: Based on the preceding AI-generated short video, corresponding rubbing carrier category features and rubbing content category features are generated; the rubbing carrier category features include sound effect features and image physical attribute features; the rubbing content features include plot semantic features and character emotional features. Based on the four-dimensional core features, multi-dimensional feature fusion is performed through a multi-head cross attention mechanism. During the fusion process, topological invariant constraints with the skeleton joint point connected graph as the object are introduced to obtain the benchmark visual-dynamic feature set and form the corresponding basic imprint layer. Based on the imprint carrier class features and the basic imprint layer, feature alignment and association fusion are performed through a cross-modal coding mechanism to establish a corresponding mapping relationship. A physical simulation mechanism is then used to verify and optimize the mapping relationship to obtain a carrier-representation joint feature representation, forming a corresponding carrier constraint layer. Based on the imprint content category features and the carrier-representation joint feature representation, a gated recurrent unit network is used to perform temporal semantic-emotion alignment fusion, and the fusion process is constrained by constructing and injecting an emotion-event causal graph model to obtain a plot-character collaborative feature sequence, forming a corresponding content control layer. Based on the basic imprinting layer, carrier constraint layer, and content control layer, they are integrated into a unified multimodal tensor structure through a hierarchical encoder to form a preliminary hierarchical imprinting map. Based on the preliminary hierarchical tracing map, the corresponding cross-level causal explanatory power index is calculated. It is determined whether the cross-level causal explanatory power index is lower than the preset index threshold. If not, the final baseline tracing map is directly output. Otherwise, the level with the weakest causal transmission is located, and the feature reconstruction process for that level is triggered. After iterative optimization until the index meets the standard, the final baseline tracing map is output.
5. The AI-based game animation frame-by-frame progressive generation method according to claim 4, characterized in that, Based on the baseline tracing map, the connection region between the target short video and the preceding short video is determined, and frame-level deviation detection is performed on the connection region based on the feature deviation detection mechanism, including: Based on the benchmark tracing map, the connection area between the target short video and the preceding AI-generated short video is determined; Based on the connection area between the target short video and the preceding AI-generated short video, the four-dimensional core features corresponding to all frames within the connection area are extracted to form a feature dataset; wherein the four-dimensional core features of each frame correspond to a data point in the dataset. Based on the isolated forest algorithm, data points in the feature dataset are mapped to a high-dimensional space. The feature dimension is randomly selected and the data points are randomly divided. An isolated forest containing multiple isolated trees is constructed iteratively. During the construction of each isolated tree, the path length from the root to the leaf node of each data point is recorded. For each data point, calculate its path length in all isolated trees and take the average as the average path length of that data point. If the average path length of a data point is less than a preset average path length threshold, it is determined that the frame corresponding to the data point has a feature deviation; if there are a preset number or more frames with feature deviations within the connection detection area, it is determined that the connection area has a deviation overall.
6. The AI-based game animation frame-by-frame progressive generation method according to claim 5, characterized in that, When a deviation is detected, local feature compensation is performed on the target short video through a compensation plan matching mechanism, including: When a deviation is detected in a frame of the connecting region or a deviation is detected in the whole, deviation pattern information is extracted and encoded. The pattern information includes the feature dimension combination where the deviation is located, the time distribution of the deviation, and the intensity vector of the deviation. The encoded deviation pattern information is matched and queried against a historical compensation strategy library, which is automatically constructed from each successful compensation. Each strategy records the correspondence between deviation pattern information and effective compensation operation sequence. If a match is successful, the compensation operation sequence in the historical strategy is directly invoked and fine-tuned to generate a compensation plan for the current deviation; otherwise, the generation of a compensation plan based on causal reasoning is initiated. Based on the generated compensation plan, compensation operations are performed on the deviation frames of the target short video.
7. The AI-based game animation frame-by-frame progressive generation method according to claim 6, characterized in that, The method further includes: Based on the final generated long-form game animation video, the story chapters of the long-form game animation video are identified and divided; for each story chapter, based on the baseline tracing map corresponding to the story chapter, it is decoupled into style genes and content genes and stored in a structured manner to form a chapter-level style gene library. When cross-chapter splicing is required, a style genetic algorithm is launched based on the chapter-level style gene library to output a set of offspring style gene sequences that are smoothly and gradually change in style and adapt to the current context in content as a reference for inheritance and transition. Using the inherited transition rubbing benchmark as a temporal constraint, a corresponding transition rubbing frame sequence is generated based on the frame-by-frame animation generation model. A pre-defined narrative coherence discriminator is run in parallel. The narrative coherence discriminator is based on an emotion-event causal graph model and determines in real time whether the generated transition frame sequence constitutes a credible causal bridge from the previous chapter to the current chapter in terms of plot logic. If not, a dynamic condition injection and local iterative regeneration process based on a causal violation heatmap is triggered to accurately correct and optimize the causal logic node with the lowest confidence in the transition frame sequence. Conversely, the transition frame sequence is inserted between chapters to complete the splicing and output.