An ai long drama automatic synthesis system based on multi-modal semantic alignment

CN122802742APending Publication Date: 2026-09-22NEUSOFT INTERNET (BEIJING) TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610943121.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-09-22

AI Technical Summary

Benefits of technology

[0015]本发明的技术效果和优点:本发明以剧情事件图和事件锚点链作为多模态素材的共同参照,使画面、对白、字幕、音效、背景音乐和镜头围绕同一剧情进程组织,必要锚点门控先排除主体、动作、对象和对白内容等关键错误,辅助锚点评分再从合格候选素材中完成选择,从而减少不同模态分别生成后出现的语义错配,事件时间轴同时考虑剧情前置关系、素材起始时间改变量、时长改变量和软性时序偏差,可降低字幕提前、音效脱离动作以及镜头切换过早等问题,复核阶段沿事件锚点链找到首个异常位置,只修正关联区间,既避免整段重新生成,也有利于保持角色、场景和声音的前后连续。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802742A_ABST
    Figure CN122802742A_ABST
Patent Text Reader

Abstract

The application relates to the field of artificial intelligence generated content and multimedia information processing technology, and discloses an AI long-form cartoon automatic synthesis system based on multi-modal semantic alignment, which analyzes a script text to form a screenplay semantic package, and establishes a plot event graph and an event anchor point chain according to the screenplay semantic package; after candidate materials of pictures, dialogues, subtitles, sound effects, background music and shots are generated, the candidate materials are subjected to necessary anchor point gating, and then target materials are determined by utilizing auxiliary anchor point scoring; a time sequence arrangement module combines an event sequence relationship and an anchor point time condition to form an event time axis, and an automatic synthesis module obtains an initial AI long-form cartoon according to the event time axis; and a result review module positions a first abnormal position along the event anchor point chain, and only modifies a material interval associated with the position, so that semantic mismatch and time dislocation of multi-modal content can be reduced, and repeated generation caused by local abnormalities can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of interdisciplinary technology of artificial intelligence-generated content and multimedia information processing, and in particular to an AI-based automatic animation synthesis system based on multimodal semantic alignment. Background Technology

[0002] AI-generated animated series typically use script text, character settings, and storyboards as a foundation. Artificial intelligence generates comic-style characters, scene images, character movements, dialogue, subtitles, sound effects, background music, and camera movements. These are then synthesized into a continuously playing dynamic narrative video according to the plot sequence. Current production processes mostly extract character, scene, dialogue, and action information from the script, generate keyframes or video clips according to the storyboard, synthesize character voiceovers, and finally write all kinds of materials into the timeline. Compared with traditional frame-by-frame drawing and manual editing, this method can significantly shorten the production cycle.

[0003] For example, Chinese patent CN119383413A discloses a method, system, and terminal for automated generation of micro-dramas based on multi-agent intelligence. It obtains a storyboard script including narration, visual information, shot information, and atmosphere information through a story generation intelligence and a storyboard generation intelligence. It generates dubbing audio and video subtitles based on the narration, generates keyframe images and video clips based on the visual information, and generates background music based on the atmosphere information. It then aligns and splices the dubbing audio, video subtitles, video clips, and background music along the timeline to obtain the target synthesized micro-drama. This technology can automatically complete the generation of storylines, storyboard design, audiovisual material generation, and micro-drama synthesis, which helps to shorten the production cycle of micro-dramas and improve the generation efficiency of micro-dramas.

[0004] The aforementioned and existing related technologies often suffer from the following drawbacks: While the visuals, dialogue, subtitles, and sound effects in AI-generated animations can play continuously in a preset order, the expression of key plot points by various content types is often not in the same place. Viewers may have already heard crucial information from the dialogue while the visuals remain focused on irrelevant actions; conflicts may have occurred on screen, but the voice-over narration and background sounds have not yet responded accordingly; subtitles may prematurely reveal subsequent content, or the sound of the previous scene may be retained after a scene change. These discrepancies disrupt the order in which viewers receive plot information, making segments that should form foreshadowing, turning points, or emotional outbursts appear bland and abrupt, even leading viewers to misjudge the current plot's focus. This is because existing automatic synthesis methods typically use the length of the material and the order of the scenes as the basis for combination, failing to determine the specific presentation timing of various materials based on the semantics of the dialogue, the progress of actions, emotional changes, and the plot's focus, making it difficult for visuals, sound, and text to develop together around the same plot point. (Invention Content) The technical problem to be solved by this invention is that the existing technology has the disadvantage that multimodal materials lack unified plot semantics and temporal constraints, which easily leads to mismatches in visuals, dialogue, subtitles, sound effects and shot content, asynchronous playback time, and the need for large-scale regeneration of local anomalies. To this end, we propose an AI-based automatic animation synthesis system based on multimodal semantic alignment.

[0005] To achieve the above objectives, this application adopts the following technical solution: an AI-based automatic animation synthesis system based on multimodal semantic alignment, comprising: a script parsing module, used to divide the script text into scenes and extract plot information to form a scene semantic package; an event anchor point construction module, used to establish plot event nodes based on the scene semantic package, construct a plot event graph based on the preconditions and causal relationships between plot event nodes, and extract an event anchor point chain arranged according to the event progression from each plot event node; a multimodal material generation module, used to generate candidate materials based on the plot event nodes and the event anchor point chain; and a semantic alignment module, used to determine the necessary anchor points of plot events, gate the candidate materials item by item according to the event anchor point chain, and, where necessary... After all anchor points pass through gating, auxiliary anchor points are used to score the candidate materials that have passed the gating to determine the target material. The timing arrangement module is used to determine the event sequence according to the plot event diagram, bind the target material to the corresponding event stage, and select the scheme with the least adjustment cost from the time adjustment schemes that satisfy the event precondition relationship to form the event timeline. The automatic synthesis module is used to combine the target material according to the event timeline to obtain the initial AI animation. The result verification module is used to extract playback features from the initial AI animation, locate the first abnormal event anchor point according to the event anchor point chain, determine the material range associated with the abnormal event anchor point and make local corrections, while keeping other material ranges that satisfy the anchor point conditions and timing constraints unchanged to obtain the final AI animation.

[0006] Preferably, the system also includes a data acquisition module, which is used to acquire script text, character setting data, scene setting data, target visual style parameters and output configuration parameters, and unify character encoding, character identification, time unit and material file format; the character setting data includes at least character name, character identification, appearance features, clothing features, voice features and character relationships; the scene setting data includes at least scene name, space type, time conditions, weather conditions, environmental sound type and scene visual features; the output configuration parameters include at least output resolution, frame rate, target duration, audio sampling rate and video encoding format.

[0007] Preferably, for sentences to be parsed where the subject of the action or the speaker of the dialogue is not clearly defined, the script parsing module extracts a set of candidate characters from the current scene and adjacent scenes, determines the location association value, semantic adaptation value, spatial association value, and dialogue relationship association value of each candidate character, and determines the subject matching value based on each association value and its corresponding weight; when the set of candidate characters includes at least two candidate characters, and the difference between the highest subject matching value and the second highest subject matching value is greater than the subject confirmation threshold, the candidate character corresponding to the highest subject matching value is determined as the event subject; when the difference is not greater than the subject confirmation threshold, at least two candidate characters with the highest subject matching values ​​are retained, and the character position, character action, and speaker characteristics are combined again for judgment during the semantic alignment process; when the set of candidate characters includes only one candidate character, that candidate character is determined as the event subject.

[0008] Preferably, the event anchor construction module configures a plot event identifier for each plot event node and divides the plot event into at least two target event stages: preparation stage, trigger stage, execution stage, completion stage, and reaction stage. Each event anchor records the plot event identifier, anchor identifier, anchor type, target semantic value, target event stage, allowed time deviation, and associated material type. Anchor types include main anchor, action anchor, object anchor, dialogue content anchor, phoneme-lip movement anchor, action-sound effect anchor, emotion anchor, scene anchor, background music anchor, and shot anchor. Multiple event anchors corresponding to the same plot event form an event anchor chain according to the target event stage and its occurrence order within the corresponding target event stage. Each candidate material records the corresponding plot event identifier and associated anchor identifier.

[0009] Preferably, the semantic alignment module extracts character identity features, character posture sequences, action object positions, action contact frames, and lip-sync category sequences from candidate scene materials; extracts speaker features, speech recognition text, phoneme time intervals, and prosodic features from candidate dialogue materials; extracts subtitle text, word-by-word display time, and full display time from candidate subtitle materials; extracts sound categories and peak sound times from candidate sound effect materials; extracts emotion categories and loudness envelopes from candidate background music materials; and extracts shot size, camera movement speed, and shot transition times from candidate shot materials. Discrete material features include character identifiers, action object identifiers, dialogue text, sound categories, and shot size, and are judged using identifier consistency. Vector material features include character identity features, dialogue semantic features, and emotion features, and are judged using vector similarity. Continuous material features include action amplitude, phoneme time, lip-sync time, peak sound times, camera movement speed, and shot transition times, and the degree of consistency is determined based on the absolute deviation between the actual values ​​and the target values ​​of the corresponding event anchor points, as well as the maximum allowable deviation.

[0010] Preferably, necessary anchors include subject anchors, action anchors, object anchors, dialogue content anchors, phoneme-lip-sync anchors, and action-sound effect anchors applicable to the current plot event. Auxiliary anchors include emotion anchors, scene anchors, background music anchors, and camera anchors. The semantic alignment module judges the consistency of each applicable necessary anchor item by item according to the order of the event anchor chain. When the consistency of any applicable necessary anchor is lower than the corresponding necessary anchor threshold, the corresponding candidate material is marked as unqualified material. When all applicable necessary anchors reach the corresponding necessary anchor threshold and there are auxiliary anchors applicable to the current material modality, the semantic alignment score of the candidate material is determined according to the consistency of each auxiliary anchor and its corresponding weight. When there are no auxiliary anchors applicable to the current material modality, the candidate materials that have passed the necessary anchor gate are sorted according to the average consistency of each applicable necessary anchor. The semantic alignment module determines the candidate material with the highest semantic alignment score or the highest average consistency of necessary anchors as the target material.

[0011] Preferably, the timing module uses video frames as time steps, generates multiple timing adjustment schemes within a preset start time movement window and a preset duration scaling range, and sets a hard timing constraint that the start time of subsequent plot events is no earlier than the completion time of their preceding plot events; for timing adjustment schemes that meet the hard timing constraint, the adjustment cost J is calculated according to the following formula: in, For the target number of materials, and The first Start and end times before and after the adjustment of project logo materials. and The first The duration of the project logo materials before and after adjustment. For the first The soft timing deviation penalty value of the project target material. , and Weights greater than 0; The soft timing deviation penalty value is determined based on the deviations between the peak of the action sound effect and the moment of action contact, the phoneme time interval and the lip-sync time interval, the moment of full subtitle display and the moment of keyword pronunciation, and the moment of camera transition and the moment of plot event completion. When the corresponding deviation does not exceed the allowable time deviation or there is no corresponding material for the current plot event, the corresponding deviation item is set to 0. The timing arrangement module selects the adjustment cost from the time adjustment schemes that meet the hard timing constraints. Minimal time adjustment scheme.

[0012] Preferably, the timing arrangement module performs dynamic planning according to the topological order of the plot event graph. The dynamic planning state includes the candidate end time of the current plot event and the cumulative adjustment cost. Among multiple dynamic planning states with the same candidate end time, the dynamic planning state with the minimum cumulative adjustment cost is retained. When there are no candidate states that meet the hard timing constraints within the preset start time movement window and preset duration scaling range, the preset start time movement window or the preset duration scaling range of the non-critical material interval is expanded, and candidate states are regenerated. When there are still no candidate states that meet the hard timing constraints within the expanded range, the first event anchor point that causes the constraint conflict is located according to the event anchor point chain. A directional regeneration instruction including plot event identifier, material modality, anchor point identifier, target semantic value, and allowed time range is generated, and the directional regeneration instruction is sent to the multimodal material generation module.

[0013] Preferably, when combining target materials, the automatic synthesis module establishes a correlation record between plot event identifiers, anchor point identifiers, material identifiers, and actual playback intervals. The result verification module queries the correlation record based on the plot event identifier and anchor point identifier of the abnormal event anchor point to determine the associated materials and actual playback intervals. When the associated material type is video material, local conditional regeneration is performed based on the area mask corresponding to the abnormal character or abnormal action object, and the previous and next frames of the abnormal material interval are used as the character identity, scene background, and lighting continuity conditions. When the associated material type is dialogue material, only the dialogue segment or phoneme interval corresponding to the abnormal event anchor point is replaced, and a cross-gradient interval is set between the replaced audio and the original audio. When the associated material type is subtitle material, sound effect material, or shot material, only the start time, end time, or key moment of the corresponding material is adjusted. During the local correction process, the video content outside the area mask and other material intervals that meet the necessary anchor point conditions and hard timing constraints remain unchanged.

[0014] Preferably, the result verification module calculates the verification deviation value of the k-th plot event according to the following formula. : ,in, To review the semantic alignment score, To normalize the time series deviation, the absolute deviation between the actual value and the corresponding target value at each applicable key moment is divided by the corresponding allowable time deviation, and the resulting value is then averaged after being restricted to the range of 0 to 1. To normalize the emotional coordination bias, the distances between the voice-over emotion, character facial expression, background music, and camera changes and the target emotion were converted to the range of 0 to 1 and then averaged. , and All weights are greater than 0 and less than 1, and When any necessary anchor point fails the review, locate the first failing necessary anchor point according to the event anchor point chain and trigger a local correction; when all necessary anchor points pass the review and the review deviation value is met, the correction is triggered. When the deviation exceeds the review threshold, the semantic bias, temporal bias, and emotional synergy bias are determined separately to affect the review deviation value. The contribution is evaluated, and the material is regenerated and corrected, or the time adjustment is adjusted or the sentiment parameter is corrected based on the deviation item with the largest contribution.

[0015] The technical effects and advantages of this invention are as follows: This invention uses a plot event diagram and an event anchor chain as a common reference for multimodal materials, so that the visuals, dialogue, subtitles, sound effects, background music, and shots are organized around the same plot process. Necessary anchor point gating first eliminates key errors such as the subject, action, object, and dialogue content. The auxiliary anchor point scoring then completes the selection from qualified candidate materials, thereby reducing semantic mismatches that occur after different modalities are generated separately. The event timeline simultaneously considers the plot's prior relationship, the amount of change in the material's start time, the amount of change in duration, and soft timing deviations, which can reduce problems such as premature subtitles, sound effects detached from actions, and premature shot transitions. In the review stage, the first abnormal position is found along the event anchor chain, and only the related interval is corrected, which avoids the entire segment being regenerated and also helps to maintain the continuity of characters, scenes, and sounds. Attached Figure Description

[0016] The disclosure of this invention is illustrated with reference to the accompanying drawings. It should be understood that the drawings are for illustrative purposes only and are not intended to limit the scope of protection of this invention. In the drawings, the same reference numerals are used to refer to the same parts: Figure 1 This is a schematic diagram of the overall system structure of the present invention; Figure 2 This is a basic operation flowchart of the AI ​​comic book automatic synthesis system of the present invention; Figure 3 This is a schematic diagram illustrating the multimodal candidate material generation and semantic alignment of the present invention; Figure 4 This is a schematic diagram of the timing arrangement and event timeline generation of the present invention; Figure 5 This is a schematic diagram illustrating the result verification and local correction of the present invention; Figure 6 This is a line graph comparing the key moments of the multimodal process in this invention before and after. Detailed Implementation

[0017] The following describes a multimodal semantic alignment-based AI-powered automatic animation synthesis system in conjunction with specific implementation methods. The following module division is mainly used to illustrate the corresponding data processing functions. In actual deployment, each module can run on the same server, or it can be deployed separately on a cloud server, edge computing device, or content generation terminal.

[0018] Example 1: Refer to Figure 1As shown, this embodiment provides an AI-based automatic animation synthesis system based on multimodal semantic alignment, including a data acquisition module, a script parsing module, an event anchor point construction module, a multimodal material generation module, a semantic alignment module, a chronological arrangement module, an automatic synthesis module, and a result verification module.

[0019] AI-generated animated series refer to dynamic narrative videos generated based on script text, character settings, and scene settings. They consist of comic-style visuals, character actions, dialogue, subtitles, environmental sound effects, action sound effects, background music, and camera movements. A storyboard refers to a video generation unit with a relatively independent scene, shot range, or narrative content within a continuous storyline. A storyboard can include one plot event or multiple plot events with a sequential relationship. A plot event refers to content that can cause changes in character actions, character language, item status, scene status, or plot relationships. A plot event node is a data node used to record plot events.

[0020] The plot event graph consists of multiple plot event nodes and the relationship information between these nodes. For the first... Each plot event has a plot event node denoted as . A plot event node includes at least an event identifier, an event subject, an event action, an event object, dialogue content, an emotional state, a scene state, a set of preceding events, and an event completion condition. Directed edges between plot event nodes are used to represent preceding or causal relationships. For plot events that are allowed to occur simultaneously, parallel relationship identifiers are used to record that there are no mandatory sequential constraints between the corresponding plot event nodes.

[0021] Event anchors are target information extracted from plot event nodes that can be detected or verified by at least one material modality. Each event anchor includes at least an anchor identifier, anchor type, target semantic value, target event stage, allowed time deviation, and associated material type. Anchor types include subject anchors, action anchors, object anchors, dialogue content anchors, emotion anchors, sound anchors, scene anchors, shot anchors, phoneme-lip movement anchors, action-sound effect anchors, and background music anchors. Multiple event anchors form an event anchor chain according to the order of occurrence within the same plot event, which is used to associate images, dialogue, subtitles, sound effects, background music, and shot materials.

[0022] Event anchors are divided into necessary anchors and auxiliary anchors. Necessary anchors determine whether candidate materials meet the basic conditions for continued use, such as whether the subject is correct, whether the action and object correspond, whether the dialogue content is consistent, and whether the audio and video are synchronized. Auxiliary anchors are used to make further selections among multiple qualified candidate materials, mainly involving emotions, background music, and camera performance. When a plot event does not contain dialogue, character actions, or action sound effects, the corresponding anchor will not be forcibly activated. The event timeline assigns start time, end time, and key moments to various materials based on the plot event diagram, event completion conditions, and event anchor chain.

[0023] The modules are divided as follows: The data acquisition module receives the script text, character settings, scene settings, visual style, and output configuration; the script parsing module extracts character, dialogue, action, object, emotion, sound events, scene, and shot information from the text; the event anchor point construction module establishes plot event nodes, event relationships, and event anchor points based on this; the multimodal material generation module is responsible for generating visuals, dialogue, subtitles, sound effects, background music, and shot candidate materials; the semantic alignment module compares material features with event anchor points; the temporal arrangement module arranges the presentation intervals of the materials; the automatic synthesis module completes video rendering, audio mixing, subtitle overlay, and shot arrangement; after the initial AI comic is generated, the result verification module re-checks the actual playback features and provides local correction instructions when there are deviations.

[0024] Data is passed between modules in the order of processing. The raw input first enters the script analysis module, and the resulting storyboard semantic package is handed over to the event anchor construction module. The plot event nodes and event anchors are then sent to the material generation, semantic alignment and timing arrangement stages respectively. Candidate materials become target materials after semantic alignment. The timing arrangement module generates an event timeline based on this. The automatic compositing module completes the initial cut. If an anomaly is found during the review, the local correction instruction will return to the material generation module or the timing arrangement module instead of restarting the entire process.

[0025] Example 2: Based on Example 1, this example further explains the basic operation process of the AI-powered automatic animation synthesis system, referring to... Figure 2 As shown, the system operation process includes: Step 1: The data acquisition module receives the script text, character setting data, scene setting data, target visual style parameters, and output configuration parameters of the AI ​​comic to be synthesized. To avoid different interpretations of the same data by subsequent modules, the system unifies the text character encoding, character identification, time unit, and the file formats of images, audio, and subtitles. Character settings can record name, identity, age range, appearance, clothing, voice, and character relationships. Scene settings can record space type, time, weather, environmental sounds, and visual features. Visual style parameters include lines, colors, brightness, texture, and aspect ratio. Output configuration includes resolution, frame rate, target duration, audio sampling rate, and encoding format. Step 2: The script analysis module divides the scenes into storyboards based on changes in scene, time, characters, and camera cues. At the same time, it checks whether the actions constitute a relatively complete narrative unit. For each storyboard, it extracts the character, predicate action, action object, dialogue content, emotional state, sound events, scene information, and camera information, and combines them into a storyboard semantic package. Step 3: The storyboard semantic package often contains continuous actions or causal narratives. The event anchor construction module breaks it down into independent plot event nodes, and then builds a plot event graph based on the preceding, causal and parallel relationships. The verifiable information in each node is organized into event anchors, and multiple anchors in the same event are connected into an event anchor chain according to the actual occurrence order. Step 4: The multimodal material generation module combines plot event nodes, event anchor chain, character settings, scene settings and target screen style to generate candidate screens, dialogue, subtitles, sound effects, background music and shot materials respectively. Each candidate material retains plot event identifiers, associated anchor identifiers, material duration and generation parameters to facilitate subsequent source tracking. Step 5: The semantic alignment module extracts character identity, character actions, action objects, dialogue content, phoneme timing, lip-sync state, sound effect peaks, emotions, background music, and shot features from the candidate materials. The system first uses necessary anchor points to eliminate key errors, then scores the remaining materials based on the consistency of auxiliary anchor points, and finally selects the target material. Step Six: The timing module binds the target material to the preparation, trigger, execution, completion, or reaction stages based on the relationships in the plot event diagram, the event completion conditions, and the target stages and allowable time deviations of each anchor point. The system generates candidate solutions by moving the start time, adjusting the duration of non-critical intervals, or moving critical moments. It also checks constraints such as event preconditions, phoneme-lip-sync, action-sound effect synchronization, subtitle display, and camera switching. Solutions that meet the hard timing constraints are compared for cost, and solutions with lower adjustment costs are written into the event timeline. Step 7: The automatic compositing module writes the target screen, dialogue, subtitles, sound effects, background music and shot parameters into the corresponding tracks according to the event timeline, and completes rendering, mixing, subtitle overlay and shot arrangement to obtain the initial AI comic series; Step 8: The result verification module re-extracts the playback features of the initial AI comic and compares them with the plot event graph, event anchor chain, and event timeline. The system finds the first position that does not meet the necessary anchor point or hard timing constraints along the event anchor chain, and only performs regeneration, time shifting, or parameter adjustment on the corresponding material interval. The corrected material re-enters the synthesis and verification process until all key conditions are met, and then outputs the final AI comic.

[0026] Example 3: This example follows the basic process described above, focusing on the methods for script text standardization, storyboard analysis, candidate subject judgment, and the creation of plot event diagrams.

[0027] Before analyzing the script, the system first standardizes character encoding, punctuation, and paragraph format, and identifies dialogue and character titles. The same character may use full name, abbreviation, alias, or pseudonym in different paragraphs. The system uses name mapping to classify these titles into the same character identifier, while preserving their position and expression in the original text.

[0028] After standardization, the system divides the storyboards based on changes in scene, time, characters, and camera cues. It then performs sentence segmentation, part-of-speech tagging, and dependency analysis on the storyboard text. Predicates, state change words, and sound description words can serve as the core of candidate events. Connective words such as "after," "following," "at the same time," "until," and "when" help determine the boundaries of events and their sequential or parallel relationships. Each candidate event contains at least one action, sound, or state change, and records the agent, patient, location, time conditions, and dialogue content.

[0029] For sentences that do not clearly define the actor or the speaker, the script analysis module extracts a set of candidate characters from the current shot and one adjacent shot before and after it. .

[0030] The determination of candidate subjects needs to consider the position of the character, the semantics of the text, the scene space and the dialogue relationship. The four indicators have different dimensions and original value ranges. Therefore, the four indicators are first converted to the range of 0 to 1, and then unified as a positive indicator that the larger the value, the more likely the candidate character is to be the subject of the event.

[0031] For multiple positive evaluation indicators, a linear weighted comprehensive evaluation relationship is adopted. The product of each normalized indicator and its corresponding weight is added to obtain the comprehensive evaluation value of the candidate. Let the normalized evaluation indicators be, in order: , , and The corresponding weights are as follows: , , and Then the comprehensive evaluation value is equal to the sum of the weighted contributions of each indicator, and the sum of each weight is 1.

[0032] Candidates Location-related values Semantic adaptation value Spatial correlation value Correlation value with dialogue Substitute the above four normalization indicators into the values ​​respectively, and... , , and Using these as corresponding weights, we obtain: Equation (1), where the weights , , and The subject identification threshold can be determined using labeled script samples. Specifically, multiple candidate weights can be set, the subject identification accuracy of each candidate can be calculated, and the weight with the highest accuracy in the validation set can be used as the preset weight for the system. The subject confirmation threshold can also be determined based on the distribution of the difference between the highest and second highest matching values ​​in correctly identified and incorrectly identified samples.

[0033] Set candidate roles Separated from the current statement to be parsed The sentence contains a positional association value. according to It is confirmed that, among them, For position attenuation parameters, semantic adaptation values The spatial correlation value is obtained by linearly transforming the cosine similarity between the feature vector of the candidate character setting text and the feature vector of the current action or dialogue text to a range of 0 to 1. Based on the normalized distance between the candidate character's position and the position where the action occurred ,according to Determined, where σ is the spatial distance decay parameter, and the dialogue relationship correlation value is determined when the candidate role is the respondent in the previous round of dialogue. Choose 1 if the candidate character appears in the two rounds of dialogue preceding the current dialogue. If the value is 0.5, and the above conditions are not met, Take 0.

[0034] In a feasible setup , , and The values ​​are set to 0.25, 0.35, 0.20, and 0.20 respectively, and the subject confirmation threshold is set to 0.15. When the difference between the highest subject matching value and the second highest subject matching value is greater than 0.15, the candidate character corresponding to the highest subject matching value is written into the plot event node; when the difference between the two is not greater than 0.15, the two candidate characters with the highest subject matching values ​​are retained, and in the subsequent semantic alignment process, the character position in the picture, the character's actions, and the speaker's features in the dialogue are judged again. The above weights and subject confirmation thresholds can be adjusted using the verification results of the labeled script samples.

[0035] Complex narratives need to be further broken down into independently processable events. For example, if "Character A hears footsteps, turns to look at the door, and whispers, 'Someone is coming,'" the system will create three plot event nodes: the sound of footsteps, Character A turning his head, and Character A's dialogue. Then, it will record the sequence and causal relationship between them.

[0036] For the Each plot event has a plot event node denoted as . A plot event node includes at least an event identifier, an event subject, an event action, an event object, dialogue content, an emotional state, a scene state, a set of preceding events, and conditions for event completion. The event anchor construction module establishes directed edges based on the preceding and causal relationships between plot event nodes. For plot events that are allowed to occur simultaneously, the parallel relationship identifier records that there are no mandatory sequential constraints between the corresponding plot event nodes, thus forming a plot event graph.

[0037] Event completion conditions are converted into states that can be detected from the generated material. For example, the completion condition for a head-turning event can be set to the angle between the character's face and the target direction being less than a preset angle; the completion condition for a dialogue event can be set to the end of the corresponding dialogue audio playback; and the completion condition for an item placement event can be set to the center of the item entering the target area and continuously maintaining it for a preset number of frames.

[0038] Example 4: Based on the plot event diagram established in Example 3, the specific organization of event anchors and event anchor chain is explained below.

[0039] The event anchor construction module extracts target information that can be detected by visuals, sound, subtitles or camera shots from plot event nodes, and establishes event anchors accordingly. Each anchor records at least the anchor identifier, anchor type, target semantic value, target event stage, allowed time deviation and associated material type.

[0040] Event anchors can cover information such as subject, action, object, dialogue content, phonemes-lip movements, action-sound effects, emotion, sound, scene, background music, and camera shots. Anchors in the same plot event are not simply listed, but arranged into an event anchor chain according to the actual plot progress.

[0041] Taking "sound triggers character action, followed by dialogue" as an example, the event anchor chain can be set sequentially as scene state, sound trigger, action start, action contact or completion, dialogue content, phoneme-lip movement, emotion and camera anchor. When a certain item does not appear in the plot, the corresponding anchor is directly omitted.

[0042] Different anchor points record content with different focuses: scene state anchor points store the character's position and environmental state before the event occurs; sound trigger anchor points record the sound type, source, and peak range; action anchor points record the subject, object, posture changes, start frame, and contact frame; dialogue content anchor points store the dialogue text, speaker, and keywords; phoneme-lip movement anchor points are responsible for the correspondence between phoneme time and lip movement type; action-sound effect anchor points focus on the contact moment and sound peak; and emotion, background music, and shot anchor points describe the intensity of emotion, music changes, shot size, focus object, and switching timing, respectively.

[0043] Phoneme-lip shape anchors can be generated by a preset mapping table. Taking Chinese speech as an example, bilipine closed sounds correspond to closed lip shapes, open vowels correspond to open lip shapes, rounded vowels and unrounded vowels correspond to rounded and unrounded lip shapes respectively, and labiodental sounds correspond to labiodental lip shapes. At the same time, neutral transitional lip shapes are set. After the speech synthesis interface provides the start and end times of phonemes or pinyin initials and finals, the system can generate lip shape category sequences and target time intervals.

[0044] Each event anchor also records the allowed time deviation. One feasible setting is: the allowed deviation for the start time of phoneme-lip movement is 80 milliseconds, and the allowed deviation for the peak sound of action-sound effect is 120 milliseconds. The complete subtitle must not be pronounced earlier than the corresponding keyword, and the camera cannot switch before the current plot event is completed. The actual values ​​can be calculated by combining the output frame rate, plot rhythm and playback terminal.

[0045] Necessary anchors and auxiliary anchors perform different tasks. Anchors such as subject, action, object, dialogue content, phoneme-lip movement, and action-sound effect are used to determine whether there are critical errors in candidate material. Anchors such as emotion, background music, and shot are more suitable for comparison among multiple qualified candidate materials. When a plot event does not contain a certain type of content, the corresponding anchor does not participate in the judgment.

[0046] Example 5: This example combines Figure 3 This section explains how candidate materials are generated, how material features are extracted, and how the consistency of continuous features is calculated.

[0047] Reference Figure 3 As shown, the multimodal material generation module receives plot event nodes, event anchor chain, character settings, scene settings and target screen style, and organizes the generation conditions for screen, dialogue, subtitles, sound effects, background music and shots respectively.

[0048] Different generation interfaces use different inputs. The screen or video interface receives character identity, scene, action subject and object, posture or action control sequence, target lip shape, screen style and target duration, and outputs screen frame sequence and metadata such as character area, posture key points and lip shape category. The dialogue interface receives dialogue text, speaker voice, target emotion, keyword emphasis and target duration, and outputs audio, phoneme sequence and time interval. The subtitle interface provides the display content and display time based on the dialogue text and word time. The sound effects, background music and camera interfaces output waveform, peak value, loudness envelope, rhythm characteristics, camera parameters and switching time, respectively.

[0049] Multiple candidate results can be generated for the same material type in the same plot event. Each candidate material saves material identifier, plot event identifier, associated anchor point identifier, material modality, duration, generation parameters, generation confidence, and key time points. This information will be passed to the semantic alignment, temporal arrangement, and result verification stages along with the material, so that the system can trace the source of the material and the actual playback interval when anomalies are subsequently detected.

[0050] After candidate frames are generated, the system obtains the character area through character detection and tracking, obtains the identity features through character recognition, obtains the normalized key point sequence through pose estimation, and then uses target detection to determine the position of the action object. The speed of key points in adjacent frames and the distance change between the subject and the object can be used to determine the action contact frame. The mouth key points or the metadata returned by the generation interface are used to form the lip shape category sequence.

[0051] The processing methods for dialogue materials differ. Speaker recognition provides speaker characteristics, speech recognition outputs the dialogue text, forced speech alignment determines the phoneme time interval, fundamental frequency, energy, speech rate, and pause positions are used to describe rhythm, subtitle materials provide word-by-word or phrase-by-word display time and full display time, sound effect materials provide sound category, peak time and duration, and background music and shot materials provide emotion, loudness, rhythm, shot size, direction of motion, zoom speed and switching time, respectively.

[0052] The raw data of different modalities are not directly compared with each other, but are matched with the event anchor points corresponding to the type respectively, and the results are unified to the range of 0 to 1. The character identification, action object identification, sound category and scene identification adopt discrete consistency judgment. The character identity, speaker, dialogue semantics and emotional semantics adopt vector similarity. Continuous features such as character posture, movement amplitude, phoneme time, lip movement time, sound peak, emotional intensity and camera movement speed are determined by the normalized absolute deviation between the actual value and the target value.

[0053] For continuous features such as emotional intensity, amplitude of movement, phoneme duration, lip-sync duration, and camera movement speed, first calculate the actual values ​​of the candidate material. target value of event anchor point The absolute deviation between them, since different features correspond to different units and permissible errors, is further determined by using the maximum permissible deviation of that feature. Normalize the absolute deviation to obtain the ratio of the actual deviation to the maximum allowable deviation.

[0054] When the actual value is the same as the target value, the normalized deviation is 0 and the consistency level should be 1. As the actual deviation gradually increases, the consistency level should decrease accordingly. When the actual deviation reaches the maximum allowable deviation, the consistency level should decrease to 0. Therefore, subtracting the normalized absolute deviation from 1 yields the consistency level that decreases linearly with the deviation.

[0055] When the actual deviation exceeds the maximum allowable deviation, the linear decay result will be less than 0. Since the consistency degree should not be negative, the lower limit is restricted to 0 using the maximum truncation relation. Therefore, we obtain: Equation (2), maximum allowable deviation The allowable time deviation for phonemes-lip movements can be calculated based on the output frame rate, while the allowable time deviation for actions-sound effects can be determined based on the action type and sound duration. The allowable deviation for emotional intensity can be determined based on the recognition error distribution of labeled emotional samples. In equation (2), Indicates the first The first plot event, the... The first material mode The candidate material and the first The consistency between event anchors is set to 1 when the actual value of the material feature is the same as the target value, and 0 when the difference between the actual value and the target value reaches or exceeds the maximum allowable deviation. The maximum allowable deviation is set for emotional intensity, action amplitude, phoneme duration, lip movement duration and camera movement speed to avoid direct comparison of data with different dimensions.

[0056] Example 6: Based on the aforementioned feature comparison, this example illustrates necessary anchor point gating, auxiliary anchor point scoring, and how to determine target material when there are no auxiliary anchor points.

[0057] The semantic alignment module first checks what the current plot event actually contains, and then determines the necessary anchor points to be checked by combining the modality of the candidate materials. Necessary anchor points are not always enabled. When there is no dialogue, character actions or action sound effects in the plot, the corresponding anchor points will not participate in gating.

[0058] For candidate video footage, applicable necessary anchors can include subject anchors, action anchors, and object anchors. Subject anchors require that the character's identity in the candidate video matches the target character. Action anchors require that the character's posture changes and action stages match the target action. Object anchors require that the action object identifier and the position of the action object relative to the action subject in the candidate video match the target value. For candidate dialogue footage, applicable necessary anchors can include subject anchors, dialogue content anchors, and phoneme-lip shape anchors. Subject anchors require that the speaker characteristics in the dialogue audio match the voice characteristics of the target character. Dialogue content anchors require that the dialogue text obtained from speech recognition achieves a preset text consistency level with the target dialogue. Phoneme-lip shape anchors require that the lip shape category corresponding to each phoneme time interval is the same as the target lip shape category, and the time deviation between the two does not exceed the allowed time deviation recorded by the corresponding event anchor.

[0059] For candidate sound effects, the action-sound effect anchor point requires that the sound category of the candidate sound effect be the same as that of the target sound effect, and that the time deviation between the peak moment of the sound effect and the moment of contact with the action does not exceed the allowable time deviation. For plot events that do not contain dialogue, character actions or action sound effects, the corresponding phoneme-lip-sync anchor point or action-sound effect anchor point is not enabled.

[0060] Each necessary anchor point is checked in the order of the event anchor point chain. If any one of them is below the corresponding threshold, the current candidate material is marked as unqualified, and subsequent auxiliary anchor point scoring will no longer be performed.

[0061] When the consistency of all applicable necessary anchor points reaches the corresponding necessary anchor point threshold, the semantic alignment module determines the set of auxiliary anchor points applicable to the current plot event and the current material modality. .

[0062] Necessary anchors are used to exclude candidate materials with critical errors, while auxiliary anchors are used to distinguish multiple candidate materials that have passed the necessary anchor gate. Since different auxiliary anchors have different levels of importance, the ordinary average of the consistency of each item cannot be directly calculated. Therefore, a weighted arithmetic average relationship is used.

[0063] For the set of auxiliary anchor points First, calculate the consistency of each auxiliary anchor point. Its weight The product of these products is summed to obtain the weighted consistent contribution of the candidate material to the auxiliary anchor points; this is then divided by the sum of the weights of all auxiliary anchor points involved in the evaluation, so that the calculation result is not affected by the number of auxiliary anchor points or the absolute size of their weights. Thus, we obtain: Equation (3), in Equation (3), Indicates participation in the The first plot event A set of auxiliary anchor points for modal scoring of various materials. Indicates the first The weights of each auxiliary anchor point This indicates that the material features of the r-th candidate material are similar to those of the r-th candidate material. The consistency between the auxiliary anchors is such that the weight of each auxiliary anchor is a value greater than 0, and is normalized according to the current material modality. When the weight of each auxiliary anchor has been normalized and the total weight is 1, equation (3) can be simplified to the weighted sum of the consistency of each auxiliary anchor. The form of the denominator is retained, which allows different numbers of auxiliary anchors to be used in different plot events.

[0064] For example, candidate scenes can be scored based on emotion, scene, and shot performance; candidate dialogue can be compared based on target emotion, speech rate, and keyword emphasis; and candidate background music focuses on musical emotion, loudness changes, and how well it matches the plot stage.

[0065] Under the same plot event and the same material modality, the system first excludes materials that have not passed the necessary anchor point gating, and then selects the semantic alignment score from the remaining candidate results. The highest-ranking item will be used as the target material.

[0066] Some plot events do not have auxiliary anchor points applicable to the current material modality. For example, when a static image only requires the display of a specified character, action, and object, there may only be anchor points for the main body, action, and object. In this case, the auxiliary anchor point scoring formula (3) is not executed. Instead, the average consistency of each candidate material on all applicable necessary anchor points is compared, and the material with the highest average value is selected.

[0067] If multiple candidate materials have the same semantic alignment score or the average value of necessary anchor points, a selection can be made by combining generation confidence, generation order, and the degree of continuity of character identity with the previous plot event.

[0068] One feasible threshold setting is as follows: the subject and object identifiers must be completely consistent, the consistency of action posture is not less than 0.85, and the consistency of dialogue text is not less than 0.90; at least 80% of the phoneme intervals need to meet the consistency of lip shape category, and the time deviation does not exceed 80 milliseconds; the deviation between the peak of the sound effect and the moment of contact with the action does not exceed 120 milliseconds, and the semantic alignment score threshold of the auxiliary anchor point can be set to 0.75.

[0069] If all candidate results for the same material modality fail the necessary anchor point gating, the system finds the first failed anchor point along the event anchor point chain and outputs the plot event identifier, material modality, abnormal anchor point identifier, target semantic value, and generation conditions that need to remain unchanged. The multimodal material generation module generates new materials accordingly, and the correct character identity, scene background, or dialogue content does not need to be modified repeatedly.

[0070] Example 7: This example combines Figure 4 This section explains the formation process of the event timeline, focusing on hard timing constraints, soft timing deviations, adjustment costs, dynamic programming, and handling situations where no feasible state exists.

[0071] Reference Figure 4 As shown, the timing arrangement module performs topological sorting of plot events based on the preceding and causal relationships in the plot event graph, and determines the order of event anchors within the same plot event based on the event anchor chain.

[0072] The timing module divides plot events into at least two target event stages from the preparation stage, trigger stage, execution stage, completion stage, and reaction stage. For dialogue events, the execution stage can be further divided into the voice start interval, keyword interval, and voice end interval.

[0073] In the visual material, the starting frame of the action is bound to the execution phase, the action contact frame or action completion frame is bound to the completion phase, the start time of the dialogue audio is bound to the sounding start interval, the sounding time of the dialogue keywords is bound to the keyword interval, the start display time of the subtitles is bound to the dialogue start time, the sound peak of the action sound effect is bound to the moment of action contact, the loudness or rhythm change position of the background music is bound to the emotional anchor point, and the moment of camera switching is bound to the completion phase or reaction phase.

[0074] The timing arrangement module sets the sequential relationship between plot events as a hard timing constraint. For two plot events with a prerequisite relationship, the start time of the subsequent plot event must not be earlier than the time when the completion condition of the prerequisite plot event is met. Time adjustment schemes that violate the hard timing constraint will not be included in the adjustment cost comparison process.

[0075] The time deviation between phonemes and lip movements, the time deviation between the moment of action contact and the peak moment of sound effects, the advance of complete subtitles relative to the moment of keyword pronunciation, and the advance of camera transitions relative to the moment of plot event completion are set as soft timing deviations. Soft timing deviations are allowed to exist within a limited range, but the larger the deviation, the greater the cost of adjusting the corresponding timing adjustment scheme.

[0076] The timing module is based on the original start time and original duration of each target material. It generates multiple timing adjustment schemes by moving the start time of the material, adjusting the duration of non-critical screen intervals, adjusting the natural pauses of non-keyword dialogue intervals, moving the subtitle display time, and moving the peak time of the sound.

[0077] The selection of the event timeline needs to meet three objectives simultaneously: reduce the amount of movement in the start time of the target material, reduce the amount of change in the duration of the target material, and reduce soft timing deviations such as phoneme-lip movement, motion-sound effects, subtitles, and shots.

[0078] First, define the total change in start time as the sum of the absolute values ​​of the difference between the start times before and after the adjustment of each target material; define the total change in material duration as the sum of the absolute values ​​of the difference between the durations before and after the adjustment of each target material; and define the total soft timing deviation as the sum of the soft timing deviation penalty values ​​of each target material.

[0079] The smaller the values ​​of the three objectives mentioned above, the closer the time adjustment scheme is to the original data and the lower the multimodal time series deviation. A linear weighted scalarization method from multi-objective optimization is used, multiplying each of the three objectives by its respective weight. , and By summing the results, the multi-objective selection problem is transformed into a single-cost minimization problem. Substituting the specific expressions of the three objectives into the linear weighted relation, we obtain: Equation (4), in Equation (4), This indicates the number of target clips in the current storyboard that are involved in time adjustments. and They represent the first The start time before and after the adjustment of the project logo materials. and They represent the first The duration of the project's logo materials before and after the adjustment. Indicates the first The soft timing deviation penalty value of the project target material. , and Weights greater than 0 , and The priority is determined based on the degree of modification to the original material and the accuracy of multimodal synchronization. When the system prioritizes timing synchronization, the accuracy is improved. When the system places greater emphasis on maintaining the original material's duration and pacing, it improves... or .

[0080] The soft timing deviation penalty value consists of motion-sound effect deviation, phoneme-lip movement deviation, subtitle advance, and shot transition advance. These four deviations are first divided by their respective allowable deviations to convert them into dimensionless normalized deviation values. Then, they are synthesized using a multi-index linear weighted relationship, resulting in: Equation (5), in Equation (5), This represents the normalized deviation exceeding the allowable time deviation between the peak of the action sound effect and the moment of action contact. This represents the normalized deviation that exceeds the allowable time deviation between the phoneme time interval and the corresponding lip shape time interval. This represents the normalized lead time of the complete subtitle preceding the moment the keyword is pronounced. This refers to the normalized lead time, which indicates that the camera transition occurs before the completion of the plot event. , , and The weights are set to be greater than 0 and the sum is set to 1, so that the soft timing deviation penalty value is kept within a uniform evaluation scale. Each weight can be determined according to the degree of influence of different deviations on viewing continuity or the user evaluation results in the labeled samples.

[0081] When a soft timing deviation does not exceed the allowable time deviation of the corresponding event anchor record, its normalized deviation is measured. When it exceeds the allowable time deviation, the excess time is divided by the corresponding allowable time deviation to obtain the normalized deviation. Thus, the closer the soft timing relationship is to the target state, the smaller the soft timing deviation penalty value.

[0082] The candidate time adjustment scheme is generated with video frames as the time step. Taking a video of 25 frames per second as an example, the initial movement range of the target material's start time can be set to 15 frames before and after the original start time. The duration scaling range of the non-critical intervals of the screen can be set to 0.85 times to 1.15 times the original duration. The duration scaling range of the non-keyword intervals of dialogue can be set to 0.90 times to 1.10 times the original duration. The keyword pronunciation intervals remain unchanged.

[0083] The timing arrangement module performs dynamic programming according to the topological order of the plot event graph. The dynamic programming state includes at least the candidate end time of the current plot event and the cumulative adjustment cost from the first plot event to the current plot event. For multiple dynamic programming states with the same candidate end time, only the state with the smallest cumulative adjustment cost is retained to reduce the number of candidate states that need to be calculated later.

[0084] After all story events have been processed, the timing module selects the cumulative adjustment cost from the termination states that satisfy the hard timing constraints. The minimum termination state is defined, and an event timeline is formed based on the start time, end time, and key time points of each target material in that termination state.

[0085] When there are no candidate states that meet the hard timing constraints within the current preset start time movement window and preset duration scaling range, the timing arrangement module expands the search range. For example, it expands the start time movement range from 15 frames before and after to 30 frames before and after, or expands the duration scaling range of static images, ambient sounds, background music, and non-keyword dialogue intervals, and regenerates candidate states.

[0086] When expanding the search scope, the preceding relationships between plot events are not changed, the keyword pronunciation range is not compressed, and candidate states are not obtained by changing the causal relationship of events.

[0087] If no candidate state that meets the hard timing constraints is found within the expanded search range, the timing orchestration module locates the first conflicting event anchor point that causes the preceding and following plot events to be disconnected, based on the plot event graph and the event anchor point chain, and generates a targeted regeneration instruction. The targeted regeneration instruction includes at least the plot event identifier, material modality, conflicting event anchor point identifier, target semantic value, target event stage, and allowed time range.

[0088] For example, when the duration of a certain action clip is too long, causing subsequent dialogue events to be unable to start after the previous event is completed, the directional regeneration instruction can require the multimodal clip generation module to regenerate a shorter action clip while keeping the action subject, action object, and action result unchanged. The newly generated clip then re-enters the necessary anchor point gating and timing arrangement process until a candidate state that meets the hard timing constraints is obtained or the preset number of regenerations is reached.

[0089] Example 8: Example 7 solved the time arrangement problem. This example specifically explains how to identify the emotions in the plot, and how voice-over, facial expressions, background music and camera work revolve around the same emotional changes.

[0090] The script analysis module combines emotional words, degree words, negation words, punctuation, character relationships, and adjacent event states in the dialogue to form target emotional information. This information includes the emotional category, the emotional intensity in the range of 0 to 1, and the direction of change of strengthening, weakening, or maintaining.

[0091] The target emotion does not directly replace the basic parameters of each material, but determines the adjustment range. For dialogue, the speech rate, pitch, volume, pauses and keyword emphasis can be adjusted; for visuals, the range of facial expressions, gaze and body posture can be adjusted; for background music, the loudness, rhythm density and intensity change speed can be changed; and for camera shots, the framing, zoom speed and dwell time can be changed.

[0092] Taking tension as an example, when the emotional intensity is between 0.6 and 0.8, the dialogue speed can be increased by 5% to 15% based on the character's basic speaking speed, ordinary pauses can be shortened by 5% to 20%, pauses before keywords can be extended by 100 to 300 milliseconds, the pitch can be increased by 0.5 to 1.5 semitones, the loudness of background music can be increased by 2 to 5 decibels, and the camera zoom speed can be increased by 5% to 20%. If the target emotion is sadness, the dialogue speed can be reduced by 5% to 15%, the pitch can be reduced by 2% to 8%, the pauses at the end of sentences can be extended by 150 to 400 milliseconds, and the camera movement speed can be reduced by 5% to 20%.

[0093] These adjustment parameters are linked to emotional anchors. Before the emotional anchors take effect, all materials continue with the parameters of the previous event. After the emotional anchors take effect, they are gradually adjusted to the target state of the current event within a preset transition range to avoid sudden changes in visuals or sound.

[0094] The range of a character's facial expressions is limited by the upper limit of emotional intensity. Low-intensity emotions will not correspond to overly exaggerated expressions. Subtitles are displayed word by word or phrase by phrase according to the timing of dialogue words. Complete subtitles containing plot keywords cannot be spoken before the keywords are spoken.

[0095] Example 9: This example combines Figure 5 This document explains the initial AI animation synthesis, necessary anchor point verification, verification deviation calculation, and local correction methods for different anomalies.

[0096] The automatic compositing module writes the target visuals, dialogue, subtitles, sound effects, background music, and shot parameters into the corresponding tracks according to the event timeline. Then, it completes video rendering, audio mixing, subtitle overlay, and shot arrangement to obtain the initial AI comic.

[0097] Each target material simultaneously saves plot event identifiers, anchor point identifiers, original material identifiers, generation parameters, actual start and end times, and key time points. When an anomaly is found during review, this information can be used to directly locate the corresponding material and playback segment.

[0098] Reference Figure 5 As shown, the result verification module segments the initial AI comic according to the plot event identifier, and uses the feature extraction method of Example 5 to re-extract character identity, character actions, action objects, dialogue text, phoneme time, lip-sync type, sound type, sound effect peak, subtitle time, emotional features and shot switching time from the actual playback interval of each event.

[0099] The review begins with the necessary anchor points and is executed strictly in the order of the event anchor point chain. When a necessary anchor point is below the threshold or exceeds the allowable time deviation, that anchor point is identified as the first abnormal position, and the system directly enters local correction without using the higher scores of subsequent auxiliary anchor points to offset this error.

[0100] After all necessary anchor points are qualified, the system will then comprehensively examine the semantic deviation, temporal deviation, and emotional coordination deviation of the current plot event.

[0101] The results review requires evaluating the semantic consistency, accuracy of key timing points, and emotional coherence of the initial AI animation, and reviewing the semantic alignment score. The larger the value, the higher the semantic consistency; therefore, we should first utilize... Semantic consistency scores are converted into semantic bias.

[0102] Normalized time series bias The normalized emotional synergy bias is obtained by normalizing the deviation between the actual key time points and the target time points in the event timeline. The result is obtained by normalizing the distance between the actual voice-over emotion, character expression, background music, and camera changes and the target emotion. Semantic bias, temporal bias, and emotional coherence bias are all limited to the range of 0 to 1, with larger values ​​indicating more severe bias.

[0103] By employing a multi-indicator weighted comprehensive evaluation relationship, the three types of deviations are weighted and summed according to their importance, thus obtaining the review deviation value of the k-th plot event: Equation (6), in Equation (6), This represents the verification semantic alignment score between the actual playback features re-extracted from the initial AI comic and the corresponding event anchor points; This represents the average normalized timing deviation of the actual key time points relative to the event timeline; This indicates the normalized coherence bias between the actual voice-over emotion, character facial expression, and background music emotion and the target emotion. , and This indicates the corresponding weights, all of which are greater than 0 and less than 1. The weights can be verified using labeled qualified and unqualified AI animation clips, and the set of weights with the highest accuracy in distinguishing qualified and unqualified clips is used as the preset weights.

[0104] In one alternative implementation, , and The values ​​are set to 0.5, 0.3, and 0.2 respectively, and the verification deviation threshold is set to 0.20. This is done when all necessary anchor points are qualified and the verification deviation value is... If the deviation value is not greater than 0.20, the current plot event passes the review; if the deviation value is less than 0.20, the event passes the review. When the value is greater than 0.20, based on semantic bias, temporal bias, and sentiment co-operation bias... The size of the contribution determines the correction type.

[0105] When all necessary anchor points have passed the verification, but the verification deviation value When the deviation exceeds the review deviation threshold, the result review module determines the semantic deviation item. Timing deviation term And emotional coordination bias For the verification deviation value The contribution is determined, and the correction type is determined based on the deviation term with the largest contribution.

[0106] When the contribution of the semantic deviation term is the largest, the result verification module determines the material to be regenerated based on the event anchor point with the lowest semantic consistency. When the material to be regenerated is a scene material, the region mask corresponding to the abnormal character or abnormal action object is obtained through character detection, object detection and semantic segmentation. Local conditional regeneration is performed only within the range defined by the region mask, and the previous frame and the next frame of the abnormal material interval are used as the continuity conditions of character identity, scene background, color and lighting.

[0107] When semantic deviations correspond to dialogue material, only the dialogue segment or phoneme range corresponding to the abnormal dialogue anchor point is regenerated, while the speaker's identity, dialogue texts without abnormalities, and other dialogue segments remain unchanged. A cross-gradual interval of 50 to 100 milliseconds is set between the regenerated dialogue audio and the original dialogue audio.

[0108] When the time sequence deviation contributes the most, the result verification module locates the material with the largest deviation between the actual key time point and the event timeline, and only adjusts the start time, end time or key time of the material. For example, if the complete subtitle is earlier than the keyword sounding time, the display time of the complete subtitle is delayed; if the peak of the action sound effect deviates from the action contact time, the time interval of the moving sound effect peak is determined; if the shot switch is earlier than the completion time of the plot event, the shot switch time is delayed.

[0109] When the contribution of the emotion coordination bias term is the greatest, the result review module readjusts the emotion control parameters of the corresponding material based on the emotion anchor point. For dialogue material, it can adjust the speech rate, pitch, volume, pause time and keyword emphasis. For visual material, it can adjust the character's facial expression range and gaze direction. For background music material, it can adjust the loudness envelope and rhythm changes. For camera material, it can adjust the camera movement speed or the duration of the image.

[0110] When the first abnormal event anchor point is associated with subtitle material, sound effect material, or shot material, only the start time, end time, or key time point of the corresponding material is adjusted. When performing local corrections, other material ranges that already meet the necessary anchor point conditions and hard timing constraints remain unchanged.

[0111] The partially corrected material is rewritten to the corresponding track and re-enters the result review process. The result review module only re-detects the plot events that have been corrected and their adjacent plot events. It does not regenerate material for other plot events that have already passed the review. After review, when all necessary anchor points meet the corresponding conditions and the review deviation value is not higher than the review deviation threshold, the final AI animation is output.

[0112] Example 10: To visually illustrate how the modules work together, this example uses the following script text: "On a rainy night, Lin Chuan stood by the window. Suddenly, the sound of breaking glass came from downstairs. Lin Chuan paused for a moment, turned to look out the window, and whispered, 'He still came.'" The input includes Lin Chuan's appearance, clothing, and voice settings, the indoor rainy night scene setting, and a visual style combining black and white line art with low-saturation coloring. The output parameters are set to 25 frames per second, and the target storyboard duration is 6 seconds.

[0113] The script analysis module extracts candidate character Lin Chuan and another present character. Regarding the judgment of the dialogue speaker, Lin Chuan's corresponding location association value, semantic fit value, spatial association value, and dialogue relationship association value are 0.95, 0.88, 0.90, and 0.80, respectively. =0.25 =0.35 =0.20 and Calculated at 0.20, Lin Chuan's subject matching value is 0.886.

[0114] The location association value, semantic adaptation value, spatial association value, and dialogue relationship association value of the other present character are 0.45, 0.40, 0.35, and 0.20, respectively, and the corresponding subject matching value is 0.363. The difference between the two subject matching values ​​is greater than the subject confirmation threshold of 0.15, so Lin Chuan is identified as the subject of the dialogue.

[0115] The event anchor construction module breaks down the storyboard content into events triggered by the sound of breaking glass, Lin Chuan's pause, Lin Chuan's head turn, and Lin Chuan's dialogue, forming corresponding plot event diagrams. The event anchor chain includes, in sequence, the rainy night scene state anchor, the sound of breaking glass anchor, the head turn action start anchor, the head turn completion anchor, the dialogue anchor, the phoneme-lip movement anchor, the tension anchor, and the shot anchor. The allowable time deviation between the peak of the breaking glass sound effect and the action-sound effect anchor is set to 120 milliseconds, and the allowable time deviation between the phoneme time interval and the lip movement time interval is set to 80 milliseconds.

[0116] The multimodal material generation module generates three candidate screen materials, three candidate dialogue materials, and corresponding subtitles, sound effects, background music, and shot materials. During the necessary anchor point gating process, the first candidate screen material is excluded because the position of the action object does not match the action object anchor point. The auxiliary anchor point scores of the second and third candidate screen materials are 0.84 and 0.78, respectively. Therefore, the second candidate screen material is determined as the target screen material.

[0117] Among the three candidate dialogue materials, the second candidate dialogue material meets the requirements in terms of speaker identity and dialogue content, phoneme-lip shape consistency ratio is 92%, and auxiliary anchor comment score is 0.86. Therefore, the second candidate dialogue material is determined as the target dialogue material.

[0118] The timing module uses one frame as the time adjustment step and generates candidate time adjustment schemes within a 15-frame movement range before and after the original time point. One candidate scheme sets the peak of the glass breaking sound at 1.55 seconds, the completion time of the head turning action at 2.65 seconds, the start time of the dialogue at 2.90 seconds, the sound of the keyword "coming" at 4.10 seconds, and the shot switching time at 5.25 seconds. This candidate scheme satisfies all necessary timing constraints and has a lower adjustment cost than other candidate schemes. Therefore, the event timeline is formed based on this candidate scheme.

[0119] After the initial synthesis, the result verification module obtains the verification semantic alignment score. The normalized time series bias is 0.82. The normalized sentiment co-operation bias is 0.35. It is 0.20, according to , and Calculate and verify the deviation value. The value is 0.235, which is higher than the review deviation threshold of 0.20.

[0120] The result verification module determined that the timing deviation contributed the most to the verification deviation value, and located the subtitle display anchor point according to the event anchor point chain. The inspection results showed that the complete subtitle was displayed at 3.55 seconds, earlier than the sound time of the keyword "arrived" at 4.10 seconds. As a result, the verification module adjusted the display time of the complete subtitle to around 4.10 seconds and adjusted the peak value of the glass breaking sound effect to 1.55 seconds.

[0121] Reference Figure 6 As shown, during the initial synthesis, the completion time of the screen action, the voice-over time of the dialogue keywords, the display time of the complete subtitle, the peak time of the action sound effect, and the shot transition time were relatively scattered relative to the corresponding event anchor points. After result review and local correction, the display time of the complete subtitle was adjusted to after the voice-over time of the keywords, the peak time of the action sound effect was moved to the vicinity of the corresponding action trigger time, and the shot transition time was located after the completion time of the action, so that the key moments of various multimodal materials were concentrated within the time range allowed by the corresponding event anchor points.

[0122] After local correction, the semantic alignment score was reviewed. The normalized time series bias is 0.91. The normalized sentiment synergy bias is 0.07. The deviation value is 0.10. The value is 0.086, which is lower than the review deviation threshold, and the final AI animation is output.

[0123] This embodiment also provides an electronic device, including a processor and a memory. The memory stores a computer program. When the processor executes the computer program, it implements the processing of the above-mentioned modules. The subject confirmation threshold, necessary anchor point threshold, allowable time deviation, candidate time window, and verification deviation threshold can be calibrated according to the labeled training samples, output frame rate, and terminal playback conditions.

[0124] The technical scope of this invention is not limited to the content described above. Those skilled in the art can make various modifications and variations to the above embodiments without departing from the technical concept of this invention, and all such modifications and variations should fall within the protection scope of this invention.

Claims

1. An AI-powered automatic animation synthesis system based on multimodal semantic alignment, characterized in that, include: The script analysis module is used to divide the script text into scenes and extract plot information to form a scene semantic package; The event anchor point construction module is used to establish plot event nodes based on the storyboard semantic package, construct a plot event graph based on the preconditions and causal relationships between plot event nodes, and extract the event anchor point chain arranged according to the event process from each plot event node. A multimodal material generation module is used to generate candidate materials based on the plot event nodes and event anchor chain; The semantic alignment module is used to determine the necessary anchor points of the plot event, gate the candidate material item by item according to the event anchor point chain, and after all the necessary anchor points have passed the gate, use auxiliary anchor points to score the candidate material that has passed the gate to determine the target material. The timing arrangement module is used to determine the event sequence according to the plot event diagram, bind the target material to the corresponding event stage, and select the scheme with the least adjustment cost from the time adjustment schemes that satisfy the event precondition relationship to form an event timeline; the automatic synthesis module is used to combine the target material according to the event timeline to obtain the initial AI animation; the result verification module is used to extract playback features from the initial AI animation, locate the first abnormal event anchor point according to the event anchor point chain, determine the material interval associated with the abnormal event anchor point and make local corrections, while keeping other material intervals that satisfy the anchor point conditions and timing constraints unchanged, to obtain the final AI animation.

2. The AI-based automatic animation synthesis system based on multimodal semantic alignment according to claim 1, characterized in that: It also includes a data acquisition module, which is used to acquire the script text, character setting data, scene setting data, target visual style parameters and output configuration parameters, and to unify character encoding, character identification, time unit and material file format; the character setting data includes at least character name, character identification, appearance features, clothing features, voice features and character relationship; the scene setting data includes at least scene name, space type, time conditions, weather conditions, environmental sound type and scene visual features; the output configuration parameters include at least output resolution, frame rate, target duration, audio sampling rate and video encoding format.

3. The AI-based automatic animation synthesis system based on multimodal semantic alignment according to claim 2, characterized in that: For sentences to be parsed where the subject of the action or the speaker of the dialogue is not clearly defined, the script parsing module extracts a set of candidate characters from the current scene and adjacent scenes, determines the location association value, semantic adaptation value, spatial association value, and dialogue relationship association value of each candidate character, and determines the subject matching value based on each association value and its corresponding weight. When the set of candidate characters includes at least two candidate characters, and the difference between the highest subject matching value and the second highest subject matching value is greater than the subject confirmation threshold, the candidate character corresponding to the highest subject matching value is determined as the event subject. When the difference is not greater than the subject confirmation threshold, at least two candidate characters with the highest subject matching values ​​are retained, and the character position, character action, and speaker characteristics are combined for further judgment during the semantic alignment process. When the set of candidate characters includes only one candidate character, that candidate character is determined as the event subject.

4. The AI-based automatic animation synthesis system based on multimodal semantic alignment according to claim 1, characterized in that: The event anchor construction module configures a plot event identifier for each plot event node and divides the plot event into at least two target event stages: preparation stage, trigger stage, execution stage, completion stage, and reaction stage. Each event anchor records the plot event identifier, anchor identifier, anchor type, target semantic value, target event stage, allowed time deviation, and associated material type. Anchor types include main anchor, action anchor, object anchor, dialogue content anchor, phoneme-lip movement anchor, action-sound effect anchor, emotion anchor, scene anchor, background music anchor, and shot anchor. Multiple event anchors corresponding to the same plot event form an event anchor chain according to the target event stage and its occurrence order within the corresponding target event stage. Each candidate material records the corresponding plot event identifier and associated anchor identifier.

5. The AI-based automatic animation synthesis system based on multimodal semantic alignment according to claim 4, characterized in that: The semantic alignment module extracts character identity features, character posture sequences, action object positions, action contact frames, and lip-sync category sequences from candidate scene materials; it extracts speaker features, speech recognition text, phoneme time intervals, and prosodic features from candidate dialogue materials; it extracts subtitle text, word-by-word display time, and complete display time from candidate subtitle materials; it extracts sound categories and peak sound times from candidate sound effect materials; it extracts emotion categories and loudness envelopes from candidate background music materials; and it extracts shot size, camera movement speed, and shot transition times from candidate shot materials. Discrete material features include character identifiers, action object identifiers, dialogue text, sound categories, and shot size, and are judged using identifier consistency. Vector material features include character identity features, dialogue semantic features, and emotion features, and are judged using vector similarity. Continuous material features include action amplitude, phoneme time, lip-sync time, peak sound times, camera movement speed, and shot transition times, and the degree of consistency is determined based on the absolute deviation between the actual values ​​and the target values ​​of the corresponding event anchor points, as well as the maximum allowable deviation.

6. The AI-based automatic animation synthesis system based on multimodal semantic alignment according to claim 5, characterized in that: The necessary anchors include main anchors, action anchors, object anchors, dialogue content anchors, phoneme-lip-sync anchors, and action-sound effect anchors applicable to the current plot event. The auxiliary anchors include emotion anchors, scene anchors, background music anchors, and camera anchors. The semantic alignment module judges the consistency of each applicable necessary anchor item by item according to the order of the event anchor chain. When the consistency of any applicable necessary anchor is lower than the corresponding necessary anchor threshold, the corresponding candidate material is marked as unqualified material. When all applicable necessary anchor points reach the corresponding necessary anchor point threshold and there are auxiliary anchor points applicable to the current material modality, the semantic alignment score of the candidate material is determined according to the consistency degree and corresponding weight of each auxiliary anchor point; when there are no auxiliary anchor points applicable to the current material modality, the candidate materials that have passed the necessary anchor point gating are sorted according to the average consistency degree of each applicable necessary anchor point. The semantic alignment module determines the candidate material with the highest semantic alignment score or the highest average consistency of necessary anchor points as the target material.

7. The AI-based automatic animation synthesis system based on multimodal semantic alignment according to claim 6, characterized in that: The timing arrangement module uses video frames as time steps and generates multiple timing adjustment schemes within a preset start time movement window and a preset duration scaling range. It sets a hard timing constraint that the start time of subsequent plot events must not be earlier than the completion time of their preceding plot events. For timing adjustment schemes that satisfy the hard timing constraint, the adjustment cost is calculated according to the following formula. : ,in, For the target number of materials, and The first Start and end times before and after the adjustment of project logo materials. and The first The duration of the project logo materials before and after adjustment. For the first The soft timing deviation penalty value of the project target material. , and The weight is greater than 0; the soft timing deviation penalty value is determined based on the deviation between the peak of the action sound effect and the moment of action contact, the time interval of phonemes and the time interval of lip movements, the moment of full subtitle display and the moment of keyword pronunciation, and the moment of camera transition and the moment of completion of the plot event; when the corresponding deviation does not exceed the allowable time deviation or there is no corresponding material for the current plot event, the corresponding deviation item is set to 0; the timing arrangement module selects the adjustment cost from the time adjustment schemes that meet the hard timing constraints. Minimal time adjustment scheme.

8. The AI-based automatic animation synthesis system based on multimodal semantic alignment according to claim 7, characterized in that: The timing arrangement module performs dynamic planning according to the topological order of the plot event graph. The dynamic planning state includes the candidate end time of the current plot event and the cumulative adjustment cost. Among multiple dynamic planning states with the same candidate end time, the dynamic planning state with the minimum cumulative adjustment cost is retained. If no candidate state that meets the hard timing constraints is found within the preset start time movement window and preset duration scaling range, the preset start time movement window or the preset duration scaling range of the non-critical material interval is expanded, and the candidate state is regenerated. If no candidate state that meets the hard timing constraints is found within the expanded range, the first event anchor point that causes the constraint conflict is located according to the event anchor point chain, and a directional regeneration instruction including plot event identifier, material modality, anchor point identifier, target semantic value and allowed time range is generated. The directional regeneration instruction is then sent to the multimodal material generation module.

9. The AI-based automatic animation synthesis system based on multimodal semantic alignment according to claim 8, characterized in that: When combining target materials, the automatic synthesis module establishes a correlation record between plot event identifiers, anchor point identifiers, material identifiers, and actual playback intervals. The result verification module queries the correlation record based on the plot event identifiers and anchor point identifiers of abnormal event anchor points to determine the associated materials and actual playback intervals. When the associated material type is video footage, local conditional regeneration is performed based on the region mask corresponding to the abnormal character or abnormal action object, and the previous and next frames of the abnormal material interval are used as continuity conditions for character identity, scene background, and lighting status; when the associated material type is video footage, local conditional regeneration is performed based on the region mask corresponding to the abnormal character or abnormal action object, and the previous and next frames of the abnormal material interval are used as continuity conditions for character identity, scene background, and lighting; when the associated material type is dialogue footage, only the dialogue segment or phoneme interval corresponding to the abnormal event anchor point is replaced, and a cross-gradient interval is set between the replaced audio and the original audio; when the associated material type is subtitle footage, sound effect footage, or shot footage, only the start time, end time, or key moment of the corresponding footage is adjusted; during the local correction process, the video content outside the region mask and other material intervals that meet the necessary anchor point conditions and hard timing constraints remain unchanged.

10. The AI-based automatic animation synthesis system based on multimodal semantic alignment according to claim 9, characterized in that: The result verification module calculates the result according to the following formula: Verification deviation of each plot event : ,in, To review the semantic alignment score, To normalize the time series deviation, the absolute deviation between the actual value and the corresponding target value at each applicable key moment is divided by the corresponding allowable time deviation, and the resulting value is then averaged after being restricted to the range of 0 to 1. To normalize the emotional coordination bias, the distances between the voice-over emotion, character facial expression, background music, and camera changes and the target emotion were converted to the range of 0 to 1 and then averaged. , and All weights are greater than 0 and less than 1, and When any necessary anchor point fails the review, locate the first failing necessary anchor point according to the event anchor point chain and trigger a local correction; when all necessary anchor points pass the review and the review deviation value is met, the correction is triggered. When the deviation exceeds the review threshold, the semantic bias, temporal bias, and emotional synergy bias are determined separately to affect the review deviation value. The contribution is evaluated, and the material is regenerated and corrected, or the time adjustment is adjusted or the sentiment parameter is corrected based on the deviation item with the largest contribution.

Citation Information

Patent Citations

  • Multi-agent-based automatic generation method, system and terminal for micro-drama

    CN119383413A