Multi-agent collaborative creation method and electronic equipment

By constructing a multi-agent system and a task-triggered collaboration mechanism, the problem of insufficient structured planning in the generation of long videos by AI video generation systems is solved, achieving efficient and professional film-level content creation, ensuring narrative coherence and character consistency, and solving the problems of dynamic collaboration, character evolution, narrative planning and film and television style adaptation in existing technologies.

CN121509779APending Publication Date: 2026-02-10BEIJING ELECTRONIC DIGITAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511794670.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-01
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing AI video generation systems lack structured planning capabilities for generating long videos, especially cinematic content. They struggle to meet the dynamic demands of complex narrative scenarios, resulting in unclear video logic, poor character consistency, and prominent audio-visual synchronization issues. Furthermore, they lack advanced narrative planning and multi-scene coordination capabilities, failing to meet the needs of professional film and television production.

Method used

A multi-agent system is constructed, including a director agent, a scene planning agent, and a shot planning agent. The functional boundaries are dynamically adjusted through a task-triggered collaboration mechanism to achieve script division, scene description, and shot decomposition. Combined with the dynamic evolution of plot-driven character characteristics and agent knowledge sharing, the consistency of narrative structure and adaptation to film and television style are ensured.

Benefits of technology

It improves the collaborative efficiency of long video generation, enables high-quality, stylized automated film production, ensures narrative coherence, character consistency, and audio-visual synchronization, and meets the complex needs of professional film and television production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121509779A_ABST
    Figure CN121509779A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-agent collaborative creation method and electronic equipment, and relates to the field of artificial intelligence, in the method, when any agent recognizes a preset task node, a cross-agent collaborative process is triggered, the collaborative process enables the function boundary of at least one agent to be changed, and the change comprises the steps of requesting data from other agents; and at least one of initiating collaboration to other agents and temporarily taking over partial functions of the other agents. According to the method, the function boundary of each agent is automatically adjusted when the task node is triggered, so that the cooperation efficiency of each agent in a complex scene is improved, and shooting of a long video is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of artificial intelligence, in particular to a multi-agent collaborative creation method and an electronic device. BACKGROUND

[0002] With the rapid development of generative artificial intelligence technology, significant breakthroughs have been made in the field of video generation. Models such as Stable Video Diffusion and Sora can generate high-quality short video content. However, in the generation of long videos, especially movie-level content, existing technologies still face major challenges.

[0003] Traditional movie production requires the collaboration of multiple professional roles such as directors, screenwriters, and storyboard artists, involving complex processes such as scriptwriting, scene planning, and shot design. However, current AI video generation systems lack this structured planning capability and are difficult to meet the dynamic needs of complex narrative scenarios.

[0004] The film and television industry has an urgent need for automated content production. According to statistics, traditional movie production requires millions of dollars and years of time, while platform video content also requires several weeks to several months of production cycle. This high-cost, long-cycle production mode severely restricts the efficiency of content production. Therefore, developing a technology that can automatically complete movie-level long video planning and generation, achieving low-cost and high-efficiency content production, has become a key problem to be solved in the field of AI video generation. SUMMARY

[0005] To overcome the problems in the related art, the present disclosure provides a multi-agent collaborative creation method, comprising: constructing a multi-agent system, the multi-agent system comprising a director agent, a scene planning agent, and a shot planning agent, the director agent being configured to divide a script into multiple structured narrative units, the scene planning agent being configured to generate a structured scene description based on the narrative units, and the shot planning agent being configured to perform shot decomposition on the scene description to obtain a shooting script; during the sequential processing of the director agent, the scene planning agent, and the shot planning agent based on the output results of the previous agent, a task-triggered collaboration mechanism is executed; wherein the task-triggered collaboration mechanism is configured to trigger a cross-agent collaboration process when any agent identifies a preset task node, the collaboration process causing the functional boundaries of at least one agent to change, the change including at least one of requesting data from other agents, initiating collaboration with other agents, and temporarily taking over part of the functions of other agents.

[0006] Optionally, the cross-agent collaboration process triggered when any agent identifies a preset task node includes: When the scene planning agent divides the scene boundary, if the scene planning agent identifies that the emotional contrast intensity of the front and rear scenes exceeds a preset threshold, the scene planning agent initiates a narrative logic analysis request to the director agent, so that the director agent provides guidance on the conversion purpose and conversion method for the scene conversion of this emotional contrast based on the management of the narrative unit.

[0007] Optionally, the cross-agent collaboration process triggered when any agent identifies a preset task node includes: When the director agent identifies that the script involves nonlinear narration, the director agent temporarily takes over the task of dividing the time-space boundary that is responsible by the scene planning agent.

[0008] Optionally, the cross-agent collaboration process triggered when any agent identifies a preset task node includes: When the shot planning agent performs shot decomposition, if the shot planning agent identifies that a one-shot-to-end operation is needed, it initiates a parameter coordination request to the scene planning agent, so that the scene planning agent adjusts the element layout and / or light parameters in the scene based on the shot motion trajectory data provided by the shot planning agent.

[0009] Optionally, the cross-agent collaboration process triggered when any agent identifies a preset task node includes: When the director agent analyzes the script, if the director agent identifies a convergence point of multi-line narration, the director agent requests the scene planning agent for time-space correlation data corresponding to the convergence point.

[0010] Optionally, the method further includes: Based on the shooting script, generating a video sequence consistent with the role through a plot-driven role feature dynamic evolution system; Through voice and lip synchronization, subtitle and voice synchronization, sound effect and action synchronization, music and editing synchronization, and emotion and audio-video synchronization, outputting a target movie with audio-video synchronization.

[0011] Optionally, the method further includes: Extracting the original features of the role, the original features including: original visual features, original voiceprint features, and original behavior features; Based on the role growth line extracted from the script, performing dynamic change processing on the original features; The feature spaces of different original features of the role are uniformly mapped through a cross-modal alignment mechanism. The video sequence consistent with the role is obtained through hierarchical consistency verification, wherein the hierarchical consistency verification includes at least one of frame-level verification, shot-level verification, scene-level verification, and global verification.

[0012] Optionally, during the sequential processing of the director agent, the scene planning agent, and the shot planning agent based on the output results of the previous agent, a forward inference process of narrative structure analysis, key element extraction, boundary definition, emotion enhancement, and technology planning is performed. A reverse verification process is performed on the structured parameters output at each stage in the forward inference process to obtain a verification result, wherein the reverse verification process includes reverse verification of the structured parameters output at a previous stage based on the structured parameters output at a subsequent stage. The structured parameters are iteratively corrected according to the verification result.

[0013] Optionally, the method further includes: An agent knowledge sharing pool is constructed for storing intermediate data generated during the processing of the director agent, the scene planning agent, and the shot planning agent. When the cross-agent collaboration process is triggered, the collaborating agent calls the intermediate data of other agents from the agent knowledge sharing pool, and cross-agent parameter sharing is achieved through a knowledge interconnection pool.

[0014] According to a second aspect of the embodiments of the present disclosure, a multi-agent collaborative creation device is provided, including: A construction module is configured to construct a multi-agent system including a director agent, a scene planning agent, and a shot planning agent. The director agent is configured to divide a script into a plurality of structured narrative units. The scene planning agent is configured to generate a structured scene description based on the narrative units. The shot planning agent is configured to perform shot decomposition on the scene description to obtain a shooting script. A collaboration module is configured to perform a task-triggered collaboration mechanism during the sequential processing of the director agent, the scene planning agent, and the shot planning agent based on the output results of the previous agent. The task-triggered collaboration mechanism is configured to trigger a cross-agent collaboration process when any agent identifies a preset task node. The collaboration process causes the functional boundaries of at least one agent to change, including at least one of requesting data from other agents, initiating collaboration with other agents, and temporarily taking over part of the functions of other agents.

[0015] According to a third aspect of the embodiments of the present disclosure, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program. The computer program is executed by a processor to implement the steps of the multi-agent collaborative creation method provided in the first aspect.

[0016] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the executable instructions to perform the steps of the multi-agent collaborative creation method provided in the first aspect.

[0017] Through the above technical solution, the director agent, the scene planning agent and the shot planning agent are no longer fixed division of labor, but when any agent identifies a preset task node, a cross-agent collaboration process is triggered. The collaboration process causes the functional boundary of at least one agent to change, and the change includes at least one of requesting data from other agents, initiating collaboration with other agents, and temporarily taking over part of the functions of other agents. The method automatically adjusts the functional boundaries of the agents when the task node is triggered, which improves the collaboration efficiency of the agents in complex scenes and helps to shoot long videos.

[0018] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF DRAWINGS

[0019] The accompanying drawings are included to provide a further understanding of the present disclosure and constitute a part of the specification, and are used together with the following detailed description to explain the present disclosure, but do not constitute a limitation of the present disclosure. In the drawings: Figure 1 FIG. 1 is a flowchart of a multi-agent collaborative creation method according to an example embodiment.

[0020] Figure 2 FIG. 2 is a schematic diagram of a cross-agent collaboration process according to an example embodiment.

[0021] Figure 3 FIG. 3 is a flowchart of a movie creation method according to an example embodiment.

[0022] Figure 4 FIG. 4 is a flowchart of a role characteristic evolution method according to an example embodiment.

[0023] Figure 5 FIG. 5 is a block diagram of a multi-agent collaborative creation device according to an example embodiment.

[0024] Figure 6is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0025] The exemplary embodiments will be described in detail with reference to the accompanying drawings. In the following description, unless otherwise denoted, the same numbers in different drawings denote the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present disclosure. Instead, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0026] It can be understood that the terms "first", "second", and the like in the present disclosure are used to describe various information, but the information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other, and do not indicate a particular order or importance.

[0027] It can be further understood that, although the operations are described in a particular order in the embodiments of the present disclosure, it should not be understood as requiring the operations to be performed in the particular order shown or in a serial order, or requiring all of the operations to be performed to obtain a desired result. In a particular environment, multi-tasking and parallel processing can be advantageous.

[0028] It should be noted that all actions of obtaining signals, information or data in the present application are performed in compliance with the corresponding data protection regulations and policies of the country where the device is located, and with the authorization of the owner of the corresponding device.

[0029] Traditional film production requires the collaboration of multiple professional roles such as directors, screenwriters, and storyboard artists, involving complex processes such as scriptwriting, scene planning, and shot design. However, current AI video generation systems lack this structured planning capability. Existing long video generation methods have the following limitations: First, in terms of narrative coherence, the generated video is difficult to maintain a logically clear storyline, and there is a lack of natural transition between scenes. Second, in multi-character scenes, the appearance and behavior of the characters are difficult to maintain consistency. Third, the audio-video synchronization problem is prominent, with dialogues and mouth movements, and subtitles and speech often out of sync. In addition, existing solutions can only achieve basic long video synthesis, lacking advanced narrative planning and multi-scene coordination capabilities, resulting in fragmented content that cannot meet professional film production needs. At the same time, the division of labor of intelligent agents in existing technologies is mainly in a static process-based collaboration mode, which is difficult to meet the dynamic needs of complex narrative scenes; the control of character features is limited to static alignment, and cannot achieve natural evolution with the development of the plot; the narrative planning process is linearly advanced, lacking reverse checking and user interaction mechanisms; and there is a lack of precise adaptation capability to different film styles, which collectively restricts the professionalism and flexibility of long video generation.

[0030] In general, the current AI video generation technology has the following bottlenecks in the key technologies of long video production: First, the problem of insufficient dynamic agent cooperation. In related technologies, the division of labor of the agent is fixed, and the functional boundaries cannot be dynamically adjusted according to the complexity of the script and the narrative style, resulting in low efficiency of cooperation in complex scenes.

[0031] Second, the problem of lack of dynamic evolution of character features. Related technologies can only realize static feature alignment of characters, and cannot cope with the natural evolution of features such as appearance, behavior, and emotion of characters with the development of the plot in long videos.

[0032] Third, the lack of narrative planning closed loop. The current linear planning process lacks a reverse checking mechanism, making it difficult to ensure the consistency of decisions at each stage, and lacks a user interaction interface, making it difficult to meet the needs of personalized creation.

[0033] Fourth, the weak ability to adapt to film styles. It is difficult to convert the unique styles of different directors into quantifiable generation parameters, making it difficult to realize stylized long video creation.

[0034] Fifth, the difficulty of coding professional film knowledge. Traditional methods are difficult to convert professional experience such as directors and photographers into computable generation rules.

[0035] The above problems result in the current AI system being able to only generate short video clips or low-quality long videos, and being unable to realize truly automated, high-quality, and stylized film production. In view of this, the present disclosure provides a multi-agent collaborative creation method and an electronic device.

[0036] Figure 1 is a flowchart of a multi-agent collaborative creation method according to an example embodiment. As Figure 1 shown, the method includes the following steps: In step S11, a multi-agent system is constructed, the multi-agent system including a director agent, a scene planning agent, and a shot planning agent, the director agent being configured to divide a script into a plurality of structured narrative units, the scene planning agent being configured to generate a structured scene description based on the narrative units, and the shot planning agent being configured to perform shot decomposition on the scene description to obtain a shooting script.

[0037] In the present disclosure, an agent refers to an intelligent decision-making module with specific domain expertise, such as a director agent referring to an intelligent decision-making module with director expertise. The division of labor of the director agent, the scene planning agent, and the shot planning agent is different.

[0038] Specifically, the director agent is responsible for deep semantic understanding and structural decomposition of the input script, so as to divide the script into multiple structured narrative units. Taking time sequence division as an example, the narrative units can include the beginning, development, climax, and ending of the story. Or the narrative units are the plot paragraphs divided based on theme continuity.

[0039] The scene planning agent generates structured scene descriptions based on the narrative units output by the director agent. Each scene description defines a specific environment in which the story takes place, such as a scene description including time and space coordinates, participating roles, emotional tone, and visual style guidelines, etc.

[0040] The shot planning agent decomposes each scene description into shots, outputting executable shooting scripts. The shooting script contains specific shot sequences, and each shot defines its visual parameters, such as: shot type (such as close-up, medium shot, long shot, etc.), camera movement (such as push, pull, pan, and move, etc.), duration, character positioning, etc.

[0041] The following will explain the process that each agent may involve in turn.

[0042] Director agent: (1) Identify the core narrative structure.

[0043] Among them, the purpose of the director agent to identify the core narrative structure is to grasp the overall framework of the script. For example, identify that the script is a three-act structure, or other multi-act narrative forms.

[0044] (2) Divide the logical plot paragraphs.

[0045] Further subdivided into paragraphs, such as dividing the three key confrontations between the main character and the villain into three plot paragraphs.

[0046] (3) Label key plot turning points.

[0047] Mark the key nodes in the story that change the fate of the characters or drive the plot development, such as plot reversal, truth revelation, etc.

[0048] (4) Maintain the role relationship graph.

[0049] Build and update the relationships between characters, such as including social relationships and emotional relationships between characters.

[0050] (5) Generate decomposition explanation documents.

[0051] This process outputs technical documents that comprehensively record the analysis results described above, providing creative guidance for subsequent agents, such as the scene planning agent can perform scene division based on the explanation documents.

[0052] Scene planning agent: (1) Scene boundary division.

[0053] The scene boundary can be divided according to the space-time conversion or the mood change. Specifically, a new scene is divided when the space-time coordinates change significantly. A new scene is divided when the scene emotional tone changes significantly.

[0054] (2) Scene element annotation.

[0055] The scene elements include a list of participating roles, an emotional tone description, a visual style description, etc.

[0056] (3) Generation of a cinematography guide.

[0057] The cinematography includes light design requirements such as light source angle, color temperature, etc., and prop arrangement requirements such as which props to arrange on the table.

[0058] (4) Output of a structured scene description.

[0059] The structured scene description is represented by a set P = {P1, P2…P N}, and each scene description can include at least one of the scene boundary, the scene element, and the cinematography guide.

[0060] Lens planning agent: (1) Determining the lens sequence.

[0061] The lens sequence is used to plan the visual narration order of the scene, such as using an establishing shot at the beginning, then switching to a main shot, inserting a reaction shot, and adding a close-up shot at a critical moment. The lens sequence is obtained by lens identification and insertion time.

[0062] (2) Annotation of lens parameters.

[0063] The lens parameters include shot type (such as close-up, medium shot, long shot, etc.), camera movement (such as push, pull, pan, and shift, etc.), duration, character positioning, etc.

[0064] (3) Generation of a time code with subtitle text.

[0065] The subtitle text is generated to match the voice rhythm accurately, and each subtitle text has an entry and exit time code to ensure sound and picture synchronization.

[0066] (4) Output of the final shooting script.

[0067] The shooting script includes the lens parameters of i scene j shot.

[0068] The above multiple agents work through division of labor to simulate the cooperation of various staff in professional film production.

[0069] The director agent, the scene planning agent, and the shot planning agent are not fixed division of labor, but adopt a task node triggered cooperation mode to automatically adjust the function boundary. Details are as follows.

[0070] In step S12, during the sequential processing of the director agent, the scene planning agent, and the shot planning agent based on the output results of the previous agent, a task triggered cooperation mechanism is executed; the task triggered cooperation mechanism is configured to trigger a cross-agent cooperation process when any agent identifies a preset task node, the cooperation process causes the function boundary of at least one agent to change, the change includes at least one of requesting data from other agents, initiating cooperation with other agents, and temporarily taking over part of the functions of other agents.

[0071] Exemplarily, in the multi-agent collaborative creation process, the director, the scene, and the shot sequentially execute their respective main processes, and the task triggered cooperation mechanism is introduced in the embodiment of the disclosure, which breaks the fixed division of labor boundary of each agent and triggers the cooperation process across the agents when a preset task node is identified.

[0072] In some feasible embodiments, the above-mentioned triggering of the cooperation process across the agents when any agent identifies a preset task node includes any one or multiple steps of steps S21 to S24 as shown in the figure. Figure 2

[0073] Step S21, when the scene planning agent divides the scene boundary, if the scene planning agent identifies that the emotional contrast intensity of the scenes before and after exceeds a preset threshold, the scene planning agent initiates a narrative logic analysis request to the director agent, so that the director agent provides guidance on the conversion purpose and conversion method for the scene conversion of this emotional contrast based on the management of the narrative unit.

[0074] In this embodiment, the emotional values of each scene can be set, and the difference between the emotional values of two consecutive scenes is taken as the emotional contrast intensity of the scenes before and after. Exemplarily, the scene planning agent processes two consecutive scenes, the scene before is a happy birthday party with an emotional value of 0.9, and the scene after is a sad farewell ceremony with an emotional value of 0.1, and the emotional contrast intensity is 0.8. The preset threshold can be set to a value greater than 0 and less than 1, for example, the preset threshold is set to 0.7. The emotional contrast intensity is calculated to be greater than the preset threshold, and the task node of strong emotional contrast is identified. The scene planning agent initiates a narrative logic analysis request to the director agent, so as to upgrade the simple technical switching problem to the narrative logic level for solving. The director agent provides guidance on the conversion purpose and conversion method based on the management of the overall narrative unit, so as to avoid harsh scene switching and help the emotional switching transition.

[0075] ​Step S22, when the director agent identifies that the script involves nonlinear narration, the director agent temporarily takes over the spatiotemporal boundary division task originally responsible by the scene planning agent.

[0076] In this embodiment, nonlinear narration is preset as a task node with high complexity. When the director agent analyzes the script, it identifies that the story structure is a nonlinear narration of “dream and reality interwoven”. To prevent the division according to the conventional rules from causing the scene to be fragmented, the director agent temporarily takes over the spatiotemporal boundary division task originally responsible by the scene planning agent. The director agent sets the scene types of reality line, dream line, and virtual-real interweaving, and sets the visual identifier and conversion rule for each scene type, thereby ensuring the coherence of the complex story structure.

[0077] Step S23, when the shot planning agent is performing shot decomposition, if the shot planning agent identifies that it needs to implement one-shot-to-end panning, it initiates a parameter coordination request to the scene planning agent, so that the scene planning agent adjusts the element layout and / or light parameters in the scene based on the shot motion trajectory data provided by the shot planning agent.

[0078] In this embodiment, one-shot-to-end panning is preset as a task node with high complexity. When one-shot-to-end panning is involved, the shot planning agent initiates a parameter coordination request to the scene planning agent and provides the shot motion trajectory data planned by the shot planning agent to the director agent. The scene planning agent dynamically adjusts the element layout and light parameters in the scene based on the shot motion trajectory data. For example, the scene planning agent finds that according to the original layout, a vase will block the main character when the shot motion reaches the 5th second, so the scene planning agent automatically moves the vase away; and the scene planning agent adjusts the light angle to ensure that the main character's face is evenly illuminated on the action path. This coordination ensures the execution effect of complex shots.

[0079] Step S24, when the director agent analyzes the script, if the director agent identifies the intersection point of multi-line narration, the director agent requests the spatiotemporal correlation data corresponding to the intersection point from the scene planning agent.

[0080] In this embodiment, the intersection point of multi-line narration is identified as a preset task node. To accurately analyze the logic of the intersection point, the director agent requests the spatiotemporal correlation data corresponding to the intersection point from the scene planning agent. The scene planning agent provides the spatiotemporal correlation data, and the director agent can use this data to accurately construct the tension and logical rigor of the intersection point.

[0081] In some embodiments, when the shot design involves a "close-up group scene depicting character relationships," the director's AI agent is simultaneously invoked to analyze the character close-up content, allowing the camera AI agent to refine the camera movement design. Permissions are automatically reset upon completion of the collaboration, and the system records the collaboration case for future optimization of improvement modules.

[0082] Through the above technical solution, the director agent, scene planning agent, and shot planning agent no longer have fixed divisions of labor. Instead, they trigger a cross-agent collaboration process when a preset task node is identified. Specifically, when any agent identifies a preset task node, it triggers a cross-agent collaboration process. This collaboration process causes a change in the functional boundaries of at least one agent. These changes include at least one of the following: requesting data from other agents, initiating collaboration with other agents, and temporarily taking over some functions of other agents. This method automatically adjusts the functional boundaries of each agent when a task node is triggered, improving the collaboration efficiency of agents in complex scenes and facilitating the shooting of long videos.

[0083] In some feasible embodiments, the method further includes: during the sequential processing of the director agent, scene planning agent, and shot planning agent based on the output of the previous agent, executing a forward reasoning process of narrative structure analysis, key element extraction, boundary definition, emotion enhancement, and technical planning; performing a reverse verification process on the structured parameters output at each stage of the forward reasoning process to obtain verification results, wherein the reverse verification process includes: reverse verification of the structured parameters output at the previous stage based on the structured parameters output at the subsequent stage; and iteratively correcting the structured parameters according to the verification results.

[0084] Specifically, the sequential processing of the aforementioned intelligent agent is designed as a cyclical iterative chain-like thinking and planning process, which includes: 1. Narrative structure analysis stage.

[0085] This stage primarily analyzes the deep structure of the input script, which may include the following four aspects.

[0086] (1) Mark the main plot points (such as the inciting event, climax, ending, etc.).

[0087] (2) Identify changes in the emotional curve.

[0088] (3) Extract the key role interaction relationships.

[0089] (4) Construct a story timeline (such as flashbacks, parallel narrative annotations, etc.).

[0090] 2. Key element extraction stage.

[0091] (1) Core character profiles (such as appearance, personality, and relationships).

[0092] (2) Important prop symbol system.

[0093] (3) Scene time and space coordinates (such as era, season, day and night).

[0094] (4) Thematic keywords (such as redemption, growth).

[0095] This stage allows for the creation of a multimodal feature library based on the extracted key elements.

[0096] 3. Boundary definition stage.

[0097] This stage yields structured narrative units, specifically comprising the following four parts.

[0098] (1) Divide the structure into acts (such as three-act plays and five-act plays).

[0099] (2) Mark the scene transition trigger conditions.

[0100] (3) Set the lens assembly rules.

[0101] (4) Generating transition effects (such as fade-out, wipe).

[0102] 4. The stage of enhanced emotion.

[0103] This stage is used to inject cinematic expressive elements, including the following four aspects.

[0104] (1) Assign visual style codes (e.g., film grain).

[0105] (2) Design color emotion mapping.

[0106] (3) Configure ambient sound effects prompts.

[0107] (4) Mark the music mood curve.

[0108] 5. Technology planning stage.

[0109] This stage generates executable parameters, including the following four aspects.

[0110] (1) Camera parameter set (such as focal length, aperture, frame rate).

[0111] (2) Lighting scheme (such as a three-point lighting diagram).

[0112] (3) Role performance guidance (such as micro-expressions and body language).

[0113] (4) Post-production special effects markers (such as CGI, compositing requirements).

[0114] After the above five stages are executed sequentially, the structured intermediate representations generated in each stage need to be reverse-verified by the verification module. Specifically, the verification module performs the following reverse verification stages.

[0115] 6. Reverse verification stage.

[0116] (1) Verify the consistency of style in the emotion enhancement stage by reverse verification of parameters in the technical planning stage: compare the lighting, camera movement parameters, etc. extracted in the technical planning stage with the "style tag library" output in the emotion enhancement stage (e.g., "suspense style" needs to match low light and handheld camera movement features). Mismatches trigger parameter correction.

[0117] (2) Verify the rationality of scene transition in the boundary definition stage by reverse verification of parameters in the emotion enhancement stage: Substitute the emotion curve data in the emotion enhancement stage into the "scene transition rule base" in the boundary definition stage (e.g., "great sorrow → great joy" requires setting a buffer scene) to detect the conflict of transition logic.

[0118] (3) Verify the completeness of elements in the key element extraction stage by reverse verification of parameters in the boundary definition stage: compare the list of scene elements in the boundary definition stage with the "core element content" in the key element extraction stage to check whether any essential elements are missing (such as whether the protagonist's token is missing).

[0119] (4) Verify the logical coherence of the narrative structure analysis stage by reverse-engineering the parameters of the key element extraction stage: Substitute the character relationship data of the key element extraction stage into the "logic chain test" of the narrative structure analysis stage to verify whether there is a relationship contradiction (such as a dead character appearing in a subsequent scene).

[0120] In addition, this disclosure also provides a user interaction module to enable an explainable and interventionable decision-making process.

[0121] 6. User interaction module.

[0122] At each stage, a "user correction interface" is set up, allowing users to adjust intermediate parameters through natural language (such as "The lighting in this scene is too dim and does not match the warm family atmosphere") or example images (upload a reference lighting style image). The system uses the CLIP model to convert user input into quantitative constraints (such as "Adjust the light brightness parameter from 0.3 to 0.7, and the color temperature from 5000K to 3000K"), and automatically synchronizes them to subsequent stages and reverse verification.

[0123] The cyclical iterative thinking and planning mechanism provided in this disclosure ensures the consistency and traceability of decisions at each stage, and ultimately outputs a complete shooting blueprint that meets the user's needs.

[0124] To support the aforementioned dynamic collaboration and iterative planning, this disclosure also constructs an agent knowledge sharing pool to store intermediate data generated by each agent, such as the "role relationship graph" of the director agent and the "emotional tone parameters" of the scene planning agent. When any agent needs to initiate collaboration, it retrieves the required data from the sharing pool through the knowledge interoperability connection pool, thereby achieving efficient and transparent cross-agent parameter sharing.

[0125] For example, after the camera planning agent calls the "character redemption plot marker" data generated by the director agent, it first identifies the "redemption trigger scene" marked in the data (such as character A blocking an attack to protect character B) and the associated "emotional transition tag" (guilt → relief). Then, it matches the "emotional symbol library" pre-stored by the scene agent and extracts the visual elements corresponding to the emotion (soft backlight, slow motion transition). Finally, it generates a camera scheme that conforms to the narrative emotional logic—using a close-up freeze-frame at the moment of the attack, switching to a backlight close-up when switching to character A's facial expression, and using a slow push-in shot to enhance the emotional impact of the redemption moment. After the scene planning agent calls the "long shot motion trajectory" data of the camera planning agent, it first analyzes the "camera focus shift path" marked in the trajectory (gradually focusing from a group panorama to character C's hand movements). Then, it combines its own scene layout data to identify possible visual obstructions in the path (such as a pillar in the foreground). Subsequently, it automatically adjusts the arrangement of scene elements—placing the pillar in the back and pre-setting prop details at the end of the focus shift (the letter in character C's hand reveals key text) to ensure that the narrative information is completely conveyed during the camera movement.

[0126] In some of these embodiments, such as Figure 3 As shown, this method also includes: In step S31, based on the shooting script, a video sequence consistent with the characters is generated through a plot-driven character feature dynamic evolution system.

[0127] In step S32, the target movie with synchronized audio and video is output by synchronizing speech with lip movements, subtitles with speech, sound effects with actions, music with editing, and emotions with audio and video.

[0128] Among them, such as Figure 4 As shown, step S31 further includes: Step S41: Extract the character's original features, which include: original visual features, original voiceprint features, and original behavioral features.

[0129] For example, and not as a limitation, the original visual features can be character faces, body shapes, etc., generated by visual models such as CLIP; the original voiceprint features can be speech spectrum features extracted by models such as ECAPA-TDNN; and the original behavioral features can be data such as distinctive gestures and gait frequencies.

[0130] In addition, emotional feature modeling can be performed to generate character emotional curves based on text analysis. A character feature mapping table is established, including original visual features, original voiceprint features, original behavioral features, and emotional features.

[0131] Step S42: Based on the character growth line extracted from the script, perform dynamic change processing on the original features.

[0132] This method is based on the character growth trajectory extracted from the script, such as from innocence to darkness, and dynamically processes these features. For example, through Temporal Generative Adversarial Network (T-GAN), the character's facial features are made to gradually enhance their eye makeup as the plot progresses; through a voice emotion conversion model, the character's voice is made to speak faster and with a higher pitch when angry, so that the character's features change with the character's growth trajectory in the long video, resulting in a natural transition.

[0133] Specifically, visual features, behavioral features, and voiceprint features can be dynamically processed from the following aspects: Visual features: Through temporal generative adversarial networks (T-GAN), the character's appearance changes naturally as the plot progresses (such as wrinkles as the character ages and skin color changes when the character is emotionally agitated), and the magnitude of the changes is strongly correlated with plot points (such as the character's eyes becoming sharper after the "betrayal incident").

[0134] Voiceprint features: Through a voice emotion conversion model, the character's voice changes with emotion, such as increasing speech speed and pitch when angry.

[0135] Behavioral characteristics: Based on changes in the character relationship graph (such as "from ally to enemy"), automatically adjust the character's typical action patterns, such as changing from "frequent handshakes" to "distancing physical distance".

[0136] Step S43: Ensure a unified mapping of the feature space of different original features of the character through a cross-modal alignment mechanism.

[0137] In this embodiment, the role of the cross-modal alignment mechanism is to ensure that the features of different modalities remain synchronized as they change, for example, to ensure that the mouth movements and the speech waveform are accurately matched in each frame.

[0138] For example, cross-modal alignment mechanisms include the following aspects: (1) Visual and speech alignment: Through cross-modal contrastive learning, ensure that mouth movements are synchronized with speech waveforms.

[0139] (2) Behavior and scene adaptation: Verify the rationality of actions based on the physics engine.

[0140] (3) Temporal continuity detection: calculate SSIM > 0.8 and PSNR > 30dB for adjacent frames.

[0141] SSIM refers to the Structural Similarity Index, where 1 indicates that two images are completely identical. PSNR refers to the Peak Signal-to-Noise Ratio, where a higher value indicates better image quality or less distortion.

[0142] (4) Dynamic feature compensation: LSTM is used for feature smoothing of long shots.

[0143] Step S44: Obtain a video sequence with consistent roles through hierarchical consistency verification, wherein the hierarchical consistency verification includes at least one of frame-level verification, shot-level verification, scene-level verification, and global verification.

[0144] Among them, the layered consistency verification constructs a multi-level quality inspection system. For example, frame-level verification ensures that the proportions of a character's facial features are not distorted in a single frame; shot-level verification ensures that the character's clothing and accessories do not change within the same shot; scene-level verification ensures that the voiceprint of a character remains recognizable and similar when the character moves between scenes; and global verification checks whether the evolution of character relationships is logical from a story perspective.

[0145] Layered consistency verification can include the following aspects: (1) Frame-level verification: The error of facial features ratio in a single frame is <7%, and the matching degree of expression and emotion value is >90%.

[0146] (2) Shot-level verification: Clothing / accessories matching degree between shots > 90%, and behavioral pattern consistency > 85%.

[0147] (3) Scene-level verification: Cross-scene voiceprint cosine similarity > 0.93, and the rationality of the change in the emotion curve > 90%.

[0148] (4) Global verification: There are no logical conflicts in the role relationship graph.

[0149] In addition, this method may also include an exception handling module, which can specifically execute the following strategies: (1) Feature drift detection: Set a threshold to trigger regeneration. If FID > 10, it indicates that the feature corresponding to the feature identifier has drifted, and the error will be corrected in time.

[0150] (2) Emergency replacement strategy: Enable alternative generation channels (such as switching to LoRA fine-tuning model).

[0151] (3) Progressive repair: GAN inpainting is applied to problematic frames.

[0152] (4) Version rollback: Automatically call the most recently valid feature cache.

[0153] This solution achieves high consistency of role IDs on the MoviePrompts test set through feature space decoupling and dynamic fusion techniques.

[0154] While generating the video sequence, the following audio-video synchronization method is performed in this embodiment: (1) Synchronization of speech and mouth shape: A 3D face mesh model is used to achieve frame-level alignment of phonemes and mouth shapes.

[0155] (2) Synchronization of subtitles and audio: The display duration of subtitles is dynamically adjusted to match the rhythm and pauses of the audio.

[0156] (3) Sound effects are synchronized with the action and music is synchronized with the editing: sound effects such as footsteps are generated based on physical simulation, and music accents are matched according to the beat points of the camera switching.

[0157] (4) Music and editing are synchronized: the camera switches are automatically matched according to the beat.

[0158] (5) Synchronization of emotion with audio and video: Ensure that the melody of the background music and the style of the sound effects are highly consistent with the emotional tone of the picture.

[0159] In some of the feasible embodiments, this disclosure also includes film and television knowledge quantification encoding for digitizing professional parameters.

[0160] (1) Shot Scale Quantification: Define 9 levels of standards (from extreme long shot 1.0 to close-up 9.0).

[0161] (2) Camera motion encoding: Establish a motion vector space (translation and rotation speed 0~100). (3) Lighting parameters: Standardized three-point lighting (the ratio of main light, fill light and backlight intensity is 6:3:1) (4) Color Science: Converting DCI-P3 color gamut to Log-C curve.

[0162] Furthermore, embodiments of this disclosure also include dynamic constraint optimization.

[0163] (1) Photographic rules check: verify the 180-degree axis principle and the rule of thirds composition.

[0164] (2) Aesthetic evaluation: Calculate the dynamic composition score (based on the golden ratio).

[0165] (3) Physical rationality: Detect the consistency of shadow direction / object proportion.

[0166] (4) Style filter: Forced to match the preset visual dictionary.

[0167] In some embodiments, this disclosure also constructs a director style knowledge graph. Specifically, for multiple director styles, a knowledge graph including the following elements is constructed: (1) Narrative structure template.

[0168] (2) Lens language parameter set.

[0169] (3) Color Mood Mapping Table.

[0170] (4) Camera movement feature library.

[0171] When a user inputs a script and specifies a reference style, the system automatically calculates the weights through a style fusion model, such as "60% for healing colors and 40% for symmetrical composition." The system then injects these style parameters into the decision-making process of the three-level agent, enabling the agent to create films based on the reference style and meet the user's personalized and stylized creative needs.

[0172] As can be seen from the above description of the method, the embodiments of this disclosure achieve automated movie generation through the following technical solutions: First, a dynamic hierarchical collaborative intelligent agent architecture: This application embodiment constructs a three-level intelligent agent that can cooperate with each other, namely director, scene planner and shot planner, to realize the dynamic evolution of intelligent agent functions and cross-intelligent agent knowledge transfer, and simulate the dynamic collaborative process of various roles in professional film production.

[0173] Second, the cyclical iterative chain thinking (CoT) planning mechanism employs a cyclical process of five stages: narrative analysis, key element extraction, boundary definition, emotional enhancement, and technical planning, along with reverse verification. A user interaction module is incorporated to achieve an explainable and interventionist decision-making process.

[0174] Third, a plot-driven dynamic evolution system for character features: through an ID-aware generation model, a temporal feature evolution algorithm, and a cross-modal constraint closed loop, it ensures the natural evolution and consistency of character image, voice, behavior, and emotions in multiple scenarios as the plot develops.

[0175] Fourth, synchronized audio and video generation: combining a visual generation model with a speech synthesis system to achieve frame-level synchronization of subtitles, speech, and video.

[0176] Fifth, the film and television style adaptive generation engine: constructs a director style knowledge graph to achieve quantitative encoding and adaptive conversion of different film and television styles.

[0177] Sixth, film and television knowledge encoding: transforming professional photography parameters (shot size, camera movement, etc.) into quantifiable generation constraints.

[0178] Figure 5 This is a block diagram of a multi-agent collaborative creation device according to an exemplary embodiment. (Refer to...)Figure 5 The device includes a construction module 510, a collaboration module 520, a feature evolution module 530, and an audio / video synchronization module 540.

[0179] Module 510 is used to build a multi-agent system, which includes a director agent, a scene planning agent, and a shot planning agent. The director agent is configured to divide the script into multiple structured narrative units, the scene planning agent is configured to generate structured scene descriptions based on the narrative units, and the shot planning agent is configured to decompose the scene descriptions into shots to obtain the shooting script.

[0180] The collaboration module 520 is used to execute a task-triggered collaboration mechanism during the sequential processing of the director agent, scene planning agent, and shot planning agent based on the output results of the previous agent. The task-triggered collaboration mechanism is configured such that when any agent identifies a preset task node, it triggers a cross-agent collaboration process. The collaboration process causes a change in the functional boundary of at least one agent. The change includes at least one of requesting data from other agents, initiating collaboration with other agents, and temporarily taking over some functions of other agents.

[0181] In some possible implementations, the collaboration module 520 is also used to trigger the scene planning agent to initiate a narrative logic analysis request to the director agent when the scene planning agent identifies that the intensity of the emotional contrast between the preceding and following scenes exceeds a preset threshold when the scene planning agent is dividing the scene boundaries. This allows the director agent to provide guidance on the purpose and method of scene transition based on the management of narrative units.

[0182] In some possible implementations, the collaboration module 520 is also used to temporarily take over the spatiotemporal boundary delineation task that was previously handled by the scene planning agent when the director agent recognizes that the script involves non-linear narrative.

[0183] In some possible implementations, the collaboration module 520 is also used to initiate a parameter collaboration request to the scene planning agent when the shot planning agent identifies that a one-shot camera movement is required during shot decomposition. This allows the scene planning agent to adjust the element layout and / or lighting parameters within the scene based on the shot motion trajectory data provided by the shot planning agent.

[0184] In some possible implementations, the collaboration module 520 is also used to request spatiotemporal correlation data corresponding to the intersection point from the scene planning agent when the director agent is parsing the script and identifies the intersection point of multiple narratives.

[0185] In some possible implementations, feature evolution module 530 is used to generate a consistent video sequence for a character based on a shooting script through a plot-driven character feature dynamic evolution system. And audio / video synchronization module 540 is used to output a target film with synchronized audio and video through synchronization of speech and lip movements, subtitles and speech, sound effects and actions, music and editing, and emotion and audio / video.

[0186] In some possible implementations, the feature evolution module 530 is also used to extract the original features of the character, including: original visual features, original voiceprint features, and original behavioral features; to dynamically change the original features based on the character growth line extracted from the script; to ensure a unified mapping of the feature space of different original features of the character through a cross-modal alignment mechanism; and to obtain a consistent video sequence of the character through hierarchical consistency verification, wherein the hierarchical consistency verification includes at least one of frame-level verification, shot-level verification, scene-level verification, and global verification.

[0187] Figure 6 This is a block diagram illustrating an electronic device 600 according to an exemplary embodiment. For example... Figure 6 As shown, the electronic device 600 may include a processor 601 and a memory 602. The electronic device 600 may also include one or more of a multimedia component 603, an input / output (I / O) interface 604, and a communication component 605.

[0188] The processor 601 controls the overall operation of the electronic device 600 to complete all or part of the steps in the multi-agent collaborative creation method described above. The memory 602 stores various types of data to support the operation of the electronic device 600. This data may include, for example, instructions for any application or method operating on the electronic device 600, and application-related data such as contact data, sent and received messages, pictures, audio, video, etc. The memory 602 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, disk, or optical disk. The multimedia component 603 may include a screen and audio components. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in memory 602 or transmitted via communication component 605. The audio component also includes at least one speaker for outputting audio signals. I / O interface 604 provides an interface between processor 601 and other interface modules, such as a keyboard, mouse, buttons, etc. These buttons may be virtual or physical buttons. Communication component 605 is used for wired or wireless communication between the electronic device 600 and other devices. Wireless communication may include Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination thereof; therefore, the corresponding communication component 605 may include a Wi-Fi module, a Bluetooth module, or an NFC module.

[0189] In an exemplary embodiment, the electronic device 600 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to execute the multi-agent collaborative creation method described above.

[0190] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implement the steps of the multi-agent collaborative creation method described above. For example, the computer-readable storage medium may be the memory 602 including the program instructions, which may be executed by the processor 601 of the electronic device 600 to complete the multi-agent collaborative creation method described above.

[0191] In another exemplary embodiment, a computer program product is also provided, which includes a computer program executable by a processor, which, when executed by the processor, implements the steps of the multi-agent collaborative creation method described above.

[0192] The preferred embodiments of this disclosure have been described in detail above with reference to the accompanying drawings. However, this disclosure is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this disclosure, various simple modifications can be made to the technical solutions of this disclosure, and these simple modifications all fall within the protection scope of this disclosure.

[0193] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, this disclosure will not describe the various possible combinations separately.

[0194] Furthermore, various different embodiments of this disclosure can be combined in any way, as long as they do not violate the spirit of this disclosure, they should also be regarded as the content disclosed in this disclosure.

Claims

1. A multi-agent collaborative creation method, characterized in that, include: A multi-agent system is constructed, comprising a director agent, a scene planning agent, and a shot planning agent. The director agent is configured to divide the script into multiple structured narrative units. The scene planning agent is configured to generate structured scene descriptions based on the narrative units. The shot planning agent is configured to decompose the scene descriptions into shots to obtain a shooting script. During the sequential processing of the director agent, the scene planning agent, and the shot planning agent based on the output of the previous agent, a task-triggered collaboration mechanism is executed. The task-triggered collaboration mechanism is configured such that when any agent identifies a preset task node, a cross-agent collaboration process is triggered. The collaboration process causes a change in the functional boundary of at least one agent. The change includes at least one of requesting data from other agents, initiating collaboration with other agents, and temporarily taking over some functions of other agents.

2. The method according to claim 1, characterized in that, When any agent identifies a preset task node, a cross-agent collaboration process is triggered, including: When the scene planning agent is dividing the scene boundaries, if the scene planning agent detects that the intensity of the emotional contrast between the preceding and following scenes exceeds a preset threshold, the scene planning agent is triggered to initiate a narrative logic analysis request to the director agent, so that the director agent can provide guidance on the purpose and method of scene transition for this emotional contrast based on the management of the narrative unit.

3. The method according to claim 1, characterized in that, When any agent identifies a preset task node, a cross-agent collaboration process is triggered, including: When the director agent recognizes that the script involves non-linear narrative, the director agent temporarily takes over the spatiotemporal boundary division task that was previously handled by the scene planning agent.

4. The method according to claim 1, characterized in that, When any agent identifies a preset task node, a cross-agent collaboration process is triggered, including: When the shot planning agent is performing shot decomposition, if the shot planning agent recognizes that a one-shot camera movement is required, it initiates a parameter coordination request to the scene planning agent, so that the scene planning agent adjusts the element layout and / or lighting parameters in the scene based on the shot motion trajectory data provided by the shot planning agent.

5. The method according to claim 1, characterized in that, When any agent identifies a preset task node, a cross-agent collaboration process is triggered, including: When the director agent is parsing the script, if the director agent identifies a point of intersection of multiple narratives, the director agent requests spatiotemporal correlation data corresponding to the point of intersection from the scene planning agent.

6. The method according to claim 1, characterized in that, The method further includes: Based on the shooting script, a video sequence with consistent characters is generated through a plot-driven dynamic evolution system of character characteristics. By synchronizing speech with lip movements, subtitles with speech, sound effects with actions, music with editing, and emotions with audio and video, the target movie is output with synchronized audio and video.

7. The method according to claim 6, characterized in that, The generation of consistent video sequences based on the shooting script and through a plot-driven character feature dynamic evolution system includes: Extract the original features of the character, including: original visual features, original voiceprint features, and original behavioral features; Based on the character growth line extracted from the script, the original features are dynamically modified. A cross-modal alignment mechanism ensures a unified mapping of the feature space for different original features of a character. By performing a layered consistency check, a video sequence with consistent roles is obtained. The layered consistency check includes at least one of frame-level verification, shot-level verification, scene-level verification, and global verification.

8. The method according to claim 1, characterized in that, During the sequential processing of the director agent, the scene planning agent, and the shot planning agent based on the output of the previous agent, a forward reasoning process is executed, including narrative structure analysis, key element extraction, boundary definition, emotion enhancement, and technical planning. For the structured parameters output at each stage of the forward reasoning process, a reverse verification process is performed to obtain the verification result. The reverse verification process includes: verifying the structured parameters output at the previous stage based on the structured parameters output at the next stage. The structured parameters are iteratively corrected based on the verification results.

9. The method according to claim 1, characterized in that, The method further includes: Construct an intelligent agent knowledge sharing pool to store intermediate data generated during the processing of the director intelligent agent, the scene planning intelligent agent, and the shot planning intelligent agent; When a cross-agent collaboration process is triggered, the agent initiating the collaboration calls intermediate data from other agents from the agent knowledge sharing pool and achieves cross-agent parameter sharing through the knowledge interoperability connection pool.

10. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to execute the executable instructions to perform the steps of the method as described in any one of claims 1 to 9.