Interactive video generation method capable of intelligent interaction based on multi-agent cooperation and related device

Through the multi-agent collaborative generation method, the problem of single traditional video interaction form is solved, personalized and diversified interactive video experience is achieved, and user satisfaction is improved.

CN120281978AActive Publication Date: 2025-07-08BEIJING BAIDU NETCOM SCI & TECH CO LTD

Patent Information

Application Number
CN202510347394.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-08
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

传统视频内容交互形式单一,用户只能被动观看,难以实现个性化、高自由度的互动体验。

Method used

Using a multi-agent collaboration method, the user needs are determined through the main agent, and copywriting, interactive options, storyboard planning and material generation agents are coordinated to generate interactive videos that match user needs.

Benefits of technology

Personalized and diversified interactive video generation is realized, improving user stickiness and interactive experience, and improving user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281978A_ABST
    Figure CN120281978A_ABST
Patent Text Reader

Abstract

The invention provides an interactive video generation method capable of intelligent interaction based on multi-agent cooperation and a related device, and relates to the technical field of artificial intelligence such as large language models, generative models, agents, interactive videos and the like. The method comprises the following steps: determining a user demand according to interaction between a target user and initial video content; controlling a preset copywriting agent to generate an interactive video copywriting matched with the user demand; controlling a preset interaction option generation agent and a split agent to respectively generate interaction options and split plans according to the interaction video copywriting; controlling a preset material generation agent to generate materials forming each video frame according to the mirror splitting plan; and generating an interactive video matched with a user demand from the interactive video copywriting, the interactive options and the materials according to a preset video rendering template. According to the method, through automation and agent cooperation between the main agent and each special agent, the generated interactive video can fully meet personalized and diversified user requirements, the user stickiness and the interactive experience of the interactive video are further improved, and the satisfaction degree of the user for such services or products is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technologies, specifically to artificial intelligence technology fields such as large language models, generative models, agents, and interactive videos. In particular, it relates to a method, device, electronic device, computer-readable storage medium, and computer program product for generating an interactive video with intelligent interaction based on multi-agent collaboration. Background Art

[0002] With the continuous development of digital media technology, videos have become an important way for users to obtain information, entertainment, and socialize. However, the traditional forms of video content interaction are relatively single. Users usually can only passively watch videos, lacking a deep interactive experience and unable to meet the user's need for exploring personalized content.

[0003] Existing interactive videos mainly rely on pre-set branch plots or specific interactive buttons. Users can only make selections within a limited range of options, making it difficult to achieve a personalized and highly free interactive experience. Summary of the Invention

[0004] Embodiments of the present disclosure propose a method, device, electronic device, computer-readable storage medium, and computer program product for generating an interactive video with intelligent interaction based on multi-agent collaboration.

[0005] In a first aspect, embodiments of the present disclosure propose a method for generating an interactive video with intelligent interaction based on multi-agent collaboration, including: determining user needs according to the interaction between a target user and initial video content; controlling a preset copywriting agent to generate interactive video copywriting that matches the user needs; controlling a preset interactive option generation agent and a storyboard agent to generate interactive options and storyboard plans respectively according to the interactive video copywriting; controlling a preset material generation agent to generate materials constituting each video frame according to the storyboard plan; and generating an interactive video that matches the user needs according to the interactive video copywriting, interactive options, and materials according to a preset video rendering template.

[0006] Second aspect, an embodiment of the present disclosure provides an interactive video generation device capable of intelligent interaction based on multi-agent collaboration, including: a user requirement determination unit configured to determine user requirements according to the interaction between a target user and initial video content; a copywriting generation unit configured to control a preset copywriting agent to generate interactive video copywriting that matches the user requirements; an interaction option and storyboard planning generation unit configured to control a preset interaction option generation agent and a storyboard agent to generate interaction options and storyboard plans respectively according to the interactive video copywriting; a material generation unit configured to control a preset material generation agent to generate materials for each video frame according to the storyboard plan; an interactive video generation unit configured to generate an interactive video that matches the user requirements according to the interactive video copywriting, interaction options, and materials according to a preset video rendering template.

[0007] Third aspect, an embodiment of the present disclosure provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to implement the method for generating an interactive video capable of intelligent interaction based on multi-agent collaboration as described in the first aspect.

[0008] Fourth aspect, an embodiment of the present disclosure provides an interactive video generation system capable of intelligent interaction based on multi-agent collaboration, the system includes: a main agent, configured to determine user requirements according to the interaction between a target user and initial video content; generate an interactive video that matches the user requirements according to the received interactive video copywriting, interaction options, and materials according to a preset video rendering template; a copywriting agent, configured to generate interactive video copywriting that matches the user requirements under the control of the main agent; an interaction option generation agent, configured to generate interaction options according to the interactive video copywriting under the control of the main agent; a storyboard agent, configured to generate a storyboard plan according to the interactive video copywriting under the control of the main agent; a material generation agent, configured to generate materials for each video frame according to the storyboard plan under the control of the main agent.

[0009] Fifth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to enable a computer to implement the method for generating an interactive video capable of intelligent interaction based on multi-agent collaboration as described in the first aspect when executed.

[0010] Sixth aspect, an embodiment of the present disclosure provides a computer program product including a computer program, and when the computer program is executed by a processor, it can implement the steps of the method for generating an interactive video capable of intelligent interaction based on multi-agent collaboration as described in the first aspect.

[0011] The interactive video generation solution based on multi-agent collaboration and capable of intelligent interaction provided by the present disclosure has a main agent responsible for determining the user requirements of the target user based on the interaction behaviors of the target user with the initial video content. Then, based on these user requirements, each intermediate stage (including copywriting, interaction options, storyboard planning, and materials) for generating a matching interactive video is dispatched to each specialized agent for execution, enabling each specialized agent to better complete the deliverables at each intermediate stage. Finally, the main agent generates an interactive video matching the user requirements according to the deliverables at each intermediate stage according to the video rendering template. Through the automation and collaboration between the main agent and each specialized agent, this solution can enable the generated interactive video to fully meet the personalized and diverse user requirements, further improving the user stickiness and interaction experience of the interactive video, and thereby enhancing the user satisfaction with such services or products.

[0012] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objectives, and advantages of the present disclosure will become more apparent:

[0014] Figure 1 is an exemplary system architecture to which the present disclosure can be applied;

[0015] Figure 2 is a flowchart of a method for generating an interactive video based on multi-agent collaboration and capable of intelligent interaction provided by an embodiment of the present disclosure;

[0016] Figure 3 is a flowchart of a method for determining user requirements provided by an embodiment of the present disclosure;

[0017] Figure 4 is a flowchart of a method for controlling an agent to generate deliverables at an intermediate stage provided by an embodiment of the present disclosure;

[0018] Figure 5 is a flowchart of a method for evaluating the quality of a generation result provided by an embodiment of the present disclosure;

[0019] Figure 6-1 is a schematic diagram of the execution process of each functional entity of the overall solution in an application scenario provided by an embodiment of the present disclosure;

[0020] Figures 6-2 to 6-7 is a schematic diagram of the effect including five interactions provided by an embodiment of the present disclosure;

[0021] Figure 7 Block diagram of an interactive video generation device based on multi-agent collaboration and capable of intelligent interaction provided by an embodiment of the present disclosure;

[0022] Figure 8 Schematic structural diagram of an electronic device suitable for executing an interactive video generation method based on multi-agent collaboration and capable of intelligent interaction provided by an embodiment of the present disclosure. Detailed implementation manners

[0023] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below. It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0024] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information and other processing are all in compliance with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0025] Figure 1 Exemplary system architecture 100 showing embodiments of an interactive video generation method, device, electronic device, and computer-readable storage medium based on multi-agent collaboration of the present disclosure that can be applied.

[0026] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0027] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various applications for realizing information communication between the two can be installed on the terminal devices 101, 102, 103 and the server 105, such as interactive video generation applications, video viewing applications, instant messaging applications, etc. A variety of intelligent agents can be installed or hosted on the server 105, such as an interactive intelligent agent dedicated to information interaction with users, and special intelligent agents for undertaking various tasks, and can also be the main intelligent agent and sub-intelligent agent, etc. according to the primary and secondary.

[0028] The terminal devices 101, 102, 103 and the server 105 can be either hardware or software. When the terminal devices 101, 102, 103 are hardware, they can be various electronic devices with a display screen, including but not limited to smartphones, tablets, laptop computers, desktop computers, etc.; when the terminal devices 101, 102, 103 are software, they can be installed in the above-listed electronic devices, and can be implemented as multiple software or software modules, or can also be implemented as a single software or software module, which is not specifically limited here. When the server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or can also be implemented as a single server; when the server is software, it can be implemented as multiple software or software modules, or can also be implemented as a single software or software module, which is not specifically limited here.

[0029] Through various built-in applications, the server 105 can provide various services. Taking the interactive video generation application that can provide interactive video generation services as an example, the main agent hosted on the server 105 can achieve the following effects when running this interactive video generation application: First, receive, through the network 104, the interactions made by the user on the initial video content viewed through the terminal devices 101, 102, 103, and then determine the user's needs; then, control the preset copywriting agent to generate an interactive video copywriting that matches the user's needs; next, control the preset interactive option generation agent and the storyboard agent to generate interactive options and storyboard plans respectively according to the interactive video copywriting; then, control the preset material generation agent to generate the materials that make up each video frame according to the storyboard plan; finally, generate an interactive video that matches the user's needs according to the preset video rendering template with the interactive video copywriting, the interactive options, and the materials.

[0030] Furthermore, the main agent hosted on the server 105 can also send the interactive video to the session where the user views the initial video content in the form of a subsequent video stream, and then present it to the user.

[0031] It should be noted that the interactions made by the user on the initial video content can be obtained not only from the terminal devices 101, 102, 103 through the network 104, but also can be pre-stored locally on the server 105 in various ways. Therefore, when the server 105 detects that these data have been stored locally (such as the pending tasks left before starting to process), it can choose to directly obtain these data from the local. In this case, the exemplary system architecture 100 may not include the terminal devices 101, 102, 103 and the network 104.

[0032] Since generating an interactive video requires a large amount of computing resources and strong computing power, the interactive video generation method based on multi-agent collaboration and capable of intelligent interaction provided in subsequent embodiments of the present disclosure is generally executed by a server 105 with strong computing power and a large amount of computing resources. Correspondingly, the interactive video generation device based on multi-agent collaboration and capable of intelligent interaction is generally also disposed in the server 105. However, it should also be noted that when the terminal devices 101, 102, and 103 also have computing power and computing resources that meet the requirements, the terminal devices 101, 102, and 103 can also complete the above operations that were originally performed by the server 105 through the interactive video generation applications installed thereon, and then output the same results as the server 105. Especially in the case where there are multiple terminal devices with different computing capabilities at the same time, when the interactive video generation application determines that the terminal device where it is located has strong computing power and a large amount of remaining computing resources, the terminal device can be allowed to execute the above operations, thereby appropriately reducing the computing pressure on the server 105. Correspondingly, the interactive video generation device based on multi-agent collaboration and capable of intelligent interaction can also be disposed in the terminal devices 101, 102, and 103. In this case, the exemplary system architecture 100 may not include the server 105 and the network 104.

[0033] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0034] Please refer to Figure 2 , Figure 2 which is a flowchart of an interactive video generation method based on multi-agent collaboration and capable of intelligent interaction provided by an embodiment of the present disclosure. The process 200 includes the following steps:

[0035] Step 201: Determine user requirements according to the interaction between the target user and the initial video content;

[0036] This step aims to accurately capture the user's interests, preferences, and requirements by the execution entity of the interactive video generation method based on multi-agent collaboration and capable of intelligent interaction (such as Figure 1 the main agent hosted on the server 105 shown) through analyzing the interaction behavior between the user and the initial video content, so as to provide a clear direction for subsequent steps such as copywriting generation, storyboard planning, and material production. That is, this step serves as the core starting point in the personalized interactive video generation solution provided in this embodiment.

[0037] Among them, the interaction behavior between the user and the initial video content may include but is not limited to the following types:

[0038] 1) Click behavior: The user clicks on specific elements in the video (such as buttons, links, options, etc.); 2) Dwell time: The length of time the user stays on certain video segments or content; 3) Selection behavior: The choices made by the user in the video (such as options, votes, answers, etc.); 4) Feedback behavior: The user's likes, comments, shares, or ratings on the video content; 5) Repeated viewing: The user's behavior of repeatedly viewing certain segments, etc. The above interaction behaviors can be collected in real time through the built-in monitoring module of the video player or third-party data analysis tools.

[0039] After obtaining the above interaction behaviors, in this step, the main intelligent agent deeply analyzes the collected interaction behavior data and extracts the key information from it. For example: Which options or elements the user clicks on indicate their interest in these contents; Which segments the user stays on for a long time may indicate that they are more interested in these contents; Whether the user's selection behavior reflects certain preferences or tendencies. Then, map the extracted key information to specific user needs. For example, if the user frequently clicks on options related to "technology", their needs may be biased towards technology-related content; If the user stays on the "humorous" segments in the video for a long time, their needs may be biased towards an entertainment style. That is, in this link, user needs can be classified according to different dimensions. For example: Content theme: The theme that the user is interested in (such as technology, education, entertainment, etc.); Content style: The style preferred by the user (such as humorous, serious, concise, etc.); Interaction depth: The degree of preference of the user for the interactive content (such as light interaction, in-depth exploration, etc.).

[0040] The above execution entity needs to analyze the behavior data in real time during the process of user-video interaction, quickly determine the user needs to support the rapid response in the subsequent links, and ensure the accuracy of the judgment of user needs as much as possible through multi-dimensional data analysis (such as behavior type, time distribution, frequency, etc.). At the same time, during the video playback process, it is necessary to dynamically adjust the judgment of user needs according to the user's latest interaction behavior to ensure that the generated content always matches the user needs.

[0041] Furthermore, in addition to the explicit interaction behaviors between the user and the initial video content, other data sources can be combined to further enrich the understanding of the user's needs. For example, the emotions or moods extracted from voice interactions, as well as the facial expressions and emotional changes of the user while watching the video captured by the camera. Additionally, historical behavior data can be combined: the user profile formed by integrating the user's historical viewing records and preferences can be used to further enhance or correct the judgment of the user's current needs. Among them, the user profile is constructed by accumulating the user's interaction behavior data over a long period, including interest tags, preferred styles, interaction habits, etc., and can provide more comprehensive support for determining the user's needs. A specific implementation method including but not limited to: First, determine the interaction information based on the interaction between the target user and the initial video content, and then determine the user's needs based on the interaction information and the profile information of the target user. The interaction information may include at least one of the following: option selection information for the interactive options provided by the initial video content, input text information, input image information, input voice information, input or shared accessible links or callable links (i.e., relevant information introduced in the form of links that is not directly presented in text or image form).

[0042] Specifically, the initial video content mentioned in this step can be presented in different forms according to different application scenarios:

[0043] 1) In the education field, the initial video can be a teaching video watched by students. Then, by having the students click on specific knowledge points or repeat watching certain segments, it is possible to determine their mastery of certain content and generate a targeted interactive video that includes review materials or practice questions for presentation.

[0044] 2) In the advertising field, the initial video can be an advertising video watched by users. Then, by having the users click on the links of certain products or services, it is possible to determine their points of interest and generate personalized advertising content that better meets their needs.

[0045] 3) In the entertainment field, the initial video can be a short video watched by users. Then, by having the users make personalized choices for plot branches or show strong interest in certain segments, it is possible to generate subsequent content that better matches their preferences.

[0046] Step 202: Control the preset copywriting agent to generate an interactive video copy that matches the user's needs;

[0047] Based on Step 201, the purpose of this step is for the above-mentioned execution entity to control the copywriting generation agent with the ability to generate copywriting to generate a video copy that highly matches it based on the determined user needs, so as to provide a basis for subsequent interactive options, storyboard planning, and material generation.

[0048] Among them, the copywriting agent can provide the following specific functions, including but not limited to:

[0049] 1) Narrative text of the video (such as narration, dialogue, etc.); 2) The theme and core information of the video; 3) Design of interactive points in the video (such as questions, options, branching plots, etc.), that is, used to design questions or options at key nodes of the video to guide user participation, and the complexity and quantity of interactive points can be adjusted according to the user's interaction depth (such as light interaction or in-depth exploration); 4) Style adaptation: The copywriting agent can adjust the copywriting style according to user preferences. For example, formal style: suitable for educational and professional videos; humorous style: suitable for entertainment and relaxing videos; concise style: suitable for advertising and fast-paced videos; 5) Language expression: The copywriting agent can generate corresponding copy according to the user's language habits (such as colloquial or written) or language types (such as Chinese or English), etc.

[0050] Specifically, this step aims to transfer the determined user requirements (such as content theme, style preference, interaction depth, etc.) by the above-mentioned execution entity to the copywriting agent as the input for copywriting generation. The copywriting agent then dynamically generates matching copy according to user requirements from a preset copywriting library or through natural language understanding (NLU, Natural Language Understanding, which aims to enable machines to understand, interpret, and generate human language. It extracts semantic information from natural language text and converts it into structured data that can be processed by machines) and natural language generation technology (NLG, Natural Language Generation, which aims to convert structured data or non-linguistic information into natural language text that can be understood by humans. Its goal is to generate language output that conforms to grammar rules, is semantically coherent, and is contextually relevant). For example, if the user is interested in the "technology" theme, the copywriting agent generates narrative text and interactive points related to technology; if the user prefers the "humorous" style, the copywriting agent adds light and humorous language elements to the copy. At the same time, when generating copy, the copywriting agent also needs to combine the overall context of the video (such as theme, target users, interaction history, etc.) to ensure the consistency and coherence of the copy.

[0051] Furthermore, the copywriting agent can generate multi-language copy according to the user's language preference to meet the needs of global users, and can also try to incorporate emotional elements into the copy. For example, adjust the tone of the copy according to the user's emotional state (such as excitement or calm), and add emotional expressions to the copy to enhance the user's sense of immersion and resonance. Even further, personalized recommendation content can be embedded in the copy. For example, recommend relevant themes or products according to the user's historical behavior, and add information or links that the user may be interested in to the copy.

[0052] Specifically, the copywriting generation mentioned in this step can be manifested in different forms according to different application scenarios:

[0053] 1) Education field: Generate targeted teaching video copy according to students' learning needs and interests. For example, for students who like hands-on practice, more case analyses and interactive exercises are embedded in the copy; for students who like theoretical learning, more concept explanations and logical reasoning are embedded in the copy.

[0054] 2) Advertising field: Generate personalized advertising copy according to users' interests and preferences. For example, for users who like fashion, fashion elements and trend trends are embedded in the copy; for users who like practicality, product functions and usage scenarios are embedded in the copy.

[0055] 3) Entertainment field: Generate personalized plot copy according to users' viewing habits and preferences. For example, for users who like suspense plots, more suspense and reversals are embedded in the copy; for users who like light-hearted plots, more humor and warm elements are embedded in the copy.

[0056] Step 203: Control the preset interaction option generation agents and storyboard agents to generate interaction options and storyboard plans respectively according to the interactive video copy.

[0057] Based on step 202, the purpose of this step is for the above-mentioned execution entity to further generate interaction options and storyboard plans based on the interactive video copy generated by the copywriting agent, providing structured support for subsequent material generation and video rendering.

[0058] Among them, the functions that the interaction option generation agent can provide include but are not limited to the following specific functions:

[0059] 1) Interaction option generation: The interaction option generation agent generates specific interaction options according to the interaction points in the interactive video copy. These options can include: Multiple-choice questions: Users can choose one answer or direction from multiple options; True or false questions: Users can answer questions through "yes / no" or "right / wrong"; Input box: Users can express their thoughts or answers through text input; Click trigger: Users can trigger subsequent content by clicking on specific elements.

[0060] 2) Subtitle generation: The interaction option generation agent generates subtitles that match the copy content, ensuring that the subtitles are synchronized with the video content and meet the reading habits and preferences of users.

[0061] 3) Option Design Optimization: The interactive option generation agent can optimize the option design according to the user's interaction habits and preferences. For example, it can adjust the number of options based on the user's interaction depth (e.g., 2 - 3 options for simple interaction and 4 - 5 options for in - depth interaction), and design more attractive option content according to the user's interests.

[0062] Among them, the storyboard agent can provide specific functions including but not limited to the following:

[0063] 1) Storyboard Planning Generation: The storyboard agent generates a detailed storyboard plan according to the interactive video copywriting, including: Shot Design: The shooting angle, scene type (such as long shot, medium shot, close - up) and movement method (such as push shot, pull shot) of each shot; Scene Transition: The transition method and effect between different scenes (such as fade - in, fade - out, quick switch); Time Allocation: The time length of each shot and the rhythm control of the overall video.

[0064] 2) Storyboard and Copywriting Matching: The storyboard agent ensures a high degree of matching between the storyboard plan and the copywriting content. For example, it designs close - up shots at the key information points of the copywriting to highlight the key points; and designs multiple storyboards at the interaction points to provide different visual experiences for users.

[0065] 3) Storyboard Style Adaptation: The storyboard agent can adjust the storyboard style according to the user's preferences. For example, Simple Style: Using simple shot transitions and clear picture layouts; Dynamic Style: Using rich shot movements and visual effects; Immersive Style: Using first - person perspective or long shots to enhance the sense of immersion.

[0066] This step aims to have the main agent control the interactive option generation agent and the storyboard agent respectively to accurately parse different parts of the interactive video copywriting to extract key information (such as interaction points, themes, styles, etc.). And the interactive option generation agent and the storyboard agent work together under the coordination of the main agent to ensure that the generated interactive options and storyboard plans match each other and are consistent with the copywriting content. Specifically, the interactive option generation agent and the storyboard agent can execute corresponding tasks in parallel, or one of them can execute the corresponding task first, and then use the interactive video copywriting and the generated task execution results as input for the other, so as to obtain a more accurate output of the second task. For example, first the interactive option generation agent generates interactive options, and then sends the interactive options and the interactive video copywriting to the storyboard agent, and the storyboard agent outputs a storyboard plan considering the interactive options. The reverse order is also possible.

[0067] Furthermore, the interaction option generation agent can design interaction options by combining multiple interaction methods. For example: Voice interaction: Users can select options or answer questions through voice; Gesture interaction: Users can trigger specific options or content through gestures. And the storyboard agent can also combine the preset material library or the capabilities of the material generation agent when generating the storyboard plan, and plan the usage method of the materials in advance. For example, select or generate matching materials such as backgrounds, characters, and props according to the storyboard plan; Design special effects or animation effects according to the storyboard plan. Moreover, the storyboard agent can also design personalized storyboard plans according to the user's viewing habits and preferences. For example, for users who like a fast pace, design storyboards with rapid switching, and for users who like details, design more close-up shots and slow-motion shots.

[0068] Specifically, the interaction option generation and storyboard plan generation mentioned in this step can be manifested in different forms according to different application scenarios:

[0069] 1) Education field: In an interactive teaching video, the interaction option generation agent generates multiple-choice questions or true / false questions related to knowledge points, and the storyboard agent designs a clear storyboard plan to highlight the teaching key points;

[0070] 2) Advertising field: In a personalized advertising video, the interaction option generation agent generates options related to the user's interests (such as selecting different product functions), and the storyboard agent designs an eye-catching storyboard plan to enhance the advertising effect;

[0071] 3) Entertainment field: In an interactive drama video, the interaction option generation agent generates options related to the plot development (such as selecting plot branches), and the storyboard agent designs an immersive storyboard plan to enhance the user's sense of immersion.

[0072] Step 204: Control the preset material generation agent to generate the materials that make up each video frame according to the storyboard plan;

[0073] Based on step 203, this step aims to dynamically generate or select the materials that make up each frame of the video by the above-mentioned execution subject according to the storyboard plan generated by the storyboard agent, ensure that the video content highly matches the user's needs, and provide high-quality material support for the final video rendering.

[0074] Among them, the specific functions that the material generation agent can provide include but are not limited to the following:

[0075] 1) Material Generation: The material generation agent generates, searches, queries, or selects materials that make up each frame of the video according to the storyboard plan, including but not limited to: Background materials: Background images or video clips that match the video theme and scene; Character materials: Character images or animations related to the video content; Prop materials: Props or decorative elements related to the video plot; Special effect materials: Special effects used to enhance the visual effect (such as lighting, particle effects, transition animations, etc.).

[0076] 2) Material Adaptation: The material generation agent can adjust the attributes of the materials, such as size, proportion, color, etc., according to the requirements of the storyboard plan to ensure a high degree of matching between the materials and the storyboard plan.

[0077] 3) Material Stylization: The material generation agent can generate or select materials that conform to a specific style according to user preferences and video styles. For example: Cartoon style: Suitable for entertainment and children's videos; Realistic style: Suitable for educational and professional videos; Minimalist style: Suitable for advertising and fast-paced videos.

[0078] Specifically, the material generation agent can specifically implement the generation of materials through the following technologies:

[0079] 1) Storyboard Analysis: The material generation agent analyzes the storyboard plan and extracts the material requirements for each frame (such as background, character, prop, special effect, etc.);

[0080] 2) Material Generation Methods: The material generation agent can generate or select materials in the following ways: Dynamic Generation: Use image generation technologies (such as generative adversarial networks, Diffusion Models, i.e., diffusion models) to generate materials that meet the requirements in real time; Material Library Selection: Select materials that match the storyboard plan from a preset material library; Material Combination: Combine or synthesize multiple materials to generate complex materials that meet the storyboard plan.

[0081] 3) Material Optimization: The material generation agent optimizes the generated materials. For example: Adjust the resolution, brightness, and contrast of the materials to ensure their visual effect in the video; Compress or convert the format of the materials to ensure their compatibility with the video rendering template.

[0082] Furthermore, the material generation agent can generate materials by combining multiple modalities. For example: 3D Material Generation: Use 3D modeling technologies to generate three-dimensional characters, scenes, and props; Audio Material Generation: Generate background music, sound effects, or voiceovers that match the video content; Personalized Material Design: The material generation agent can generate customized materials according to the user's needs. For example: Generate the user's preferred characters or scenes according to the user's historical behavior; Generate relevant props or decorative elements according to the user's interest points.

[0083] Specifically, the material generation mentioned in this step can take different forms according to different application scenarios:

[0084] 1) Education field: In interactive teaching videos, the material generation agent generates backgrounds, characters, and props related to knowledge points. For example, when explaining historical events, it generates scenes and character images that conform to the historical background; when explaining scientific principles, it generates dynamic 3D models or animations.

[0085] 2) Advertising field: In personalized advertising videos, the material generation agent generates materials related to user interests. For example, for users who like fashion, it generates the background and characters of fashion brands; for users who like technology, it generates 3D models and special effects of technology products.

[0086] 3) Entertainment field: In interactive plot videos, the material generation agent generates materials related to the plot development. For example, it generates different scenes and characters according to the plot branches selected by the user; it generates matching background music and special effects according to the user's emotional state.

[0087] Step 205: Generate an interactive video that matches the user's needs by embedding the interactive video copywriting, interaction options, and materials into a preset video rendering template.

[0088] Based on Steps 202 to 204, this step aims to integrate the copywriting, interaction options, and materials generated in the previous steps by the above-mentioned execution entity into a complete and interactive video through a preset video rendering template, ensuring that the finally output video highly matches the user's needs.

[0089] Among them, the video rendering template provides a standardized framework for integrating the interactive video copywriting, interaction options, and materials into a complete video. Specifically, it includes: Copywriting rendering: Embedding the copywriting content into the video in the form of subtitles, voiceovers, or dialog boxes; Interaction option rendering: Embedding the interaction options into the video in the form of buttons, selection boxes, or click areas; Material rendering: Embedding the generated materials (such as backgrounds, characters, props, special effects) into the video frames according to the storyboard plan. And the video rendering template can control the timeline of each frame according to the storyboard plan to ensure the rhythm and smoothness of the video, and the video rendering template embeds the logic of the interaction options (such as jumps and trigger events after selection) into the video to ensure the realization of the interaction function.

[0090] The specific integration process can be shown in the following order (only as an example and does not reflect a fixed or mandatory integration order):

[0091] 1) Copywriting embedding: Embedding the copywriting content into the video frames in the form of text or voice.

[0092] 2) Option embedding: Embed interactive options into video frames in a visual form and set up interaction logic;

[0093] 3) Material embedding: Embed materials into video frames according to the storyboard plan and apply special effects or animations.

[0094] That is, according to the preset rules and logic, the video rendering template renders the input content into video frames, stitches them into a complete video in chronological order, and then outputs the rendered video frames as an interactive video file, supporting multiple formats (such as MP4, WebM, etc.) and playback platforms (such as web pages, mobile devices).

[0095] It should be noted that during the video generation process in this step, it is necessary to rely on the video rendering template to ensure the consistency and coordination of the text, interactive options, and materials in the video, avoid content conflicts or visual incongruities, and ensure that the functional logic of the interactive options is correctly implemented in the video. For example, after the user selects an option, the video can jump to the corresponding segment or trigger a specific event, and the user's interaction behavior can be reflected in real time and affect the subsequent video content.

[0096] Furthermore, the video rendering template can generate adapted video formats and resolutions according to different playback platforms (such as web pages, mobile devices, TVs) to ensure the playback effect of the video on different devices. And during the video playback process, the video rendering template can also be required to dynamically adjust the subsequent video content according to the user's real-time feedback. For example, render different plot branches according to the user's selection results, or adjust the rhythm or content of the video according to the user's interaction behavior. Even further, the video rendering template can select different rendering styles according to the user's needs, such as a simple style with a clear layout and simple special effects, a dynamic style with rich animations and special effects, and an immersive style with a full-screen layout and 3D effects, etc.

[0097] Specifically, the generation of the interactive video mentioned in this step can be manifested in different forms according to different application scenarios:

[0098] 1) Education field: In interactive teaching videos, the video rendering template integrates knowledge point explanations, interactive exercises, and teaching materials into a complete video, and students can enter the next stage of learning by selecting answers or clicking on links;

[0099] 2) Advertising field: In personalized advertising videos, the video rendering template integrates product introductions, user interest points, and interactive options into a video, and users can view product functions or purchase links of interest by selecting different options;

[0100] 3) Entertainment field: In an interactive drama video, the video rendering template integrates the plot development, character dialogues, and interaction options into a single video. Users can experience different plot branches by selecting different options.

[0101] The interactive video generation method based on multi-agent collaboration and intelligent interaction provided by the embodiments of the present disclosure. The main agent is responsible for determining the user needs of the target user based on the interaction behaviors of the target user with the initial video content. Then, based on these user needs, each intermediate stage (including copywriting, interaction options, storyboard planning, and materials) for generating a matching interactive video is dispatched to each specialized agent for execution, so as to better complete the deliverables of each intermediate stage with the help of each specialized agent. Finally, the main agent generates an interactive video that matches the user needs according to the video rendering template from the deliverables of each intermediate stage. Through the automation and agent collaboration between the main agent and each specialized agent, this solution enables the generated interactive video to fully meet the personalized and diverse user needs, further improving the user stickiness and interaction experience of the interactive video, and thus enhancing the user satisfaction with such services or products.

[0102] Please refer to Figure 3 , Figure 3 It is a flowchart of a method for determining user needs provided by the embodiments of the present disclosure, aiming to give a more specific implementation for the solution of determining user needs based on interaction information and portrait information mentioned in step 201. The process 300 includes the following steps:

[0103] Step 301: Determine the interaction information according to the interaction of the target user with the initial video content;

[0104] This step is the same as that described in the above embodiments and will not be elaborated here.

[0105] Step 302: Determine the video direction requirements according to the initial video content and the interaction information;

[0106] Among them, the analysis of the initial video content may include the theme, style, structure, and interaction points, so as to be used as the basis for determining the video direction requirements. Combining the interaction information and the initial video content is used to deduce the user's requirements for the video direction. For example, if the user shows interest in "technology" related content, the video direction requirements may tend to be technology-related content; if the user stays on a "humorous" segment for a long time, the video direction requirements may tend to be an entertainment style; if the user frequently selects a certain plot branch, the video direction requirements may tend to be the subsequent development of this branch.

[0107] Among them, the video trend requirements can be classified according to different dimensions. For example: Content theme: The themes that users are interested in (such as technology, education, entertainment, etc.); Content style: The styles that users prefer (such as humorous, serious, concise, etc.); Interaction design: The preferences of users for interactive content (such as light interaction, in-depth exploration, etc.).

[0108] Step 303: Use the portrait information to correct the video trend requirements to obtain user requirements.

[0109] Among them, the user portrait is usually constructed based on the user's historical behavior data (such as viewing records, interaction habits, preference tags, etc.), including: Interest tags: The fields or themes that users are interested in (such as technology, fashion, travel, etc.); Behavior habits: The interaction habits of users (such as preferring to click, select, repeat viewing, etc.); Style preferences: The video styles that users prefer (such as humorous, serious, concise, etc.), etc. And in the case of insufficient personal portrait data, the group portrait information of similar users can be referred to in order to better assist in determining user requirements.

[0110] The process of using the user portrait information to correct the video trend requirements can be illustrated by examples: If the user portrait shows that they have a long-term interest in "travel" related content, then the video trend requirements are corrected to travel-related content; if the user portrait shows that they prefer the "concise" style, then the style of the video trend requirements is corrected to be more concise.

[0111] This embodiment aims to accurately determine the user requirements of users by analyzing the interaction behavior between users and the initial video, combined with the user portrait information, so as to provide a clear direction for subsequent video generation.

[0112] Specifically, the user requirement determination solution provided in this embodiment can be presented in different forms according to different application scenarios:

[0113] 1) Education field: When students watch teaching videos, by clicking on specific knowledge points or repeating certain segments, the main intelligent agent determines their interaction information, and combines their learning history (such as preferring practice or theory) to correct the video trend requirements and generate targeted teaching content.

[0114] 2) Advertising field: When users watch advertising videos, by clicking on the links of certain products or services, the main intelligent agent determines their interaction information, and combines their shopping history (such as preferring fashion or technology) to correct the video trend requirements and generate advertising content that better suits their interests.

[0115] 3) Entertainment field: When users watch short videos, by selecting different plot branches or showing strong interest in certain segments, the main intelligent agent determines their interaction information, and combines their viewing history (such as preferring suspense or comedy) to correct the video trend requirements and generate subsequent content that better suits their preferences.

[0116] Further, in the process of determining user needs, environmental information such as the user's geographical location and time can also be combined to further optimize the need judgment. And during the video playback, the main intelligent agent can dynamically adjust the user needs according to the user's real-time interaction behavior. If the user shows new points of interest during the playback, the video direction needs will be adjusted in real time. And if the user's interaction behavior deviates from the expectation, the user needs will be corrected again.

[0117] Please refer to Figure 4 , Figure 4 which is a flowchart of a method for controlling an intelligent agent to generate a deliverable in an intermediate stage provided by an embodiment of the present disclosure, to better reflect the interaction behavior between the main intelligent agent and each special intelligent agent in the above embodiment, so as to more clearly specify how to obtain the required deliverable. The process 400 includes the following steps:

[0118] Step 401: Generate a copywriting generation instruction corresponding to the user needs;

[0119] Step 402: Send the copywriting generation instruction to the copywriting intelligent agent and receive the returned interactive video copywriting;

[0120] The above two steps are intended to first generate a copywriting generation instruction corresponding to the determined user needs by the main intelligent agent, then send the copywriting generation instruction to the copywriting intelligent agent to trigger the copywriting generation process, and finally the copywriting intelligent agent returns the generated interactive video copywriting to the main intelligent agent.

[0121] Among them, the instruction may include: a need description for clarifying the user's needs (such as theme, style, interaction depth, etc.), and a generation rule for specifying the specific rules of copywriting generation (such as word count limit, language style, interaction point design, etc.).

[0122] Step 403: Generate an interaction option generation instruction and a storyboard planning instruction respectively according to the interactive video copywriting;

[0123] Step 404: Correspondingly send the interaction option generation instruction and the storyboard planning instruction to the interaction option generation intelligent agent and the storyboard intelligent agent respectively, and receive the returned interaction options and storyboard planning respectively;

[0124] The above two steps are intended to first generate an interaction option generation instruction and a storyboard planning instruction respectively by the main intelligent agent according to the received interactive video copywriting, then send the interaction option generation instruction to the interaction option generation intelligent agent and send the storyboard planning instruction to the storyboard intelligent agent, thereby triggering their respective generation processes. Finally, the interaction option generation intelligent agent generates interaction options, the storyboard intelligent agent generates a storyboard planning, and returns the results to the main intelligent agent.

[0125] Among them, the interactive option generation instruction is used to clarify the interactive points in the copywriting, specify the types of interactive options (such as multiple-choice questions, true or false questions, input boxes, etc.) and design rules. The storyboard planning instruction is used to clarify the content structure of the copywriting and specify the generation rules of the storyboard planning (such as shot design, scene switching, time allocation, etc.).

[0126] Step 405: Generate a material generation instruction corresponding to the storyboard planning;

[0127] Step 406: Send the material generation instruction to the material generation agent and receive the returned materials.

[0128] The above two steps are intended to have the main agent first generate a corresponding material generation instruction according to the received storyboard planning, then send the material generation instruction to the material generation agent to trigger the material generation process. Finally, the material generation agent generates or selects matching materials according to the instruction and returns them to the main agent.

[0129] Among them, the instruction may include: material requirements for clarifying the types of materials required for each frame of the video (such as backgrounds, characters, props, special effects, etc.), and generation rules for specifying the specific rules of material generation (such as style, resolution, format, etc.).

[0130] The steps 401 - 406 provided in this embodiment are intended to reflect that through the coordination of the main agent, the user requirements are transformed into specific instructions and distributed to each special agent for execution, and finally an interactive video copywriting, interactive options, storyboard planning, and materials are generated.

[0131] Specifically, the control generation scheme provided in this embodiment may be presented in different forms according to different application scenarios:

[0132] 1) Education field: The main agent generates a copywriting generation instruction according to the user requirements of the students, and the copywriting agent generates a teaching video copywriting; the main agent generates an interactive option generation instruction and a storyboard planning instruction according to the copywriting, the interactive option generation agent generates exercise options, and the storyboard agent plans the teaching scene; the main agent generates a material generation instruction according to the storyboard planning, and the material generation agent generates teaching materials.

[0133] 2) Advertising field: The main agent generates a copywriting generation instruction according to the user's interests, and the copywriting agent generates an advertising copywriting; the main agent generates an interactive option generation instruction and a storyboard planning instruction according to the copywriting, the interactive option generation agent generates product options, and the storyboard agent plans the advertising scene; the main agent generates a material generation instruction according to the storyboard planning, and the material generation agent generates advertising materials.

[0134] 3) Entertainment field: The main intelligent agent generates a copywriting generation instruction according to the user's preference, and the copywriting intelligent agent generates a plot copywriting; the main intelligent agent generates an interactive option generation instruction and a storyboard planning instruction according to the copywriting, the interactive option generation intelligent agent generates plot options, and the storyboard intelligent agent plans the plot scene; the main intelligent agent generates a material generation instruction according to the storyboard planning, and the material generation intelligent agent generates plot materials.

[0135] Furthermore, during the video generation process, the main intelligent agent can also dynamically adjust subsequent instructions according to the user's real-time feedback or the execution results of the special intelligent agents. For example, if the user's feedback on a certain interaction point is not ideal, the main intelligent agent can adjust the interactive option generation instruction and redesign the options; if the material generation intelligent agent cannot generate materials that meet the requirements, the main intelligent agent can adjust the material generation instruction and select an alternative solution.

[0136] Moreover, the main intelligent agent can set priorities for the instructions according to the urgency or importance of the tasks. For example, for the material generation instruction of key frames, a high priority is set to ensure its priority execution; for the non-core interactive option generation instruction, a low priority is set to optimize resource allocation. At the same time, the main intelligent agent can establish an instruction feedback mechanism to monitor the execution status of each special intelligent agent in real time. For example, if a certain intelligent agent times out or fails to execute, the main intelligent agent can reissue the instruction or switch to a standby intelligent agent; if the deliverable generated by a certain intelligent agent does not meet the expectations, the main intelligent agent can adjust the instruction and trigger the generation process again.

[0137] Considering that the user requirements determined by the main intelligent agent are more comprehensive and can more accurately describe the actual needs of the user, but often cannot be all sent to each special intelligent agent, only part of the information related to the specific task type can be extracted from them to form task execution instructions, and even if the special intelligent agents are based on the same instruction information, the results obtained in different executions are not the same. Therefore, in order to ensure that the deliverables output by each special intelligent agent can meet the actual needs of the user, it is also possible to perform quality evaluation on the products delivered by the above special intelligent agents and establish a quality evaluation and feedback mechanism to better control the deliverables of each link.

[0138] See Figure 5 , Figure 5 which is a flowchart of a method for quality evaluation of generation results provided by an embodiment of the present disclosure, and its process 500 includes the following steps:

[0139] Step 501: Perform quality evaluation on any one of the received interactive video copywriting, interactive options, storyboard planning, and materials based on user requirements;

[0140] Among them, the quality evaluation can be understood as that the above-mentioned execution entity formulates specific quality evaluation criteria for each type of deliverable based on user requirements. For example:

[0141] 1) For the interactive video copy: Does it accurately reflect the user's needs? Does it conform to the preset style and theme? Is the interactive point design reasonable?

[0142] 2) For the interaction options: Do they match the copy content? Do they conform to the user's interaction habits? Is the option design clear and easy to understand?

[0143] 3) For the storyboard planning: Is it consistent with the copy content? Does it conform to the user's visual preferences? Are the shot design and scene transitions reasonable?

[0144] 4) For the materials: Do they match the storyboard planning? Do they conform to the user's style preferences? Does the material quality meet the requirements (such as resolution, clarity, etc.)?

[0145] And the specific evaluation methods can also include: 1) Rule matching: Check whether the deliverables conform to the preset rules and standards; 2) Semantic analysis: Analyze whether the semantics of the copy and options match the user's needs through natural language processing technology; 3) Visual analysis: Analyze whether the visual effects of the materials meet the requirements through image recognition technology; 4) User simulation: Evaluate the actual effects of the interaction options and storyboard planning by simulating the user's interaction behavior, etc.

[0146] Step 502: In response to the quality evaluation result being unqualified, determine the adjustment instruction information according to the user's needs;

[0147] Among them, the adjustment instruction information is used to locate the specific problems that do not meet the requirements in the deliverables. For example: The interactive point design in the copy is unreasonable, the expression of the interaction options is not clear enough, the shot design of the storyboard planning does not conform to the user's preferences, and the style of the materials does not match the user's needs, etc.

[0148] And after locating the specific problems, the adjustment instruction information can be further manifested as:

[0149] 1) For the copy agent: Adjust the interactive point design and add content that the user is interested in;

[0150] 2) For the interaction option generation agent: Optimize the expression of the options to make them clearer and easier to understand;

[0151] 3) For the storyboard agent: Adjust the shot design to make it more in line with the user's visual preferences;

[0152] 4) For the material generation agent: Replace the material style to make it match the user's needs.

[0153] Step 503: Issue a regeneration instruction containing the adjustment instruction information to the agent with an unqualified quality evaluation result until a qualified quality evaluation result is obtained.

[0154] This step aims to generate a regeneration instruction by the above-mentioned execution entity according to the adjustment instruction information, clarify the content and specific requirements to be adjusted, and then send the regenerated instruction to the corresponding special intelligent agent to trigger the regeneration process. After that, the special intelligent agent regenerates the deliverable according to the instruction, and then the main intelligent agent conducts a quality evaluation on the new deliverable again until it passes the evaluation.

[0155] Furthermore, if the regenerated deliverable still fails to pass the quality evaluation, the main intelligent agent can initiate a multi-round adjustment mechanism to gradually optimize the deliverable. For example, the first round of adjustment: optimize the design of the interaction points in the copywriting; the second round of adjustment: further adjust the language style of the copywriting; the third round of adjustment: optimize the expression of the interaction options. And if the quality problems of a certain deliverable involve multiple intelligent agents, the main intelligent agent can coordinate multiple intelligent agents for collaborative adjustment. For example, if the storyboard planning does not match the materials, the main intelligent agent can simultaneously adjust the instructions of the storyboard intelligent agent and the material generation intelligent agent.

[0156] Even further, to reduce the probability of multiple rework and modification behaviors caused by failing to pass the quality evaluation, at least one of the copywriting intelligent agent, the interaction option generation intelligent agent, the storyboard intelligent agent, and the material generation intelligent agent can be controlled to provide at least two alternative deliverables for the received generation instruction at the same time. That is, the deliverable correspondingly includes at least one of the interactive video copywriting, interaction options, storyboard planning, and materials. That is, by providing multiple alternative deliverables at the same time, the main intelligent agent can select from them, thereby increasing the probability of passing the quality evaluation at one time.

[0157] Steps 501 - 503 provided in this embodiment aim to conduct a quality evaluation on the deliverables (such as interactive video copywriting, interaction options, storyboard planning, and materials) generated by each special intelligent agent, and make dynamic adjustments according to the evaluation results until all deliverables meet the quality requirements.

[0158] The quality evaluation scheme provided in this embodiment can be presented in different forms according to different application scenarios:

[0159] 1) Education field: The main intelligent agent conducts a quality evaluation on the teaching video copywriting, finds that the design of the interaction points is unreasonable, and sends adjustment instruction information to require the copywriting intelligent agent to regenerate the copywriting; evaluates the regenerated copywriting again until it passes.

[0160] 2) Advertising field: The main intelligent agent conducts a quality evaluation on the advertising materials, finds that the style does not match the user needs, and sends adjustment instruction information to require the material generation intelligent agent to regenerate the materials; evaluates the regenerated materials again until it passes.

[0161] 3) Entertainment field: The main intelligent agent evaluates the quality of the plot storyboard planning. If it finds that the shot design does not meet the user's preferences, it issues adjustment instruction information to require the storyboard intelligent agent to regenerate the storyboard planning. It evaluates the regenerated storyboard planning again until it passes.

[0162] Based on any of the above embodiments, if the initial video content includes a real image or a virtual image, then it should be controlled that the generated interactive video includes the same real image or virtual image, and it should be controlled that the postures of the real image or virtual image appearing in the interactive video match the video content of the interactive video, such as hand postures, face postures, and body postures, etc.

[0163] Based on any of the above embodiments, after the generation of the interactive video based on the current user requirements is completed, it is also possible to further control each intelligent agent to jointly generate a predicted interactive video corresponding to the future (relative to the current interaction moment, or relative to the currently generated interactive video) user requirements based on the current user requirements.

[0164] That is, the above execution entity predicts the user's future user requirements based on the current user requirements and the user's historical behavior data. That is, it can control each special intelligent agent (such as the copywriting intelligent agent, the interactive option generation intelligent agent, the storyboard intelligent agent, the material generation intelligent agent) to jointly generate an interactive video that matches the predicted requirements. For example, generate video content related to "new technology product evaluation"; generate video content related to "travel route planning".

[0165] During the subsequent interaction process, if there is a target preset interactive video whose matching degree with the future real user requirements exceeds the preset degree, then the target preset interactive video is directly provided to the target user. The number of the preset interactive videos and the number of future nodes involved are determined based on the available performance and / or the priority of the target user. Among them, in terms of available performance, it mainly dynamically adjusts the number of predicted videos and the number of future nodes according to the computing power and resource occupancy of the system. For example, if the system performance is sufficient, more predicted videos can be generated and more future nodes can be covered. If the system performance is limited, the number of predicted videos and the number of future nodes are reduced. In terms of user priority, it mainly adjusts the generation strategy of the predicted videos according to the importance and priority of the user. For example, for high-priority users, more predicted videos are generated for them and more future nodes are covered, while for ordinary users, only fewer predicted videos are generated and fewer future nodes are covered.

[0166] When the user's future real needs are clear, the above-mentioned execution entity evaluates the matching degree between the predicted video and the real needs. If there is a predicted video whose matching degree with the future real needs exceeds the preset threshold (i.e., the target preset interactive video), the main intelligent agent directly provides it to the user, avoiding the waiting time for regenerating the video. For example, if the user actually shows interest in "new technology product reviews" in the future and the content of the predicted video highly matches it, the matching degree exceeds the preset threshold, and then the pre-generated video content related to new technology product reviews that has been pre-generated is provided for it to interact with the user in a new round.

[0167] Based on any of the above embodiments, considering that multiple interactive videos (fragments) sorted by time axis will be gradually generated, at the interaction level, this embodiment can also provide the following rollback mechanism:

[0168] That is, the above-mentioned execution entity can receive the rollback selection information of the target user input for the historically generated interactive video. If the target user has a new interaction with the historical node corresponding to the rollback selection information, new interactive videos can also be jointly generated by each intelligent agent based on the new user needs corresponding to the new interaction.

[0169] Among them, this rollback function is used to allow the user to return to the historically generated interactive video node. For example, when the user is watching a video, they can choose to return to a certain historical node by clicking the "rollback" button, or the user can select a specific historical node through the time axis or node list. Specifically, the rollback information can come from the user's selection of the target node for rollback (such as a specific time point or interaction point), and the user's dissatisfaction with the current content or the willingness to reselect.

[0170] After the user rolls back to the historical node, the main intelligent agent real-time collects the new interaction behaviors of the user at this node. For example, the user reselects an interaction option, the user clicks on an element in the video, the user inputs new feedback through voice or text, etc. Then the main intelligent agent analyzes the corresponding new user needs based on the new interaction behaviors of the user. For example, if the user reselects an interaction option, the new needs may be related to the content theme or style corresponding to this option. If the user inputs new feedback, the new needs may be related to the keywords or emotional tendencies in the feedback. Finally, the main intelligent agent generates new instructions based on the new user needs and issues them to each special intelligent agent to obtain the newly generated interactive video, and provides the newly generated interactive video to the user to continue the subsequent interaction process.

[0171] Furthermore, it can also support users to roll back to multiple historical nodes and perform new interactions at different nodes. For example, users can roll back to the beginning of the video and reselect different plot branches; users can roll back to a certain key node and reselect different interaction options. In addition, it can record the user's rollback path and new interaction behaviors for optimizing subsequent video generation. For example, if a user frequently rolls back to a certain node and selects the same interaction option, the main intelligent agent can enhance the optimization of the content of that node. At the same time, after the user rolls back to a historical node, the main intelligent agent can also provide guiding information to help the user perform new interactions. For example, it can prompt the user to select different interaction options and provide background information or suggestions related to the historical node.

[0172] The rollback solution provided in this embodiment can be presented in different forms according to different application scenarios:

[0173] 1) Education field: When students are watching teaching videos, they roll back to a certain knowledge point node and reselect exercise options; the main intelligent agent generates a new teaching video based on the new selection and provides targeted explanations.

[0174] 2) Advertising field: When users are watching advertising videos, they roll back to a certain product introduction node and reselect function options; the main intelligent agent generates a new advertising video based on the new selection and shows the product functions that the user is interested in.

[0175] 3) Entertainment field: When users are watching plot videos, they roll back to a certain plot branch node and reselect options; the main intelligent agent generates a new plot video based on the new selection and shows different story developments.

[0176] To deepen the understanding, in view of the problems existing in the existing interactive video interaction methods, such as insufficient intelligence in interaction, difficulty in meeting user needs, and single interaction content, this embodiment specifically proposes an interactive video interaction method based on multi-agent collaboration, aiming to efficiently and accurately understand the user's interaction needs and realize the dynamic generation and continuous optimization of interactive video content through the collaboration of intelligent agents. The key technical solutions are as follows:

[0177] As Figure 6-1 shown, this embodiment designs an interactive video generation system based on the collaborative work of a main intelligent agent and multiple sub-intelligent agents, specifically including the following intelligent agents:

[0178] 1) Main agent: responsible for overall interactive process control, accurately understands user needs, coordinates and dispatches other agents in real time, and implements unified evaluation feedback and quality monitoring of the output content of sub-agents such as copywriting, subtitles, interactive options, storyboards, pictures or short film generation to ensure the consistency, accuracy and quality of the overall interactive video generation. In addition, the main agent can also perform intelligent automatic rendering based on video templates, quickly generate the final video presentation effect, and improve the interactive response speed.

[0179] 2) Copywriting Agent: Responsible for generating video copy based on user needs, including video theme planning, content organization and expression design, taking into full consideration user preferences and interactive context. The generated copy will be uniformly evaluated and optimized by the main agent to ensure the quality, relevance and attractiveness of the copy.

[0180] 3) Subtitle and option agent (i.e., the interactive option generation agent mentioned in the above embodiment): Generate corresponding subtitles and user interactive options according to the video text, where the subtitles support highlighting of key content to facilitate users to quickly capture key information; the interactive options are used to guide users to further interact and enhance user participation.

[0181] 4) Storyboard Agent: It intelligently plans the visual presentation requirements of the video based on the text and subtitles, including the type and theme of the illustrations or short films, determines the visual expression method and the method of obtaining material resources (such as creation or search), and interacts with the main agent in real time for quality feedback and adjustment.

[0182] 5) Image or short film generation agent: Based on the planning of the storyboard agent, it intelligently creates or obtains the required image or short film materials through search engines, realizes efficient acquisition and accurate generation of materials, and the main agent implements strict quality evaluation and control.

[0183] A specific execution process may include the following steps in order:

[0184] 1) The user first interacts with the default initial interactive video and inputs his or her needs, such as clicking on options, entering text, or voice input;

[0185] 2) The main agent analyzes and dispatches the copywriting agent in real time based on the user's needs and user profile to generate copywriting that meets the user's needs, and implements strict quality control and feedback;

[0186] 3) After the copy is evaluated, the main agent continues to dispatch the subtitle and option agents to generate corresponding subtitle and option interactive content to further improve the user experience. The main agent evaluates and gives feedback on the subtitles and options. If the subtitles and options are of poor quality, the main agent will give feedback to the subtitle and option agents, and the subtitle and option agents will regenerate subtitles and options.

[0187] 4) After the subtitle and option evaluation is passed, the main agent schedules the storyboard agent to generate a reply storyboard. The main agent evaluates and gives feedback on the storyboard. If the storyboard quality is not good, the main agent will give feedback to the storyboard agent, and the storyboard agent will regenerate the storyboard;

[0188] 5) After the storyboard evaluation is passed, the main agent schedules the picture or short video generation agent to generate a reply picture or short video. Specifically, if it is a creation and generation plan for pictures or short videos, the picture or short video generation agent is scheduled to generate pictures or short videos; if it is a search engine acquisition plan for pictures or short videos, the picture or short video search agent is scheduled to acquire pictures or short videos. Whether it is through creation and generation or search engine acquisition, the output of the picture or short video generation agent needs to go through the evaluation and feedback of the main agent. If the picture or short video quality is not good, the main agent will give feedback to the picture or short video generation agent, and the picture or short video generation agent will regenerate the picture or short video;

[0189] 6) After the picture or short video evaluation is passed, the main agent integrates the content such as the copywriting, subtitles, options, storyboards, pictures or short videos, etc., and automatically generates the template content according to the front-end video rendering template, and renders it into an interactive video to be displayed to the user;

[0190] 7) After that, the user can continue to interact with the generated interactive video and input their own requirements.

[0191] A specific example can be seen in Figures 6-2 to 6-6 the content of each interactive video segment shown and the corresponding interactive options, Figure 6-7 which shows a comprehensive strategy given after the selection of the previous interactive options. Among them, Figure 6-2 what is shown is the first interaction, which is used to provide the selection of the specific number of travelers, Figure 6-3 what is shown is the second interaction, which is used to display graphic content, Figure 6-4 what is shown is the third interaction, which is used to collect the number of days of play, Figure 6-5 what is shown is the fourth interaction, which is used to collect the budget, Figure 6-6 what is shown is the fifth interaction, which is used to collect the play methods, Figure 6-7 and is used to return the result of the strategy customization.

[0192] Compared with the prior art, the solution provided in this embodiment has the following invention innovation points:

[0193] 1) An interactive video generation mechanism based on multi-agent collaboration is proposed. Through the accurate and real-time understanding of user needs by the main agent, the unified scheduling and quality feedback of content generation by sub-agents, the personalization, intelligence, and automation of interactive video generation are comprehensively realized, significantly improving the overall user interaction experience and satisfaction.

[0194] 2) An intelligent video template rendering method is designed. The main agent can dynamically and intelligently select and generate video rendering template content according to user needs and interaction characteristics, significantly improving the flexibility, visual unity, and response speed of interactive video generation.

[0195] 3) Based on the multi-agent collaboration mode of the end-to-end generative large model, accurate and efficient task coordination, real-time feedback, and continuous optimization among agents are achieved, ensuring the global optimality of the entire interactive video task process, fully meeting the personalized and diverse user needs, and further improving the user stickiness and interaction efficiency of interactive videos.

[0196] The solution provided in this embodiment can be used in various interactive video application scenarios, such as online education, e-commerce promotion ( Figures 6-2 to 6-7 taking the promotion of tourism projects as an example), social interaction, live broadcast interaction, etc., and has broad application prospects and market potential. At the same time, this embodiment can also be deeply integrated with existing interactive video platforms, content generation tools, etc., further enhancing the intelligence, personalization, and user experience of interactive videos, and having high commercial value and social benefits.

[0197] For further reference Figure 7 , as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of an interactive video generation device capable of intelligent interaction based on multi-agent collaboration. This device embodiment corresponds to the Figure 2 method embodiment shown, and this device can be specifically applied to various electronic devices.

[0198] Such as Figure 7As shown in the figure, the interactive video generation device 500 based on multi-agent collaboration and capable of intelligent interaction in this embodiment may include: a user demand determination unit 701, a copywriting generation unit 702, an interaction option and storyboard planning generation unit 703, a material generation unit 704, and an interactive video generation unit 705. Among them, the user demand determination unit 701 is configured to determine user demands according to the interaction between the target user and the initial video content; the copywriting generation unit 702 is configured to control a preset copywriting agent to generate interactive video copywriting that matches the user demands; the interaction option and storyboard planning generation unit 703 is configured to control a preset interaction option generation agent and a storyboard agent to generate interaction options and storyboard plans respectively according to the interactive video copywriting; the material generation unit 704 is configured to control a preset material generation agent to generate materials constituting each video frame according to the storyboard plan; the interactive video generation unit 705 is configured to generate an interactive video that matches the user demands according to the interactive video copywriting, interaction options, and materials according to a preset video rendering template.

[0199] In this embodiment, in the interactive video generation device 700 based on multi-agent collaboration: the specific processing of the user demand determination unit 701, the copywriting generation unit 702, the interaction option and storyboard planning generation unit 703, the material generation unit 704, and the interactive video generation unit 705 and the technical effects brought by them can respectively refer to Figure 2 the relevant descriptions of steps 201-205 in the corresponding embodiment, which will not be elaborated here.

[0200] In some other optional implementation manners of this embodiment, the user demand determination unit 701 may include:

[0201] An interaction information determination subunit, configured to determine interaction information according to the interaction between the target user and the initial video content;

[0202] A user demand determination subunit, configured to determine user demands based on the interaction information and the portrait information of the target user.

[0203] In some other optional implementation manners of this embodiment, the interaction information may include at least one of the following:

[0204] Option selection information, input text information, input image information, input voice information, input or shared accessible link or callable link for the interactive options provided by the initial video content.

[0205] In some other optional implementation manners of this embodiment, the user demand determination subunit may be further configured to:

[0206] Determine the video direction demand according to the initial video content and the interaction information;

[0207] Modify the video direction requirement using the image information to obtain the user requirement.

[0208] In some other alternative implementation manners of this embodiment, the interaction option and storyboard planning generation unit 703 may be further configured to:

[0209] Control the interaction option generation agent to generate interaction options according to the interactive video copywriting;

[0210] Control the storyboard agent to generate a storyboard plan according to the interactive video copywriting and interaction options.

[0211] In some other alternative implementation manners of this embodiment, the copywriting generation unit 702 may be further configured to:

[0212] Generate a copywriting generation instruction corresponding to the user requirement;

[0213] Send the copywriting generation instruction to the copywriting agent and receive the returned interactive video copywriting;

[0214] Correspondingly, the interaction option and storyboard planning generation unit 703 may be further configured to:

[0215] Generate an interaction option generation instruction and a storyboard planning instruction respectively according to the interactive video copywriting;

[0216] Send the interaction option generation instruction and the storyboard planning instruction to the interaction option generation agent and the storyboard agent respectively, and receive the returned interaction options and storyboard plan respectively;

[0217] Correspondingly, the material generation unit 704 may be further configured to:

[0218] Generate a material generation instruction corresponding to the storyboard plan;

[0219] Send the material generation instruction to the material generation agent and receive the returned materials.

[0220] In some other alternative implementation manners of this embodiment, the interactive video generation device 700 capable of intelligent interaction based on multi-agent cooperation may further include:

[0221] A quality evaluation unit, configured to perform a quality evaluation on any one of the received interactive video copywriting, interaction options, storyboard plan, and materials based on the user requirement;

[0222] An adjustment instruction information determination unit, configured to determine adjustment instruction information according to the user requirement in response to the quality evaluation result being not passed;

[0223] A regeneration instruction sending unit, configured to send a regeneration instruction including adjustment indication information to an agent with a failed quality evaluation result until a quality evaluation result with a passing status is obtained.

[0224] In some other alternative implementation manners of this embodiment, the interactive video generation device 700 capable of intelligent interaction based on multi-agent collaboration may further include:

[0225] A multi-alternative deliverable providing unit, configured to control at least one of a copywriting agent, an interaction option generation agent, a storyboard agent, and a material generation agent to provide at least two alternative deliverables for a received generation instruction at the same time; wherein, the deliverables correspondingly include at least one of an interactive video copywriting, interaction options, a storyboard plan, and materials.

[0226] In some other alternative implementation manners of this embodiment, the interactive video generation device 700 capable of intelligent interaction based on multi-agent collaboration may further include:

[0227] An image and posture control unit, configured to, in response to the initial video content including a real image or a virtual image, control the generated interactive video to include the same real image or virtual image, and control the posture of the real image or virtual image appearing in the interactive video to match the video content of the interactive video.

[0228] In some other alternative implementation manners of this embodiment, the interactive video generation device 700 capable of intelligent interaction based on multi-agent collaboration may further include:

[0229] A predicted interactive video generation unit, configured to, in response to the generation of an interactive video based on the current user requirements being completed, control each agent to jointly generate a predicted interactive video corresponding to future user requirements based on the current user requirements.

[0230] In some other alternative implementation manners of this embodiment, the interactive video generation device 700 capable of intelligent interaction based on multi-agent collaboration may further include:

[0231] A direct use unit, configured to, in response to the existence of a target preset interactive video whose matching degree with future real user requirements exceeds a preset degree, directly provide the target preset interactive video to the target user.

[0232] In some other alternative implementation manners of this embodiment, the number of preset interactive videos and the number of future nodes involved are determined based on available performance and / or the priority of the target user.

[0233] In some other alternative implementations of this embodiment, the interactive video generation device 700 capable of intelligent interaction based on multi-agent collaboration may further include:

[0234] A fallback selection information receiving unit, configured to receive fallback selection information of an interactive video generated historically input by a target user;

[0235] A new interactive video generation unit, configured to, in response to the target user making a new interaction with a historical node corresponding to the fallback selection information, control each agent to jointly generate a new interactive video based on the new user requirements corresponding to the new interaction.

[0236] This embodiment exists as a device embodiment corresponding to the above method embodiment. The interactive video generation device capable of intelligent interaction based on multi-agent collaboration provided in this embodiment has a main agent responsible for determining the user requirements of the target user according to the interaction behavior of the target user with the initial video content, and then based on the user requirements, dispatching each intermediate stage (including copywriting, interaction options, storyboard planning, and materials) for generating a matching interactive video to each special agent for execution, so as to better complete the deliverables of each intermediate stage with the help of each special agent. Finally, the main agent generates an interactive video matching the user requirements according to the video rendering template from the deliverables of each intermediate stage. Through the automation and agent collaboration between the main agent and each special agent, this solution can enable the generated interactive video to fully meet the personalized and diverse user requirements, further improve the user stickiness and interaction experience of the interactive video, and thus enhance the user's satisfaction with such services or products.

[0237] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to implement the method for generating an interactive video capable of intelligent interaction based on multi-agent collaboration described in any of the above embodiments.

[0238] According to an embodiment of the present disclosure, the present disclosure also provides a readable storage medium, which stores computer instructions for enabling a computer to implement the method for generating an interactive video capable of intelligent interaction based on multi-agent collaboration described in any of the above embodiments when executed.

[0239] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, which can implement the method for generating an interactive video capable of intelligent interaction based on multi-agent collaboration described in any of the above embodiments when executed by a processor.

[0240] Figure 8 FIG. 1 shows a schematic block diagram of an example electronic device 800 that may be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementations of the present disclosure described and / or claimed herein.

[0241] As Figure 8 shown, the device 800 includes a computing unit 801 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0242] A plurality of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0243] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the intelligent interactive video generation method based on multi-agent collaboration. For example, in some embodiments, the intelligent interactive video generation method based on multi-agent collaboration can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the intelligent interactive video generation method based on multi-agent collaboration described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the intelligent interactive video generation method based on multi-agent collaboration by any other suitable means (e.g., by means of firmware).

[0244] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0245] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0246] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0247] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0248] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0249] A computer system can include clients and servers. The clients and servers are generally far from each other and usually interact through a communication network. The client-server relationship is generated by computer programs running on respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability existing in traditional physical hosts and virtual private server (VPS, Virtual Private Server) services.

[0250] According to the technical solution of the embodiment of the present disclosure, the main agent is responsible for determining the user needs of the target user according to the interaction behavior of the target user with the initial video content, and then based on the user needs, dispatching each intermediate stage (including copywriting, interaction options, storyboard planning, and materials) for generating a matching interactive video to each special agent for execution, so as to better complete the deliverables of each intermediate stage with the help of each special agent. Finally, the main agent generates an interactive video matching the user needs according to the deliverables of each intermediate stage according to the video rendering template. Through the automation and agent collaboration between the main agent and each special agent, this solution can enable the generated interactive video to fully meet the personalized and diverse user needs, further improve the user stickiness and interaction experience of the interactive video, and then enhance the user satisfaction with such services or products.

[0251] It should be understood that various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in the present disclosure can be achieved, and no limitation is made herein.

[0252] The above specific embodiments do not constitute a limitation to the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. An interactive video generation method based on multi-agent collaboration and capable of intelligent interaction, comprising: Determining user requirements according to the interaction between the target user and the initial video content; Controlling a preset copywriting agent to generate an interactive video copy that matches the user requirements; Controlling a preset interaction option generation agent and a storyboard agent to generate interaction options and storyboard plans respectively according to the interactive video copy; Controlling a preset material generation agent to generate materials constituting each video frame according to the storyboard plan; Generating an interactive video that matches the user requirements according to the interactive video copy, the interaction options, and the materials according to a preset video rendering template.

2. The method according to claim 1, wherein The determining user requirements according to the interaction between the target user and the initial video content includes: Determining interaction information according to the interaction between the target user and the initial video content; Determining the user requirements based on the interaction information and the portrait information of the target user.

3. The method according to claim 2, wherein The interaction information includes at least one of the following: Option selection information, input text information, input image information, input voice information, input or shared accessible link or callable link for the interactive options provided by the initial video content.

4. The method according to claim 2, wherein The determining the user requirements based on the interaction information and the portrait information of the target user includes: Determining video direction requirements according to the initial video content and the interaction information; Correcting the video direction requirements by using the portrait information to obtain the user requirements.

5. The method according to claim 1, wherein The controlling the preset interaction option generation agent and the storyboard agent to generate interaction options and storyboard plans respectively according to the interactive video copy includes: Controlling the interaction option generation agent to generate the interaction options according to the interactive video copy; Controlling the storyboard agent to generate the storyboard plan according to the interactive video copy and the interaction options.

6. The method according to claim 1, wherein The controlling the preset copywriting agent to generate an interactive video copy that matches the user requirements includes: Generating a copywriting generation instruction corresponding to the user requirements; Sending the copywriting generation instruction to the copywriting agent and receiving the returned interactive video copy; Correspondingly, the controlling the preset interaction option generation agent and the storyboard agent to generate interaction options and storyboard plans respectively according to the interactive video copy includes: Generating an interaction option generation instruction and a storyboard plan instruction respectively according to the interactive video copy; Correspondingly sending the interaction option generation instruction and the storyboard plan instruction to the interaction option generation agent and the storyboard agent respectively, and receiving the returned interaction options and storyboard plan; Correspondingly, the controlling the preset material generation agent to generate materials constituting each video frame according to the storyboard plan includes: Generating a material generation instruction corresponding to the storyboard plan; Sending the material generation instruction to the material generation agent and receiving the returned materials.

7. The method according to claim 6, further comprising: Performing a quality evaluation on any one of the received interactive video copy, interaction options, storyboard plan, and materials based on the user requirements; In response to the quality evaluation result being unqualified, determine adjustment indication information according to the user requirements; For the agent with the unqualified quality evaluation result, issue a regeneration instruction containing the adjustment indication information until an agent with a qualified quality evaluation result is obtained.

8. The method according to claim 7, further comprising: Controlling at least one of the copywriting agent, the interaction option generation agent, the storyboard agent, and the material generation agent to provide at least two alternative deliverables for the received generation instruction at the same time; wherein the deliverables correspondingly include at least one of an interactive video copywriting, interaction options, a storyboard plan, and materials.

9. The method according to claim 1, wherein, In response to the initial video content including a real image or a virtual image, control the generated interactive video to include the same real image or virtual image, and control the posture of the real image or virtual image appearing in the interactive video to match the video content of the interactive video.

10. The method according to any one of claims 1-9, further comprising: In response to the completion of generating the interactive video based on the current user requirements, control each agent to jointly generate a predicted interactive video corresponding to future user requirements based on the current user requirements.

11. The method according to claim 10, further comprising: In response to the existence of a target preset interactive video whose matching degree with future real user requirements exceeds a preset degree, directly provide the target preset interactive video to the target user.

12. The method according to claim 10, wherein, The number of the preset interactive videos and the number of future nodes involved are determined based on available performance and / or the priority of the target user.

13. The method according to claim 10, further comprising: Receiving the fallback selection information input by the target user for the historically generated interactive video; In response to the target user having a new interaction with the historical node corresponding to the fallback selection information, control each agent to jointly generate a new interactive video based on the new user requirements corresponding to the new interaction.

14. An interactive video generation device capable of intelligent interaction based on multi-agent collaboration, comprising: A user requirement determination unit configured to determine user requirements according to the interaction between the target user and the initial video content; A copywriting generation unit configured to control a preset copywriting agent to generate an interactive video copywriting matching the user requirements; An interaction option and storyboard plan generation unit configured to control a preset interaction option generation agent and a storyboard agent to generate interaction options and a storyboard plan respectively according to the interactive video copywriting; A material generation unit configured to control a preset material generation agent to generate materials constituting each video frame according to the storyboard plan; An interactive video generation unit configured to generate an interactive video matching the user requirements by using the interactive video copywriting, the interaction options, and the materials according to a preset video rendering template.

15. An interactive video generation system capable of intelligent interaction based on multi-agent collaboration, comprising: A main agent for determining user requirements according to the interaction between the target user and the initial video content; Generate an interactive video that matches the user's needs by rendering the received interactive video copywriting, interaction options, and materials according to a preset video rendering template; A copywriting agent for generating interactive video copywriting that matches the user's needs under the control of the main agent; An interaction option generation agent for generating interaction options according to the interactive video copywriting under the control of the main agent; A storyboard agent for generating a storyboard plan according to the interactive video copywriting under the control of the main agent; A material generation agent for generating materials that make up each video frame according to the storyboard plan under the control of the main agent.

16. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method for generating an intelligently interactive interactive video based on multi-agent collaboration according to any one of claims 1-13.

17. A non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the method for generating an intelligently interactive interactive video based on multi-agent collaboration according to any one of claims 1-13.

18. A computer program product comprising a computer program, the computer program when executed by a processor implements the steps of the method for generating an intelligently interactive interactive video based on multi-agent collaboration according to any one of claims 1-13.

Citation Information

Patent Citations

  • Method for creating interactive video file and method and device for playing interactive video

    CN115460468A

  • Video recommendation method and device, electronic equipment and storage medium

    CN116405736A

  • Short video production method and system based on artificial intelligence

    CN117714797A

  • Video generation method and device based on large model, electronic equipment, storage medium and program product

    CN118870146A

  • AIGC-based aisle deduction system and use method thereof

    CN119152743A

Cited By

  • Reading understanding test question generation method based on large language model multi-agent

    CN120471183A

  • Multi-agent figure video generation method and device based on potential diffusion model, equipment and storage medium

    CN121486635A

  • Short video editing method and device, electronic equipment, storage medium and program product

    CN121509775A

  • Multi-modal content generation method and device, intelligent agent and electronic equipment

    CN121541807A

  • Video generation method, device and system based on multi-agent cooperation

    CN121691839A