A system and method for real-time interactive video stream generation and intelligent interaction

By combining AI Agent with a real-time interactive video stream generation system for multimodal data, the problems of static content, single interactive form, data isolation, poor scene adaptability, and delay in existing technologies are solved, achieving efficient, dynamic, and intelligent content generation, and improving the effect of real-time interactive live broadcast and user experience.

CN120186376BActive Publication Date: 2025-09-16SANSHENG ZHILIAN TECHNOLOGY (HANGZHOU) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510645116.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-09-16
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

Existing interactive video stream generation technologies have problems such as static content, single interactive form, data isolation and poor scene adaptability, as well as delay and real-time issues. They are unable to achieve efficient, dynamic and intelligent content generation based on real-time data and user interaction.

Method used

A real-time interactive video stream generation system is adopted, including a control center, a real-time video stream module, a knowledge base, a content generation module and a feedback module. Through the combination of AI Agent and multimodal data, dynamic content generation, intelligent interactive closed loop, full-scene compatibility and low-latency transmission are realized. The organic combination of multimodal data acquisition module, content generation engine and feedback optimization module is utilized to ensure the efficient flow of the entire process from audience interaction to video content generation, transmission and feedback.

Benefits of technology

It significantly improves the real-time, interactivity and data accuracy of interactive video stream generation technology, and provides an innovative, efficient, flexible and low-latency solution to meet the real-time interactive live broadcast needs of various industries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120186376B_ABST
    Figure CN120186376B_ABST
Patent Text Reader

Abstract

The present invention discloses a system and method for real-time interactive video stream generation and intelligent interaction, which belongs to the field of artificial intelligence and data fusion technology. The control center generates instructions based on the preset video stream theme, interactive information and the video stream content of the real-time video stream module. The content generation module obtains corresponding data from the knowledge base according to the instructions, generates new video stream content in combination with the video stream content of the real-time video stream module, and the feedback module adjusts data acquisition and content generation in real time according to the interactive information. By building a full-link intelligent interactive ecology, the present invention realizes highly automated processing from data collection, intelligent decision-making to content generation, realizes real-time interactive video live broadcast, ensures the efficient flow of the entire process from audience interaction to video content generation, transmission and feedback, and improves the real-time, interactivity and data accuracy of the interactive video stream generation technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence and data fusion, and specifically relates to a system and method for real-time interactive video stream generation and intelligent interaction. Background Art

[0002] With the rapid development of artificial intelligence (AI), particularly breakthroughs in large language models, computer vision, speech recognition, and multimodal data fusion, real-time interactive video stream generation technology has become a hot application across many industries. In scenarios such as interactive live streaming, distance learning, corporate training, virtual customer service, and online shopping, the demand for real-time, intelligent, and personalized interactive content is increasing. Traditional interactive video stream generation methods typically rely on static scripts or manual input, lacking an automated, real-time data-driven mechanism. This results in limited interactive effects and high latency.

[0003] The combination of AI agents and multimodal data opens up new possibilities for generating real-time interactive video streams. AI agents are AI agents driven by large language models, used to interpret user commands, control data collection, and drive content generation. Multimodal AI technologies encompass the collection and analysis of data from vision (image recognition, cameras), speech (audience, natural language processing), olfaction, touch, and sensors (temperature, humidity, and location). Through natural language processing, big data analytics, and machine learning, AI agents can efficiently interpret viewer or user commands and automatically control data collection and content generation. By integrating multimodal data (such as sensor data, images, speech, and video), AI agents can drive content generation engines to dynamically generate real-time interactive video streams tailored to the current environment, scenario, and audience needs. This brings a new level of intelligence and automation to various real-time interactive application scenarios.

[0004] However, existing technical implementations still face some core bottlenecks, especially in terms of real-time performance, interactivity, and data fusion. Specifically, they include the following aspects:

[0005] (1) Static content and low responsiveness

[0006] Traditional interactive video stream generation methods, such as scripted digital human live broadcasts and manually driven live broadcasts, typically generate static content and lack dynamic responses based on real-time data and user interaction. While some advanced digital human technologies can generate and interact with virtual characters, these interactions still rely primarily on pre-set scripts or manual input and cannot adaptively adjust content based on real-time data or user needs. Digital human technology is used to generate video streams of virtual characters, typically using pre-trained image models, speech synthesis, and motion-driven animations. Live broadcasts, on the other hand, are based on the real host's footage and enhance interactivity through real-time data overlay and dynamic content generation.

[0007] For example, in digital human live broadcasts, while the avatars can complete certain interactive tasks, they are unable to adjust their presentation content in real time based on environmental changes (such as logistics progress, inventory changes, and production line status). Live broadcasts with real people face similar challenges, as the host must rely on manual operations and lacks intelligent automation support.

[0008] (2) Single and inefficient interactive form

[0009] Existing interactive video generation systems still rely on a relatively simple interaction model, often relying on a feedback loop between user commands and manual processing. Traditionally, user questions often require manual analysis and response. This inefficient approach cannot meet the demands of high-frequency, high-precision, real-time interaction.

[0010] For example, when viewers ask questions during a live broadcast, the host needs to check real-time data or call external system interfaces to obtain answers. This approach not only has a high response delay, but is also prone to inaccurate information due to human omissions or technical bottlenecks.

[0011] (3) Data isolation and poor scenario adaptability

[0012] Currently, most existing interactive video streaming systems lack effective linkage between data collection and content generation. Many systems are unable to effectively integrate multimodal data (such as sensor data, image data, and voice data) into live content. This leads to a disconnect between content generation and the real-world, affecting the credibility of live content and the interactive experience.

[0013] Furthermore, existing technologies have limited adaptability to complex scenarios. In many practical applications, real-time data needs to be flexibly adjusted based on different scenarios, user needs, and external conditions, but traditional systems are unable to achieve this. For example, in fields such as industrial production and intelligent manufacturing, rapid acquisition and accurate feedback of real-time data are crucial, and existing technology systems are unable to effectively support efficient data integration and content generation in such scenarios.

[0014] (4) Delay and real-time issues

[0015] When generating real-time interactive video streams, low latency is crucial for ensuring a positive user experience. However, existing systems often experience significant delays between data processing and content generation. This delay, particularly during the acquisition and processing of multimodal data, can reach several seconds or even longer, making it difficult to meet the demands of real-time interaction.

[0016] For example, when transmitting sensor data, video streams, or voice streams in real time, bottlenecks in data processing and transmission may cause viewers or users to perceive significant delays, affecting the interactive experience.

[0017] In summary, existing interactive video stream generation technologies have significant shortcomings in many areas, making them incapable of achieving efficient, dynamic, and intelligent content generation based on real-time data and user interaction. Therefore, a new technical solution is urgently needed that can combine AI agents with multimodal data to generate interactive content in real time, while also possessing high adaptability and flexibility to enhance interactive effects and optimize the user experience. Summary of the Invention

[0018] To address the shortcomings of existing technologies, achieve dynamic content generation, closed-loop intelligent interaction, full-scenario compatibility, real-time data-driven and automated control, and reduce system latency and improve response speed, the present invention adopts the following technical solutions:

[0019] A real-time interactive video stream generation system includes a control center, a real-time video stream module, a knowledge base, a content generation module, and a feedback module. The control center generates instructions based on the video stream theme, interactive information, and the video stream content of the real-time video stream module. The content generation module obtains corresponding data from the knowledge base according to the instructions, and generates new video stream content based on the video stream content of the real-time video stream module. The feedback module adjusts data acquisition and content generation in real time according to the interactive information to ensure that key information can be presented to the audience with minimal delay.

[0020] Furthermore, the real-time video stream module includes generating real-time real-person video stream content, and the content generation module includes a digital human content generation module, a real-person video stream content integration module and a multimedia content mixing module. The real-person video stream content integration module obtains corresponding data through the instructions and integrates it into the real-person video stream content. The digital human content generation module generates digital human video stream content that matches the real-person video stream content according to the instructions, and integrates it with the real-person video stream content through the multimedia mixing module. The integrated content is displayed to obtain the interactive information, and the integrated content is fed back to the real-time video stream module to assist in generating the real-person video stream content.

[0021] Furthermore, the interactive information includes interactive information based on the content of the video stream shown, and / or interactive information between real people and digital humans, so that instructions can be generated based on the feedback of viewers watching the video stream to adjust the video stream content in real time, and the adjustment of the feedback module can also adjust the data collection strategy and optimize the content generation efficiency in real time based on the feedback of viewers watching the video stream.

[0022] Furthermore, the content generation module is also used to generate unmanned video stream content without the real person part and / or the digital human part, and the unmanned video stream content is set in conjunction with the real person video stream content and / or the digital human video stream to generate new video stream content.

[0023] Furthermore, the system also includes a streaming media server module, and the control center is an artificial intelligence control center, which generates content generation instructions based on the video stream theme, the interactive information, the real-time video stream content, the knowledge base information, and the system status; the content generation module obtains data from the knowledge base to generate corresponding content based on the content generation instructions, the real-time video stream content, the video stream theme and the analysis of the interactive information. The content is encoded by the streaming media server module to generate new video stream content, which is pushed to the display end to obtain the interactive information.

[0024] Furthermore, the corresponding content includes but is not limited to video content and / or text content, and the content is encoded by a streaming media server module to generate new video stream content and / or text content, which is pushed to the display end.

[0025] Furthermore, the system also includes a multimodal data acquisition module, and the control center triggers data acquisition instructions based on the video stream theme and the interactive information , the multimodal data acquisition module collects real-time multimodal data that is not retrieved in the knowledge base. The content generation module searches the knowledge base according to the video stream theme and the interactive information to generate corresponding content. The formula is as follows:

[0026]

[0027]

[0028]

[0029] in, Represents the generated content, Represents the content generation engine function (i.e. multimodal fusion), Represents content generation instructions (such as scripts, charts, digital human actions), represents multimodal real-time data (such as data obtained by sensors, cameras, and external system interfaces), t represents time, represents the knowledge base retrieval function (such as vector similarity matching), Represents a knowledge base (including static knowledge + Dynamic Knowledge ), Indicates the preset subject of operation configuration (such as preset topic, knowledge base related data), Represents the natural language NLP parsing function (used for intent recognition and entity extraction), Indicates interactive input (such as text / voice requests, click actions), Represents real-time video streaming content, Represents the video stream content, Indicates streaming media encoding and distribution, Indicates the upper bound of streaming media delay. Streaming media encoding and distribution need to dynamically adjust the bit rate (i.e., adaptive bit rate) according to the upper bound of delay to ensure that the generated video stream can guarantee the audience's viewing experience. represents the control function, Indicates system status (such as current live broadcast mode, content generation strategy, and data collection priority);

[0030] The feedback module adjusts the system status in real time based on the video stream content and interactive information to ensure that the live broadcast content always meets the audience's needs and maintains efficient operation. The formula is as follows:

[0031]

[0032]

[0033] in, Indicates interactive feedback (such as likes, question frequency, latency indicators), represents the feedback function, Indicates that the content is based on the video stream and interactive feedback The loss function constructed to measure user satisfaction and system performance, Represents the learning rate, which is used to feedback the adjustment strength of the optimization module. Gradient descent dynamically adjusts the system state (e.g. data collection priorities, content generation patterns).

[0034] Furthermore, the control center classifies the content generated subsequently based on the acquired topics, driving the content generation module to sequentially generate the video streams according to the class. The feedback module adjusts the class of the content to be generated in real time based on the match between the interactive information and the subsequent content to be generated, providing early feedback on the interactive information. Furthermore, the feedback module further acquires corresponding data based on the interactive information to generate inserted content, integrating the inserted content into the content to be generated. Traditional live broadcasts often emphasize timely, low-latency feedback of interactive information, such as answering user questions. However, overly rapid responses can interrupt the current live broadcast. For example, while explaining product performance, a question about product discounts is suddenly answered, disrupting the overall plan of the live broadcast. By dividing the entire live broadcast into multiple segments and adjusting the order of each segment based on real-time user questions, the pre-selected content is relevant to the question to be answered, or even the content itself can answer the user's question. This allows users to receive timely answers while seamlessly inserting the answers, without making the response process appear too abrupt. At the same time, while waiting for answers, users can also learn about other performance features of the product, and the content of the reply is included in the content related to the answer, so that users can understand the answer to the question and the content related to the answer as a whole. Because sometimes users ask questions that may not solve the problem itself, but to understand another implicit problem related to the problem, or there is another related aspect of the problem that the user has not considered. Through relevant replies, users can have a more comprehensive understanding of the product and solutions to related problems.

[0035] For the inserted content, an anchor point is set in the video stream, and the time step of the anchor point interval is set based on the number of similar interactive information after the anchor point. According to the time step, in the subsequent video stream, the inserted content is regenerated by the feedback module to respond to the previous inserted content. For users who are new to the live broadcast room and ask the same or similar questions, the answers to the questions can be answered in a timely manner, and the impression of the previous users on the solution to the problem is also deepened. Moreover, the content fed back again is the content adjusted and updated by the feedback module, which is a further description or correction of the previous feedback content. On the other hand, the content generated by the interactive feedback includes information such as figures and tables displayed on the live broadcast interface. For malicious and frequent questions in a short period of time, the figures, tables and other information will always occupy the screen. The more frequent the malicious questions are, the shorter the time step of the corresponding anchor point interval will be. The figures and tables corresponding to the malicious questions will soon be replaced by the figures and tables corresponding to the next hot frequency.

[0036] The inserted contents are sorted according to the number of identical or similar inserted contents. The content generation level of the corresponding content to be generated is adjusted according to the ranking of multiple inserted contents that are ranked high and correspond to the same content to be generated, so as to provide feedback on interactive information in advance. On the one hand, high-frequency similar questions are answered first. On the other hand, multiple answers to different questions corresponding to the same content segment are grouped together and answered first along with the content segment, thereby improving the answer efficiency, reducing the frequency of switching anchor point answers at each time step, and improving the user experience.

[0037] A system for real-time interactive video stream generation and intelligent interaction, based on the real-time interactive video stream generation system, also includes an operation configuration module and an interaction module. The operation configuration module is connected to the AI ​​Agent control center and is used for video stream theme setting, data source configuration and interactive content planning, so as to set the key topics of the video stream and the expected data sources for display, generate video stream themes, and provide a basis for the subsequent data acquisition and automatic execution of content generation instructions, ensuring that the live broadcast content meets the operation goals and can be flexibly adjusted to meet the needs of the audience; the interaction module is connected to the AI ​​Agent control center and is used for real-time interaction between the system and users, obtaining user requests in text and / or voice, and preprocessing and / or parsing to generate interactive information, thereby ensuring that the system can accurately capture and understand the intention of the request, providing a basis for subsequent decision-making.

[0038] A method for generating a real-time interactive video stream and intelligent interaction, which performs real-time interactive video stream generation and intelligent interaction according to a system for generating a real-time interactive video stream and intelligent interaction, comprises the following steps:

[0039] Step 1: Preset the video streaming theme;

[0040] Step 2: Obtain data corresponding to the topic from the knowledge base based on the video stream topic to generate real-time video stream content;

[0041] Step 3: Obtain interactive information of the video stream content;

[0042] Step 4: Analyze the interactive information and trigger data collection based on the generated real-time video stream content to obtain data corresponding to the interactive information for generating new video stream content;

[0043] Step 5: Content generation and video streaming;

[0044] Step 6: Adjust data acquisition and content generation in real time based on interactive information to ensure that key information can be presented to the audience with minimal latency.

[0045] The advantages and beneficial effects of the present invention are:

[0046] The present invention is composed of a number of closely cooperating core modules to ensure efficient flow of the entire process from audience interaction to video content generation, transmission and feedback. Among them, through the organic combination of multimodal data acquisition, AI Agent intelligent control, content generation engine and feedback optimization module, highly automated and real-time interactive video live broadcast is achieved. By building a full-link intelligent interactive ecosystem, the present invention realizes automated processing from data acquisition, intelligent decision-making to content generation, significantly improving the real-time, interactivity and data accuracy of interactive video stream generation technology, and providing various industries with an innovative, efficient, flexible and low-latency solution. Compared with the existing technology, it has obvious advantages in content dynamics, interactive efficiency, technical scalability, data fusion capability, real-time performance and other aspects. Through intelligent AI Agent control, real-time acquisition and fusion of multimodal data, low-latency real-time transmission and flexible live broadcast mode switching, it provides a more efficient and personalized solution for various real-time interactive live broadcast applications, bringing a new live broadcast experience to enterprises and users. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 Schematic diagram of the system structure of an embodiment of the present invention.

[0048] Figure 2 This is one of the pure digital human live broadcast interface diagrams in the system of an embodiment of the present invention.

[0049] Figure 3 This is the second pure digital human live broadcast interface diagram in the system of an embodiment of the present invention.

[0050] Figure 4 This is one of the live broadcast interface diagrams in the system of an embodiment of the present invention.

[0051] Figure 5 This is the second live broadcast interface diagram of the system in an embodiment of the present invention.

[0052] Figure 6 This is one of the mixed live broadcast interface diagrams in the system of an embodiment of the present invention.

[0053] Figure 7 This is the second hybrid live broadcast interface diagram in the system of an embodiment of the present invention.

[0054] Figure 8 This is one of the unmanned live broadcast interface diagrams in the system of an embodiment of the present invention.

[0055] Figure 9 This is the second unmanned live broadcast interface diagram of the system in an embodiment of the present invention.

[0056] Figure 10 It is a flow chart of a method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0057] The following describes the specific embodiments of the present invention in detail with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present invention and are not intended to limit the present invention.

[0058] like Figure 1 As shown, this invention proposes a system for real-time interactive video stream generation and intelligent interaction, combining AI Agents with multimodal data. This system aims to address existing issues such as static content, insufficient interactivity, and lack of data fusion, further promoting the intelligent and automated generation of interactive video streams. By combining advanced technologies such as artificial intelligence, big data analysis, computer vision, speech processing, and digital humans, this system can generate dynamic interactive content driven by real-time data. It supports live broadcasts of digital humans, live broadcasts of real people, and hybrid modes of digital and real people, meeting the needs of various scenarios.

[0059] The system includes an operator configuration module, an audience interaction module, an AI agent control center, a live broadcast room module, a multimodal data acquisition module, a knowledge base, a content generation engine, a feedback optimization module, and a streaming server module. Under this architecture, each module maintains close collaboration through standardized interfaces and instruction flows. The multimodal data acquisition module provides real-time data, the content generation engine generates corresponding content according to instructions, and the streaming server module distributes the content to the audience. The execution process is as follows:

[0060]

[0061]

[0062] in, Represents the generated content, Represents the content generation engine function (i.e. multimodal fusion), Represents content generation instructions (such as scripts, charts, digital human actions), represents multimodal real-time data (such as data obtained by sensors, cameras, and external system interfaces), t represents time, represents the knowledge base retrieval function (such as vector similarity matching), Represents a knowledge base (including static knowledge + Dynamic Knowledge ), Indicates the preset subject of operation configuration (such as preset topic, knowledge base related data), Represents NLP parsing function (for intent recognition and entity extraction), Indicates audience interaction input (such as text / voice requests, click actions), Indicates the content of real-time broadcast by a live anchor (such as video stream, audio stream, etc.). Represents the video stream content, Indicates streaming media encoding and distribution, It represents the upper bound of streaming media delay. Streaming media encoding and distribution need to dynamically adjust the bit rate (i.e., adaptive bit rate) based on the delay upper bound to ensure that the generated video stream can guarantee the audience's viewing experience.

[0063] As shown in formula (3), the operator configuration module provides preset themes for the system , AI Agent Control Center (AI_Agent) analyzes audience demand input Or the content explained by a real anchor in real time And generate task instructions, including data collection instructions and content generation instructions :

[0064]

[0065] in, represents the control function, Indicates the system status (such as the current live broadcast mode, content generation strategy, and data collection priority).

[0066] As shown in formula (4) and formula (5), the feedback optimization module optimizes and adjusts the live content and system status in real time , which is used to dynamically adjust the system status based on audience interaction data and system performance, ensuring that the live content always meets audience needs and maintains efficient operation. The formula is as follows:

[0067]

[0068]

[0069] in, Indicates audience feedback (such as likes, question frequency, latency indicators), represents the feedback function, Indicates that the content is based on the video stream and audience feedback The loss function constructed to measure audience satisfaction and system performance, Represents the learning rate (used to feed back the adjustment strength of the optimization module), which dynamically adjusts the system state (such as data collection priority and content generation mode) through gradient descent.

[0070] The modular design of the system architecture enables each part to operate independently while collaborating flexibly. In the future, the system can easily integrate new data sources or live broadcast platforms to support more diverse live broadcast modes or platform access. For example, in addition to supporting third-party platforms such as specific social media video accounts, the system can also access other social media or customized platforms as needed to provide personalized interactive live broadcast solutions for different users. Overall, this system architecture not only meets the current needs of real-time interactive live broadcasts, but also has a high degree of scalability and flexibility, which can provide strong support for future technological innovations and changes in business scenarios.

[0071] The operator configuration module serves as the system's configuration center, primarily responsible for pre-live theme setting, data source configuration, and interactive content planning. Operators use this module to define key topics for the live stream and the data sources to be displayed. These configured topics and themes are transmitted to the AI ​​Agent control center, providing a basis for subsequent data collection and automated execution of content generation instructions. This module ensures that live content meets operational objectives and can be flexibly adjusted to meet audience needs.

[0072] The audience interaction module bridges the gap between the system and audiences, supporting both text and voice interaction. Audiences can submit requests to the system by typing questions, selecting options, or using voice commands. Leveraging speech recognition and natural language processing (NLP) technologies, the system converts speech into text and pre-processes the text to ensure accurate understanding of audience intent. After processing, audience requests are communicated to the system's AI Agent control center through a user-friendly interactive interface, providing a basis for subsequent decision-making.

[0073] The AI ​​Agent Control Center is the system's intelligent brain, responsible for core decision-making and automated operations. Its functions extend beyond responding to live audience input and live stream content from live streamers. It also automatically generates content and triggers data collection based on pre-set topics and information from a knowledge base. This enables the system to proactively and intelligently guide the live stream, automatically generating interactive content based on pre-set topics or scenarios without human intervention.

[0074] First, the AI ​​Agent Control Center automatically schedules content generation and data collection tasks based on the themes and topics preset by operators in the operator configuration module. Preset topics might include product showcases, industry trends, and real-time data presentations. In these cases, the AI ​​Agent automatically retrieves relevant information from the knowledge base based on the topic and generates content tailored to the current topic. For example, when showcasing a new product, the AI ​​Agent extracts information such as the product's introduction, features, and price from the knowledge base. It also activates the multimodal data collection module to obtain relevant real-time inventory data. It then combines all this information to generate real-time interactive content, which is then displayed through the content generation engine.

[0075] This automated process, based on pre-set topics, significantly reduces the need for manual intervention and improves the efficiency and accuracy of content generation. Whether preparing content before a live broadcast or presenting real-time data during the broadcast, the AI ​​Agent Control Center automatically executes tasks based on topic settings, ensuring the real-time and relevance of live content.

[0076] During the live broadcast, when viewers ask questions or request interaction, the AI ​​Agent control center dynamically responds based on their input. Viewers can interact with the system via text or voice. The AI ​​Agent uses natural language processing technology to analyze the viewer's needs, identify the type of question, and determine whether further data collection, knowledge base querying, or new content generation is necessary. For example, if a viewer asks, "What's the current inventory?" the AI ​​Agent retrieves inventory data from the data source in real time and uses the content generation engine to update the live broadcast content, providing feedback to the viewer via graphics, text, animation, or voice.

[0077] Therefore, the AI ​​Agent Control Center is more than just the system's response to viewer input; it also serves as the core for proactive, knowledge-driven content generation and data collection based on pre-set topics. This intelligent mechanism enables the system to prepare relevant content before audience interaction and flexibly respond to audience needs during live broadcasts, ensuring consistent, accurate and rapid interaction and content generation.

[0078] Through its dual functions of being driven by preset topics and audience input, the AI ​​Agent Control Center promotes a high degree of automation in the system's content generation and data collection, achieving a more personalized, real-time interactive experience and providing flexibility and efficiency for live broadcasts.

[0079] The multimodal data acquisition module is a crucial component of the system's real-time interaction. The visual data acquisition component uses cameras or image recognition devices to capture real-time footage related to the live broadcast. This not only supports standard video stream acquisition but also extracts key information through image recognition and analysis, providing image data for subsequent content generation. Simultaneously, the sensor network utilizes temperature and humidity sensors, GPS positioning devices, and other sensors to collect real-time data on the environment, location, inventory, and production status. The system also interconnects data through interfaces with third-party systems such as ERP and MES, acquiring key information such as production, inventory, and logistics from external systems in real time. All collected data undergoes preprocessing and is then uniformly transmitted to the knowledge base.

[0080] The content generation engine flexibly generates video content based on multimodal data and scene requirements, and can also decide whether to output subtitles as needed.

[0081] like Figure 2 、 Figure 3 As shown in the figure, in the pure digital human mode, the system uses real-time data to drive the digital human to generate a virtual character video stream, enabling the digital human to dynamically display environmental data, such as the breeding status and inventory information of agricultural products, and combine it with speech synthesis technology for real-time commentary.

[0082] like Figure 4 、 Figure 5 As shown in the figure, in live-action mode, the system uses the AI ​​Agent control center to identify the content of the real-life anchor's real-time broadcast, and then collects data and generates content, so as to realize the real-time superposition of charts, animations or other data onto the anchor's screen, and realize synchronous display with the on-site commentary.

[0083] For the hybrid mode, based on the live broadcast mode, the system not only generates data display-related content, but also drives the generation of digital human content. Then, using image synthesis technology, it realizes seamless collaboration between the digital human and the live broadcaster, making the two content display methods complement each other and further enriching the interactive effect. The process of hybrid live broadcast is as follows:

[0084] 1. The live broadcast room is equipped with real anchors and digital anchors, and the anchors can have a face-to-face effect at a certain angle, such as Figure 6 、 Figure 7 As shown;

[0085] 2. Live broadcast by real anchors to create digital human anchor content , the content is input into the AI ​​Agent control center and content generation engine;

[0086] 3. The AI ​​Agent control center will analyze the anchor's content to obtain data collection instructions and content generation instructions to drive the content generation engine to generate content;

[0087] 4. The content generated by the content generation engine includes the live broadcast content of the digital human, the accompanying explanation materials, and the arrangement between the contents (such as the position between the digital human and the real anchor, etc.). In addition, the video stream of the real anchor will be integrated into the content.

[0088] 5. The content generation engine mixes multimedia content of the live broadcast of real people and the live broadcast of digital people, pushes the synthesized content to the streaming media server, and finally presents it to the audience in the live broadcast room and the real anchor.

[0089] In addition to the live broadcast module for real-person anchors and / or digital human anchors, an unmanned mode is also provided, that is, a mode without digital humans or real people appearing on screen, such as Figure 8 、 Figure 9 As shown in Figure 2, this model can be used in conjunction with the previous model to complete a live broadcast. The feedback optimization module continuously monitors audience interaction data, such as question frequency, likes, and comments, during system operation. Based on this feedback, the system dynamically adjusts its data collection strategy and content generation methods. For hot topics of interest to viewers, the system prioritizes data collection and optimizes the content generation process to ensure that key information is presented to viewers with minimal latency. Through this closed-loop optimization, the system continuously improves data response speed and the overall interactive experience.

[0090] The streaming server module is responsible for encoding, distributing, and transmitting generated content in real time. This module supports access to both self-built streaming servers and third-party streaming platforms. Self-built streaming servers are used for distributing local streaming content. For those who wish to push live content to other platforms, the system also supports pushing live video streams to third-party streaming platforms such as social media video accounts. Through efficient video encoding technology, the system ensures low-latency, high-definition video streaming, while also supporting adaptive bitrate adjustment to accommodate diverse network environments and viewer devices.

[0091] The design of the streaming media server module ensures that the system can broadcast live on multiple platforms simultaneously. Whether it is real-time distribution in the self-built streaming media server or pushing content to a third-party platform, the system can guarantee a smooth and stable viewing experience.

[0092] In the scenario of agricultural product live streaming, the present invention proposes a method for generating real-time interactive video streams and intelligent interaction. The system based on AI Agent and multimodal data is used to provide full-process automated operations from content generation, interactive response to data collection and feedback optimization through the collaboration of various modules. Figure 10 As shown, the details are as follows:

[0093] Step 1: Operation staff configuration and preset topics;

[0094] Before the live broadcast begins, the operator configures the module to set the theme and topic of the live broadcast. For agricultural product live broadcasts, the operator will preset the theme of the live broadcast based on actual needs, such as "The advantages and characteristics of organic apples" or "The picking and distribution process of fresh vegetables." At the same time, the operator can configure relevant live broadcast points, such as product introductions, promotional information, inventory quantities, etc., and link them to existing content in the knowledge base. For example, when configuring the theme P:

[0095] P='Advantages and characteristics of organic apples' (6)

[0096] Preset topics and live broadcast targets are automatically transmitted to the AI ​​Agent Control Center. The AI ​​Agent Control Center then initiates subsequent content generation and data collection tasks based on these preset topics, ensuring that the system automatically prepares relevant content and data before the live broadcast begins.

[0097] Step 2: Based on the video stream topic, the data corresponding to the topic is obtained from the knowledge base, and the data collection and content generation are automatically driven to generate real-time video stream content;

[0098] During the preparation phase of the live broadcast, the AI ​​Agent control center begins automated operations based on the preset live broadcast topic and the content in the knowledge base:

[0099] 1) Automated data collection: The AI ​​Agent control center automatically triggers data collection instructions based on preset topics and requests relevant data from the multimodal data collection module. For example, if collection instruction A is triggered:

[0100] A='Collect real-time apple inventory data' (7)

[0101] For example, if the live broadcast topic is about apples, the AI ​​Agent control center will automatically obtain data such as apples' real-time inventory, production progress, and logistics status from external systems or sensors.

[0102] 2) Automated content generation: The AI ​​Agent control center not only relies on real-time data collection but also extracts relevant information about agricultural products from a knowledge base, such as product characteristics, origin, and health benefits. This information is then combined with charts, videos, or animations to generate content. This generated content is processed by a content generation engine and can be a digital human voice commentary, dynamic chart presentations, or a combination of text and images.

[0103] C instruction ='Generate content containing real-time inventory data' (8)

[0104]

[0105]

[0106] The generated content includes an introduction to the product's features, the relationship between the product and health, the product's production / distribution process, etc., and is pushed to the live broadcast platform through the streaming server module so that viewers can watch it in real time.

[0107] Step 3: Obtain interactive information about the video stream content through audience viewing and interaction;

[0108] After the live broadcast begins, viewers can watch the agricultural product live broadcast content through the streaming server module. The live broadcast content is transmitted to the audience in real time through the video stream. The audience's viewing experience includes the digital human anchor's commentary, product display, interactive Q&A, etc. For example:

[0109] I='How many apples are left' (11)

[0110] At the same time, viewers interact with the host through text or voice. They may ask questions about product details, inquire about pricing and availability, or inquire about product delivery. These interactive messages are transmitted in real time to the AI ​​Agent control center for processing.

[0111] Step 4: The AI ​​Agent control center parses the interaction request and generates a response based on the generated real-time video stream content. This triggers data collection to obtain data corresponding to the interaction information and use it to generate new video stream content.

[0112] The audience's interactive information is transmitted to the AI ​​Agent Control Center, which uses natural language processing (NLP) technology to analyze the audience's input and generate corresponding operation instructions based on the interactive request:

[0113]

[0114] A='Query Inventory' (13)

[0115] C='Generate inventory trend chart' (14)

[0116] 1) Data collection instructions: If a viewer asks about the inventory level of a certain product, the AI ​​Agent will automatically generate data collection instructions and obtain the latest inventory data through the multimodal data collection module.

[0117] 2) Content Generation Instructions: If a viewer asks about a product's health benefits or production process, the AI ​​Agent will query the knowledge base and generate relevant content based on the viewer's needs. The AI ​​Agent control center will pass this information to the content generation engine, which will generate new content.

[0118] Step 5: Content generation and video streaming;

[0119] The generated content (including text, graphics, and digital human commentary) is processed by the content generation engine. Once generated, it is encoded and pushed to the viewer via the streaming server module. This content can be presented to the viewer via the live video stream V, responding to their needs in an interactive manner.

[0120] For example, if a viewer asks about the health benefits of a product, the AI ​​Agent will automatically generate an explanation containing information such as the product's nutritional content and health effects, and present it to the viewer through a digital human anchor or chart.

[0121] Step 6: Audience feedback and optimization: adjust data acquisition and content generation in real time based on interactive information to ensure that key information is presented to viewers with minimal latency.

[0122] During the live broadcast, the system continuously collects interactive feedback from viewers, including likes, frequency of questions, content of comments, and video playback status (such as latency and image quality). This information is analyzed in real time by the feedback optimization module and used to adjust subsequent interaction and content generation strategies.

[0123] R='The audience is very interested in the eco-certification information of a product' (15)

[0124] S(t+1)='Generate more content related to eco-product certification' (16)

[0125] 1) Interaction frequency analysis: If certain questions are asked frequently by the audience, the system will automatically prioritize collecting relevant data or displaying more information.

[0126] 2) Content Optimization: If viewers report that a presentation is not ideal or needs improvement, the system will adjust the presentation through the feedback optimization module to improve the live broadcast experience.

[0127] Through this real-time feedback and optimization mechanism, the system can continuously adjust the live broadcast content and interaction methods based on the audience's needs and feedback, thereby improving the interaction effect and audience satisfaction.

[0128] After the live broadcast, all generated content, audience interaction data, and system logs are stored for subsequent analysis and optimization. Operations personnel can review the live broadcast results, analyze audience engagement, interaction, and feedback, and help optimize future live broadcast topics, content, and interaction methods.

[0129] Through the above process, this system achieves full automation of the live broadcast process of agricultural products. First, the operator configures the module to preset the live broadcast topic. Then, the AI ​​Agent Control Center automatically generates content and collects data based on the preset topic and knowledge base. During the live broadcast, the audience drives the AI ​​Agent Control Center to generate new content through interaction. This content is pushed to the audience through the video stream, and the interactive content is optimized in real time through the feedback optimization module. The entire system ensures the efficiency, real-time nature, and interactivity of the live broadcast process through intelligent and automated workflows, greatly improving the effect of agricultural products live broadcasts and the audience experience.

[0130] The purpose of this invention is to achieve dynamic content generation, intelligent interactive closed loop, full scene compatibility, real-time data drive and automatic control, reduce system latency and improve response speed, as follows:

[0131] 1) Dynamic content generation

[0132] As shown in formula (17), the content of real-time video streams changes over time, with data and interaction. The core innovation of this invention in "dynamic content generation" lies in the real-time integration of multimodal data, and the formation of highly flexible live broadcast content with the help of knowledge bases and preset topics. Before the live broadcast, the system first obtains the theme or product information configured by the operator and combines it with existing text and multimedia materials to prepare preliminary content. During the live broadcast, real-time acquisition equipment (such as cameras, sensors, external system interfaces, etc.) continuously obtains new environmental, business, or user interaction data. By analyzing and filtering this dynamic data, the system can update the commentary script, superimpose visual charts, or generate animation effects in real time on the digital human or real-person live broadcast screen, so that the content is closely aligned with the current business status and audience needs.

[0133]

[0134] For example, livestreaming agricultural products can instantly integrate data from the system into livestreams, providing the latest growing environment data (such as soil moisture or temperature) as well as inventory and logistics progress. The digital human host can adjust their commentary based on this information, demonstrating the ripeness of fruits and vegetables or delivery speeds to viewers. Livestreamers can also overlay visual charts or dynamic subtitles on the livestream to enhance the intuitive presentation of information. This real-time generation and dynamic presentation mechanism not only significantly enhances the authenticity and interactivity of livestream content, but also allows viewers to stay informed about the latest product status, thereby increasing their trust in and engagement with the content.

[0135] 2) Intelligent interactive closed loop

[0136] In order to ensure high efficiency and high-quality interaction during the live broadcast process, the present invention sets up an intelligent interactive closed loop driven by AI Agent in the architecture. The AI ​​Agent parses the input (text or voice) of the audience or user through natural language processing (NLP) technology, automatically determines its intention and decides how to respond in combination with the data decision engine within the system. If the problem involves real-time data query (such as inventory, production status or agricultural product growth information), the AI ​​Agent will call the corresponding interface or trigger multimodal data collection; if new commentary or pictures need to be generated, the AI ​​Agent will schedule the content generation module to update the live broadcast content. The generated results can be presented in various forms such as digital humans, charts or subtitles to help viewers obtain the required information more intuitively. Based on the current system state S(t) and audience feedback R(t), an update function F is constructed to obtain the updated system state S(t+1):

[0137]

[0138] This closed-loop interactive process allows the system to promptly respond to audience needs: when viewers ask questions or provide instructions, the AI ​​Agent makes decisions and provides feedback in real time, forming a complete loop of "viewer input - system decision - content update - feedback display." This cycle can be triggered multiple times and dynamically iterated, ensuring that live content remains synchronized with audience needs and external data, making the entire live broadcast more vivid, accurate, and deeply interactive.

[0139] 3) Full-scenario compatibility

[0140] Another important goal of this invention is to achieve full compatibility across all scenarios, supporting flexible switching between different live broadcast modes, including pure digital human live broadcast, pure human live broadcast, unmanned live broadcast, and collaborative live broadcast between digital humans and real people. The system can automatically select the most appropriate live broadcast mode based on the scenario requirements, ensuring an optimal interactive experience in different application scenarios.

[0141]

[0142] in, 、 、 The weight parameters of vision, speech, and text are expressed respectively. The content engine generates a function G(·) that performs multimodal fusion of video, speech, and text through the weight parameters.

[0143] 3.1) Pure digital human live broadcast mode

[0144] Suitable for completely virtual scenes or content presentation without human intervention, digital humans can generate dynamic explanations and interactions driven by real-time data.

[0145] 3.2) Pure live broadcast mode

[0146] This system is suitable for situations where a real person is needed for hosting or where digital humans cannot be fully relied upon in certain scenarios. The system will enhance the host's interactive capabilities and data presentation effects through technologies such as data overlay.

[0147] 3.3) Hybrid Live Streaming Mode

[0148] Suitable for scenarios that require collaboration between digital humans and real people. In the same live broadcast screen, digital humans and real people can cooperate with each other to jointly display data, provide explanations, or conduct multi-angle interactive presentations.

[0149] 3.4) Unmanned live streaming mode

[0150] The unmanned live broadcast mode is a mode without digital humans or real people appearing on screen. This mode can be combined with the aforementioned pure digital human live broadcast mode, pure real person live broadcast mode, or mixed real person and digital human live broadcast mode to complete a live broadcast.

[0151] This full-scenario compatibility ensures that the system can be flexibly adjusted according to specific needs, adapt to different application scenarios, improve live broadcast effects and meet audience needs.

[0152] 4) Real-time data-driven and automated control

[0153] The core of this invention lies in achieving real-time data-driven content generation and automated control. The system collects various multimodal data (such as sensor data, video streams, and audio data) in real time and automatically analyzes, processes, and provides feedback through AI agents. Based on this, the system can integrate real-time external data (such as logistics progress, production line data, and inventory information) with live content, achieving highly automated interactive content generation.

[0154]

[0155] in, Represents the data acquisition instruction execution function (such as sensor triggering and API calling). The data acquisition instruction execution function updates the current multimodal real-time data D(t) based on the data acquisition instruction A(t) to obtain the multimodal real-time data D(t+1) at the next moment.

[0156] For example, in the field of intelligent manufacturing, the system can adjust live broadcast content based on real-time data from the production line, displaying the production line status in real time, avoiding the inefficiencies of traditional live broadcasts that rely on manual operations. Furthermore, the system can automatically trigger data collection and content generation based on audience needs, reducing the need for manual operations and improving efficiency and interactivity.

[0157] In the field of agricultural production and processing, the system can also use smart sensors and real-time video monitoring to obtain key information such as crop growth environment information, harvesting progress, and agricultural product processing procedures. Based on this data, the AI ​​agent automatically generates content, whether showcasing the freshness of fruits and vegetables, the production environment, or the processing and packaging process, allowing viewers to intuitively understand the entire agricultural product chain from field to packaging during the live broadcast. This allows the host or digital human to clearly present the product's origin advantages, nutritional value, and fast delivery status to the audience through real-time data overlay and multimedia presentation, enhancing the credibility and interactivity of the live broadcast, thereby effectively improving the overall effectiveness of the live broadcast.

[0158] 5) Reduce system latency and improve response speed

[0159] In order to ensure real-time interactive effects, the present invention pays special attention to the low-latency architecture of the system. Through edge computing and optimized streaming media transmission technology, the system can maintain low latency in all aspects of data collection, processing and feedback, ensuring that all interactive content can be presented in real time. The streaming media server needs to meet the upper limit of latency. , achieved through adaptive bitrate technology. Whether it’s digital human-generated content or hybrid mode switching between real and digital humans, the system ensures that latency remains at an industry-leading level, ensuring viewers have a seamless interactive experience.

[0160]

[0161] in, Represents the streaming media delay measurement function, which is used to measure the actual delay of the streaming media server. It represents the delay upper bound generation function, which calculates the delay upper bound required by the streaming media server based on the network bandwidth and edge computing capabilities. ,only The audience's viewing experience can only be guaranteed if the measured actual delay is less than the upper limit of the delay.

[0162] In summary, this invention aims to solve the main problems in current real-time interactive video stream generation technology through the combination of AI Agent and multimodal data, and promote the development of interactive content generation technology towards intelligence, automation, and real-time, thereby providing more efficient, accurate and interactive content generation solutions for various application scenarios.

[0163] This invention achieves full-link automated processing from user command acquisition to real-time content generation, effectively overcoming the shortcomings of traditional interactive video generation technology in terms of real-time performance, interactivity, and data fusion. By organically combining multiple innovative technologies, it provides an efficient, flexible, and low-latency solution for various real-time interactive scenarios. Specifically:

[0164] First, in terms of technological innovation, this invention fully leverages advanced AI agent-driven technology to create an intelligent interactive closed loop. The system's embedded AI agent accurately interprets user or viewer voice and text commands, automatically deciding and scheduling relevant data collection and content generation tasks, automating the entire process from command reception to feedback generation. This process not only reduces the need for manual intervention in traditional live broadcasts and interactive video production, but also significantly improves response speed and interaction efficiency.

[0165] At the same time, this invention utilizes multimodal data deep fusion technology to organically integrate visual, voice, sensor data, and database information. Through data preprocessing, calibration, and fusion algorithms, the system ensures that the generated video stream content is highly consistent with the real-time status of the scene. This collaborative processing of multi-source data ensures that the generated interactive content is both highly real-time and accurately reflects the various dynamic changes in the actual environment, providing users with a more realistic and reliable information display.

[0166] Furthermore, the system supports seamless switching between purely digital humans, purely human humans, and mixed modes, fully demonstrating its flexible content generation strategy. Thanks to scene-aware algorithms and automatic mode adaptation technology, the system automatically selects the most appropriate interaction mode based on actual application needs, meeting the requirements of various scenarios. This flexibility not only enhances the system's applicability but also leaves room for future functional expansion.

[0167] In terms of system architecture design, this invention utilizes edge computing nodes and advanced streaming media transmission technology to ensure low-latency operation in all aspects of data collection, processing, and content generation. Adaptive bitrate technology maintains stable and smooth video transmission even in poor network conditions, significantly enhancing the real-time interactive experience.

[0168] From a technical perspective, this invention is particularly effective in improving interaction efficiency and user experience. Users can directly trigger system responses through simple voice or text commands, avoiding the delays and information errors caused by manual responses in traditional interaction models. At the same time, the deep integration of multimodal data and real-time processing mechanisms ensure that video streams can accurately reflect the real-time status of production, logistics, inventory, and other on-site processes, greatly enhancing the timeliness and accuracy of information.

[0169] This invention also boasts flexible scalability for all scenarios. Whether in a purely digital environment, a live interaction scenario, or a hybrid model combining both, it can provide optimal interactive solutions. This not only enables the system to demonstrate excellent performance in a variety of application scenarios, but also effectively reduces enterprise operation and maintenance costs, promoting the development of digital transformation and automated operations.

[0170] Compared with existing interactive video stream generation technologies, the present invention has significant advantages, especially in terms of content dynamics, interactive efficiency, technical scalability, and data fusion capabilities, as follows:

[0171] 1) Content dynamism:

[0172] In traditional digital human live broadcasts, content generation often relies on preset scripts or fixed data, making it highly static and difficult to dynamically adjust to audience needs or changes in the external environment. Traditional live broadcasts with real people rely on manual operations, resulting in slow content updates and data overlays, making instant responsiveness impossible.

[0173] This invention achieves dynamic content generation based on real-time data through intelligent control from an AI agent control center and multimodal data integration. The AI ​​agent can interpret viewer commands in real time and dynamically update content based on external data sources (such as sensors, databases, and video streams). Whether broadcasting in digital human, live human, or hybrid modes, the system can adjust content presentation in real time, ensuring that live content closely matches the real-time environment and viewer needs.

[0174] 2) Interaction efficiency:

[0175] In existing live streaming technologies, especially live broadcasts with real people, audience interaction often relies on manual responses from the host, resulting in low interaction efficiency and delays in real-time data responses. This not only impacts the viewer experience but also increases the host's workload, leading to longer delays in answering viewer questions.

[0176] This invention utilizes an AI-powered, intelligent, interactive closed loop that instantly interprets audience input and triggers data collection and content generation through automated processes. The system automatically responds to audience needs with virtually no latency and adjusts content presentation in real time. This intelligent interactive mechanism significantly improves audience interaction efficiency and reduces manual intervention and response time.

[0177] 3) Technical scalability:

[0178] Traditional interactive video streaming systems are typically limited to a single live streaming mode, such as a simple digital human live stream or a purely real human live stream. These systems lack flexible scalability and mode switching capabilities, and often require redevelopment or adjustments to meet diverse scenario requirements.

[0179] Through its modular design, this system supports seamless switching between digital humans, real humans, and hybrid modes. Whether in education and training, corporate presentations, retail e-commerce, or smart manufacturing, the system can be flexibly adapted to meet specific needs. Each module of the system is independent and scalable, allowing for easy integration of new data sources and functional modules as new technologies develop and requirements change.

[0180] 4) Data fusion capability:

[0181] Existing technologies have significant limitations in data integration, particularly in traditional livestreams, where data collection and content presentation are often isolated and lack deep integration with external data sources (such as production line status, logistics information, and inventory levels). This results in livestream content failing to reflect real-world dynamics, impacting its credibility and interactivity.

[0182] Through multimodal data acquisition and intelligent data fusion, the present invention effectively integrates multiple information sources, including visual, sensor, voice, and external system data. Through the intelligent scheduling and decision-making engine of the AI ​​Agent, the present invention can integrate various data in real time across different live broadcast modes, ensuring that the live broadcast content is highly consistent with the external environment. For example, the system of the present invention can provide a more accurate and real-time viewer experience by displaying data such as temperature changes on a production line, inventory status, or logistics progress in real time.

[0183] 5) Real-time and low latency:

[0184] Traditional live streaming technologies often face high latency, especially when it comes to data collection, processing, and content generation. Due to system processing and network transmission limitations, latency can reach several seconds or even longer. This delay impacts the viewer experience, especially in scenarios that require instant interactivity and real-time data feedback.

[0185] This system leverages edge computing and streaming optimization technologies to maintain low latency across all stages of data acquisition, content generation, and transmission. By employing the WebRTC protocol and adaptive bitrate technology, the system maintains smooth video transmission in a variety of network environments, ensuring the real-time and interactivity of live content even in bandwidth-constrained environments.

[0186] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A real-time interactive video stream generation system, comprising a control center, a real-time video stream module, a knowledge base, a content generation module, and a feedback module, characterized in that: The control center generates instructions based on the video stream theme, interactive information, and the video stream content of the real-time video stream module. The content generation module obtains corresponding data from the knowledge base according to the instructions and generates new video stream content based on the video stream content of the real-time video stream module. The feedback module adjusts data acquisition and content generation in real time according to the interactive information. The system also includes a multimodal data acquisition module. The control center triggers a data acquisition instruction based on the video stream theme and the interactive information. The multimodal data acquisition module collects real-time multimodal data of products that are not retrieved from the knowledge base. The multimodal data acquisition module extracts key information through identification and analysis to provide key information for subsequent content generation. The content generation module extracts product information based on the video stream theme and the interactive information from the knowledge base to generate corresponding content, combining real-time data from the outside world with the live broadcast content. The control center analyzes the content of the anchor, identifies the content of the anchor's real-time broadcast, and then performs data acquisition and content generation. The collected data is superimposed on the anchor's screen in real time, or the commentary lines are adjusted based on the collected data and displayed to the audience. The feedback module adjusts system strategies, including data collection strategies, in real time based on video stream content and interactive information. A measure of user satisfaction and system performance loss, built based on video stream content and interactive feedback, is used to adjust the feedback module's intensity and dynamically adjust the system state, including data collection priority. The control center classifies content to be generated subsequently according to the acquired theme, and drives the content generation module to sequentially generate the video streams according to the levels. The feedback module adjusts the content generation level of the content to be generated in real time according to the matching degree between the interactive information and the content to be generated subsequently, so as to provide feedback on the interactive information in advance, and further acquires corresponding data based on the interactive information to generate inserted content, and integrates the inserted content into the content to be generated. For the inserted content, anchor points are set in the video stream, and a time step between anchor points is set based on the number of similar interactive messages after the anchor points. In subsequent video streams, the inserted content is regenerated by the feedback module according to the time step to respond to the previously inserted content; The inserted contents are sorted according to the number of identical or similar inserted contents, and the content generation levels of the corresponding contents to be generated are adjusted according to the rankings of multiple inserted contents that are ranked high and correspond to the same contents to be generated, so as to provide feedback on interactive information in advance.

2. A real-time interactive video stream generation system according to claim 1, characterized in that: The real-time video stream module generates real-time live video stream content. The content generation module includes a digital human content generation module, a live video stream content integration module, and a multimedia content mixing module. The live video stream content integration module obtains corresponding data through the instructions and integrates it into the live video stream content. The digital human content generation module generates digital human video stream content that matches the live video stream content according to the instructions, and integrates it with the live video stream content through the multimedia mixing module. The integrated content is displayed to obtain the interactive information, and the integrated content is fed back to the real-time video stream module to assist in generating live video stream content.

3. A real-time interactive video stream generation system according to claim 2, characterized in that: The interactive information includes interactive information based on the content of the video stream shown, and / or interactive information between a real person and a digital human.

4. A real-time interactive video stream generation system according to claim 1 or 2, characterized in that: The content generation module is further configured to generate unmanned video stream content without the real person part and / or the digital human part. The unmanned video stream content is coordinated with the real person video stream content and / or the digital human video stream to generate new video stream content.

5. A real-time interactive video stream generation system according to claim 1, characterized in that: The system further includes a streaming media server module, and the control center is an artificial intelligence control center, which generates content generation instructions based on the video stream theme, the interactive information, the real-time video stream content, the knowledge base information, and the system status; The content generation module obtains data from the knowledge base to generate corresponding content based on content generation instructions, real-time video stream content, video stream topics and analysis of interactive information. The content is encoded by the streaming server module to generate new video stream content, which is pushed to the display end to obtain the interactive information.

6. A real-time interactive video stream generation system according to claim 5, characterized in that: The corresponding content includes video content and / or text content, which is encoded by a streaming media server module to generate new video stream content and / or text content, and pushed to the display end.

7. A real-time interactive video stream generation system according to claim 1, characterized in that: The system further includes a multimodal data acquisition module. The control center triggers a data acquisition instruction A(t) based on the video stream theme and the interactive information. The multimodal data acquisition module acquires real-time multimodal data that is not retrieved from the knowledge base. The content generation module searches the knowledge base based on the video stream theme and the interactive information to generate corresponding content. The formula is as follows: C content (t)=G(C instruction (t)+D(t)+Φ(K,P(t))+Ψ(I(t))+C reallife (t)) V(t)=Stream(G(t),τ) A(t),C instruction (t)=AI_Agent(P(t),I(t),D(t),K(t),S(t),C reallife (t)) Among them, C content (t) represents the generated content, G(·) represents the content generation engine function, and C instruction (t) represents content generation instructions, D(t) represents multimodal real-time data, t represents time, Φ(·) represents knowledge base retrieval function, K represents knowledge base, P(t) represents topic, Ψ(·) represents natural language parsing function, I(t) represents interactive input, C reallife (t) represents the real-time video stream content, V(t) represents the video stream content, Stream(·) represents streaming media encoding and distribution, τ represents the upper bound of streaming media delay, and streaming media encoding and distribution dynamically adjust the bit rate according to the upper bound of delay, AI_Agent(·) represents the control function, and S(t) represents the system status; The feedback module adjusts the system strategy in real time based on the video stream content and interactive information. The formula is as follows: R(t)=Feedback(V(t),I(t)) Among them, R(t) represents interactive feedback, Feedback(·) represents feedback function, represents the loss function that measures user satisfaction and system performance based on video stream content V(t) and interactive feedback R(t), η represents the learning rate, which is used to adjust the intensity of the feedback module. Gradient descent dynamically adjusts the system strategy S(t).

8. A system for generating real-time interactive video streams and intelligent interaction, characterized by: A real-time interactive video stream generation system according to any one of claims 1 to 7, further comprising an operation configuration module and an interaction module, wherein the operation configuration module is connected to a control center and is used for setting themes of video streams, configuring data sources, and planning interactive content, so as to set key topics of the video streams and the data sources to be displayed to generate video stream themes; The interactive module is connected to the control center and is used for the audience to conduct real-time intelligent interaction with the video stream content generated by the system. The system interacts with the user in real time, obtains user requests in text and / or voice, and pre-processes and / or analyzes them to generate interactive information.

9. A method for generating real-time interactive video streams and intelligent interaction, characterized in that: The system for generating real-time interactive video streams and intelligent interaction according to claim 8 performs real-time interactive video stream generation and intelligent interaction, comprising the following steps: Step 1: Preset the video streaming theme; Step 2: Obtain data corresponding to the topic from the knowledge base based on the video stream topic to generate real-time video stream content; Step 3: Obtain interactive information of the video stream content; Step 4: Analyze the interactive information and trigger data collection based on the generated real-time video stream content to obtain data corresponding to the interactive information for generating new video stream content; Step 5: Content generation and video streaming; Step 6: Adjust data acquisition and content generation in real time based on interaction information.

Citation Information

Patent Citations

  • Video rendering method and device for live broadcast scene and electronic equipment

    CN117812375A

  • Live broadcast interaction method and system based on AI digital human

    CN119071521A

  • Self-learning digital human live broadcast method

    CN119967196A