Live broadcast interaction method and apparatus, and live broadcast device
By collecting multimedia information in real time through live streaming equipment and generating a list of interactive candidate operations using a multimodal model, the problem of the lack of intelligent decision-making in online live streaming interactive strategies is solved, realizing intelligent and real-time interaction between the anchor and the audience, and improving the flexibility and effectiveness of the interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-04-07
AI Technical Summary
Existing online live streaming interactive strategies lack dynamic analysis and intelligent decision-making capabilities, resulting in a mechanical interactive process that fails to meet viewers' demands for intelligent interactive experiences.
Multimedia information is collected in real time through live streaming equipment, understood using a multimodal model, and a list of interactive candidate operations is generated. Intelligent interaction is achieved through the anchor's virtual assistant and audience voting.
It enhances the flexibility and effectiveness of live streaming interaction, providing viewers with a smoother and richer participation experience while reducing the operational burden on broadcasters.
Smart Images

Figure CN121815022A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of virtual anchor, and particularly to a live broadcast interaction method and device and live broadcast equipment. BACKGROUND
[0002] As a real-time interactive content form, network live broadcast has been widely applied in entertainment, e-commerce, education and other fields. In the current network live broadcast scene, especially the live broadcast application based on embedded devices, the interaction strategy of anchors and audiences is usually fixed. In such network live broadcast scenes, the generation of the interaction strategy mostly depends on fixed rules, lacks dynamic analysis and intelligent decision-making ability for complex live broadcast scenes, and is difficult to flexibly generate and adjust the interaction strategy, which leads to poor interaction flexibility, making the interaction process appear mechanical and rigid, and failing to meet the demand of audiences for intelligent interaction experience. SUMMARY
[0003] The present application provides a live broadcast interaction method and device and live broadcast equipment, which performs multi-modal understanding on multimedia information in a live broadcast process through a multi-modal model, dynamically generates an interaction candidate operation list highly related to live broadcast content, and thus realizes intelligent interaction between anchors and audiences and improves the flexibility of live broadcast interaction.
[0004] In a first aspect, the present application provides a live broadcast interaction method. The live broadcast interaction method comprises: collecting, by a live broadcast equipment, multimedia information of a live broadcast video in a live broadcast process; performing, by a multi-modal model in the live broadcast equipment, understanding on the multimedia information to generate an interaction candidate operation list; and distributing, by the live broadcast equipment, the interaction candidate operation list to an anchor live broadcast network for live broadcast interaction of audiences and anchors.
[0005] In an implementation manner of the first aspect, performing, by the multi-modal model in the live broadcast equipment, understanding on the multimedia information to generate the interaction candidate operation list comprises: performing, by the multi-modal model, basic understanding on the multimedia information associated with live broadcast content to generate a first candidate operation list; performing, by the multi-modal model, collaborative understanding on the multimedia information associated with multi-modal to generate a second candidate operation list; and generating the interaction candidate operation list according to the first candidate operation list and the second candidate operation list.
[0006] In an implementation form of the first aspect, the generating the first candidate operation list by utilizing the multi-modal model to perform a basic understanding associated with the live content on the multimedia information comprises: encoding the multimedia information into multi-modal tokens by a multi-modal token encoder; rearranging the multi-modal tokens according to a time / space context, and splitting the rearranged multi-modal tokens into several sub-tasks and performing multi-turn dialogue understanding; performing semantic understanding and reasoning on a result of the multi-turn dialogue understanding by a large language model to generate interactive operations; and converting the interactive operations into a structured instruction form to generate the first candidate operation list.
[0007] In an implementation form of the first aspect, the generating the second candidate operation list by utilizing the multi-modal model to perform a collaborative understanding associated with the multi-modal on the multimedia information comprises: understanding multi-modal information in the live video by an agent biased to objectivity, and extracting content understanding biased to objectivity in the multi-modal information; understanding the multi-modal information in the live video by an agent biased to emotion, and evaluating an emotional state of the audience in the multi-modal information; understanding the multi-modal information in the live video by an agent biased to interaction, and analyzing interactive behaviors between the audience and the host in the multi-modal information; and generating the second candidate operation list by a decision-making agent according to the content understanding biased to objectivity, the emotional state, and the interactive behaviors.
[0008] In an implementation form of the first aspect, the distributing, by the live device, the interactive candidate operation list to the host live network for live interaction of the audience and the host comprises: distributing the interactive candidate operation list through the host live network for voting of the audience; determining an audience voting result, and feeding back the audience voting result to the host, so that the host performs corresponding interactive operations in the live process according to the audience voting result; and distributing a live interaction picture of the live process to the live platform.
[0009] In an implementation form of the first aspect, the live interaction method further comprises: evaluating a live atmosphere after the host performs the interactive operation; and updating and optimizing the interactive candidate operation list according to the live atmosphere, and distributing the updated and optimized interactive candidate operation list to the host live network for live interaction of the audience and the host.
[0010] In an implementation form of the first aspect, the live interaction method further comprises: when receiving a host interactive operation initiated by the host, responding to the host interactive operation preferentially.
[0011] In an implementation form of the first aspect, the preferentially responding to the anchor interaction operation comprises: obtaining a currently generated interaction candidate operation list, adding the anchor interaction operation to the interaction candidate operation list with the highest priority, updating the interaction candidate operation list, and distributing the updated interaction candidate operation list to an anchor live streaming network for live streaming interaction of the audience and the anchor.
[0012] In an implementation form of the first aspect, the multimedia information comprises at least one of an image, audio and text.
[0013] In a second aspect, the present application provides a live streaming interaction device. The live streaming interaction device comprises: an anchor virtual assistant module configured to collect multimedia information of a live streaming video in a live streaming process; and a virtual assistant intelligent agent module configured to understand the multimedia information by using a multi-modal model to generate an interaction candidate operation list, wherein the anchor virtual assistant module is further configured to distribute the interaction candidate operation list to an anchor live streaming network for live streaming interaction of the audience and the anchor.
[0014] In a third aspect, the present application provides a live streaming device. The anchor device comprises: a memory configured to store an executable program; and a processor configured to execute the program to enable the live streaming device to perform any of the live streaming interaction methods described above.
[0015] According to embodiments of the present application, the live streaming device collects multimedia information corresponding to a live streaming video in a live streaming process in real time, and calls a multi-modal model to understand and infer the multimedia information, thereby dynamically generating an interaction candidate operation list highly related to the multimedia information in the live streaming video in real time, so that the anchor and the audience can perform live streaming interaction based on the interaction candidate operation list, and realize intelligent and real-time interaction between the anchor and the audience. In summary, the present application changes the generation of interaction strategies from the traditional preset rules or simple trigger mode to an intelligent decision mode based on real-time live streaming content understanding, so that the anchor can interact with the audience more easily, reduces the operation pressure of the anchor in real-time interaction, improves the interaction flexibility and interaction effect, and provides a smoother and richer participation experience for the audience. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 is a flowchart illustrating a live streaming interaction method according to an embodiment of the present application.
[0017] Figure 2 is a system architecture diagram illustrating a live streaming interaction system according to an embodiment of the present application.
[0018] Figure 3 is a block diagram illustrating a live streaming interaction device according to an embodiment of the present application. DETAILED DESCRIPTION
[0019] To describe the technical solutions in the present application in detail, the purposes achieved and effects are explained below in conjunction with the embodiments and the accompanying drawings.
[0020] An embedded network live streaming device is a special hardware designed for real-time video streaming, which usually integrates core functions such as video / audio acquisition, encoding compression, network transmission, etc. With the popularization of 5G network and the progress of related technologies, such devices have been widely used in content creation, online education, enterprise applications, live streaming e-commerce, and security monitoring. However, embedded network live streaming devices still face several technical challenges in practical applications: first, the limited computing resources of the device itself make it difficult to support deep understanding and analysis of live streaming content; second, the device's task scheduling ability is weak, and its interactive strategy generation mostly relies on fixed rules, lacking dynamic analysis and intelligent decision-making capabilities for complex live streaming scenarios, making it difficult to flexibly generate and adjust interactive strategies, with poor interactive flexibility; third, the live streaming scenario has very high real-time requirements, further putting strict requirements on the efficiency and response speed of the end-side processing.
[0021] To solve at least the above technical problems, the present disclosure provides a live streaming interaction method. According to the present disclosure, the live streaming device collects multimedia information corresponding to the live streaming video in the live streaming process in real time, and calls a multi-modal model to understand and infer the multimedia information, thereby generating a list of interactive candidate operations highly related to the multimedia information in the live streaming video in real time and dynamically, so that the host and the audience can interact based on the list of interactive candidate operations, realizing intelligent and real-time interaction between the host and the audience. In summary, the present application changes the generation of interactive strategies from traditional preset rules or simple trigger mode to intelligent decision-making mode based on real-time live streaming content understanding, making it easier for the host to interact with the audience, reducing the host's operation pressure in real-time interaction, improving the interactive flexibility and interactive effect, and providing a smoother and richer participation experience for the audience.
[0022] Hereinafter, the technical solutions according to the present disclosure will be described with reference to specific embodiments and in conjunction with the accompanying drawings.
[0023] Figure 1 is a flowchart showing a live streaming interaction method 100 according to an embodiment of the present disclosure. Referring to Figure 1 , the live streaming interaction method 100 includes the following steps 102 to 106.
[0024] Step 102, collecting multimedia information of a live streaming video in a live streaming process by a live streaming device.
[0025] The multimedia information refers to multi-modal information such as images, audio, and text generated in real time during live streaming, which can include live video streams (such as live picture images), audio streams (such as anchor voice), and accompanying text information (such as real-time comments, bullet screens, gift messages, etc.).
[0026] Step 104: using a multi-modal model in the live streaming device to understand the multimedia information to generate a list of interactive candidate operations.
[0027] The multi-modal model refers to a model that can perform unified encoding, associated understanding, and semantic reasoning on various data such as images, audio, and text. In actual applications, the multi-modal model can include multiple specialized models with different functions and capable of working together. For example, the multi-modal model can include an LLM (Large Language Model) model that supports image, audio, and text understanding, and an ASR (Automatic Speech Recognition) model and a TTS (Text-to-Speech) model that support voice interaction. In this embodiment, the ASR model can be used to convert voice data to text data, and the TTS model can be used to convert text data to voice data.
[0028] The understanding is an intelligent cognitive process that comprehensively extracts semantic context, user intent, and interactive opportunities of live streaming content by performing unified encoding, associated analysis, and semantic reasoning on multimedia information such as images, audio, and text collected in real time during live streaming, and finally generating structured interactive options.
[0029] The list of interactive candidate operations refers to a structured instruction set of a series of specific interactive operations or strategies that can be selected and executed by the anchor or the audience after multi-modal understanding by the multi-modal model, such as a list of product recommendations, a list of game strategy selection, etc. In actual applications, each interactive candidate operation in the list of interactive candidate operations corresponds to a specific interactive intent highly related to the current live streaming content, such as initiating a vote on a specific theme, recommending the next item to be displayed, or proposing a game tactic.
[0030] In practical applications, the understanding of live content in a live video and interaction support can be achieved by deploying the anchor virtual assistant, the virtual assistant agent, and the end-side multi-modal model. Specifically, at the beginning of the live broadcast, the live broadcast device first loads the virtual assistant agent and the end-side multi-modal model, and then the live broadcast device collects multimedia information in the live video during the live broadcast in real time. After the multimedia information is collected, the multimedia information can be transmitted to the virtual assistant agent by the anchor virtual assistant. Based on this, the virtual assistant agent can understand the multimedia information by using the understanding ability of the multi-modal model and combining the live broadcast interaction domain knowledge, so as to generate a list of interactive candidate operations. In this embodiment, the live broadcast interaction domain knowledge is partially included in the multi-modal model, partially (i.e., persistent knowledge) included in the live broadcast interaction knowledge base, and partially (i.e., temporary knowledge) included in the dialogue context. In addition, it should be noted that after the virtual assistant agent calls the multi-modal model, the multi-modal model can understand the multimedia information by performing a series of processes such as splitting, embedding, position encoding, LLM understanding, and task mapping on the multimedia information.
[0031] Step 106, the live broadcast device distributes the list of interactive candidate operations to the anchor live broadcast network for live broadcast interaction between the anchor and the audience.
[0032] The anchor live broadcast network refers to a live broadcast transmission system connecting the anchor end, the live broadcast server, and the audience end, responsible for distributing the list of interactive candidate operations to the live broadcast server and receiving audience voting feedback to realize real-time transmission and response of the list of interactive candidate operations. In practical applications, after the anchor virtual assistant in the live broadcast device generates the list of interactive candidate operations by using the virtual assistant agent module, it can distribute the list of interactive candidate operations to the live broadcast server through the anchor live broadcast network for live broadcast interaction between the anchor and the audience.
[0033] In some embodiments, the understanding of the multimedia information by using the multi-modal model in the live broadcast device to generate the list of interactive candidate operations includes: performing basic understanding of the multimedia information associated with live content by using the multi-modal model to generate a first list of candidate operations; performing collaborative understanding of the multimedia information associated with multi-modal by using the multi-modal model to generate a second list of candidate operations; and generating the list of interactive candidate operations according to the first list of candidate operations and the second list of candidate operations. It should be noted that after the first list of candidate operations and the second list of candidate operations are obtained, the first list of candidate operations and the second list of candidate operations can be input into the multi-modal model for reasoning to obtain the final list of interactive candidate operations.
[0034] The basic understanding refers to the preliminary and structured semantic analysis and information extraction process of the multi-modal model on multimedia information, aiming to identify key entities, events and basic intentions from the multimedia information, and form preliminary interactive options.
[0035] The collaborative understanding refers to the collaborative analysis and comprehensive judgment of the live broadcast situation from different dimensions (such as objective content, subjective emotion, interactive behavior) on the basis of the basic understanding, to generate more detailed and more real-time atmosphere fitting interactive options.
[0036] As can be seen from the above description, by dividing the understanding task into two stages of basic understanding and collaborative understanding, the basic understanding is responsible for quickly extracting key information and direct intentions, and the collaborative understanding is for comprehensive judgment from multiple dimensions (such as objective content, emotion, interactive behavior). Based on this, the final interactive candidate operation list is obtained by synthesizing the interactive options of the two stages, thereby improving the quality and situation fitting degree of the interactive candidate operation, improving the scene adaptation accuracy of the interactive recommendation, and making the generated candidate list more decision-making valuable.
[0037] In some embodiments, utilizing the multi-modal model to perform basic understanding associated with live broadcast content on the multimedia information to generate a first candidate operation list comprises: encoding the multimedia information into multi-modal Tokens by a multi-modal Token encoder; rearranging the multi-modal Tokens according to time / space context, and splitting the rearranged multi-modal Tokens into several sub-tasks and performing multi-round dialogue understanding; utilizing a large language model to perform semantic understanding and reasoning on the results of the multi-round dialogue understanding to generate interactive operations; and converting the interactive operations into structured instruction form to generate the first candidate operation list.
[0038] The multi-modal Token encoder is a component in the multi-modal model responsible for mapping raw input data of different modalities (such as image blocks, audio frames, text words) into a unified, machine-processable vector sequence. The multi-modal Token is a unified embedding vector converted by the multi-modal Token encoder.
[0039] The time context refers to the time sequence dependency relationship of the video frame sequence, and the space context refers to the spatial position relationship of different regions (such as faces, objects) in a single frame picture.
[0040] The rearrangement refers to reordering the multi-modal Tokens according to timestamps and spatial coordinates to ensure that the LLM can understand the content in the order of live event development.
[0041] The multi-round dialogue understanding refers to simulating a multi-round question and answer mechanism, guiding the model to gradually deepen the understanding of the live scene through an iterative prompting engineering, and focusing on different sub-tasks in each round. In actual application, the multi-round dialogue understanding simulates the thinking way of humans clarifying complex problems through continuous questioning, so that the model can perform more fine and coherent semantic analysis on multimedia information in the live broadcast.
[0042] The structured instruction form refers to converting the interactive operation intention obtained after understanding and reasoning into a standardized, explicit data format or command that can be directly parsed and executed by the anchor virtual assistant, etc.
[0043] In actual application, the virtual assistant agent uses the understanding ability of the multi-modal model and combines the live interactive field knowledge to perform basic understanding on multimedia information, and can generate a first candidate operation list. The specific operation is as follows: the virtual assistant agent first loads an end-side multi-modal model, for example, loads a visual and voice model VLM (such as Qwen2.5-VL) supporting a dual-mode of text and image, or loads an omni-modal LLM model (such as Qwen2.5-omni); then the virtual assistant agent obtains multimedia information of a live video collected by a live device; next, the multi-modal model encodes the multimedia information into a unified embedding representation (i.e., multi-modal Token) through a multi-modal Token encoder to support cross-modal data understanding, and rearranges the multi-modal Token according to the time / space context; then, the multi-modal model splits the Token into several sub-tasks and performs multi-round dialogue understanding, and performs semantic understanding and reasoning on the result of the multi-round dialogue understanding through the LLM model to infer the interactive operation; finally, the virtual assistant agent converts the interactive operation into a structured instruction form that can be understood by the anchor virtual assistant to comprehensively generate an interactive candidate operation list.
[0044] As can be known from the above description, by unified encoding of the multi-modal Token, the heterogeneous data is mapped to the same embedding space, solving the problem of cross-modal alignment. Then the time / space context rearrangement mechanism is introduced, and the Token sequence is reorganized according to the time continuity and spatial layout of the live scene, effectively suppressing understanding bias. On this basis, further sub-task splitting and multi-round dialogue mechanism are adopted, the complex problem is decomposed into a progressive reasoning chain, and the single reasoning complexity is reduced, so that the end-side chip can also smoothly run the large language model. Finally, the interactive operation is converted into a structured instruction form that can be directly understood by the anchor virtual assistant, improving the overall interaction efficiency. Based on the above steps, the deep fusion understanding of the multimedia information is realized, thereby effectively generating a first candidate operation list highly related to the live content.
[0045] In some embodiments, the utilizing the multi-modal model to perform multi-modal associated collaborative understanding on the multimedia information to generate a second candidate operation list comprises: performing objective-biased agent understanding on the multi-modal information in the live video to extract objective-biased content understanding in the multi-modal information; performing emotion-biased agent understanding on the multi-modal information in the live video to evaluate an emotional state of an audience in the multi-modal information; performing interaction-biased agent understanding on the multi-modal information in the live video to analyze an interaction behavior between the audience and the host in the multi-modal information; and performing decision-biased agent understanding to generate the second candidate operation list according to the objective-biased content understanding, the emotional state, and the interaction behavior.
[0046] The virtual assistant agent refers to a lightweight artificial intelligence module with a specific function orientation, which can include an objective-biased agent, an emotion-biased agent, an interaction-biased agent, and a decision-biased agent. The objective-biased agent focuses on identifying and extracting objectively existing and factual information in the live content, such as objects, scenes, product parameters, data statistics, etc. The emotion-biased agent focuses on analyzing and evaluating the emotional tendency and emotional state of the audience reflected in the live content (such as audience comment text, voice tone, emoticon package). The interaction-biased agent focuses on identifying and analyzing specific interaction behaviors between the audience and the host, such as the frequency and content of likes, comments, gifts, and questions. The decision-biased agent is responsible for receiving and fusing the analysis results of the above agents, performing comprehensive calculation and judgment according to the pre-set rules, strategies or models, and finally generating a second candidate operation list. It should be noted that the virtual assistant agent can be an end-side agent deployed and running locally on a live device (i.e., an end-side agent).
[0047] In practical applications, after completing the basic understanding of the multimedia information, the collaborative understanding of the multimedia information will be further implemented. To this end, the virtual assistant agent will split the interactive question into multiple analysis dimensions to improve the generation quality of interactive operations in collaborative understanding. In order to facilitate the virtual assistant agent to understand the multimedia information from multiple analysis dimensions, multiple agent sub-modules with different specialities (such as objective-oriented agent, emotion-oriented agent, and interaction-oriented agent) are arranged in the virtual assistant agent, each of which is responsible for an analysis dimension, that is, three types of agents with different divisions of labor analyze the live content in the live video in cooperation. Specifically, the objective-oriented agent can extract objective content in the multimedia information by using a multi-modal model, such as identifying product information, statistical data, and other objective information; the emotion-oriented agent can evaluate the emotional state of the audience in the multimedia information by using a multi-modal model; and the interaction-oriented agent can analyze the interactive behavior between the audience and the host, such as likes, comments, and the like, by using a multi-modal model. On this basis, the host virtual assistant comprehensively considers the multi-dimensional information such as the aforementioned objective information, emotional state, and interactive behavior, and other factors such as user interest and current trend, calculates and decides to generate a series of reasonable second candidate operation list. Of course, when calculating and deciding to generate a series of reasonable second candidate operation list by the decision-oriented agent, it can fill the aforementioned multi-dimensional information into the multi-model VLM conversation prompt word template, and then perform VLM reasoning to generate a reasonable second candidate operation list. Finally, the second candidate operation list will be passed to the host virtual assistant for further action or screening execution according to the second candidate operation list.
[0048] As can be known from the foregoing description, by designing multiple function-specific agents to process information of different dimensions respectively, and by the decision-oriented agent to make comprehensive decisions, the multi-dimensional information such as the objective facts of the live content, the emotional state of the audience, and the real-time interactive behavior can be simultaneously and efficiently considered, thereby improving the intelligent level of interactive recommendation, making the recommended candidate interactive operation more in line with the live atmosphere, and effectively improving the user experience.
[0049] In some embodiments, the distribution of the interactive candidate operation list to the host live network by the live device for live interaction between the audience and the host includes: distributing the interactive candidate operation list through the host live network for the audience to vote; determining the audience voting result and feeding back the audience voting result to the host, so that the host performs corresponding interactive operation in the live process according to the audience voting result; and distributing the live interaction picture of the live process to the live platform.
[0050] In practical applications, by distributing the interactive candidate operation list to the anchor live network, the audience and the anchor can jointly complete a series of intelligent interactive operations based on the interactive candidate operation list. Typical scenarios include: in game live, the audience can vote to select the next tactical strategy; in shopping live, the next displayed goods can be jointly determined; in the chat room live room, the audience for on-mic interaction can be interactively screened; in the talent live (such as singing and dancing), the next performance song, dance action or clothing matching can be selected; and in the “watch together” live room, the subsequent program content can be jointly selected.
[0051] As can be known from the above description, by distributing the interactive candidate operation list to the anchor live network and collecting the audience voting results, and feeding back the results to the anchor to guide the live broadcast, the interactive form is greatly enriched, the live interaction is changed from one-way or simple feedback to a dynamic decision-making process assisted by an intelligent system and participated by the audience, the audience selection instantaneously affects the live content, enhances the audience participation, and improves the user experience.
[0052] In some embodiments, the live interaction method can further include: evaluating the live atmosphere after the anchor performs the interactive operation; and updating and optimizing the interactive candidate operation list according to the live atmosphere, and distributing the updated and optimized interactive candidate operation list to the anchor live network for live interaction by the audience and the anchor.
[0053] The live atmosphere refers to the atmosphere state formed in the live room in real time after performing a specific interactive operation. It should be noted that the live atmosphere can be comprehensively evaluated by various atmosphere information such as audience interaction enthusiasm (such as comment / like growth rate, barrage density, gift value), emotional feedback tendency (positive / negative comment ratio), and user retention rate. In practical applications, after collecting and extracting the atmosphere information related to the atmosphere (such as user retention rate), the atmosphere information representing the live atmosphere can be input into a multi-modal model together with the currently generated interactive candidate operation list for reasoning. Based on this, the multi-modal model can understand the live atmosphere in real time based on the atmosphere information, and then reorder, filter or adjust the candidate operations, and finally output an interactive candidate operation list optimized by the atmosphere. For example, in the game live scenario, if the audience is in high spirits and interacts frequently after voting for a certain tactic, the system can accordingly increase similar tactic options or improve their priority in the next round of recommendation, thereby continuously adapting to audience preferences and improving live interaction effect.
[0054] As can be seen from the above description, by evaluating the live atmosphere after the anchor performs the interactive operation, and dynamically adjusting the subsequent generated interactive candidate operation according to the live atmosphere, the interactive candidate operation is made to be more suitable for the current live situation and audience preference, and the optimized interactive candidate operation is put into a new round of interaction, thereby improving the accuracy, attractiveness and user stickiness of live interaction.
[0055] In some embodiments, the live interaction method can further include: when receiving an anchor-initiated anchor interaction operation, responding to the anchor interaction operation preferentially.
[0056] The anchor interaction operation refers to an instruction or action that is actively triggered or input by an anchor and explicitly expresses the anchor's interaction intention, such as an interaction operation initiated by the anchor through physical buttons, voice commands, or screen touch, etc. The anchor interaction operation has the highest priority.
[0057] As can be seen from the above description, by responding to the anchor interaction operation preferentially, the anchor always has the highest decision-making power and control, thereby guaranteeing the anchor's use experience.
[0058] In some embodiments, responding to the anchor interaction operation preferentially includes: obtaining the current generated interactive candidate operation list, adding the anchor interaction operation to the interactive candidate operation list with the highest priority to update the interactive candidate operation list; and distributing the updated interactive candidate operation list to the anchor live network for live interaction by the audience and the anchor.
[0059] In actual application, adding the anchor interaction operation to the interactive candidate operation list with the highest priority can mean placing the anchor interaction operation in the most prominent, most likely to be selected or default execution position in the interactive candidate operation list, for example, can be placed at the top, marked as "anchor recommendation" or set as the default option.
[0060] Illustratively, the anchor interacts with the user more easily by using the anchor virtual assistant and the virtual assistant intelligent agent behind it to understand the live scene and generate an interactive candidate operation list. The specific process is as follows: first, the anchor initiates an anchor interaction operation actively through a specific way such as voice command, button click, etc. The anchor interaction operation can be in the form of asking questions, voting or game invitation, etc. Then, the anchor virtual assistant obtains the interactive candidate operation list generated by the virtual assistant intelligent agent, and adds the anchor-initiated highest-priority active interaction operation to the interactive candidate operation list to update the interactive candidate operation list. Next, the anchor virtual assistant distributes the updated interactive candidate operation list to the live network. The video anchor continues to live according to the interactive operation selected by the user through voting, and the live device synchronously distributes the interactive picture to the live platform, completing the entire interactive closed loop.
[0061] From the above description, when the host initiates the interaction operation actively, the host interaction operation is dynamically integrated into the generated candidate list with the highest priority, establishing an arbitration mechanism with the highest priority of the host instruction, guaranteeing the control right of the host creation, realizing the organic integration of artificial intervention and intelligent recommendation, and improving the flexibility and practicality of the system.
[0062] Figure 2 is a system architecture diagram illustrating a live interaction system according to an embodiment of the present disclosure. The live interaction system is arranged in a live device, and with reference to Figure 2 , the system includes a host virtual assistant network and a multi-modal intelligent agent on the terminal side.
[0063] The host virtual assistant network includes a host virtual assistant (i.e., a host virtual assistant module) and a host live network and other core components. The host virtual assistant undertakes part of the intelligence of multimedia interaction of the network host and participates in the host live network. In actual application, the host virtual assistant can be used to acquire multimedia information of a live video collected by the live device and send it to the virtual assistant intelligent agent to generate an interaction candidate operation list (i.e., multimedia interaction information) through the virtual assistant intelligent agent, and at the same time, it can also interact with the user preliminarily, such as sending the interaction candidate operation list to the audience for voting through the host live network.
[0064] The multi-modal intelligent agent on the terminal side includes a virtual assistant intelligent agent (i.e., a virtual assistant intelligent agent module) and a multi-modal model and other core components. The virtual assistant intelligent agent understands the multimedia information by calling the multi-modal model, and generates an interaction candidate operation list based on the multimedia information. In this embodiment, the understanding process of calling the multi-modal model to understand the multimedia information includes two levels of basic understanding and collaborative understanding. In addition, the multi-modal model can include various models such as a YOLO model for image target detection, an ASR model for speech recognition, a multi-modal LLM for cross-modal understanding, and a TTS model for speech synthesis.
[0065] In summary, the live interaction system includes four components: a host virtual assistant, a host live network, a virtual assistant intelligent agent, and a multi-modal model. The host virtual assistant undertakes part of the intelligence of multimedia interaction of the network host and participates in the host live network. The virtual assistant intelligent agent utilizes the understanding ability of the multi-modal model and combines with the knowledge in the field of live interaction to generate multimedia interaction information. Based on the above components, the typical data flow of the live interaction system is: host virtual assistant (acquire multimedia information) -> virtual assistant intelligent agent (multimedia information -> multi-modal model -> interaction candidate operation list) -> host virtual assistant (distribute interaction candidate operation list) -> host live network (respond to interaction candidate operation list).
[0066] To make the application easier to understand, an example application is provided below, in which the live interaction method application flow of the application can be as follows.
[0067] (1) Multimedia information collection: the live device collects multimedia information in the live process in real time and sends it to the anchor virtual assistant. The multimedia information can include the anchor's explanation of the picture, audio (such as anchor voice, background music), and real-time scrolling of audience comment barrage, and other different modal information.
[0068] (2) Multi-modal understanding and candidate list generation First, basic understanding: the anchor virtual assistant sends the obtained multimedia information to the virtual assistant agent, and the virtual assistant agent loads the multi-modal model to understand the multimedia information. The specific understanding process of the multimedia information can be as follows: the multi-modal model encodes the multimedia information into multi-modal tokens through a multi-modal token encoder, and rearranges them according to the time sequence and picture spatial relationship; then, the multi-modal model splits the rearranged multi-modal tokens into several sub-tasks and performs multi-round dialogue understanding; then, a large language model is used to understand the semantics of the multi-round dialogue understanding result and reasoning, generating some interactive operations, which will be converted into structured instructions to form a first candidate operation list.
[0069] Second, collaborative understanding: the collaborative understanding can be jointly worked by multiple special agents, specifically: the objective agent can use the multi-modal model to extract the objective content understanding in the multimedia information, such as identifying product information, statistical data, and other objective information; the emotion-oriented agent can use the multi-modal model to evaluate the emotional state of the audience in the multimedia information; the interactive agent can use the multi-modal model to analyze the interactive behavior between the audience and the anchor in the multimedia information, such as likes, comments, etc.; and the decision-making agent can comprehensively consider the above information to generate a second candidate operation list.
[0070] It should be noted that the first candidate operation list and the second candidate operation list are both a set of interactive candidate operations. For example, in the game live scene, the first / second candidate operation list can be a vote for the next team action, such as: A. Gather to open the shadow master; B. Huddle to push the middle tower; C. Invade the enemy's wild area to plunder resources; in the shopping live scene, the first / second candidate operation list can be a vote for the next displayed product, such as: A. Mini humidifier; B. Air fryer; C. Smart fragrance machine.
[0071] (3) List integration and anchor priority: the anchor virtual assistant obtains the first candidate operation list and the second candidate operation list generated above and integrates them based on the two lists to obtain the final interactive candidate operation list.
[0072] It needs to be particularly pointed out that when the interactive candidate operation list is integrated, if the host suddenly initiates a host interactive operation, for example, in a shopping live broadcast scene, the host initiates to say: “I see that many friends ask about the small dehumidifier. Do we need to explain the small dehumidifier first?” In a game live broadcast, the host initiates to say: “Wait, the opposite side just exchanged a flash, and their auxiliary position is out of sync. I think this wave can try to directly open two towers and play a time difference. Do you want to try a bet?” This host interactive operation (such as explaining the small dehumidifier / strongly pushing the two towers in the middle road) is immediately added to the interactive candidate operation list by the system with the highest priority to finally integrate the interactive candidate operation list.
[0073] (4) List distribution and audience voting: the host virtual assistant in the live broadcast device distributes the integrated interactive candidate operation list to the live broadcast server through the host live broadcast network to present it to the audience for voting.
[0074] (5) Voting result execution and live broadcast promotion: after the audience voting result is obtained, the host performs subsequent live broadcast for the interactive operation with the highest number of votes in the audience voting result, for example, in a game live broadcast scene, if “huddle to push the two towers in the middle road” leads the voting, the host will next preferentially and decisively command the teammates to huddle to push the two towers in the middle road; in a shopping live broadcast scene, if “small dehumidifier” leads the voting, the host will next start to focus on introducing the small dehumidifier.
[0075] (6) Atmosphere evaluation and strategy optimization: the system evaluates the live broadcast atmosphere after this round of interaction (i.e., the host performs the corresponding interactive operation according to the voting result), updates and optimizes the interactive candidate operation list according to the live broadcast atmosphere, and distributes the updated and optimized interactive candidate operation list to the host live broadcast network for the audience and the host to re-perform live broadcast interaction.
[0076] Figure 3 is a block diagram illustrating a live broadcast interaction device according to an embodiment of the present disclosure. Referring to Figure 3 , the live broadcast interaction device 200 includes a host virtual assistant module 210 and a virtual assistant intelligent agent module 220.
[0077] The host virtual assistant module 210 is configured to collect multimedia information of a live broadcast video in a live broadcast process.
[0078] The virtual assistant intelligent agent module 220 is configured to understand the multimedia information by using a multi-modal model to generate an interactive candidate operation list.
[0079] The host virtual assistant module 210 is further configured to distribute the interactive candidate operation list to a host live broadcast network for live broadcast interaction by the audience and the host.
[0080] In some embodiments, the anchor virtual assistant module 210 can be further configured to evaluate a live streaming atmosphere after the anchor performs the interactive operation, and update the list of interactive candidate operations according to the live streaming atmosphere, and distribute the updated list of interactive candidate operations to the anchor live streaming network for live streaming interaction by the audience and the anchor.
[0081] In some embodiments, the anchor virtual assistant module 210 can be further configured to prioritize responding to the anchor-initiated anchor interactive operation when the anchor-initiated anchor interactive operation is received.
[0082] It should be understood that the anchor virtual assistant module 210 and the virtual assistant intelligent agent module 220 can be further configured to perform the respective corresponding steps or actions in the live streaming interaction method described in the above embodiments, which will not be repeated here.
[0083] According to another aspect of the present application, the present disclosure also provides a live streaming device. The live streaming device comprises a memory and a processor. The memory is configured to store an executable program. The processor is communicatively connected with the memory and is configured to execute the program to enable the live streaming device to perform the live streaming interaction method as described above.
[0084] In summary, according to the live streaming interaction method and device and the live streaming device provided by the present application, on the one hand, the present application proposes an anchor virtual assistant scheme based on a multi-modal intelligent agent. The technical architecture of the present application mainly includes four components: an anchor live streaming network, an anchor virtual assistant, a virtual assistant intelligent agent, and a multi-modal model. The anchor virtual assistant is responsible for undertaking part of the multimedia interaction functions of the network anchor and accessing the anchor live streaming network. The virtual assistant intelligent agent generates a list of interactive candidate operations relying on the content understanding ability of the multi-modal model and combining professional knowledge in the live streaming interaction field. On the other hand, the present application can support the end-side deployment of LLM-VL and other multi-modal models by introducing a special application chip with medium NPU computing power, thereby providing the necessary computing basis for the end-side deployment of LLM-VL and other multi-modal models, expanding the depth understanding ability of the device for video content, and effectively supporting the depth understanding and analysis of live streaming content. Moreover, the present application also designs and introduces a collaborative architecture composed of an anchor virtual assistant and a virtual assistant intelligent agent. Through the division of labor and cooperation among the virtual assistant intelligent agents, the dynamic arrangement and execution ability of complex interactive tasks is significantly enhanced. In addition, the present application can realize multi-modal real-time understanding and low-latency response for live streaming content in a resource-constrained end-side environment through key technologies such as model lightweight, process optimization (such as multi-modal Token efficient encoding, sub-task splitting, and multi-round dialogue understanding), which helps to alleviate the bottlenecks of the device in computing power, task scheduling, and real-time performance.
[0085] It should be understood that the terms "first", "second" and so on are used to describe various information in the present application, but these information should not be limited to these terms, and these terms are only used to distinguish the same type of information from each other. For example, the "first" information can also be referred to as "second" information, and similarly, the "second" information can also be referred to as "first" information without departing from the scope of the present application.
[0086] The above description is only an embodiment of the present application, and does not limit the patent scope of the present application, and any equivalent transformation or direct or indirect application in the related technical field using the content of the present application specification and drawings is also included in the patent protection scope of the present application.
Claims
1. A live streaming interactive method, characterized in that, include: Multimedia information from the live video collected by the live streaming equipment during the live streaming process; The multimedia information is understood using the multimodal model in the live streaming device to generate a list of interactive candidate operations; as well as The live streaming device distributes the list of interactive candidate operations to the live streaming network for viewers and streamers to interact during the live stream.
2. The live streaming interactive method according to claim 1, characterized in that, The process of understanding the multimedia information using a multimodal model in the live streaming device to generate a list of interactive candidate operations includes: The multimodal model is used to perform a basic understanding of the multimedia information in relation to the live broadcast content in order to generate a first candidate operation list; The multimedia information is subjected to multimodal collaborative understanding using the multimodal model to generate a second candidate operation list; and The interactive candidate operation list is generated based on the first candidate operation list and the second candidate operation list.
3. The live streaming interactive method according to claim 2, characterized in that, Using the multimodal model to perform a basic understanding of the multimedia information in relation to the live broadcast content, and generating a first candidate operation list, includes: The multimedia information is encoded into a multimodal token using a multimodal token encoder; The multimodal tokens are rearranged according to the temporal / spatial context, and the rearranged multimodal tokens are divided into several sub-tasks for multi-round dialogue understanding. Using a large language model, semantic understanding and reasoning are performed on the results of the multi-turn dialogue understanding to generate interactive operations; and The interactive operation is converted into a structured instruction form to generate the first candidate operation list.
4. The live streaming interactive method according to claim 2, characterized in that, Utilizing the multimodal model to perform collaborative understanding of the multimedia information in a multimodal context to generate a second candidate operation list includes: The system uses a biased intelligent agent to understand the multimodal information in the live video and extracts the biased content from the multimodal information. The system uses an emotion-biased agent to understand the multimodal information in the live video and assesses the emotional state of the audience within the multimodal information. By using a semi-interactive intelligent agent to understand the multimodal information in the live video, the interactive behavior between the viewer and the broadcaster in the multimodal information is analyzed; and The second candidate operation list is generated by a biased decision-making agent based on the biased objective content understanding, the emotional state, and the interactive behavior decision.
5. The live streaming interactive method according to claim 1, characterized in that, The live streaming device distributes the list of interactive candidate operations to the live streaming network for viewers and streamers to interact during the live stream, including: The list of interactive candidate operations will be distributed through a live streaming network for viewers to vote on. Determine the audience voting results and feed them back to the streamer so that the streamer can perform corresponding interactive actions during the live broadcast based on the audience voting results; and Distribute the live interactive footage of the live stream to the live streaming platform.
6. The live interactive method according to claim 5, characterized in that, Also includes: Assess the live stream atmosphere after the host performs the interactive operation; as well as Based on the live broadcast atmosphere, the interactive candidate operation list is updated and optimized, and the updated and optimized interactive candidate operation list is distributed to the live broadcast network for viewers and broadcasters to interact during the live broadcast.
7. The live streaming interactive method according to claim 1, characterized in that, Also includes: When a host interaction operation is received from the host, the host interaction operation is responded to first.
8. The live streaming interactive method according to claim 7, characterized in that, Prioritizing responses to the broadcaster's interactive actions includes: Obtain the currently generated list of candidate interactive operations, add the broadcaster's interactive operation to the list of candidate interactive operations with the highest priority, and update the list of candidate interactive operations; and The updated list of interactive candidate operations will be distributed to the live streaming network for viewers and streamers to interact during the live stream.
9. The live streaming interactive method according to claim 1, characterized in that, The multimedia information includes at least one of images, audio, and text.
10. A live interactive device, characterized in that, include: The virtual assistant module for broadcasters is configured to collect multimedia information from the live video during the broadcast. as well as The virtual assistant intelligent agent module is configured to understand the multimedia information using a multimodal model to generate a list of interactive candidate operations. The virtual assistant module for broadcasters is further configured to distribute the list of interactive candidate operations to the broadcaster live streaming network for viewers and broadcasters to interact during the live stream.
11. A live streaming device, characterized in that, include: The memory is configured to store executable programs; as well as A processor is configured to execute the program to cause the live streaming device to perform the live interactive method according to any one of claims 1 to 9.