Live broadcast interaction management method and device, equipment, storage medium and product
By analyzing and processing the video, audio and text data in the live broadcast room, and using the trained model to obtain interactive information, the problem of low accuracy of live broadcast assistant information is solved, and the interaction effect between the anchor and the audience is improved.
Patent Information
- Application Number
- CN202510029700.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-06-06
AI Technical Summary
The information feedback from existing live broadcast assistants is poor, which makes it difficult for anchors to interact with the audience accurately and efficiently, and the interaction effect of live videos is poor.
By obtaining video, audio and text data in the live broadcast room, the trained live broadcast assistant model and callback function model are used for analysis and processing, the live broadcast data requirement instructions are obtained and external data is obtained from the external system, and live broadcast interactive information is generated to improve information accuracy and usability.
It improves the accuracy and usability of the information feedback from the live broadcast assistant, helps the anchor to have more effective live broadcast interactions and improves the interactive effect of live videos.
Smart Images

Figure CN120111259A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technology, and in particular to a live interactive management method, device, equipment, storage medium and product. Background Art
[0002] With the development of the Internet and computer technology, the application of live video broadcasting has become more and more widespread, and more and more anchors and viewers have participated in live video broadcasting. When the anchor starts the live broadcast in the live broadcast room, he can interact with the users who enter the live broadcast room or participate in the live broadcast through voice, video, and text. The audience can also enter text in the chat window of the live broadcast room to interact with the anchor.
[0003] However, during the live broadcast, the host (for example, a new host) has an inaccurate grasp of the live broadcast rhythm and finds it difficult to interact with the audience accurately and efficiently. The live broadcast assistant provided in the live broadcast room generally only provides the host with information such as the number of viewers in the current live broadcast room. The information provided to the host is limited, and the information feedback from the live broadcast assistant is less accurate. It cannot effectively help the host to interact with the live broadcast, and the interactive effect of the live video broadcast is poor. Summary of the invention
[0004] The embodiments of the present application provide a live broadcast interaction management method, device, equipment, storage medium and product to solve the problem that the information feedback from the live broadcast assistant in the related technology is poor in accuracy, which can effectively improve the accuracy and availability of the information feedback from the live broadcast assistant, help the anchor to interact with the live broadcast, and improve the interactive effect of the live video broadcast.
[0005] In a first aspect, an embodiment of the present application provides a live interactive management method, including:
[0006] Acquire live video data, live audio data and live text data of the live broadcast room, and determine video text description data according to the live video data, and determine audio text description data according to the live audio data;
[0007] The video text description data, the audio text description data and the live text data are sent to the trained live assistant model, and the video text description data, the audio text description data and the live text data are analyzed and processed by the live assistant model to obtain a live data demand instruction;
[0008] Sending the live broadcast data demand instruction to the trained callback function model, analyzing and processing the live broadcast data demand instruction through the callback function model to obtain an external data call instruction;
[0009] According to the external data calling instruction, external data is obtained from a preset external system, and the external data is sent to the live broadcast assistant model. The video text description data, the audio text description data, the live broadcast text data and the external data are analyzed and processed by the live broadcast assistant model to obtain live broadcast interaction information, and live broadcast interaction is performed in the live broadcast room according to the live broadcast interaction information.
[0010] In a second aspect, an embodiment of the present application provides a live broadcast interaction management device, including a data acquisition module, a data analysis module, an instruction determination module and an interaction management module, wherein:
[0011] The data acquisition module is configured to acquire live video data, live audio data and live text data of the live broadcast room, and determine video text description data according to the live video data, and determine audio text description data according to the live audio data;
[0012] The data analysis module is configured to send the video text description data, the audio text description data and the live text data to the trained live assistant model, and analyze and process the video text description data, the audio text description data and the live text data through the live assistant model to obtain a live data demand instruction;
[0013] The instruction determination module is configured to send the live data demand instruction to the trained callback function model, analyze and process the live data demand instruction through the callback function model, and obtain an external data call instruction;
[0014] The interaction management module is configured to obtain external data from a preset external system according to the external data call instruction, and send the external data to the live broadcast assistant model, analyze and process the video text description data, the audio text description data, the live broadcast text data and the external data through the live broadcast assistant model to obtain live broadcast interaction information, and perform live broadcast interaction in the live broadcast room according to the live broadcast interaction information.
[0015] In a third aspect, an embodiment of the present application provides a live interactive management device, including: a memory and one or more processors;
[0016] The memory is used to store one or more programs;
[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the live interactive management method as described in the first aspect.
[0018] In a fourth aspect, an embodiment of the present application provides a non-volatile storage medium storing computer executable instructions, which, when executed by a computer processor, are used to execute the live interactive management method as described in the first aspect.
[0019] In the fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor of the device reads and executes the computer program from the computer-readable storage medium, so that the device executes the live interactive management method as described in the first aspect.
[0020] The embodiment of the present application obtains video text description data, audio text description data and live text data of the live broadcast room, sends the video text description data, audio text description data and live text data to a trained live broadcast assistant model, analyzes and processes the video text description data, audio text description data and live broadcast text data through the live broadcast assistant model to obtain a live broadcast data demand instruction, sends the live broadcast data demand instruction to the trained callback function model, analyzes and processes the live broadcast data demand instruction through the callback function model to obtain an external data call instruction, obtains external data from a preset external system according to the external data call instruction, and sends the external data to the live broadcast assistant model, analyzes and processes the video text description data, audio text description data, live broadcast text data and external data through the live broadcast assistant model to obtain live broadcast interaction information, and performs live broadcast interaction in the live broadcast room according to the live broadcast interaction information, and manages the live broadcast interaction of the live broadcast room according to the video, voice, text data and external data of the live broadcast room, thereby effectively improving the accuracy and availability of information feedback from the live broadcast assistant, helping the anchor to perform live broadcast interaction, and improving the interactive effect of video live broadcast. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a flow chart of a live interactive management method provided by an embodiment of the present application;
[0022] Figure 2 is a flow chart of another live interactive management method provided by an embodiment of the present application;
[0023] Figure 3 It is a structural diagram of a live interactive management device provided in an embodiment of the present application;
[0024] Figure 4 It is a structural diagram of a live interactive management device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical scheme and advantages of the present application clearer, the specific embodiments of the present application are further described in detail below in conjunction with the accompanying drawings. It is understood that the specific embodiments described herein are only used to explain the present application, rather than to limit the present application. It should also be noted that, for the convenience of description, only the part related to the present application but not all the contents are shown in the accompanying drawings. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow chart describes each operation (or step) as a sequential process, many of the operations therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of each operation can be rearranged. The above process can be terminated when its operation is completed, but it can also have additional steps not included in the accompanying drawings. The above process can correspond to a method, a function, a procedure, a subroutine, a subprogram, etc.
[0026] The live broadcast interaction management method provided in this application can be applied to the interaction management of video live broadcast rooms. It aims to manage the live broadcast interaction of the live broadcast room based on the video, voice, text data and external data of the live broadcast room, improve the accuracy and availability of information feedback from the live broadcast assistant, help the anchor to interact with the live broadcast, and improve the interactive effect of the video live broadcast.
[0027] In the existing live interactive management, the anchor usually uses the anchor assistant to report the relevant situation of the live broadcast room to the anchor, such as providing the anchor with information such as the number of viewers in the current live broadcast room. However, the information provided by the live broadcast assistant to the anchor is limited, and the accuracy of the information fed back by the live broadcast assistant is poor, which leads to the anchor's inaccurate grasp of the live broadcast rhythm, making it difficult to interact with the audience accurately and efficiently, and failing to effectively help the anchor to interact with the live broadcast, resulting in poor interactive effect of the live video broadcast. Based on this, a live interactive management method according to an embodiment of the present application is provided to solve the technical problem of poor accuracy of the information fed back by the live broadcast assistant used in the existing live interactive management solution.
[0028] Figure 1 A flow chart of a live broadcast interaction management method provided in an embodiment of the present application is given. The live broadcast interaction management method provided in an embodiment of the present application can be executed by a live broadcast interaction management device, which can be implemented by hardware and / or software and integrated in a live broadcast interaction management device.
[0029] The following description is made by taking the live broadcast interaction management device executing the live broadcast interaction management method as an example. Figure 1 , the live interactive management method includes:
[0030] S110: Acquire live video data, live audio data, and live text data of the live broadcast room, and determine video text description data according to the live video data, and determine audio text description data according to the live audio data.
[0031] Exemplarily, the live video data, live audio data and live text data of the live broadcast room are obtained in real time. Among them, the live video data is the screen data of the live broadcast room, and the live video data can be provided by the anchor end, for example, the video data collected by the anchor end through a video acquisition device (such as a camera) is used as the live video data, or the shared data screen (such as a desktop, a preset file, etc.) is used as the live video data. The live video data may also include the screen data of other connected microphone users (such as other anchors or viewers) in the current live broadcast room. The live audio data is the audio data of the live broadcast room, and the live audio data can be the audio data provided by the anchor end, or it can be the audio data provided by other connected microphone users in the current live broadcast room. The live text data is the text data that appears in the live broadcast room, and the live text data can be the text data entered by the anchor and the audience in the live broadcast room in the chat window.
[0032] In one embodiment, when acquiring live video data and live audio data, text description data may be extracted from the live video data (live video data may be all or multiple video frames of a live video stream) and the live audio data, respectively, wherein the text description data may be used to reflect the information of the live video data or the live audio data. For example, the video text description data is determined based on the live video data, and the audio text description data is determined based on the live audio data. Among them, the video text description data may be used to reflect the picture information in the live broadcast room (for example, describing the corresponding scene in the live video data (indoor and outdoor, lighting, weather, background, etc.), what is included (characters, objects, etc.), character features (long or short hair, hair color, gender, etc.), what happened (for example, someone is singing, dancing, exercising, etc.)), etc., and the audio text description data may be used to reflect the sound in the live broadcast room (for example, lyrics, what the characters say, etc.).
[0033] In a possible embodiment, the live video data can be input into a trained video text description data extraction model, and the video text description data of the live video data can be extracted through the video text description data extraction model. Among them, the data extraction model can be a stableDiffusion model, and the video screen of the stableDiffusion model host or guest is subjected to BLIP (Bootstrapping Language-Image Pretraining, a visual-language model framework) and DeepBooru (a Chinese word segmentation tool based on Deep learning, which can perform back-inference of prompt words) analysis, and the video screen is converted into an output label mode or text content described in natural language, and these text contents can be used as video text description data. Optionally, the video text description data can also be translated, and then these text contents are converted into voice through speech synthesis (TTS, Text-To-Speech) technology, and sent to the audience end for playback, so that visually impaired people can "watch" the live content by listening to natural language or the voice corresponding to the prompt words, thereby improving the user coverage and the live experience of special groups of users.
[0034] In a possible embodiment, the live audio data can be sent to a trained audio text description data extraction model (such as an ASR speech model), and the audio text description data of the live audio data can be extracted through the audio text description data. Optionally, the audio text description data extraction model can be obtained by secondary development using the Whisper module.
[0035] S120: Send the video text description data, audio text description data and live broadcast text data to the trained live broadcast assistant model, and analyze and process the video text description data, audio text description data and live broadcast text data through the live broadcast assistant model to obtain live broadcast data demand instructions.
[0036] Exemplarily, the video text description data, audio text description data and live text data obtained above can be sent to the trained live assistant model, and the video text description data, audio text description data and live text data can be analyzed and processed by the live assistant model, and the live data requirement instruction can be output. The live assistant module provided by this solution can be obtained by fine-tuning the large language model (LLM). For example, the live assistant module can be obtained by fine-tuning the large language model using the data related to the live broadcast room, so that the live assistant module can accurately handle the tasks related to the live broadcast room.
[0037] In one embodiment, the live broadcast assistant model can analyze the video text description data, audio text description data and live broadcast text data based on preset instructions, and output corresponding live broadcast data demand instructions. The preset instructions can be described in natural language. For example, the preset instructions can be "analyze the current live broadcast room interaction ratio", "the current live broadcast room information flow is: xxx..., analyze the core focus of the discussion and whether it is interesting". "What is the behavior of most viewers in the current live broadcast room", "analyze the gender and regional distribution of the audience in the current live broadcast room, are there any good chat suggestions", "What is the style of live broadcast that the audience in the live broadcast room are used to watching? How should I adjust it?". Among them, the live broadcast data demand instruction can be understood as an instruction to obtain live broadcast related data. The live broadcast data demand instruction can be a data acquisition request described in natural language. For example, the live broadcast data demand instruction can be "obtain the audience composition information of the live broadcast room from the data background", "query the gender and regional distribution of the audience in the live broadcast room", "query the viewing style of the audience in the live broadcast room", etc.
[0038] S130: Send the live broadcast data demand instruction to the trained callback function model, analyze and process the live broadcast data demand instruction through the callback function model, and obtain the external data call instruction.
[0039] Exemplarily, the live broadcast data demand instruction output by the live broadcast assistant model is sent to the trained callback function model, and the live broadcast data demand instruction is analyzed and processed by the callback function model to output the corresponding external data call instruction.
[0040] Among them, the external data call instruction is an instruction that can be understood by the computer. The computer can obtain the required data from the location corresponding to the external data call instruction according to the external data call instruction. The callback function (FunctionCall) model can generate an external data call instruction to call an external API or function during the conversation to obtain specific information or perform operations. This mechanism converts the request into a standard function call format, allowing the model to automatically complete tasks or answer questions that require real-time data. Optionally, the callback function model can be constructed based on llama3.1-Instruct (Instruct-tuned models, that is, models that have been fine-tuned with specific task instructions).
[0041] In one embodiment, the callback function model provided by the present solution receives training data with task instructions as input during the training process, and gives appropriate responses based on these instructions. In the callback function model, the input can describe the problem or task in a clearer way, for example: "Please summarize this text" or "Explain the meaning of X". Compared with the traditional language model, the callback function model can better understand and follow the user's instructions, give answers that better meet the user's needs, and more accurately map natural language to a specific function call, which is then handed over to the live broadcast assistant for processing.
[0042] S140: Obtain external data from a preset external system according to an external data call instruction, and send the external data to the live broadcast assistant model, analyze and process the video text description data, audio text description data, live broadcast text data and external data through the live broadcast assistant model to obtain live broadcast interaction information, and perform live broadcast interaction in the live broadcast room according to the live broadcast interaction information.
[0043] Exemplarily, after receiving the external data call instruction, the external data can be obtained from the preset external system pointed to by the external data call instruction according to the external data call instruction, and the external data can be sent to the live broadcast assistant model. After receiving the external data, the live broadcast assistant model can analyze and process the video text description data, audio text description data, live broadcast text data and external data, and output live broadcast interaction information in the form of natural language description, and conduct live broadcast interaction in the live broadcast room according to the live broadcast interaction information.
[0044] Among them, the live broadcast interaction in the live broadcast room according to the live broadcast interaction information can be that the live broadcast assistant model displays the content related to the live broadcast interaction information to the anchor (for example, when the live broadcast interaction information is a live broadcast suggestion to the anchor), or it can be sent in the chat window of the live broadcast room. Content related to the live broadcast interaction information (for example, when the live broadcast interaction information is a topic guide content for the live broadcast room). The live broadcast assistant model can speak in the live broadcast room according to the live broadcast interaction information to interact with the anchor or the audience, show suggestions for live broadcast interaction to the anchor, etc. For example, when the video text description data, audio text description data, and live broadcast text data involve a place, the live broadcast interaction information can be a live broadcast suggestion to the anchor: "The live broadcast room talked about XX place, why not chat with everyone about the food in this place", at this time, the live broadcast assistant model can show the anchor a chat bubble of "The live broadcast room talked about a certain place, why not chat with everyone about the food in this place", and guide the anchor and the audience in the live broadcast room to talk about topics that are more likely to arouse interest; the live broadcast interaction information can also be the topic guidance content of the live broadcast room: "Chat with the audience in the live broadcast room about the food in XX place", at this time, the live broadcast assistant model can input "Wow, I have heard of XX place, is there anything delicious in this place", and guide the audience in the live broadcast room and the anchor to participate in topics that are more likely to arouse interest. Among them, the live broadcast interaction information refers to the video text description data, audio text description data, and live broadcast text data of the live broadcast room, as well as the external data of the external system. The live broadcast interaction information is more suitable for the current interactive situation of the live broadcast room, and the information reflected by the live broadcast interaction information is more intuitive and comprehensive.
[0045] In the above, by obtaining the video text description data, audio text description data and live text data of the live broadcast room, the video text description data, audio text description data and live text data are sent to the trained live broadcast assistant model, the video text description data, audio text description data and live broadcast text data are analyzed and processed by the live broadcast assistant model to obtain the live broadcast data demand instruction, and the live broadcast data demand instruction is sent to the trained callback function model, and the live broadcast data demand instruction is analyzed and processed by the callback function model to obtain the external data call instruction, and the external data is obtained from the preset external system according to the external data call instruction, and the external data is sent to the live broadcast assistant model, and the video text description data, audio text description data, live broadcast text data and external data are analyzed and processed by the live broadcast assistant model to obtain the live broadcast interaction information, and the live broadcast interaction is performed in the live broadcast room according to the live broadcast interaction information, and the live broadcast interaction is managed in the live broadcast room according to the video, voice, text data and external data of the live broadcast room, which can effectively improve the accuracy and availability of the information feedback by the live broadcast assistant, help the anchor to interact with the live broadcast, and improve the interactive effect of the video live broadcast.
[0046] Based on the above embodiments, Figure 2A flowchart of another live interactive management method provided in an embodiment of the present application is given, which is a specific implementation of the above live interactive management method. Figure 2 , the live interactive management method includes:
[0047] S210: Acquire live video data, live audio data and live text data of the live broadcast room, and determine video text description data according to the live video data, and determine audio text description data according to the live audio data.
[0048] S220: Send the video text description data, audio text description data and live broadcast text data to the trained live broadcast assistant model, and analyze and process the video text description data, audio text description data and live broadcast text data through the live broadcast assistant model to obtain the live broadcast data demand instruction.
[0049] S230: Send the live broadcast data demand instruction to the trained callback function model, analyze and process the live broadcast data demand instruction through the callback function model, and obtain the external data call instruction.
[0050] S240: Obtain external data from a preset external system according to an external data call instruction, and send the external data to the live broadcast assistant model, analyze and process the video text description data, audio text description data, live broadcast text data and external data through the live broadcast assistant model to obtain live broadcast interaction information, and perform live broadcast interaction in the live broadcast room according to the live broadcast interaction information.
[0051] S250: Sending live broadcast interaction information to the anchor terminal corresponding to the live broadcast room, so that the anchor terminal can display the live broadcast interaction information.
[0052] Exemplarily, after the live broadcast assistant model analyzes and processes the video text description data, audio text description data, live broadcast text data and external data to obtain the live broadcast interaction information, the live broadcast interaction information can be sent to the anchor end corresponding to the live broadcast room, so that the anchor end can display the live broadcast interaction information. The anchor can interact with the audience according to the live broadcast interaction suggestions corresponding to the live broadcast interaction information, thereby improving the level of interaction in the live broadcast room.
[0053] For the anchor, the live broadcast assistant model can analyze the feedback of the audience in the current live broadcast room on the live broadcast content (such as whether it is interesting, what they want to discuss, and what is the current focus of discussion) by analyzing the information flow (including video text description data, audio text description data, live broadcast text data and external data). Audiences in different regions and different time periods prefer different live broadcast styles.
[0054] The live assistant model can analyze indicators such as the user ratio, audience preferences, and fan attributes in the current live broadcast room, analyze the audience participation, emotions in the live broadcast room, the proportion of interactive behaviors, etc., and output live broadcast interaction information, so that the anchor can specifically fill in the lack of interactive points in the live broadcast room and enrich the atmosphere of the live broadcast room. The live assistant model can also help anchors expand more live broadcast topics. For example, it can generate deeper discussion topics based on the audience's speeches, or extract hot events of the day from the Internet, terminal forums, and terminal hot searches to extend the live broadcast content, and propose new discussion topics in the live broadcast room (for example, show new discussion topics to the anchor or enter new discussion topics in the live chat box (public screen)), further improving the participation of live broadcast room topics.
[0055] In one embodiment, each user can set his or her own system language (for example, determined according to the language code of the terminal device, or determined according to the language manually set by the user). When the audience enters the live broadcast room, the audience subscribes to the language broadcast channel of the live broadcast room (for example, the Chinese broadcast channel: live broadcast room id11111_Chinese, the English broadcast channel: live broadcast room id11111_English). After obtaining the video text description data, audio text description data and live text data in the live broadcast data stream, the video text description data, audio text description data and live text data are converted into multiple languages in the live broadcast room respectively, and then sent to the broadcast channels of different languages in the live broadcast room to provide the live broadcast content to the visually impaired (video content recognition and translation), the hearing impaired (audio content recognition and translation) and cross-regional users (public screen content translation), thereby improving the usability of the product for different types of users. Among them, the translation model can use the built-in translation of the live broadcast assistant, or a third-party translation service.
[0056] In a possible embodiment, the live interactive management method provided by the present solution further includes, after sending the video text description data, the audio text description data and the live text data to the trained live assistant model:
[0057] S260: Analyze and process the video text description data, the audio text description data, and the live text data through the live assistant model to obtain live text description information.
[0058] S270: Determine the live broadcast voice description information according to the live broadcast text description information, and send the live broadcast voice description information to the first preset user terminal so that the first preset user terminal plays the live broadcast voice description information.
[0059] Exemplarily, after receiving the video text description data, audio text description data and live text data, the live assistant model can generate live text description information based on the video text description data, audio text description data and live text data. For example, send the instruction "describe what happened in the live broadcast room based on the video text description data, audio text description data and live text data" to the live assistant model, and the live assistant model will output the live text description information in the form of natural language description. For example, the live text description information can be: "This is a description of a photo. The subject of the photo is a person sitting in the front row and wearing headphones, holding a microphone in his hand and looking at the camera", "This picture describes a girl with long blonde hair. She is sitting on a chair with her hand on her chin, looking at the camera. This is a very calm and thoughtful scene". Optionally, the live text description information can also be translated into live text description information corresponding to different languages according to the languages subscribed by the audience of the live broadcast room.
[0060] In one embodiment, after obtaining the live text description information, the live text description information can be voice-to-text converted to obtain the corresponding voice information (i.e., live voice description information), and can also be converted into live voice description information in the corresponding language according to the language corresponding to the live voice description information. Optionally, the live voice description information can be sent to the first preset user terminal (for example, the live voice description information in the corresponding language can be sent to the audience terminal) for the first preset user terminal to play the live voice description information, so that the user can understand what is happening in the live broadcast room by listening to the voice corresponding to the live voice description information, and achieve barrier-free interaction with the live broadcast room. Visually impaired users can also understand the content of the current live broadcast room communication in real time, thereby improving the user experience. In one embodiment, the live assistant model provided by this solution can provide different functions (such as real-time subtitle recognition and translation, live broadcast effect evaluation, real-time story creation, audience atmosphere group, diversified live content generation, live broadcast room order management, etc.) based on different fine-tuned large language models, enriching the interactive management function of the live broadcast room.
[0061] In one embodiment, the video text description data, audio text description data, live text data, etc. provided by the live assistant model can be translated and processed through the live studio multi-language stream subscription system and sent to the audience for display. After the user enters the live studio, the desired language set by the user (i.e., the preset language) is sent to the preset live studio multi-language stream subscription system. If there is no translation stream of the user's language for the live studio in the live studio multi-language stream subscription system, a new translation workflow is created to translate the text messages of the live studio in real time; if there is a translation workflow of the preset language (other users have triggered the translation before), no new translation workflow is created, and the user is then added to the translation stream corresponding to the language. At this time, the translation stream will continuously send translated subtitles to users of the same language in the live studio, so that users in different regions can also communicate in real time without obstacles.
[0062] In a possible embodiment, the live broadcast interaction management method provided by the present scheme can also analyze and process the video text description data, audio text description data and live broadcast text data through the live broadcast assistant model after sending the video text description data, audio text description data and live broadcast text data to the trained live broadcast assistant model, so as to obtain subtitle text information in a preset language, and send the subtitle text information to a second preset user terminal so that the second preset user terminal can display the subtitle text information.
[0063] Exemplarily, after obtaining the video text description data, audio text description data and live text data, the live assistant model can generate subtitle text information reflecting the content of the live broadcast room communication based on the video text description data, audio text description data and live text data based on one or more preset languages subscribed by the user, and send the subtitle text information to a second preset user terminal (such as an audience terminal) so that the second preset user terminal can display the subtitle text information, so that the user can understand what is happening in the live broadcast room by viewing the subtitle text information, and realize barrier-free interaction with the live broadcast room. Hearing-impaired users can also understand the content of the current live broadcast room communication in real time, thereby improving the user experience and allowing users in different language regions to communicate in real time without barriers.
[0064] In a possible embodiment, the live interactive management method provided by the present solution can, after sending the video text description data, audio text description data and live text data to the trained live assistant model, also analyze and process the video text description data, audio text description data and live text data through the live assistant model to obtain story recording data when the real-time story creation function is enabled.
[0065] Among them, the real-time story creation function can be understood as the function of creating stories based on the interactive content of the live broadcast room (such as the actions, sounds, texts generated by the anchor in the live broadcast room, and the comments of the audience in the live broadcast room, etc.). The real-time story creation function can be turned on by the anchor.
[0066] Exemplarily, when the real-time story creation function is enabled, the video text description data, audio text description data and live text data generated by the device that enables the real-time story creation function are collected. When the real-time story creation function is turned off, or the host determines that the real-time story creation is completed, the collected video text description data, audio text description data and live text data are sent to the live assistant model, and instructions for creating a corresponding story based on the provided video text description data, audio text description data and live text data are input into the live assistant model. The video text description data, audio text description data and live text data are analyzed and processed by the live assistant model to obtain story recording data.
[0067] Optionally, the story record data can be recorded in a combination of one or more of text, pictures and videos. The story record data can be sent to the anchor, or a post can be generated based on the story record data and posted in the live broadcast post bar to divert traffic to the anchor's live broadcast room. For example, the anchor can activate the real-time story creation function of the live broadcast assistant model through preset instructions. Then, the anchor and the audience in the live broadcast room can use the instructions together to create a special story belonging to the live broadcast room. The direction of the plot can be controlled by the live broadcast room, thereby improving the interactivity between the anchor and the audience. After the live broadcast, the generated story can be exported as story record data, and then the story record data can be sent to the end post bar to achieve a richer live interactive management effect.
[0068] In one embodiment, when it is detected that the atmosphere of the live broadcast room has decreased (the heat value of the live broadcast room is less than the preset heat threshold), the live broadcast assistant model can construct dynamic speeches based on the most recent live broadcast stream data, and act as a chat robot to guide the topic in the live broadcast room and liven up the atmosphere. The live broadcast room module regularly makes judgments based on the heat value of the live broadcast room. When the heat value of a live broadcast room is continuously less than a certain threshold preset time length, the chat robot audience atmosphere group is triggered, and the public screen is automatically sent in the live broadcast room in combination with the historical information of the live broadcast room to enhance the interactive experience of the anchor. After the introduction of the speech-to-text model, the speeches of the anchor and the guests can also be passed into the live broadcast assistant model, realizing seamless chat between the anchor and the atmosphere group.
[0069] In a possible embodiment, the live broadcast interaction management method provided by the present scheme can also analyze and process the video text description data, audio text description data and live broadcast text data through the live broadcast assistant model after sending the video text description data, audio text description data and live broadcast text data to the trained live broadcast assistant model to obtain the live broadcast room heat value, and when the live broadcast room heat value is less than a preset heat threshold, determine the interactive text information based on the video text description data, audio text description data and live broadcast text data, and display the interactive text information in the chat window of the live broadcast room and / or the anchor end.
[0070] Exemplarily, the video text description data, audio text description data and live text data collected within a certain period of time are sent to the live assistant model, and the live assistant model is required to analyze and process these video text description data, audio text description data and live text data to obtain the live room heat value. After the live assistant model determines the live room heat value, if the live room heat value is less than the preset heat threshold, an instruction to provide anchor interaction suggestions based on the above information can be input to the live assistant model, so that the live assistant model analyzes and processes the above-provided video text description data, audio text description data and live text data and outputs interactive text information, and displays the interactive text information in the chat window of the live room or the anchor end, and the audience and / or the anchor can interact based on the interactive text information, mobilize the initiative of the anchor and the audience to interact, and increase the heat of the live room.
[0071] Optionally, the live data stream of the live broadcast room can be received in real time and the preset order interface (order API) can be used to judge the violation of the live data stream. If malicious behavior is judged in the live data stream, the anchor or administrator will be notified. If the anchor (administrator) agrees to the disposal, the live order behavior (kicking people, banning) will be executed. When the popularity and traffic of the live broadcast room are high, the anchor's control over the live broadcast room is improved. For the content security of the live broadcast room, another dimension of security is added in addition to manual review. All interactions in the live broadcast room can first be judged by the original preset order interface in the terminal. If there is a strong violation, it can be directly intercepted and processed. If the data is normal or not a strong violation, it is added to the cache. When the cached data volume reaches the preset threshold, the cached message can be pushed to the live broadcast assistant model for judgment. The live broadcast assistant model can judge marginal violations, thereby further maintaining the healthy environment of the live broadcast room, helping to increase the live broadcast time of the anchor and protecting the viewing experience of other normal viewers.
[0072] In a possible embodiment, the live interactive management method provided by the present solution further includes, after sending the video text description data, the audio text description data and the live text data to the trained live assistant model:
[0073] S280: Perform violation judgment on the live broadcast room data according to the preset order interface, and determine the live broadcast room data that is judged to be non-violated as data to be judged according to the violation judgment result, and add the data to be judged to the violation judgment cache. The live broadcast room data includes video text description data, audio text description data and live broadcast text data.
[0074] S290: When the amount of data in the violation judgment cache reaches a preset data amount threshold, the data to be judged in the violation judgment cache is sent to the live broadcast assistant model, and the live broadcast assistant model performs edge violation judgment on the data to be judged to obtain an edge violation judgment result.
[0075] Exemplarily, video text description data, audio text description data, and live text data are collected in real time, and the video text description data, audio text description data, and live text data are sent as live room data to a preset order management system through a preset order interface. The live room data is judged for violations (e.g., manual review and / or machine review) through the preset order management system to obtain violation judgment results for each data in the live room data. In one embodiment, for live room data with violation judgment results of violations (strong violations) (e.g., insults, offensive actions, voices, texts, etc.), the violating viewers or anchors can be processed according to the preset violation processing method (e.g., kicking, silencing, banning, etc.).
[0076] In one embodiment, for the live broadcast room data whose violation judgment result is no violation, these live broadcast room data that are not in violation can be used as data to be judged, and the data to be judged is added to the violation judgment cache. When the amount of data in the violation judgment cache reaches the preset data amount threshold, the data to be judged in the violation judgment cache is taken out and sent to the live broadcast assistant model, and the live broadcast assistant model is sent to judge whether there is a marginal violation in the data to be judged, so as to make a marginal violation judgment on the data to be judged through the live broadcast assistant model, and obtain the marginal violation judgment result. Among them, marginal violation data can be understood as non-strong violation data, and its maliciousness is lower than that of strong violation data, but there are malicious actions (such as communicating with negative words). For example, the chat text corresponding to the audience in the live broadcast text data "The anchor is so annoying, too much nonsense", which contains dissatisfaction and malicious emotions towards the anchor, and the live broadcast text data can be determined as marginal violation live broadcast text data.
[0077] For data to be judged with a marginal violation judgment result of marginal violation, the marginally violating audience or anchor can be processed according to the preset violation processing method (such as kicking, silencing, banning, etc.), improving the comprehensiveness of the violation review of the live broadcast room and effectively maintaining a healthy environment for the live broadcast room. This solution determines the live broadcast room data that is judged to be non-violated as data to be judged based on the violation judgment result of the live broadcast room data and adds it to the violation judgment cache. When the amount of data in the violation judgment cache reaches the preset data amount threshold, the data to be judged in the violation judgment cache is sent to the live broadcast assistant model for marginal violation judgment, improving the comprehensiveness of the violation inspection of the relevant data of the live broadcast room and improving the interactive management effect of the live broadcast room.
[0078] In the above, by obtaining the video text description data, audio text description data and live text data of the live broadcast room, the video text description data, audio text description data and live text data are sent to the trained live broadcast assistant model, the video text description data, audio text description data and live broadcast text data are analyzed and processed by the live broadcast assistant model to obtain the live broadcast data demand instruction, and the live broadcast data demand instruction is sent to the trained callback function model, and the live broadcast data demand instruction is analyzed and processed by the callback function model to obtain the external data call instruction, and the external data is obtained from the preset external system according to the external data call instruction, and the external data is sent to the live broadcast assistant model, and the video text description data, audio text description data, live broadcast text data and external data are analyzed and processed by the live broadcast assistant model to obtain the live broadcast interaction information, and the live broadcast interaction is performed in the live broadcast room according to the live broadcast interaction information, and the live broadcast interaction is managed in the live broadcast room according to the video, voice, text data and external data of the live broadcast room, which can effectively improve the accuracy and availability of the information feedback by the live broadcast assistant, help the anchor to interact with the live broadcast, and improve the interactive effect of the video live broadcast. At the same time, live broadcast interaction information is sent to the anchor end corresponding to the live broadcast room, so that the anchor end can display the live broadcast interaction information. The anchor can interact with the audience according to the live broadcast interaction suggestions corresponding to the live broadcast interaction information, thereby increasing the level of interaction in the live broadcast room, improving the interactivity of the anchor's live broadcast room, and improving the interactive stickiness of users within the end.
[0079] Figure 3 Schematic diagram of a live interactive management device provided by an embodiment of the present application. Figure 3 The live interactive management device includes a data acquisition module 31, a data analysis module 32, an instruction determination module 33 and an interactive management module 34.
[0080] Among them, the data acquisition module 31 is configured to acquire the live video data, live audio data and live text data of the live broadcast room, and determine the video text description data according to the live video data, and determine the audio text description data according to the live audio data; the data analysis module 32 is configured to send the video text description data, audio text description data and live text data to the trained live broadcast assistant model, analyze and process the video text description data, audio text description data and live text data through the live broadcast assistant model, and obtain the live data demand instruction; the instruction determination module 33 is configured to send the live broadcast data demand instruction to the trained callback function model, analyze and process the live broadcast data demand instruction through the callback function model, and obtain the external data call instruction; the interaction management module 34 is configured to obtain external data from a preset external system according to the external data call instruction, and send the external data to the live broadcast assistant model, analyze and process the video text description data, audio text description data, live text data and external data through the live broadcast assistant model, obtain the live broadcast interaction information, and perform live broadcast interaction in the live broadcast room according to the live broadcast interaction information.
[0081] In the above, by obtaining the video text description data, audio text description data and live text data of the live broadcast room, the video text description data, audio text description data and live text data are sent to the trained live broadcast assistant model, the video text description data, audio text description data and live broadcast text data are analyzed and processed by the live broadcast assistant model to obtain the live broadcast data demand instruction, and the live broadcast data demand instruction is sent to the trained callback function model, and the live broadcast data demand instruction is analyzed and processed by the callback function model to obtain the external data call instruction, and the external data is obtained from the preset external system according to the external data call instruction, and the external data is sent to the live broadcast assistant model, and the video text description data, audio text description data, live broadcast text data and external data are analyzed and processed by the live broadcast assistant model to obtain the live broadcast interaction information, and the live broadcast interaction is performed in the live broadcast room according to the live broadcast interaction information, and the live broadcast interaction is managed in the live broadcast room according to the video, voice, text data and external data of the live broadcast room, which can effectively improve the accuracy and availability of the information feedback by the live broadcast assistant, help the anchor to interact with the live broadcast, and improve the interactive effect of the video live broadcast.
[0082] In a possible embodiment, the live broadcast interaction management device also includes an information push module, and the information push module is configured to: send live broadcast interaction information to the anchor end corresponding to the live broadcast room, so that the anchor end can display the live broadcast interaction information.
[0083] In a possible embodiment, the live interactive management device further includes a voice feedback module, and the voice feedback module is configured as follows:
[0084] The video text description data, the audio text description data and the live text data are analyzed and processed by the live assistant model to obtain the live text description information;
[0085] The live broadcast voice description information is determined according to the live broadcast text description information, and the live broadcast voice description information is sent to the first preset user terminal so that the first preset user terminal plays the live broadcast voice description information.
[0086] In a possible embodiment, the live interactive management device further includes a subtitle feedback module, and the subtitle feedback module is configured as follows:
[0087] The video text description data, audio text description data and live text data are analyzed and processed through the live assistant model to obtain subtitle text information in a preset language, and the subtitle text information is sent to a second preset user terminal so that the second preset user terminal can display the subtitle text information.
[0088] In a possible embodiment, the live interactive management device further includes a story creation module, and the story creation module is configured as follows:
[0089] When the real-time story creation function is enabled, the video text description data, audio text description data and live text data are analyzed and processed through the live assistant model to obtain story recording data.
[0090] In a possible embodiment, the live broadcast interaction management device further includes an interaction feedback module, and the interaction feedback module is configured as follows:
[0091] The video text description data, audio text description data and live text data are analyzed and processed through the live assistant model to obtain the live room heat value. When the live room heat value is less than a preset heat threshold, the interactive text information is determined based on the video text description data, audio text description data and live text data, and the interactive text information is displayed in the chat window of the live room and / or the anchor end.
[0092] In a possible embodiment, the live broadcast interaction management device further includes a violation feedback module, and the violation feedback module is configured as follows:
[0093] The live broadcast room data is judged as being in violation of a rule according to a preset order interface, and the live broadcast room data that is judged as not in violation of a rule is determined as data to be judged according to the violation judgment result, and the data to be judged is added to the violation judgment cache, wherein the live broadcast room data includes video text description data, audio text description data, and live broadcast text data;
[0094] When the amount of data in the violation judgment cache reaches a preset data amount threshold, the data to be judged in the violation judgment cache is sent to the live broadcast assistant model, and the live broadcast assistant model performs edge violation judgment on the data to be judged to obtain an edge violation judgment result.
[0095] It is worth noting that in the embodiment of the above-mentioned live interactive management device, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present application.
[0096] The embodiment of the present application also provides a live broadcast interaction management device, which can integrate the live broadcast interaction management device provided by the embodiment of the present application. Figure 4 Schematic diagram of a live interactive management device provided by an embodiment of the present application. Figure 4 The live interactive management device includes: an input device 43, an output device 44, a memory 42 and one or more processors 41; the memory 42 is used to store one or more programs; when one or more programs are executed by one or more processors 41, the one or more processors 41 implement the live interactive management method provided in the above embodiment. The live interactive management device, equipment and computer provided above can be used to execute the live interactive management method provided in any of the above embodiments, and have corresponding functions and beneficial effects.
[0097] The embodiments of the present application also provide a non-volatile storage medium storing computer executable instructions, which are used to execute the live interactive management method provided in the above embodiments when executed by a computer processor. Of course, the non-volatile storage medium storing computer executable instructions provided in the embodiments of the present application, whose computer executable instructions are not limited to the live interactive management method provided above, can also execute the related operations in the live interactive management method provided in any embodiment of the present application. The live interactive management apparatus, equipment and storage medium provided in the above embodiments can execute the live interactive management method provided in any embodiment of the present application. For technical details not described in detail in the above embodiments, please refer to the live interactive management method provided in any embodiment of the present application.
[0098] On the basis of the above embodiments, the embodiments of the present application also provide a computer program product. The essence of the technical solution of the present application or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer program product is stored in a storage medium, including a number of instructions for enabling a computer device, a mobile terminal or a processor therein to execute all or part of the steps of the live interactive management method provided in each embodiment of the present application.
Claims
1. A live interactive management method, characterized in that: include: Acquire live video data, live audio data and live text data of the live broadcast room, and determine video text description data according to the live video data, and determine audio text description data according to the live audio data; The video text description data, the audio text description data and the live text data are sent to the trained live assistant model, and the video text description data, the audio text description data and the live text data are analyzed and processed by the live assistant model to obtain a live data demand instruction; Sending the live broadcast data demand instruction to the trained callback function model, analyzing and processing the live broadcast data demand instruction through the callback function model to obtain an external data call instruction; According to the external data calling instruction, external data is obtained from a preset external system, and the external data is sent to the live broadcast assistant model. The video text description data, the audio text description data, the live broadcast text data and the external data are analyzed and processed by the live broadcast assistant model to obtain live broadcast interaction information, and live broadcast interaction is performed in the live broadcast room according to the live broadcast interaction information.
2. The live interactive management method according to claim 1, characterized in that: After analyzing and processing the video text description data, the audio text description data, the live text data and the external data through the live assistant model to obtain the live interactive information, the method further includes: The live broadcast interaction information is sent to the anchor terminal corresponding to the live broadcast room, so that the anchor terminal displays the live broadcast interaction information.
3. The live interactive management method according to claim 1, characterized in that: After sending the video text description data, the audio text description data, and the live text data to the trained live assistant model, the method further includes: The video text description data, the audio text description data and the live text data are analyzed and processed by the live assistant model to obtain live text description information; The live broadcast voice description information is determined according to the live broadcast text description information, and the live broadcast voice description information is sent to a first preset user terminal so that the first preset user terminal plays the live broadcast voice description information.
4. The live interactive management method according to claim 1, characterized in that: After sending the video text description data, the audio text description data, and the live text data to the trained live assistant model, the method further includes: The video text description data, the audio text description data and the live text data are analyzed and processed by the live assistant model to obtain subtitle text information in a preset language, and the subtitle text information is sent to a second preset user terminal so that the second preset user terminal can display the subtitle text information.
5. The live interactive management method according to claim 1, characterized in that: After sending the video text description data, the audio text description data, and the live text data to the trained live assistant model, the method further includes: When the real-time story creation function is enabled, the video text description data, the audio text description data and the live text data are analyzed and processed by the live assistant model to obtain story recording data.
6. The live interactive management method according to claim 1, characterized in that: After sending the video text description data, the audio text description data, and the live text data to the trained live assistant model, the method further includes: The video text description data, the audio text description data and the live text data are analyzed and processed by the live assistant model to obtain the live room heat value, and when the live room heat value is less than a preset heat threshold, the interactive text information is determined according to the video text description data, the audio text description data and the live text data, and the interactive text information is displayed in the chat window of the live room and / or the anchor end.
7. The live interactive management method according to claim 1, characterized in that: After sending the video text description data, the audio text description data, and the live text data to the trained live assistant model, the method further includes: Performing violation judgment on the live broadcast room data according to a preset order interface, and determining the live broadcast room data that is determined to be non-violated as data to be judged according to the violation judgment result, and adding the data to be judged to the violation judgment cache, wherein the live broadcast room data includes the video text description data, the audio text description data and the live broadcast text data; When the amount of data in the violation judgment cache reaches a preset data amount threshold, the data to be judged in the violation judgment cache is sent to the live broadcast assistant model, and the live broadcast assistant model performs edge violation judgment on the data to be judged to obtain an edge violation judgment result.
8. A live interactive management device, characterized in that: It includes a data acquisition module, a data analysis module, an instruction determination module and an interaction management module, wherein: The data acquisition module is configured to acquire live video data, live audio data and live text data of the live broadcast room, and determine video text description data according to the live video data, and determine audio text description data according to the live audio data; The data analysis module is configured to send the video text description data, the audio text description data and the live text data to the trained live assistant model, and analyze and process the video text description data, the audio text description data and the live text data through the live assistant model to obtain a live data demand instruction; The instruction determination module is configured to send the live data demand instruction to the trained callback function model, analyze and process the live data demand instruction through the callback function model, and obtain an external data call instruction; The interaction management module is configured to obtain external data from a preset external system according to the external data call instruction, and send the external data to the live broadcast assistant model, analyze and process the video text description data, the audio text description data, the live broadcast text data and the external data through the live broadcast assistant model to obtain live broadcast interaction information, and perform live broadcast interaction in the live broadcast room according to the live broadcast interaction information.
9. A live interactive management device, characterized in that: include: memory and one or more processors; The memory is used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the live interactive management method as described in any one of claims 1-7.
10. A non-volatile storage medium storing computer executable instructions, characterized in that: The computer executable instructions, when executed by a computer processor, are used to execute the live interactive management method as described in any one of claims 1-7.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the live broadcast interactive management method described in any one of claims 1-7 is implemented.