Multi-modal live broadcast interaction system based on artificial intelligence

By designing a multimodal live broadcast interactive system based on artificial intelligence, the problem of difficulty in collecting and processing of multimodal data in existing systems is solved, and multimodal intelligent live broadcast and more diverse interaction methods are realized, which improves user experience and sense of participation.

CN120017873APending Publication Date: 2025-05-16合肥冬葵网络科技有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510218612.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing artificial intelligence live broadcast interactive system is difficult to collect and process multimodal data, and cannot provide targeted interactive suggestions, affecting the user experience.

Method used

Design a multimodal live broadcast interaction system based on artificial intelligence, including multimodal data acquisition module, multimodal fusion processing module, intelligent interaction response module, virtual image driver engine, module training and optimization module and storage module. Through deep learning models, feature extraction and semantic association of multiple modal data will be generated, unified semantic representations will be realized, and real-time interaction and virtual anchor rendering will be realized.

Benefits of technology

Multimodal intelligent live broadcast is realized, which can collect and process a variety of different forms of data, provide more diverse interactive channels, enhance the audience's sense of participation and immersion, and provide anchors with more targeted interactive suggestions to improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017873A_ABST
    Figure CN120017873A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to a multi-modal live broadcast interaction system based on artificial intelligence, comprising a multi-modal data acquisition module; a multi-modal fusion processing module; an intelligent interaction response module; a virtual image driving engine; a module training and optimizing module; and a storage module. According to the invention, through cooperative use of the multi-modal data acquisition module and the multi-modal fusion processing module, multi-modal intelligent live broadcast can be realized; through cooperative use of the intelligent interaction response module, the virtual image driving engine and the module training and optimization module, analysis and fusion of different data can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a multimodal live broadcast interactive system based on artificial intelligence. Background Art

[0002] In recent years, the live streaming industry has shown explosive growth and has become a highly influential Internet application model. From entertainment live streaming to e-commerce live streaming, from education live streaming to corporate conference live streaming, the application scenarios of live streaming have been continuously expanded and the user scale has continued to expand. Live streaming has become one of the important ways for people to obtain information, entertainment, business transactions and knowledge learning.

[0003] Most of the existing artificial intelligence live interactive systems have relatively single and limited data dimensions and interaction methods, which makes it difficult to collect and process diversified and multi-format data, and difficult to conduct multi-modal live interaction. At the same time, some artificial intelligence live interactive systems find it difficult to accurately analyze and integrate different data, and are not convenient to provide targeted interaction suggestions, which affects the user experience.

[0004] To this end, we proposed an artificial intelligence-based multimodal live interactive system to solve the above problems. Summary of the invention

[0005] The purpose of the present invention is to provide a multimodal live interactive system based on artificial intelligence to solve the problems raised in the above-mentioned background technology.

[0006] To achieve the above object, the present invention provides the following technical solution: a multimodal live interactive system based on artificial intelligence, comprising a multimodal data acquisition module, The multimodal data acquisition module is used to collect multiple modal data during the live broadcast process, including video image data, audio data, and user interaction behavior data on the live broadcast platform, and can collect voice signals, visual images, text input, and user biometric data in the live broadcast scene in real time; A multimodal fusion processing module, which extracts features and associates semantics of collected data through a deep learning model to generate a unified semantic representation, and can pre-process, extract features and analyze the collected multimodal data to identify key information and user intent in live content; An intelligent interactive response module, which dynamically generates voice, text and avatar feedback based on semantic analysis results, interactively generates corresponding interactive instructions according to the analysis structure and data processing results, and optimizes the response strategy through reinforcement learning, pushes interactive content to users on the live broadcast platform, and realizes real-time interaction with users; A virtual image driving engine, which renders a 3D virtual anchor in real time according to the interactive content, and synchronizes lip movements, facial expressions and body movements; A module training and optimization module, which is used to collect and organize historical live broadcast data, train and optimize artificial intelligence models, so as to improve the system's adaptability and accuracy to different live broadcast scenarios and user needs; A storage module is used to store multimodal data generated during the live broadcast, processed data, and trained artificial intelligence models.

[0007] Preferably, the multimodal data acquisition module includes: A video image acquisition unit, used to obtain a video image frame sequence from a live stream; Audio acquisition unit, used to record audio information in live broadcast, and extract and collect broken jade video and audio at the same time.

[0008] Preferably, the multimodal data acquisition module is equipped with an acoustic acquisition unit of a ring array microphone, which supports sound source localization and noise suppression. The multimodal data acquisition module is equipped with a multispectral camera group, integrating RGB, depth and infrared sensors, so as to improve the acquisition efficiency of the video image acquisition unit and the audio acquisition unit.

[0009] Preferably, the multimodal fusion processing module is connected to the multimodal data acquisition module, and the multimodal fusion processing module includes: Perform pre-processing operations such as denoising, cropping, and normalization on the collected video image data, and perform denoising and feature extraction on the audio data; Multimodal feature fusion unit, which fuses the pre-processed video image features and audio features to form a comprehensive feature vector; Intent recognition unit, based on the fused feature vector, uses a pre-trained machine learning model to recognize the user's intent, such as asking questions, seeking advice, expressing emotions, etc. Hierarchical decision-making network, which processes instant interaction needs and long-term user portrait analysis respectively, makes the multimodal live interactive system more suitable for different usage needs and can provide targeted and personalized choices. For example, viewers can participate in live interaction by showing their gestures through the camera, and the system can recognize these gestures and give corresponding feedback; the host's facial expression changes can also be captured and analyzed by the system to better adjust the live atmosphere.

[0010] Preferably, the intelligent interactive response module is connected to the multimodal fusion processing module, and the intelligent interactive response module is used to monitor the user's operating behaviors on the live broadcast platform, such as likes, comments, shares, and attention, and record these behavioral data. The intelligent interactive response module includes: Generate targeted text responses, voice prompts or video clips as interactive content based on intent recognition results and preset interaction strategies; The interactive push unit sends the generated interactive content to the corresponding users through the message push interface of the live broadcast platform to ensure that the users can receive the interactive information in time to improve the accuracy and naturalness of the interaction.

[0011] Preferably, the virtual image driving engine includes: · Parametric facial binding system, supporting mixed control of multiple basic expressions; Physical simulation skeleton driving algorithm to achieve natural limb movement transition; Real-time speech lip-syncing neural network, audio to lip-sync animation delay is less than 80ms.

[0012] Preferably, the module training and optimization module includes: Data annotation unit, which manually annotates historical live broadcast data to mark key information, user intent, and corresponding interaction results; Model training unit, which uses labeled data to train deep learning models, such as recurrent neural networks, convolutional neural networks, and long short-term memory networks. Model evaluation and optimization unit, which regularly evaluates the trained model and adjusts the model parameters according to the evaluation indicators to improve the performance and accuracy of the model.

[0013] Preferably, the storage module adopts a distributed storage architecture, including multiple storage nodes, for storing different types of data respectively, thereby improving data storage efficiency and security.

[0014] The present invention provides a multi-modal live interactive system based on artificial intelligence, which has the following beneficial effects: Through the coordinated use of the multimodal data acquisition module and the multimodal fusion processing module, the purpose of multimodal intelligent live broadcast can be achieved. It can collect and process data in various forms, covering multiple dimensions such as vision, hearing, and touch. In addition to traditional text and voice interaction, it can also process images, videos, gestures, expressions and other information, and provide more diverse interaction channels. In addition to basic text and voice communication, it also supports image-based interaction, which can realize real-time action interaction and enhance the audience's sense of participation and immersion.

[0015] Through the coordinated use of the intelligent interactive response module, the virtual image driving engine and the module training and optimization module, the purpose of analyzing and integrating different data can be achieved, and richer and more accurate information can be excavated. Combined with the audience's voice intonation, facial expressions, body movements and other information, the system can more accurately judge the audience's emotions, interests and participation. After the judgment, the system can integrate this information to more comprehensively understand the audience's status and provide more targeted interactive suggestions for the anchor. Through rich interactive methods and precise intelligent analysis, it is more suitable for scenarios that require high interactivity and immersive experience, enhances the user's sense of participation and immersion, enables users to obtain a more real and vivid live broadcast experience, and enhances the emotional connection between users and anchors. BRIEF DESCRIPTION OF THE DRAWINGS In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0017] Figure 1 It is a schematic diagram of the structural module flow of the present invention.

[0018] In the figure: 1. Multimodal data acquisition module; 2. Multimodal fusion processing module; 3. Intelligent interactive response module; 4. Virtual image driving engine; 5. Module training and optimization module; 6. Storage module. DETAILED DESCRIPTION The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0020] The present invention proposes a multimodal live interactive system based on artificial intelligence, comprising a multimodal data acquisition module 1, The multimodal live interactive system based on artificial intelligence mainly includes a multimodal data acquisition module 1, a multimodal fusion processing module 2, an intelligent interactive response module 3, a virtual image driving engine 4, a module training and optimization module 5 and a storage module 6. These modules work together to realize multimodal interaction between users and the live system.

[0021] The system consists of a three-level architecture consisting of a front-end collection layer, an edge computing layer, and a cloud processing layer: Front-end layer: deploy multimodal sensor arrays and use hardware acceleration to implement data preprocessing; Edge layer: Run lightweight AI models to complete preliminary feature extraction and reduce cloud load; Cloud layer: performs complex model reasoning and big data analysis, and maintains user interaction memory.

[0022] The speech data in the present invention uses WaveNet to extract acoustic features, the visual data uses EfficientNet-V2 to extract spatiotemporal features, the text data is encoded by BERT, and the cross-modal association matrix is ​​constructed through the cross-attention mechanism, which is expressed as follows: Among them, Q, K, VQ, K, and V represent the query and key-value vectors of different modalities, respectively, to achieve semantic-level fusion.

[0023] Multimodal data acquisition module 1, which is used to collect multiple modal data during the live broadcast process, including video image data, audio data, and user interaction behavior data on the live broadcast platform, and can collect voice signals, visual images, text input, and user biometric data in the live broadcast scene in real time; Specifically, the multimodal data acquisition module 1 includes: A video image acquisition unit, used to obtain a video image frame sequence from a live stream; Audio acquisition unit, used to record audio information during live broadcast.

[0024] Specifically, the multimodal data acquisition module 1 is equipped with an acoustic acquisition unit with a circular array microphone, which supports sound source localization and noise suppression, and the multimodal data acquisition module 1 is equipped with a multispectral camera group, which integrates RGB, depth and infrared sensors; Specifically, the voice signal component uses a high-sensitivity microphone to ensure accurate capture of the user's voice input; the visual image component uses a high-definition camera to capture the user's gestures and facial expressions; the text input and user biometric data component receives the user's text input through devices such as keyboards or touch screens.

[0025] Multimodal fusion processing module 2: Multimodal fusion processing module 2 uses a deep learning model to extract features and semantically associate the collected data to generate a unified semantic representation. It can pre-process, extract features and analyze the collected multimodal data to identify key information and user intentions in the live content. Specifically, the multimodal fusion processing module 2 is connected to the multimodal data acquisition module 1, and the multimodal fusion processing module 2 includes: Perform pre-processing operations such as denoising, cropping, and normalization on the collected video image data, and perform denoising and feature extraction on the audio data; Multimodal feature fusion unit, which fuses the pre-processed video image features and audio features to form a comprehensive feature vector; Intent recognition unit, based on the fused feature vector, uses a pre-trained machine learning model to recognize the user's intent, such as asking questions, seeking advice, expressing emotions, etc. Hierarchical decision network, which handles immediate interaction needs and long-term user profile analysis respectively.

[0026] Thus, the multimodal fusion processing module 2 receives the information collected by the multimodal information collection module, and uses deep learning algorithms and natural language processing technology to understand and analyze it. The module can recognize the user's voice commands, gestures and text inputs, and infer the user's intentions based on context information. In addition, the multimodal fusion processing module 2 also includes an emotion analysis submodule for analyzing the user's emotional state in order to more accurately understand the user's interaction needs.

[0027] Intelligent interactive response module 3, which dynamically generates voice, text and virtual image feedback based on the semantic analysis results, interactively generates corresponding interactive instructions according to the analysis structure and data processing results, and optimizes the response strategy through reinforcement learning, pushes interactive content to users on the live broadcast platform, and realizes real-time interaction with users; Specifically, the intelligent interactive response module 3 is connected to the multimodal fusion processing module 2. The intelligent interactive response module 3 is used to monitor the user's operation behavior on the live broadcast platform, such as likes, comments, shares, and attention, and record these behavior data. The intelligent interactive response module 3 includes: Generate targeted text responses, voice prompts or video clips as interactive content based on intent recognition results and preset interaction strategies; ·Interaction push unit, which sends the generated interactive content to the corresponding users through the message push interface of the live broadcast platform to ensure that users can receive the interactive information in time; Therefore, the intelligent interactive response module 3 is used to monitor the user's operating behaviors on the live broadcast platform, such as likes, comments, shares, and follows, and record these behavioral data. The intelligent interactive response module 3 generates corresponding interactive processing instructions based on the analysis results of the multimodal fusion processing module 2. These instructions include updating the live broadcast screen, adjusting the live broadcast settings, playing audio content, etc. The instruction generation process fully considers the user's intentions and emotional state.

[0028] Avatar driving engine 4, which renders a 3D virtual anchor in real time according to the interactive content, and synchronizes lip movements, facial expressions and body movements; Specifically, the virtual image driving engine 4 includes: · Parametric facial binding system, supporting mixed control of multiple basic expressions; Physical simulation skeleton driving algorithm to achieve natural limb movement transition; Real-time speech lip-syncing neural network, audio to lip-sync animation delay is less than 80ms.

[0029] Therefore, the virtual image driving engine 4 promotes the application of the system in the present invention in live e-commerce, online education and telemedicine. The virtual anchor answers product inquiries in real time, supports gestures to guide users to pay attention to details, and simultaneously analyzes the patient's voice tremor and facial micro-expressions to assist in disease assessment.

[0030] Module training and optimization module 5: Module training and optimization module 5 is used to collect and organize historical live broadcast data, train and optimize artificial intelligence models, so as to improve the adaptability and accuracy of the system to different live broadcast scenarios and user needs; Specifically, module training and optimization module 5 includes: Data annotation unit, which manually annotates historical live broadcast data to mark key information, user intent, and corresponding interaction results; Model training unit, which uses labeled data to train deep learning models, such as recurrent neural networks, convolutional neural networks, and long short-term memory networks. Model evaluation and optimization unit: Regularly evaluate the trained model and adjust the model parameters according to the evaluation indicators.

[0031] Therefore, by comparing and optimizing the historical data of live broadcasts in different periods, a multimodal live broadcast system that is more adaptable to market demand is provided.

[0032] Storage module 6, which is used to store multimodal data generated during the live broadcast, processed data, and trained artificial intelligence models; Specifically, the storage module 6 adopts a distributed storage architecture, including multiple storage nodes, which are used to store different types of data respectively, so that the stored data can be retrieved more quickly when it is needed.

[0033] Working principle: The user opens the client of the live broadcast platform and enters a live broadcast room to watch the live broadcast. At this time, the multimodal data acquisition module 1 starts working. The video image acquisition unit and the audio acquisition unit respectively extract the image frame sequence and record the live audio data stream for the live video stream.

[0034] The collected data is transmitted to the multimodal fusion processing module 2, which first pre-processes the video image data, and then fuses the video image features and audio features to form a comprehensive feature vector. The machine learning model in the intention recognition unit is then used for analysis to accurately identify the user's intention.

[0035] After the intelligent interactive response module 3 receives the intention recognition result, the interactive content generation unit retrieves the relevant answer from the knowledge base according to the recognized user intention, and generates a detailed text as a targeted interactive content reply, and then sends the text reply push interface to the user in the form of a bullet screen. At the same time, the intelligent interactive response module 3 monitors the user's operation behavior on the live broadcast platform; The virtual image driving engine 4 renders according to the content of the intelligent interactive response module 3, generates corresponding facial expressions and images, and makes the interactive content more vivid; The module training and optimization module 5 collects and organizes historical live broadcast data, and continuously optimizes and trains it. The storage module 6 adopts a distributed storage architecture to classify and store different modules.

[0036] The above is the entire working principle of the present invention.

[0037] Finally, a few points should be explained: First, in the description of the present application, it should be noted that, unless otherwise specified and limited, the terms "installed", "connected" and "connected" should be understood in a broad sense, and can be mechanical or electrical connections, or internal connectivity between two components, or direct connections. "Up", "down", "left" and "right" are only used to indicate relative position relationships. When the absolute position of the described object changes, the relative position relationship may change; secondly, in the drawings of the embodiments disclosed in the present invention, only the structures involved in the embodiments disclosed in the present invention are involved. Other structures can refer to the general design. In the absence of conflict, the same embodiment and different embodiments of the present invention can be combined with each other; finally, the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A multimodal live interactive system based on artificial intelligence, comprising a multimodal data acquisition module (1), characterized in that: A multimodal data acquisition module (1), the multimodal data acquisition module (1) being used to acquire multiple modal data during a live broadcast process, including video image data, audio data, and interactive behavior data of users on a live broadcast platform, and being capable of acquiring voice signals, visual images, text input, and user biometric data in a live broadcast scene in real time; A multimodal fusion processing module (2), wherein the multimodal fusion processing module (2) performs feature extraction and semantic association on the collected data through a deep learning model to generate a unified semantic representation, and can perform preprocessing, feature extraction and analysis on the collected multimodal data to identify key information and user intentions in the live broadcast content; An intelligent interactive response module (3), the intelligent interactive response module (3) dynamically generates voice, text and virtual image feedback based on the semantic analysis results, interactively generates corresponding interactive instructions according to the analysis structure and data processing results, and optimizes the response strategy through reinforcement learning, pushes interactive content to the live broadcast platform to the user, and realizes real-time interaction with the user; A virtual image driving engine (4), wherein the virtual image driving engine (4) renders a 3D virtual anchor in real time according to the interactive content, and synchronizes lip movements, facial expressions and body movements; A module training and optimization module (5), wherein the module training and optimization module (5) is used to collect and organize historical live broadcast data, train and optimize the artificial intelligence model, so as to improve the adaptability and accuracy of the system to different live broadcast scenarios and user needs; A storage module (6), wherein the storage module (6) is used to store multimodal data generated during the live broadcast process, processed data, and a trained artificial intelligence model.

2. According to claim 1, a multi-modal live interactive system based on artificial intelligence is characterized by: The multimodal data acquisition module (1) comprises: A video image acquisition unit, used to obtain a video image frame sequence from a live stream; Audio acquisition unit, used to record audio information during live broadcast.

3. The multimodal live interactive system based on artificial intelligence according to claim 1, characterized in that: The multimodal data acquisition module (1) is equipped with an acoustic acquisition unit of a circular array microphone, supporting sound source localization and noise suppression. The multimodal data acquisition module (1) is equipped with a multispectral camera group, integrating RGB, depth and infrared sensors.

4. The multimodal live interactive system based on artificial intelligence according to claim 1, characterized in that: The multimodal fusion processing module (2) is connected to the multimodal data acquisition module (1), and the multimodal fusion processing module (2) comprises: Perform pre-processing operations such as denoising, cropping, and normalization on the collected video image data, and perform denoising and feature extraction on the audio data; Multimodal feature fusion unit, which fuses the pre-processed video image features and audio features to form a comprehensive feature vector; Intent recognition unit, based on the fused feature vector, uses a pre-trained machine learning model to recognize the user's intent, such as asking questions, seeking advice, expressing emotions, etc. Hierarchical decision network, which handles immediate interaction needs and long-term user profile analysis respectively.

5. The multimodal live interactive system based on artificial intelligence according to claim 1, characterized in that: The intelligent interactive response module (3) is connected to the multimodal fusion processing module (2), and the intelligent interactive response module (3) comprises: Generate targeted text responses, voice prompts or video clips as interactive content based on intent recognition results and preset interaction strategies; The interactive push unit sends the generated interactive content to the corresponding users through the message push interface of the live broadcast platform to ensure that the users can receive the interactive information in time.

6. The multimodal live interactive system based on artificial intelligence according to claim 1, characterized in that: The virtual image driving engine (4) comprises: · Parametric facial binding system, supporting mixed control of multiple basic expressions; Physical simulation skeleton driving algorithm to achieve natural limb movement transition; Real-time speech lip-syncing neural network, audio to lip-sync animation delay is less than 80ms.

7. The multimodal live interactive system based on artificial intelligence according to claim 1, characterized in that: The module training and optimization module (5) includes: Data annotation unit, which manually annotates historical live broadcast data to mark key information, user intent, and corresponding interaction results; Model training unit, which uses labeled data to train deep learning models, such as recurrent neural networks, convolutional neural networks, and long short-term memory networks. Model evaluation and optimization unit: Regularly evaluate the trained model and adjust the model parameters according to the evaluation indicators.

8. The multi-modal live interactive system based on artificial intelligence according to claim 1, characterized in that: The storage module (6) adopts a distributed storage architecture and includes multiple storage nodes for respectively storing different types of data.

Citation Information

Cited By

  • Interface interaction method and device, equipment and storage medium

    CN120812334A