A Smart Live Streaming Interaction Method and System Based on AI Virtual Human

CN120786085BActive Publication Date: 2026-08-14SHANGHAI CHEWEISHI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0003]然而,在实际的直播交互过程中,AI虚拟人所面临的挑战也日益严峻

Benefits of technology

[0005]本发明为解决上述技术问题,提出了一种基于AI虚拟人的智能直播交互方法及系统,以解决至少一个上述技术问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120786085B_ABST
    Figure CN120786085B_ABST
Patent Text Reader

Abstract

This invention relates to the field of live streaming interaction technology, and more particularly to an intelligent live streaming interaction method and system based on AI virtual humans. The method includes the following steps: identifying real-time live streaming interaction data streams, performing intelligent bullet screen filtering and multimodal user interaction perception, and constructing a multimodal user interaction perception map; extracting key bullet screens based on the multimodal interaction perception map and predicting user needs to generate user interaction need features; performing deep semantic space mapping on the user interaction need features, followed by reverse live streaming semantic gap analysis to obtain need gap filling information; and performing dynamic analysis of user behavior based on the multimodal user interaction perception map, and fitting the global evolution of real-time interactive emotions to construct a real-time interactive emotion map. This invention improves the emotional matching degree of user interaction and the satisfaction of interactive content through a deep understanding and analysis of user needs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of live streaming interaction technology, and in particular to an intelligent live streaming interaction method and system based on AI virtual humans. Background Technology

[0002] With the continuous evolution of artificial intelligence and virtual reality technologies, AI virtual humans, as intelligent interactive entities integrating speech synthesis, motion-driven processing, emotional computing, and natural language processing, have been widely applied in various fields such as education and training, digital marketing, entertainment live streaming, and virtual customer service, showing great development potential, especially in live interactive scenarios. Compared to the traditional live streaming model dominated by human anchors, AI virtual humans have the characteristics of strong sustainability, accurate emotional simulation, and flexible content generation, which can effectively reduce labor costs, enhance audience immersion, and achieve 24 / 7 uninterrupted operation. Therefore, against the backdrop of the digital and intelligent upgrading of the live streaming industry, live interactive systems based on AI virtual humans have become a key development direction.

[0003] However, in actual live streaming interactions, AI virtual humans face increasingly severe challenges. On the one hand, user behavior in real-time live streaming environments is highly diverse, including forms such as bullet screen input, likes, and gift-giving, with their emotional states and interests exhibiting complex, dynamically changing characteristics. On the other hand, traditional scripted or rule-driven virtual human response methods struggle to cope with high-frequency, multimodal, and highly perceptive interactive needs, resulting in poor user interaction experiences and decreased engagement. Furthermore, virtual human systems often lack a deep understanding of the semantic flow of live streaming and user intent, failing to achieve high-level emotional feedback modeling and precise content control, further limiting their intelligence level and practical application value.

[0004] Most existing AI virtual human live streaming systems rely on script templates or predefined behavior triggering mechanisms. While they possess some interactive capabilities, they frequently encounter problems such as delayed response, context mismatch, and rigid interaction when facing complex scenarios involving fluctuating user needs, emotional shifts, and semantic anomalies. These issues severely impact system stability and user satisfaction. Especially when the number of viewers grows rapidly or the live streaming content changes frequently, existing systems often struggle to achieve real-time perception and intelligent response to multi-source interactive data. Their ability to predict user emotional trends and adapt to content generation urgently needs improvement. Summary of the Invention

[0005] To address the aforementioned technical problems, this invention proposes an intelligent live-streaming interaction method and system based on AI virtual humans, thereby resolving at least one of the aforementioned technical issues.

[0006] To achieve the above objectives, this invention provides an intelligent live streaming interaction method based on AI virtual humans, comprising the following steps: Identify real-time live streaming interactive data streams, perform intelligent bullet screen filtering and multimodal user interaction perception, and construct a multimodal user interaction perception map; Key bullet comments are extracted based on the multimodal interaction perception map, and user needs are predicted to generate user interaction needs features. Deep semantic space mapping is performed on the characteristics of user interaction needs, and then reverse live streaming semantic gap analysis is performed to obtain information to fill gaps in the needs. Based on the multimodal user interaction perception map, dynamic analysis of user behavior is performed, and real-time interaction emotion global evolution is fitted to construct a real-time interaction emotion map. Based on the information to fill in the gaps in the demand and the real-time interactive emotion map, the virtual human makes interaction decisions and dynamically adjusts the live broadcast interaction to execute intelligent live broadcast interaction tasks.

[0007] This specification provides an intelligent live streaming interaction system based on AI virtual humans, used to execute the intelligent live streaming interaction method based on AI virtual humans as described above, including: The interaction perception module is used to identify real-time live interactive data streams, perform intelligent bullet screen filtering and multimodal user interaction perception, and construct a multimodal user interaction perception map. The demand prediction module is used to extract key bullet comments based on the multimodal interaction perception map and predict user demand, thereby generating user interaction demand features. The demand gap filling module is used to perform deep semantic space mapping on user interaction demand features, and then perform reverse live broadcast semantic gap analysis to obtain demand gap filling information. The emotion perception module is used to perform dynamic analysis of user behavior based on a multimodal user interaction perception map, and to fit the global evolution of real-time interactive emotions to construct a real-time interactive emotion map. The interactive control module is used to make virtual human interaction decisions based on the blank information to fill in the gaps and the real-time interactive emotion map, and to perform dynamic live broadcast interaction control in order to execute intelligent live broadcast interaction operations.

[0008] The beneficial effects of this invention are specifically as follows: By capturing multimodal interactive behavior data of viewers during live streaming in real time, including bullet screen text, like frequency, and gift-giving rhythm, a fusion data stream covering language behavior, economic behavior, and emotional behavior is constructed. Natural language processing technology is used to identify and filter abnormal bullet screens with extreme emotions, repetitive spamming, and irrelevant content, improving the quality and density of text information. Combined with temporal modeling of like and gift-giving behavior, periodic popularity fluctuations are extracted and quantified as behavioral spectra. Using a multimodal interaction perception map as the original input source, high-density, emotionally stimulating language patterns in bullet screens are identified and grouped through key bullet screen clustering and word frequency-emotion weight mapping. Combining behavioral popularity nodes, such as sudden increases in like curves and gift-giving peaks, the system uses graph neural networks to achieve semantic linkage modeling of behavior-language-emotion signals, thereby generating predictions of user interaction intentions. Furthermore, by combining bullet screen semantic similarity, tone tendency, and user group behavior consistency, the system can accurately infer the content direction, communication rhythm, or emotional resonance needs currently expected by viewers. By comparing the semantic tensor of the information range conveyed by current AI virtual humans with the characteristics of user needs, the system accurately identifies "expression gaps" or "information breaks" in the live streaming context, i.e., semantic blanks in the content layer. Based on this, a language content completion model is constructed to predictively deduce the topic motivations or emotional expressions that viewers may expect but have not yet expressed. A full-process emotional state time-series graph is constructed based on user behavior dynamics and emotional signal evolution. Through high-frequency sampling of like behavior, gift-triggered fluctuations, and the rhythm of bullet screen interactions, the system analyzes the emotional evolution trajectory of user groups using long-term dependency models such as Transformer. Combining real-time facial expression recognition and speech acoustic change trends (such as sudden pitch increases, accelerated speech speed, and other emotional outburst characteristics), the system can achieve phased identification and dynamic curve fitting of user emotions during the live stream. The system further integrates individual emotional states and group emotional aggregation, dynamically constructing a real-time interactive emotion graph, marking typical emotional cycles such as "intense—weak—reignited" wavebands. The system-perceived emotion graph and content completion information are jointly input into the interaction strategy generator to drive the AI ​​virtual human to conduct multi-dimensional decision-making behavior. This decision-making process extends beyond the language output itself, encompassing the coordinated control of multimodal output signals such as speech rate, tone, intonation, pauses, eye contact, and facial expression switching frequency. The system calculates emotional fit to select the most suitable response style across different emotional cycles and uses an emotion-behavior mapping matrix to determine the virtual human's micro-expression sequence and intonation curve. Real-time interactive data streams are continuously fused and optimized with feedback from the virtual human's behavioral responses. The interaction strategy model is continuously trained using user feedback signals (such as intonation recognition in speech, smile markings, and a surge in likes) within multi-turn dialogue structures, achieving dynamic self-optimization and learning updates.Simultaneously, the system records individual users' emotional response patterns and interaction style characteristics during high-frequency interactions, gradually accumulating personalized profiles and achieving long-term adaptability by "remembering who you are and what kind of interaction you prefer." Ultimately, through the integrated capabilities of accurate prediction, flexible generation, and personalized control, the AI ​​virtual human completes the transformation from "role-playing" to "digital personality symbiosis" in live streaming scenarios, forming a transferable and evolvable underlying capability framework for an intelligent interaction platform, applicable to various scenarios ranging from entertainment live streaming, virtual hosting, and educational interaction to business customer service. Attached Figure Description

[0009] Figure 1 This is a flowchart illustrating the steps of an intelligent live-streaming interaction method based on an AI virtual human, as described in this invention. Figure 2 This is a detailed flowchart illustrating the implementation steps of step S1. Figure 3 This is a detailed flowchart illustrating the implementation steps of step S2; Figure 4 This is a flowchart illustrating the detailed implementation steps of step S3. Detailed Implementation

[0010] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0011] This application provides an intelligent live streaming interaction method and system based on AI virtual humans. The executing entities of the intelligent live streaming interaction method and system based on AI virtual humans include, but are not limited to, mechanical equipment, data processing platforms, cloud server nodes, network upload devices, etc., which can be considered as general computing nodes in this application. The data processing platform includes, but is not limited to, at least one of an audio-visual management system, an information management system, and a cloud data management system.

[0012] Please see Figures 1 to 4 This invention provides an intelligent live streaming interaction method based on AI virtual humans, comprising the following steps: Identify real-time live streaming interactive data streams, perform intelligent bullet screen filtering and multimodal user interaction perception, and construct a multimodal user interaction perception map; Key bullet comments are extracted based on the multimodal interaction perception map, and user needs are predicted to generate user interaction needs features. Deep semantic space mapping is performed on the characteristics of user interaction needs, and then reverse live streaming semantic gap analysis is performed to obtain information to fill gaps in the needs. Based on the multimodal user interaction perception map, dynamic analysis of user behavior is performed, and real-time interaction emotion global evolution is fitted to construct a real-time interaction emotion map. Based on the information to fill in the gaps in the demand and the real-time interactive emotion map, the virtual human makes interaction decisions and dynamically adjusts the live broadcast interaction to execute intelligent live broadcast interaction tasks.

[0013] In the embodiments of the present invention, see Figure 1 The diagram below illustrates the steps of an intelligent live-streaming interaction method based on an AI virtual human according to the present invention. In this example, the steps of the method include: Identify real-time live streaming interactive data streams, perform intelligent bullet screen filtering and multimodal user interaction perception, and construct a multimodal user interaction perception map; In this embodiment, in the AI ​​virtual human-driven intelligent live streaming interaction scenario, the system first needs to perform real-time identification, structured processing, and semantic analysis on the large-scale, multimodal interactive data generated during the live streaming process, so as to support the virtual human to achieve intelligent response, personalized dialogue, and interactive content control. This step focuses on identifying and integrating three core user behavior data streams, namely bullet comments, likes, and gifts, performing intelligent bullet comment filtering processing, and integrating and constructing a multimodal user interaction perception map to provide a foundation for subsequent user intent recognition and virtual human response strategy generation. First, the system receives live streaming interactive data streams in real time at the live streaming platform interface layer, including three key input types: (1) text-based bullet comment streams; (2) like behavior data (including user ID and like timestamp); (3) gift-giving data (including gift type, value, and gift-giving time). All data is synchronized at the millisecond level, a unified timeline index is established for data of different dimensions, and real-time data stream frameworks such as Kafka are used for data caching and transmission scheduling. In the experimental scenario, a high-concurrency environment of 3000 bullet comments per second, 6000 likes per second, and 200 gifts per second was simulated to ensure the system's real-time processing performance. Next, intelligent filtering of the bullet comment stream was performed. The system first used natural language processing tools (such as Jieba or a custom BERT-based tokenizer) to segment the bullet comment text and remove stop words, special symbols, and other noise data. Subsequently, the system performed two-way recognition processing based on a deep text content analysis model: first, identifying illegal bullet comments by training a supervised multi-label classification model to identify bullet comments containing sensitive words, insulting language, advertisements, and vulgar content; second, detecting emotionally extreme bullet comments by using sentiment analysis (emotional polarity scoring) to identify highly negative, irritable, and aggressive language. In addition, a frequency analysis mechanism was introduced to monitor users' repeated bullet comment sending behavior and mark abnormal spammers. In the validation experiment, by introducing the RoBERTa text classification model for training, the recognition accuracy reached 92.3% on a dataset of 100,000 labeled bullet comments, achieving efficient and accurate filtering capabilities. After cleaning and labeling, the bullet screen data, along with behavioral data such as likes and gifts, are input into the multimodal perception fusion module. In this module, the system aggregates and models the three types of behavioral data within a 10-second sliding window period, constructing a "user interaction performance vector" for that time segment. Specifically: bullet screen behavior is represented by text embedding vectors (such as BERT embedding) and sentiment polarity vectors; like behavior is represented by frequency and temporal density distribution vectors; and gift behavior is calculated based on gift level weights to determine gift frequency and value, generating a gift feature vector. These three types of vectors are concatenated, normalized, and then input into the perception map construction model.Finally, the system utilizes graph structure modeling methods (such as Dynamic Graph Convolutional Network, D-GCN) to construct a "multimodal user interaction perception graph," where nodes represent users, interaction content, interaction types, time windows, etc., and edges represent behavioral relationships (such as the same user sending bullet comments, likes, and gifts within a short period of time). Node representations are learned through graph embedding methods. This graph possesses time-series characteristics and multimodal semantic labeling capabilities, effectively reflecting each user's interaction depth, interaction preferences, emotional state, and behavioral tendencies across different time periods.

[0014] Key bullet comments are extracted based on the multimodal interaction perception map, and user needs are predicted to generate user interaction needs features. In this embodiment, in an AI-driven intelligent live streaming environment, accurately identifying the core expressive content and potential intentions in user interactions is crucial for enabling the virtual human to proactively respond, guide interactions, and generate content. This step involves in-depth mining of the multimodal interaction perception map generated in the previous step to extract key bullet screen information. Combined with user behavior frequency, content preferences, and time series patterns, it models and predicts user interaction needs, thereby forming high-dimensional, dynamically changing user interaction demand characteristics. This provides important input for the virtual human's next intelligent decision-making and interactive behavior driving. The system analyzes the bullet screen nodes and their edge weights in the multimodal interaction perception map, using bullet screen semantic weight metrics (e.g., multiplying TF-IDF word weights by contextual interaction frequency) to filter out "key bullet screens" within each time segment. To ensure the representativeness and diversity of key bullet screens, the system sets a minimum threshold for semantic clustering (e.g., no fewer than 5 bullet screens per category) and introduces a BERT semantic embedding model based on the Transformer architecture to convert bullet screens into high-dimensional semantic vectors before performing K-Means or HDBSCAN clustering. In the experiment, 2000 bullet comments were clustered within a 10-second sliding window, ultimately extracting 10 to 15 categories of high-frequency and semantically independent key bullet comment topics. Each category represents a user's focus, such as "discussion of game character skills," "feedback on streamer behavior," and "urging for benefits." The system aligns the extracted key bullet comment semantic tags with user behavior nodes in terms of time and matches them with behavioral responses. For example, if bullet comments related to "requesting to connect" frequently appear within a certain time period, accompanied by a surge in likes or the gifting of specific items (such as virtual gifts like "microphone"), the system uses a graph neural network model (such as Temporal Graph Attention Network, T-GAT) to construct the alignment relationship between user behavior trajectories and semantic trajectories, thereby evaluating the strength of user group demand intent from two dimensions: interaction frequency and semantic consistency. The system introduces a behavior co-occurrence score (COS) to calculate the co-occurrence probability between semantic events and behavioral events. In the experimental environment, topics with a COS higher than 0.8 are marked as "candidate demand instructions." The system then performs user demand prediction modeling on these candidate demand instructions. This process constructs a deep prediction model based on a multi-task learning framework. The main task of the model is to predict the type of demand (such as "Q&A", "interaction", "welfare", "emotional reassurance", etc.), and the auxiliary task is to classify user response intentions (positive, neutral, negative). The model inputs are: user interaction sequence vectors (including bullet screen frequency, sentiment polarity, click behavior, and gift frequency), key bullet screen semantic feature vectors, and user node degree and edge weight features in the perceptual graph.During the training phase, a multi-label cross-entropy loss function was used, and 200,000 sets of user interaction sequence data were used for training and validation. The final prediction accuracy exceeded 80%, meeting the response requirements for demand recognition in live streaming scenarios. The prediction results were structured into a set of user interaction demand feature vectors. This set includes dimensions such as timestamp, user ID, main demand type, demand intensity level, latent semantic tags (e.g., "want to listen to music," "want to answer questions," "call for live chat"), and expected interaction path. These features are provided as input to the AI ​​virtual human response strategy generation module to dynamically generate multimodal response behaviors such as facial expressions, voice responses, content guidance, or program flow control.

[0015] Deep semantic space mapping is performed on the characteristics of user interaction needs, and then reverse live streaming semantic gap analysis is performed to obtain information to fill gaps in the needs. In this embodiment, in an AI-driven intelligent live-streaming interactive system, simply identifying the user's current explicit needs is insufficient to form a complete interactive loop. The key lies in uncovering potential semantic gaps to achieve proactive content generation, thereby enhancing the breadth and depth of the virtual human's intelligent response. This step revolves around this goal, and its core components include: performing deep semantic space mapping on user interaction needs features, conducting reverse semantic alignment analysis on the live-streaming content trajectory, and ultimately deriving information to fill gaps in the needs, providing feedforward support for subsequent intelligent virtual human-driven operations. The system receives a set of user interaction needs feature vectors generated in the previous stage, including the main need type, need semantic tags, behavioral preference parameters, and historical response paths. To represent these features within a unified semantic coordinate system, the system introduces a deep semantic mapping model (using a Sentence-BERT variant based on the Transformer architecture) to embed the user need vectors, projecting them into a 768-dimensional semantic space. During pre-training, the system incorporated a live-stream corpus (approximately 1.2 million semantic response dialogues) and trained using contrastive loss to obtain embeddings with good semantic neighbor discrimination capabilities. After embedding, the system models the semantic content trajectory of the current live-stream content. Specifically, within each 10-second time window, the system constructs a time-series semantic vector trajectory for the virtual human's output content (including speech-to-speech transcription, facial expression descriptions, and screen text display), models temporal relevance using Bi-GRU, and simultaneously embeds it into the same semantic space as the user's demand vector. At this point, the system holds two sets of semantic trajectories: one for user demand semantic trajectories and the other for the current live-stream content semantic trajectories. These two sets will be reverse-matched in the next step. The system performs semantic space reverse alignment analysis to identify "semantic gaps." This analysis calculates the cosine similarity matrix of the two semantic trajectories in the global space and constructs a semantic coverage matrix. For each user request, if the number of response points with a similarity higher than 0.8 in the current live stream semantic trajectory is less than a threshold (set to 3 in actual testing), the request is considered not to have been effectively met, and the system marks it as a potential semantic blank. Simultaneously, a comprehensive priority score is calculated by combining user interaction time weights (e.g., high-frequency behavior nodes have a weight of 1.0 for earlier times and 0.6 for later times) to rank the blanks. Experimental results show that, on average, 25-30 high-priority blank candidate points can be identified in 1500 live stream user request trajectories. After obtaining the semantic blanks, the system performs semantic expansion and blanking modeling based on their original semantic labels.A GPT-like large model (based on fine-tuned Prompt Learning technology) is used to generate "possible filler content prompts" for each semantic gap. If there is no interactive response to the "gift draw" within a certain period, structured language templates such as "Should we launch the next round of the draw?" or "Please flood the comments section with a certain keyword" are generated. The filler content is output in the form of a structured vector, including the target semantics, suggested content format (voice / action / text), expected interaction path, priority, etc. The set of demand gap filler information output in this step will serve as the feedforward input of the virtual human behavior planning module, enabling the AI ​​virtual human to have semantic self-drive when content saturation is insufficient. It can proactively initiate guided interactions, rhythmic control, or emotional responses, effectively filling content gaps and maintaining the continuity of the live broadcast context and audience engagement.

[0016] Based on the multimodal user interaction perception map, dynamic analysis of user behavior is performed, and real-time interaction emotion global evolution is fitted to construct a real-time interaction emotion map. In this embodiment, a multimodal interaction perception graph is invoked, which is composed of modal data fused from bullet screen text content, like sequence, gift-giving behavior, voice tone and emotional cues, and image expression recognition. Each data stream exists as a time-labeled node in the graph, and the node features include behavior type (text / action / emotion), emotion intensity, timestamp, user ID, weight parameters, etc. In the experiment, the system uses a 5-second time window to process all user interaction events within that time segment, aggregating an average of about 42 to 68 modal interaction data points per window. The system then enters the dynamic analysis stage of user behavior. By introducing a temporal modeling structure (Bi-LSTM combined with an attention mechanism), the system performs temporal context modeling on the above multimodal node sequences to identify the user behavior evolution path. For example, if a user repeatedly posts "So cute," likes frequently, and gives virtual gifts in the first two time windows, the model considers them to be on the "positive participation - emotional incentive" path through feature aggregation. Behavioral evolution labels are further categorized into five types: emotional participation, event-driven, cooling-down, noisy behavior, and abnormal reversal, with each type mapping to a corresponding emotion curve. Through dynamic behavioral labels, the system can predict the likelihood of user interactions and the direction of emotional shifts in the coming seconds. The system performs real-time interactive emotion evolution fitting on the global user's behavioral emotional state. This fitting employs an emotion vector field modeling method, mapping each user's emotional state (represented in two dimensions using a Valence-Arousal model) within each time window as vector points, and combining user interaction frequency and temporal intensity to construct a dynamic emotional flow field. The entire live stream forms a time-emotion two-dimensional vector field. The system then uses clustering (DBSCAN) and fitting (Gaussian Mixture Model, GMM) to dynamically divide this vector field into regions, thereby identifying current emotional hotspots, emotional contraction regions, and emotional reversal nodes. In experiments, the system modeled a sample of 500 live streamers, processing an average of 3400 emotion points per second, ultimately generating a real-time emotion map region structure with high aggregation. The system constructs a real-time effective interaction map (AIM), which serves as the direct input for scheduling the AI ​​virtual human's live-streaming behavior. Each node in the AIM represents an emotional unit, containing the following attributes: source user ID group, current emotional value (e.g., excitement, anger, calmness), emotional evolution trend vector, node time span, and position information within the map. By reading the emotional change trends in this map, the virtual human can adjust its language style, response speed, facial expressions, and even interactive topics, thereby achieving dynamic content adaptive generation driven by emotions.

[0017] Based on the information to fill in the gaps in the demand and the real-time interactive emotion map, the virtual human makes interaction decisions and dynamically adjusts the live broadcast interaction to execute intelligent live broadcast interaction tasks.

[0018] In this embodiment, in the AI-based intelligent live streaming system, the virtual human's interactive decisions need to fully integrate the unmet needs of the audience (needs gap filling information) and the real-time monitored group emotional state (real-time interactive sentiment graph) to achieve accurate, efficient, and emotionally appropriate live streaming interactive responses. This step aims to complete the intelligent decision generation and interactive control of the virtual human through multi-dimensional information fusion and dynamic strategy adjustment, ensuring continuous optimization of the interactive experience in the live streaming environment. The system receives needs gap filling information from the previous step, which reveals the user's currently unmet interests or topic gaps, typically including keyword sets, potential intent tags, and missing content dimensions. Simultaneously, the system acquires the real-time interactive sentiment graph, which reflects the spatial distribution and temporal evolution trend of the overall emotional state of users in the current live streaming room, including important indicators such as emotional intensity, clustering areas, and emotional flow direction. In the experimental environment, needs gap filling information is generally updated in real time through a deep semantic matching model (such as Transformer semantic embedding comparison), with the update interval controlled within 3 seconds; the real-time interactive sentiment graph is refreshed every 5 seconds to ensure information timeliness. The system uses a multi-modal fusion algorithm to perform deep correlation analysis between needs gaps and emotional state. The specific method is as follows: A weighted fusion approach is used to weight and superimpose the priority of user needs (e.g., determined by interest popularity and content novelty) with the emotional clustering and positive / negative emotion ratio in the sentiment graph, forming a comprehensive interaction decision scoring matrix. Each calculation of this matrix includes hundreds of user behavior and emotional node parameters, and a non-linear optimization algorithm (such as genetic algorithm or reinforcement learning strategy optimization) is used to find the optimal interaction response path. Experimental parameters show that in a live broadcast room with 500 participants, the calculation time for this decision scoring is controlled within 200 milliseconds, meeting the requirements of real-time interaction. Subsequently, based on the interaction decision scoring, the system generates a virtual human interaction behavior plan, covering multiple dimensions such as language output content, tone adjustment, facial expressions, and topic switching rhythm. Specifically, the virtual human first generates preliminary dialogue text through a text generation model (based on a GPT variant of the pre-trained language model Fine-tune), and then adjusts the text's emotional color using the emotion enhancement module to match the dominant emotion in the current sentiment graph; for facial expressions, the action library calls micro-expressions and gesture combinations related to the current emotional intensity; the topic control module promptly introduces new topics based on the need gap prompts, avoiding content gaps or repetitions. Each round of interaction is generated within one second to ensure smooth flow. After the interaction behavior plan is finalized, the system performs dynamic live interaction control, monitoring interaction feedback in real time (including changes in bullet comments, fluctuations in likes, and gift giving), and fine-tuning the virtual human's behavior strategy. An adaptive feedback control model is introduced during the control process, adjusting the frequency of topic introduction, the emotional tone of language, and the interaction rhythm by comparing the difference between actual interaction feedback and expected emotional trends.For example, when feedback indicates that audience emotions have cooled, the system increases the intensity of emotional incentives and the frequency of interaction to enhance engagement. Experimental data shows that this feedback control mechanism can increase user engagement metrics by 12%-18%.

[0019] In this embodiment, see Figure 2 The specific steps for identifying real-time live interactive data streams, performing intelligent bullet screen filtering and multimodal user interaction perception, and constructing a multimodal user interaction perception map are as follows: Identify real-time live interactive data streams, including bullet screen streams, like behavior data, and gift-giving data; Intelligent bullet comment filtering is applied to the bullet comment stream to construct an intelligent filtered bullet comment stream; Extract the timestamp of the like behavior data; Based on the timestamp, a time-series behavior analysis is performed to construct a time-series curve of the "like" behavior. Analyze the gift-giving data by the number of times gifts are given and the types of gifts to obtain real-time gift characteristics; Based on real-time gift characteristics, time-series curves of like behavior, and intelligent filtering of bullet screen streams, multimodal user interaction perception is performed, and a multimodal user interaction perception map is constructed.

[0020] In this embodiment, in the AI ​​virtual human intelligent live streaming scenario, the identification of real-time live interactive data streams is the starting point for system data input, mainly including three core data types: bullet screen streams, like behavior data, and gift-giving data. First, after obtaining authorization from the relevant platform and users, it is necessary to connect to the live streaming platform's data interface or SDK, collecting user interaction behavior through the WebSocket protocol or RTMP data stream interface. Each type of data stream has a clear structural format. Bullet screen streams are typically in structured JSON format, containing fields such as user ID, sent content, timestamp, and emotion tags (such as emoticons); like behavior data often only contains timestamps and user IDs, reflecting high-frequency click events; while gift-giving data includes attributes such as gift ID, gift name, quantity, user level, and gifting time. To cope with high-concurrency scenarios, the system uses Kafka as a data access buffer queue, combined with Spark Streaming or Flink for real-time streaming processing. In the experimental environment, the data stream access speed reaches 3000 records per second, and the latency can still be kept within 100ms during peak live streaming periods. Bullet comments may contain spam, invalid interactions, advertisements, duplicate content, or malicious attack text, thus requiring bullet comment filtering. This step employs a multi-layered filtering mechanism. First, regular expressions and keyword matching methods are used for initial cleaning to remove bullet comments containing blacklisted keywords or violating content guidelines. Then, deep learning text classification models (such as BERT or ERNIE) are used for contextual semantic analysis to determine whether the bullet comments have interactive value or are low-quality content. In the experiment, a multi-label classification model based on BERT was selected. The input was the bullet comment text, and the output was three labels: "high interactivity," "neutral," and "noise." The model was trained on a training set (containing 100,000 manually annotated bullet comments), achieving a classification accuracy of over 93%. The processed bullet comments are semantically labeled, and "high interactivity" bullet comments are retained based on time sequence, ultimately forming an intelligent filtered bullet comment stream. Furthermore, to improve semantic understanding, the system also introduces a sentiment analysis model to score the sentiment tendency (positive, neutral, negative) of the bullet comments, supporting subsequent user profiling and sentiment feedback. Liking is a form of instant user feedback, typically a lightweight interaction without textual expression. In practice, a "timestamp" field is extracted from the like data structure, and a high-precision time synchronization mechanism (such as NTP protocol to synchronize with the server clock) ensures data accuracy. When processing like data, the system uses a sliding window mechanism to count the frequency of like events per second for subsequent time series modeling. In the experiment, the user click frequency was controlled to no more than 20 times per second to prevent like manipulation, and a sliding window length of 5 seconds with a step size of 1 second was added during data processing to dynamically aggregate the number of likes.All "like" events are extracted and uniformly formatted into a time-series data structure (such as a time-series DataFrame) and time-aligned to provide input data for subsequent behavior analysis. The construction of the time-series curve for "like" behavior primarily employs a combination of statistical modeling and signal analysis. Using the extracted timestamp sequence of "likes," a sliding window weighted average and exponential moving average (EMA) are used to smooth the data, preventing local anomalies from interfering with the overall curve trend. Furthermore, Fourier transform (FFT) is used to analyze the frequency domain characteristics of "like" behavior, identifying periodic patterns and sudden events. The system also introduces mutation detection algorithms (such as CUSUM or Z-score methods) to annotate drastic changes in "like" frequency, helping to identify high-impact triggering events in the interaction between the AI ​​virtual human and the audience. In the experimental setup, a 60-second analysis period is used; the time interval where the cumulative number of "likes" exceeds twice the average is defined as the "high-interaction window" and serves as an important marker interval for subsequent multimodal perception. The final constructed time-series curve of like behavior was standardized and transformed into a feature vector sequence for alignment analysis with bullet screen and gift data. Gift-giving behavior is not only a core indicator of live stream revenue but also reflects users' emotions and recognition of the virtual avatar. In this stage, the system first parses the gift data, extracting the timestamp, gift ID, gift value (unit price), number of gifts, and user level for each gift. Gifts are categorized and coded by type, such as by value level into "ordinary gifts," "medium gifts," and "high-value gifts," and the frequency, cumulative value, and gift density of each type of gift within a unit of time are statistically analyzed. In the experiment, the total value of gifts and the number of gifts were analyzed in 15-second time windows, and a gift intensity feature vector was constructed (e.g., {ordinary: 15 times, medium: 7 times, high-value: 2 times, total value: 450 yuan}). Furthermore, user profiling analysis is introduced, and users who give gifts are tagged (e.g., "high-frequency gifter," "first-time gifter"), and their behavioral change trajectories are used as additional features. The analysis results provide a basis for identifying core user groups and interaction preferences. The three types of interactive data are time-aligned and standardized to second- or sub-second precision. During the fusion phase, the system employs a multimodal attention mechanism to weight different modalities, identifying the factors most influential on user emotional fluctuations and behavioral changes within specific time periods. For example, if a large number of encouraging comments appear in the bullet screen at a certain moment, accompanied by a surge in high-value gifts and likes, this event will be marked as an "emotional resonance peak."Subsequently, the system constructs a user interaction graph based on a graph neural network (GNN), treating each user as a node in the graph and interaction events (such as simultaneous likes, joint gift-giving, or activity within the same time period) as edges, with the edge weights reflecting the interaction strength. In the experiment, a graph convolutional network (GCN) was used for training and embedding learning of the graph, enabling visualization of the user relationship network and optimization of AI virtual human interaction strategies. This interaction perception graph can be updated in real time and serves as an important input for generating virtual human language, emotions, and performance strategies, improving the immersiveness of the live streaming experience and the naturalness of human-computer interaction.

[0021] In this embodiment, the specific steps for performing intelligent bullet comment filtering on the bullet comment stream to construct an intelligent filtered bullet comment stream are as follows: The bullet screen stream is preprocessed by word segmentation and stop word removal, and the preprocessed bullet screen stream is extracted. Perform deep content recognition on the pre-processed bullet comment stream to extract abnormal and illegal bullet comments; Detect negative sentiment based on pre-processed bullet screen stream and label bullet screens with negative sentiment. Calculate the frequency of the negative emotional comments and identify the users who posted them; Based on the frequency of the barrage users, abnormal barrage inferences are made, and abnormal barrage barrages and user IDs are marked. Intelligent bullet screen filtering is performed based on abnormal and inappropriate bullet screen comments, abnormal spam bullet screen comments, and user IDs to build an intelligent filtered bullet screen stream.

[0022] In this embodiment, bullet screen preprocessing is an important prerequisite step for text analysis. Its purpose is to standardize and simplify the text content, and improve the efficiency and accuracy of subsequent model analysis. First, the system connects to the real-time bullet screen stream of the live broadcast platform, extracts the core content fields of each bullet screen (such as the text field), and inputs it into the Chinese word segmentation module for word segmentation. Considering the strong oral language characteristics, concise meanings, and a large number of Internet buzzwords in bullet screens, a deep learning-based word segmentation tool "LAC (Lexical Analysis of Chinese)" and the "jieba" extended vocabulary are used in combination to ensure the effective recognition of new words and new memes. Secondly, a custom Chinese stop word list (containing more than 3,000 terms) is used to clean the stop words in the bullet screen text, removing words without substantial meaning (such as "de", "le", "a"). In addition, to adapt to the characteristics of bullet screens, regular processing of Emoji, punctuation marks, and repeated characters is also added, making the text content more compact and the semantics clearer. In the experiment, after cleaning 50,000 bullet screens, the average length of each bullet screen text was compressed from the original 12.3 terms to 6.7 effective terms, significantly improving the processing efficiency of the subsequent analysis model. Abnormal and违规 bullet screens refer to those involving inappropriate remarks, offensive content, advertising spamming, politically sensitive information, etc., which need to be accurately identified and removed. In this step, a Chinese text classification model based on BERT is used to perform in-depth semantic understanding and classification judgment on the preprocessed bullet screen content. The model is trained as a multi-label recognizer, and the output labels include categories such as "politically sensitive", "personal attack", "vulgar and pornographic", "advertising promotion", "normal", etc. The training data uses 200,000 manually labeled samples, among which the total number of various违规 bullet screens is about 55,000, and the F1 value of the model on the validation set reaches 0.92. During the recognition process, to improve robustness, the system also introduces a rule filtering mechanism as a supplement, such as regular matching of commercial promotion codes like "QQ number", "VX number", "lowest price across the network", etc., to cover edge cases that the model fails to recognize. The bullet screens identified as abnormal content will be marked with a违规 label and the user ID and timestamp will be output, serving as the basis for subsequent behavior analysis. Experimental data shows that in the actual operating environment, the违规 recognition model can process about 1,500 bullet screens per minute, and the processing delay is less than 80ms, meeting the real-time requirements of the live broadcast. Negative emotion bullet screens have a significant impact on the live broadcast atmosphere, especially in AI virtual human live broadcasts, which may cause problems such as "virtual human emotional breakdown" or decreased interaction. Therefore, the system needs to perform sentiment analysis to timely identify the user's emotional state. In this step, a Chinese sentiment analysis model (SKEP - Sentiment Knowledge Enhanced Pre-training Model) is used to perform binary classification on the preprocessed bullet screens, marked as "positive" and "negative" respectively. At the same time, to enhance the model's understanding of the context, the system combines the bullet screen with the context bullet screens within its time period (±10 seconds) and inputs them together to simulate human cognition of the local context.During training, transfer learning was performed using public sentiment datasets from Douban comments and Weibo, and fine-tuned by incorporating 30,000 manually labeled bullet screen samples from live streaming scenarios. The final model achieved an accuracy of 91.6% on the live streaming corpus. Negative sentiment tags included categories such as sarcasm, complaint, criticism, and anxiety. The system labeled these bullet screens in the data stream and bound them to user IDs to support subsequent frequency analysis and user identification. To measure the density and source of negative emotions, it was necessary to count the frequency of negative bullet screens sent by the same user within a specific time window. The system used a sliding time window mechanism to count the number of negative bullet screens sent by each user in the past 30 seconds. If a user's accumulated negative bullet screens exceeded a set threshold (5 in the experiment), they were labeled as a "high-frequency negative user." In addition, the system also calculated the overall negative bullet screen ratio in the live stream (e.g., triggering an alert mechanism if the negative bullet screen ratio exceeds 15% per minute) to judge changes in the overall atmosphere of the live stream. Experimental data shows that in a live stream containing approximately 100,000 bullet comments, negative emotions accounted for an average of 8.3%. However, this percentage rose to over 18% during periods when the AI ​​virtual human exhibited abnormal behaviors such as "mistakes" or "interaction pauses." Therefore, the frequency of negative bullet comments is significantly correlated with the quality of virtual human interaction, serving as a key reference for designing subsequent intervention mechanisms. Besides negative emotions, abnormal spamming behavior also severely impacts bullet comment readability and interaction rhythm. The system determines whether a user's spamming is abnormal based on the frequency and repetitiveness of their messages. In terms of time, the system defines "short-term spamming" as a user sending more than 7 bullet comments within 10 seconds, and "continuous spamming" as maintaining a high frequency (>5 comments / 10 seconds) for five consecutive time windows. In terms of content, the system uses cosine similarity and Jaccard similarity coefficients to calculate the repetition of consecutive bullet comments; a similarity higher than 0.9 is considered repetitive spamming. Furthermore, user level information is introduced to distinguish between "normal active users" and "low-level spamming users," with the latter receiving a higher warning level. Users identified by the system as engaging in abnormal spamming will have their ID, comment content, and timestamp recorded, and will be uniformly marked as "spamming risk users." Based on the identified anomaly information, the system performs comprehensive filtering of the comment stream, generating an "intelligent filtered comment stream" as input to the virtual human interaction system. The filtering logic consists of two layers: the first layer is mandatory blocking, directly deleting all comments marked as "violation content" and "highly repetitive spamming," while simultaneously adding the associated users to a "temporary mute list"; the second layer is strategic weak filtering, imposing a display probability penalty on comments posted by users with "mildly negative emotions" or "moderate spamming behavior," such as displaying them only on the interface of 5% of viewers, reducing their impact. Furthermore, the system provides feedback signals to the virtual human behavior generation module. When the proportion of negative comments exceeds a set threshold, the AI ​​virtual human will automatically generate language strategies to alleviate emotions, such as apologies, jokes, and self-regulation, improving the interactive experience.The entire filtering module is deployed based on the Flink stream processing framework. In the experimental environment, it can process 8,000 bullet comments per second with a filtering latency of less than 60ms, supporting the application requirements of real-time intelligent bullet comment management in large-scale live streaming environments.

[0023] In this embodiment, see Figure 3 The specific steps for extracting key bullet comments based on the multimodal interaction perception map and predicting user needs to generate user interaction demand features are as follows: Key bullet comments are extracted based on a multimodal interaction perception map and processed in multiple time segments to obtain bullet comment streams for different time periods. Deep semantic analysis was performed on the bullet screen stream to obtain the semantic features of bullet screens at different time periods; The frequency of the bullet comments is calculated and high-frequency bullet comments are identified. Based on the high-frequency bullet comments and the semantic features of the bullet comments, user interaction intent analysis is performed to obtain user interaction intent patterns. Predict user needs based on user interaction intent patterns to generate user interaction need characteristics.

[0024] In this embodiment, based on a multimodal interaction perception map, "key bullet comments" that significantly impact user interaction within different time periods are extracted. First, based on high-interaction nodes marked in the perception map (e.g., sudden increases in likes, gift surges, peaks in emotional resonance), the system identifies the bullet comment content that triggers the behavior and constructs a key bullet comment index pool. To accurately extract the causal relationship between bullet comments and behavior, a cross-modal association model based on the Attention mechanism is adopted. Behavioral signals (e.g., like frequency change curves) are used as the main input, and bullet comment text as the secondary input. Their relevance scores are calculated, and the Top-N bullet comments are selected as the "key bullet comment group" for that time period. Then, the system divides the entire live stream into multiple time periods along the timeline, employing a strategy combining fixed windows and adaptive adjustments. For example, the base time period is set to 60 seconds, but if there are significant interaction fluctuations within that period (e.g., user activity increases exceeding twice the normal value), it is automatically shortened to 30 seconds for refined processing. After processing, each time period has an independent key bullet comment stream, forming a multi-segment bullet comment sequence with a time index, facilitating subsequent semantic and frequency analysis. To understand the deeper meaning expressed by bullet comments, the system needs to perform deep semantic modeling on key bullet comment streams. This stage uses a pre-trained Chinese language model based on the Transformer architecture (such as RoBERTa-wwm-ext or ChatGLM) to vectorize each bullet comment and combines it with contextual semantic understanding. Each bullet comment stream is concatenated into temporal segments and fed into a bidirectional encoder to extract high-dimensional semantic features (usually 768-dimensional vectors). Before model processing, word segmentation and entity recognition are performed to extract key elements such as people, actions, and events. Sentiment analysis (SKEP) is used to supplement the semantic vectors with emotional dimensions, forming complete semantic description units. For example, if a large number of bullet comments focus on "dancing," "so cute," and "one more time" within a certain time period, the system can automatically classify them into the "content expectation" semantic cluster. Through semantic clustering algorithms (such as K-means and HDBSCAN), the system performs cluster analysis on the bullet comment feature vectors within each time period, outputting semantic topics and their proportions. In the experiment, the bullet comments from 12 AI virtual human live streams were clustered, identifying an average of 3-5 semantic topics per segment, with a topic consistency score (Silhouette Coefficient) of 0.67, demonstrating stable performance of the semantic model in interactive language understanding. Bullet comment frequency is an important dimension for measuring user interaction intensity and focus. In this step, the system uses a sliding window mechanism (window size of 10 seconds, sliding step size of 1 second) to dynamically count the number of bullet comments within each time period. For bullet comment texts with high recurrence frequency, a hash mapping + string normalization method is used for standardized counting.Text normalization includes synonym replacement (e.g., classifying "fantastic" and "awesome" as synonyms), punctuation removal, and emoji character mapping (e.g., classifying [doge] as "teasing"). The system then jointly analyzes the frequency of bullet comments (danmu) and user-sent messages to identify high-frequency bullet comments that are repeated by multiple users. A frequency threshold is set (e.g., if a bullet comment appears more than three times the average frequency during that period), and it is marked as a high-frequency bullet comment. High-frequency bullet comments not only reveal current user focus but may also reflect potential collective emotional trends or interaction needs. Experiments showed that high-frequency bullet comments accounted for less than 10% of the total live stream duration, but almost all were concentrated in content climaxes or segments featuring special AI virtual human performances, exhibiting high interactivity. Interaction intentions include, but are not limited to, categories such as "requesting a response," "expecting a performance," "expressing liking," and "questioning / criticizing." The system uses a FusionNet neural network structure, concatenating semantic topic vectors and high-frequency word distribution vectors as input, and outputting multiple interaction intention labels with accompanying confidence scores. For example, if the bullet comments (danmaku) focus on "singing," "one more song," and "it sounds good" during a certain period, and these phrases appear with an unusually high frequency, the system can infer that the user intends to "request content continuation." To increase robustness, the system also introduces a multi-label learning strategy and uses 50,000 manually labeled intent samples for supervised training. On the test set, the macro-average F1 score for intent recognition is 0.87, indicating that the system performs relatively evenly across multiple intent types. Ultimately, each time period outputs one or more "interaction intent patterns" to characterize the main needs or emotional orientation of the current audience. To achieve intelligent response and proactive interaction capabilities for the AI ​​virtual human, the system needs to further predict user needs based on intent patterns. This module is built on an LSTM+Attention framework based on temporal prediction. The input is the sequence of interaction intents from the most recent N time periods (N is generally set to 5~10), and the output is the dominant user need category that may appear in the next time period. Combining live stream content features (such as the current behavioral state of the AI ​​virtual human, plot points, etc.) and historical audience behavior data (such as common demand trajectories in previous similar live streams), the model performs context-aware prediction. The system also incorporates a confidence interval mechanism to add a range of possibilities to the prediction results. For example, if the current system predicts a 72% probability that "viewers will ask the virtual human to dance," then it can generate relevant language scripts and action preparations for the AI ​​virtual human in advance. In experiments, the demand prediction model achieved an accuracy rate of 81% in the middle of the live stream, and the accuracy rate improved to over 86% during the peak of the live stream. Finally, the system provides the output user interaction demand characteristics (such as demand type, popularity level, and keywords) as input to the virtual human engine, enabling automatic content generation and personalized responses.

[0025] In this embodiment, reference Figure 4 The specific steps for performing deep semantic space mapping on user interaction demand features and then performing reverse live streaming semantic gap analysis to obtain demand gap filling information are as follows: Deep semantic space mapping is performed on the user interaction demand features to obtain the user demand feature space; Deeply mine potential user needs from the user demand feature space to extract potential user demand features; Predict future bullet comments based on users' potential demand characteristics, and generate bullet comment streams for future time periods; High-frequency danmaku semantic clustering is performed on the danmaku stream in the future time period, and reverse live streaming semantic gap analysis is performed to generate demand gaps; Based on the gaps in demand, virtual human interaction instructions are defined to fill in the gaps in demand information.

[0026] In this embodiment, user interaction needs are typically represented by discrete category labels (such as "expecting a performance," "requesting a response," and "expressing approval") or low-dimensional vectors, which are insufficient to reflect their complex semantic relationships and evolutionary trends. Therefore, it is necessary to map these features into a high-dimensional semantic space with contextual linkage and semantic distribution structure. This step uses a semantic embedding model based on the Transformer architecture (such as SimCSE-BERT) to encode user interaction needs. The input consists of interaction intent labels and related bullet screen content descriptions across multiple time periods, and the output is a unified semantic vector representation. The model is optimized through contrastive learning, aggregating semantically similar needs in the feature space. Subsequently, dimensionality reduction methods such as t-SNE and PCA are used to visualize and analyze the need space, revealing that it can be divided into multiple semantic clusters, such as "content expectations," "interaction requests," and "emotional expression." In the experiment, semantic space embedding modeling was performed on 2,400 need features generated in 8 live streams, ultimately constructing a high-order semantic space with a dimension of 1024 to carry future bullet screen predictions and potential need inferences. Explicit needs typically only cover a portion of the true motivations behind user interactions; therefore, the system needs to further mine latent need features through semantic space analysis. This step employs an autoencoder model to perform nonlinear dimensionality reduction and reconstruction of the need feature space, identifying latent semantic trends hidden behind the semantic distribution. By training a deep encoder-decoder network with a bottleneck structure, explicit needs are compressed into a low-dimensional latent variable, and then unexpressed but potentially existing need directions are extracted from the reconstruction error. To improve the model's interpretability, a VAE (Variational Autoencoder) structure is introduced to model the need space as a continuous probability distribution, outputting a "probability score" for latent needs. For example, when users frequently request "singing" but no "dancing" content appears, the system discovers that the latent need probability of "dance-related interaction" is as high as 62% through the reconstruction of adjacent semantic clusters. In experiments, the model was trained on a sample set containing 90,000 need-interaction data pairs, with an average KL divergence of less than 0.05, demonstrating stable latent need generation. Finally, the system outputs 1-3 high-confidence "user latent need features" for each time period for subsequent bullet screen prediction. Once user latent needs are identified, they can be used to predict the content of bullet comments that users may post in the following time period. This step employs a Chinese bullet comment generation model based on the GPT architecture (such as ChatGLM-tuned for live comment prediction) to comprehensively model the input latent need features, historical bullet comment context, and live broadcast behavior nodes to generate a temporally correlated sequence of future bullet comments. The model learns the language generation rules in the live broadcast context through fine-tuning instructions and incorporates a reinforcement learning mechanism (RLHF) to optimize the naturalness and interactivity of the content.In practice, the input features include the interaction trend within the current 5 minutes, the distribution of user semantic clusters, and the behavioral state of the AI ​​virtual human. The output is a predicted Top-100 list of bullet comments for the next 30-60 seconds, covering possible user feedback, requests, and evaluations. To control linguistic diversity and realism, the system introduces a temperature sampling mechanism (Temperature=0.85) and Top-k filtering (k=40) to ensure that the generated content is representative without becoming homogeneous. In the test scenario, the average similarity between the system-generated bullet comments and real bullet comments reached 0.74 (based on BLEU-2 scoring), demonstrating that the prediction model has strong text reconstruction capabilities. After the bullet comment prediction is completed, it is necessary to further identify high-frequency semantic themes to infer semantic gaps in the current live content that have not yet been covered or ignored. This step first performs word segmentation and embedding processing on the predicted bullet comments, and then uses the HDBSCAN density clustering algorithm to semantically aggregate high-frequency words and phrases, forming several "pseudo-interaction themes". After clustering, these themes are compared with the current AI virtual human's behavior logs (such as spoken content, performance content, and visual expression). A semantic alignment model (such as SBERT similarity matching) is used to identify which themes are not yet reflected in the current live stream content or have "weak expression." The system defines these discrepancies as "semantic gaps" or "demand gaps" and outputs them in order of importance. For example, if potential user comments such as "please dance" or "want to see emojis" are clustered, but the virtual human does not provide any related response in the current segment, the system determines that there is an "interactive content gap" in that time period and suggests implementing a fill-in strategy. In experiments, the average accuracy of reverse semantic coverage calculation reached 88.5%, and the success rate of gap recognition combined with behavior log matching was 76%, demonstrating stable and reliable performance in real-time feedback scenarios. After identifying demand gaps, the system needs to convert them into interactive fill-in instructions that the virtual human can execute. This step defines a "semantic template system for interactive fill-in instructions" to structurally fill different types of gaps. For example, if the gap is "action not covered," an action-related instruction is generated (e.g., "trigger dance action package A"); if the gap is "emotional response missing," a language-related supplementary instruction is generated (e.g., "generate comforting statements"). The instruction generation process uses a rule-based + model hybrid approach, matching the corresponding template structure based on the gap type and calling a semantic generation module (e.g., the T5 text generation model) to complete the instruction content. At the execution level, the instruction is sent to the AI ​​virtual human control module, triggering the corresponding performance logic or language module. The system also sets a confidence threshold (default 0.7), performing the supplementary operation only when both prediction accuracy and demand intensity are met, to avoid false triggering.Experimental data shows that the average latency of the fill instruction response is 280ms, the trigger accuracy reaches 91%, and the user interaction satisfaction is improved by 13.6% (based on user survey feedback), proving that the module has significant effectiveness in enhancing the interactive performance of AI virtual humans.

[0027] In this embodiment, the specific steps for performing dynamic analysis of user behavior based on a multimodal user interaction perception map, fitting the global evolution of real-time interactive emotions, and constructing a real-time interactive emotion map are as follows: Perform dynamic analysis of user behavior on the time-series curve of "like" behavior to extract user behavior trajectory; Identify the peak intervals, durations, and fluctuation frequencies of user behavior trajectories to obtain dynamic user behavior characteristics; The user interaction density fluctuation evolution is analyzed based on the dynamic behavior characteristics of users to generate a user interaction density fluctuation map. By fitting the user interaction density fluctuation map and real-time gift features to the global evolution of real-time interaction sentiment, a real-time interaction sentiment map is constructed.

[0028] In this embodiment, the time-series curve of "like" behavior is an important data sequence reflecting the users' willingness and participation in the live broadcast room. First, the system extracts timestamps and corresponding like counts from the real-time collected like event data to form a high-frequency time series. The time granularity is usually set to 1 second or finer to capture user dynamics. A sliding window (window size 30 seconds, step size 5 seconds) is used to smooth the like data, removing occasional noise and obtaining a stable time-series curve. Subsequently, combined with user ID association analysis, the trajectory of the same user's like behavior on the time axis is extracted to obtain the time distribution and frequency characteristics of a single user's likes. The trajectory extraction not only includes the like time points but also combines the user's active time periods to analyze the coherence and concentration of user like behavior. To achieve higher accuracy, the Dynamic Time Warping (DTW) algorithm is used to match and cluster user behavior trajectories to distinguish different types of like behavior patterns, such as "explosive likes" and "uniform likes". In the experiment, based on the like data of 10,000 users, the DTW clustering results showed four significant behavioral patterns, with a cluster silhouette coefficient of 0.72, verifying the effectiveness of the behavioral trajectory analysis. For the like behavior trajectory extracted in the first step, the system further focused on the "peak intervals" in the trajectory, i.e., the time periods when the number of likes significantly increases and reaches a local maximum. Peak detection algorithms (such as Continuous Wavelet Transform (CWT)) were used to analyze the like time series curve, accurately identifying the start and end times and peak amplitude of the like surge intervals. For each peak interval, the duration (the time span between the peak and the threshold when the number of likes exceeds the threshold) and fluctuation frequency (the number of times the number of likes changes per unit time) were calculated. Through these indicators, the system summarized the dynamic characteristics of user like behavior, such as behavioral labels like "short-term strong bursts," "medium-to-long-term maintenance," and "frequent fluctuations." For example, during the peak period of a live stream, the system detected a like peak lasting 3 minutes, with the peak like rate more than 5 times the normal rate and a fluctuation frequency of approximately 12 times per minute, indicating very active audience interaction. Experimental verification shows that the peak detection accuracy of this step exceeds 92%, and dynamic behavioral characteristics can effectively distinguish user response types for different live streaming content. User dynamic behavioral characteristics are an important indicator of live streaming interaction intensity, but a single indicator cannot fully reflect the overall interaction. Therefore, the system integrates the intensity, duration, and volatility of like behavior to construct a user interaction density index. Specifically, after standardizing the number of likes, it combines the volatility frequency and calculates the interaction density score using a weighted function (the weights are optimized experimentally, such as a like intensity weight of 0.6 and a volatility frequency weight of 0.4). Using time-series data visualization technology, the evolution of this density over time is plotted as an interaction density fluctuation graph, with peak areas representing the most intensive user interactions. This graph not only visually displays changes in interaction intensity but can also be used for real-time monitoring and trend prediction.To further enhance the expressiveness of graphical information, the system introduces multi-scale time window analysis (e.g., 1-minute, 5-minute, and 10-minute windows) to obtain interaction density fluctuation maps at different granularities, supporting a comprehensive judgment of short-term responses and long-term behavioral trends. Experimental data shows that the interaction density fluctuation map is highly correlated with live stream audience activity and gift-giving behavior, with a Pearson correlation coefficient reaching 0.81. Liking behavior reflects positive feedback, while gift-giving reflects a higher level of user emotional expression. This step integrates and analyzes user interaction density fluctuation maps with real-time gift-giving data (including gift type, gift frequency, and gift user activity) to construct a global interaction sentiment model. Specifically, the system employs a multivariate time series fitting method (such as a multivariate long short-term memory network LSTM-MVN), simultaneously inputting interaction density and gift features to learn the dynamic coupling relationship between the two. The model output is a user sentiment intensity curve and sentiment evolution trajectory, reflecting the rise and fall and turning points of the user's overall emotion during the live stream. This sentiment curve is used to adjust the performance strategy of the AI ​​virtual human in real time to enhance emotional resonance. To improve model robustness, an emotion label correction mechanism is introduced, combined with the sentiment analysis results in the bullet comments for multimodal validation. In the experiment, the model was trained using data from nearly 50 live streams. The correlation coefficient between the model's sentiment prediction and actual user feedback (based on questionnaire and bullet screen sentiment statistics) reached 0.79, indicating that the fitted model can effectively capture the overall trend of sentiment evolution. Finally, the constructed real-time interactive sentiment map is presented in the form of a visual dashboard, supporting real-time monitoring and response by operations personnel and AI systems.

[0029] In this embodiment, the specific steps for detecting emotional abrupt changes in the real-time interactive sentiment graph and performing multi-time-point sentiment state modeling to construct the user sentiment state graph are as follows: Multi-dimensional spatiotemporal analysis is performed on the real-time interactive sentiment graph, and sentiment fluctuation tracking analysis is conducted to extract the user's sentiment fluctuation tracking path; Detect abrupt changes in user emotions by tracking the user's emotional fluctuations and mark the time nodes of user emotional transitions; Based on the aforementioned user emotional shift time points, high-probability triggering of bullet comments is identified; The user sentiment shift was mined from the high-probability triggered bullet comments to obtain the deep-seated patterns of user sentiment shifts. Based on the deep emotional transformation patterns of users, multi-time point emotional state modeling is performed to construct a user emotional state map; Based on the information to fill in the gaps in demand and the user's emotional state, the system makes virtual human interaction decisions and dynamically adjusts the live broadcast interaction to execute intelligent live broadcast interaction tasks.

[0030] In this embodiment, the real-time interactive sentiment graph is an expression of emotional state after integrating multimodal data from likes, gifts, and bullet comments, presenting the dynamic characteristics of users' overall emotions changing over time and space (different live streaming content modules, different user groups). To gain a deeper understanding of emotional dynamics, the system introduces a multidimensional spatiotemporal analysis framework, extending the sentiment graph to the time series and spatial levels (such as audience regional distribution and user attribute groups). In the time dimension, a sliding window technique (window length of 1 minute, sliding step of 10 seconds) is used to capture the frequency and amplitude of emotional fluctuations, identifying fine-grained time nodes of emotional ups and downs. In the spatial dimension, clustering algorithms (such as K-means clustering based on user geographical and social attributes) are used to analyze the emotional heterogeneity among different user groups. Based on this spatiotemporal data, a Hidden Markov Model (HMM) is used to track emotional state transitions, obtaining emotional fluctuation tracking paths and depicting the evolution trajectory of emotional states. In the experimental phase, spatiotemporal analysis was performed on the interaction data of more than 30,000 users in 20 live streams. The emotional fluctuation tracking paths can accurately reflect the changes in user emotions caused by important speeches and performances by the anchor, with a path detection accuracy rate of 85%. Emotional fluctuations are often accompanied by sudden emotional shifts, which typically correspond to key events in the live stream content or drastic changes in the user's psychological state. The system employs a combination of CUSUM (Cumulative and Control Graph) and Bayesian change point detection to detect abrupt shifts in the emotional fluctuation path. CUSUM sensitively captures significant deviations in the mean by analyzing cumulative differences in emotional values; Bayesian change point detection dynamically determines the probability distribution of change points using a probabilistic model, enhancing robustness. Detected abrupt shift time points are marked as "emotional transition points," and the system provides semantic annotations based on live stream event logs (e.g., "explosive user anger," "emotional stimulation," etc.). In the experimental parameter settings, the CUSUM threshold was set to 0.3, and the Bayesian change point detection sampling was 1000 times to ensure a low false positive rate (<7%) and high sensitivity in capturing emotional shifts. Verification showed that the detection of emotional transition time points matched the climax of the live stream with 89% accuracy, effectively supporting subsequent adjustments to interaction strategies. User emotional shifts are often accompanied by a surge in bullet comments or the appearance of bullet comments with specific emotions. The system uses emotional transition time points as anchors, selecting bullet comment streams within ±30 seconds before and after each point. It employs statistical learning methods (such as bullet comment density calculation within the time window and TF-IDF keyword extraction) to identify the set of triggering bullet comments. For the bullet comment content, the system applies a multi-label sentiment classification model (based on a sentiment classifier combining TextCNN and BiLSTM) for fine-grained sentiment annotation (such as anger, joy, and anticipation), selecting bullet comments with high emotional resonance as "high-probability triggering bullet comments." Experiments show that in 15,000 bullet comments during emotional transition periods, the model achieves an accuracy rate of 83% in identifying high-probability triggering bullet comments, effectively reflecting emotional fluctuations at key moments in the live stream. To reveal the deeper mechanisms of user emotional transitions, the system performs sequence pattern mining on high-probability triggering bullet comments.This study employs sequence pattern mining algorithms (such as PrefixSpan) to identify frequent sequences of different emotional tag combinations. It then combines these with association rule mining to analyze the transition probabilities between different emotional states, constructing a model of user emotional transition patterns. Furthermore, by combining user profile information (age, activity level, interest tags) and the semantics of live stream content, causal inference methods (such as Granger causality analysis) are used to verify the triggering factors of emotional transitions. Results show that different user groups exhibit significant differences in their reaction paths to emotional events. Younger users are more likely to rapidly switch between "expectation-incentive-anger," and specific content categories (such as e-sports) are more likely to trigger strong emotional fluctuations. In the experiment, this analysis of data from eight live streams yielded an emotional transition rule coverage rate of 75%, providing a solid foundation for subsequent emotional state modeling. Based on the aforementioned emotional transition patterns and multi-time-point emotional data, the system constructs a multi-time-point emotional state model. A multi-layer temporal graph convolutional network (ST-GCN) combined with an attention mechanism is used to dynamically capture the propagation and evolution of emotional states across time and user groups. The model input includes emotional time series, user group relationship graphs, and interaction behavior features; the output is emotional state scores for different time points and user groups. This model effectively reflects the multi-layered and multi-dimensional evolution of emotions during live streaming, providing emotional state map support for virtual human decision-making. The model training used 50,000 time-series emotional data points, achieving an 82% accuracy rate in emotion prediction and a 78% recall rate on the validation set, demonstrating strong generalization ability. Combining demand gap filling information with user emotional state maps, the system designs a multi-strategy decision engine to dynamically adjust the AI ​​virtual human's behavior. The decision engine uses a reinforcement learning framework (such as Deep Q-Learning), taking emotional state scores and demand gap priority as state inputs, to learn the optimal interactive actions (verbal responses, facial expressions, content switching, etc.) to maximize user emotional satisfaction and engagement. During the decision-making process, emotional feedback is monitored in real time, and action strategies are adjusted to form a closed-loop control. In experiments, strategy training was based on 100 live streaming interaction histories. After offline simulation and online A / B testing, intelligent interaction tasks improved user interaction satisfaction by 14%, increased bullet screen activity by 18%, and increased gift-giving by 12%. This dynamic adjustment mechanism effectively achieves emotional resonance and intelligent collaboration between the AI ​​virtual human and users, significantly optimizing the live streaming interactive experience.

[0031] In this embodiment, the specific steps for making virtual human interaction decisions based on demand gap filling information and user emotional state map, and for dynamically adjusting live broadcast interaction to execute intelligent live broadcast interaction operations are as follows: Decision-making based on virtual human emotional resonance based on user emotional state profile, generating virtual human emotional resonance features; Based on the characteristics of virtual human emotional resonance, the real-time action frequency, tone speed and facial micro-expression parameters of virtual human are defined to obtain the virtual human emotional resonance mode. Based on the blank information to fill in the gaps in the requirements, make virtual human interaction decisions and generate real-time information on virtual human interaction. Dynamic live-stream interaction control is implemented to regulate real-time information and emotional resonance patterns of virtual human interactions in order to execute intelligent live-stream interaction tasks.

[0032] In this embodiment, the user emotion state map is a dynamic expression of emotional states across multiple time points and groups, directly reflecting the overall emotional atmosphere of the live stream and specific user groups. The system achieves synchronization and resonance between the virtual human and user emotions by constructing an emotion resonance decision model. Specifically, the emotion state map data is input into a multimodal emotion matching network based on an attention mechanism. This network integrates user emotion distribution, interaction frequency, and content context to assess the emotional state the virtual human should respond to in the current live stream emotional environment. The output virtual human emotion resonance features include emotional intensity (e.g., happiness, surprise, concern), emotional color (positive, neutral, negative), and emotional tone (warm, excited, calm). Model training utilizes a historical live stream emotion dataset (covering over 100 live streams, including millions of bullet comments and likes), and cross-validation ensures an emotion resonance prediction accuracy of over 85%. This feature lays the foundation for the subsequent generation of virtual human actions and expressions. Converting the emotion resonance features into specific virtual human expressions is a crucial step in achieving natural interaction. The system is designed with an emotion mapping module, mapping resonance features to virtual human action frequencies (such as gesture frequency and nod frequency), speech rate (adjustable to 120-180 words per minute), and facial micro-expression parameters (eyebrow raising, mouth corner raising, etc., parameters ranging from 0 to 1). This mapping process uses a multilayer perceptron (MLP) combined with a rule engine. The rule engine defines threshold ranges based on different emotion categories. For example, "happiness" corresponds to an action frequency higher than 0.7, a fast speech rate, and a bright facial expression; "concern" corresponds to a lower action frequency, a calmer speech rate, and a delicate facial expression. In the experimental settings, the adjustment accuracy of speech rate was controlled within ±5 words per minute, and the adjustment accuracy of micro-expression parameters reached 0.05, ensuring that the virtual human's expressions were natural and smooth. Through user subjective evaluation questionnaires, the adjusted emotional resonance mode was considered more approachable and realistic by 90% of the test users. Need gap filling information refers to unmet needs or emotional gaps in user interactions during live streaming. The system uses this information as decision input, combining it with the current virtual human emotional resonance mode, and employs a reinforcement learning decision model (based on Deep Q-Network, DQN) to generate real-time virtual human interaction information, including language content suggestions (such as questions and encouragement), action prompts, and emotional expression suggestions. This model is trained in a simulated environment with the goal of maximizing user satisfaction and interaction activity. The latency of real-time interaction information generation is controlled within 100 milliseconds to ensure smooth live streaming. During the experiment, the model was fine-tuned online in multiple live streaming environments, improving the relevance and adaptability of interactive responses by more than 15%, significantly enhancing user experience and interaction depth. Finally, the system integrates real-time virtual human interaction information and the emotional resonance mode to construct a dynamic interaction control module, continuously monitoring user feedback and changes in live streaming emotions, and adjusting virtual human behavior in real time. The control module adopts a closed-loop control system architecture, with feedback signals including real-time bullet screen sentiment analysis, changes in likes and gifts, and other indicators.Based on multi-objective optimization algorithms (such as genetic algorithms), the system dynamically balances action diversity, emotional matching, and user satisfaction, adjusting the virtual human's action combinations, tone of voice, and facial expressions. The adjustment cycle is typically set to 500 milliseconds to ensure real-time and natural responses. Online A / B testing results show that this dynamic adjustment mechanism increased average user dwell time by 12%, interaction frequency by 18%, and gift conversion rate by 10%, indicating that intelligent interactive operations significantly improve the quality and commercial value of live streaming interactions.

[0033] In this embodiment, an intelligent live streaming interaction system based on AI virtual humans is provided, used to execute the intelligent live streaming interaction method based on AI virtual humans as described above, including: The interaction perception module is used to identify real-time live interactive data streams, perform intelligent bullet screen filtering and multimodal user interaction perception, and construct a multimodal user interaction perception map. The demand prediction module is used to extract key bullet comments based on the multimodal interaction perception map and predict user demand, thereby generating user interaction demand features. The demand gap filling module is used to perform deep semantic space mapping on user interaction demand features, and then perform reverse live broadcast semantic gap analysis to obtain demand gap filling information. The emotion perception module is used to perform dynamic analysis of user behavior based on a multimodal user interaction perception map, and to fit the global evolution of real-time interactive emotions to construct a real-time interactive emotion map. The interactive control module is used to make virtual human interaction decisions based on the blank information to fill in the gaps and the real-time interactive emotion map, and to perform dynamic live broadcast interaction control in order to execute intelligent live broadcast interaction operations.

[0034] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.

[0035] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein are implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

Claims

1. A method for intelligent live streaming interaction based on AI virtual humans, characterized in that, Includes the following steps: Identify real-time live streaming interactive data streams, perform intelligent bullet screen filtering and multimodal user interaction perception, and construct a multimodal user interaction perception map; Key bullet comments are extracted based on the multimodal interaction perception map, and user needs are predicted to generate user interaction needs features. Deep semantic space mapping is performed on the characteristics of user interaction needs, and then reverse live streaming semantic gap analysis is performed to obtain information to fill gaps in the needs. Based on the multimodal user interaction perception map, dynamic analysis of user behavior is performed, and real-time interaction emotion global evolution is fitted to construct a real-time interaction emotion map. Based on the blank information and real-time interactive emotion map, the virtual human interaction decision is made, and dynamic live interaction control is carried out to execute intelligent live interaction tasks. The specific steps for constructing a real-time interactive sentiment map by performing dynamic analysis of user behavior based on a multimodal user interaction perception map and fitting the global evolution of real-time interactive sentiment are as follows: Perform dynamic analysis of user behavior on the time-series curve of "like" behavior to extract user behavior trajectory; Identify the peak intervals, durations, and fluctuation frequencies of user behavior trajectories to obtain dynamic user behavior characteristics; The user interaction density fluctuation evolution is analyzed based on the dynamic behavior characteristics of users to generate a user interaction density fluctuation map. By fitting the user interaction density fluctuation map and real-time gift features to the global evolution of real-time interaction sentiment, a real-time interaction sentiment map is constructed.

2. The intelligent live streaming interaction method based on AI virtual human according to claim 1, characterized in that, The specific steps for identifying real-time live interactive data streams, performing intelligent bullet screen filtering and multimodal user interaction perception, and constructing a multimodal user interaction perception map are as follows: Identify real-time live interactive data streams, including bullet screen streams, like behavior data, and gift-giving data; Intelligent bullet comment filtering is applied to the bullet comment stream to construct an intelligent filtered bullet comment stream; Extract the timestamp of the like behavior data; Based on the timestamp, a time-series behavior analysis is performed to construct a time-series curve of the "like" behavior. Analyze the gift-giving data by the number of times gifts are given and the types of gifts to obtain real-time gift characteristics; Based on real-time gift characteristics, time-series curves of like behavior, and intelligent filtering of bullet screen streams, multimodal user interaction perception is performed, and a multimodal user interaction perception map is constructed.

3. The intelligent live streaming interaction method based on AI virtual human according to claim 2, characterized in that, The specific steps for performing intelligent bullet comment filtering on the bullet comment stream to construct an intelligent filtered bullet comment stream are as follows: The bullet screen stream is preprocessed by word segmentation and stop word removal, and the preprocessed bullet screen stream is extracted. Perform deep content recognition on the pre-processed bullet comment stream to extract abnormal and illegal bullet comments; Detect negative sentiment based on pre-processed bullet screen stream and label bullet screens with negative sentiment. Calculate the frequency of the negative emotional comments and identify the users who posted them; Based on the frequency of the barrage users, abnormal barrage inferences are made, and abnormal barrage barrages and user IDs are marked. Intelligent bullet screen filtering is performed based on abnormal and inappropriate bullet screen comments, abnormal spam bullet screen comments, and user IDs to build an intelligent filtered bullet screen stream.

4. The intelligent live streaming interaction method based on AI virtual human according to claim 1, characterized in that, The specific steps for extracting key bullet comments based on a multimodal interaction perception map and predicting user needs to generate user interaction demand features are as follows: Key bullet comments are extracted based on a multimodal interaction perception map and processed in multiple time segments to obtain bullet comment streams for different time periods. Deep semantic analysis was performed on the bullet screen stream to obtain the semantic features of bullet screens at different time periods; The frequency of the bullet comments is calculated and high-frequency bullet comments are identified. Based on the high-frequency bullet comments and the semantic features of the bullet comments, user interaction intent analysis is performed to obtain user interaction intent patterns. Predict user needs based on user interaction intent patterns to generate user interaction need characteristics.

5. The intelligent live streaming interaction method based on AI virtual human according to claim 1, characterized in that, The specific steps for performing deep semantic space mapping on user interaction demand features and then performing reverse live streaming semantic gap analysis to obtain demand gap filling information are as follows: Deep semantic space mapping is performed on the user interaction demand features to obtain the user demand feature space; Deeply mine potential user needs from the user demand feature space to extract potential user demand features; Predict future bullet comments based on users' potential demand characteristics, and generate bullet comment streams for future time periods; High-frequency danmaku semantic clustering is performed on the danmaku stream in the future time period, and reverse live streaming semantic gap analysis is performed to generate demand gaps; Based on the gaps in demand, virtual human interaction instructions are defined to fill in the gaps in demand, so as to obtain the information for filling in the gaps in demand.

6. The intelligent live streaming interaction method based on AI virtual human according to claim 1, characterized in that, The specific steps for making virtual human interaction decisions based on demand-filling information and real-time interactive sentiment maps, and for dynamically adjusting live-stream interaction to execute intelligent live-stream interaction tasks are as follows: Multi-dimensional spatiotemporal analysis is performed on the real-time interactive sentiment graph, and sentiment fluctuation tracking analysis is conducted to extract the user's sentiment fluctuation tracking path; Detect abrupt changes in user emotions by tracking the user's emotional fluctuations and mark the time nodes of user emotional transitions; Based on the aforementioned user emotional shift time points, high-probability triggering of bullet comments is identified; The user sentiment shift was mined from the high-probability triggered bullet comments to obtain the deep-seated patterns of user sentiment shifts. Based on the deep emotional transformation patterns of users, multi-time point emotional state modeling is performed to construct a user emotional state map; Based on the information to fill in the gaps in demand and the user's emotional state, the system makes virtual human interaction decisions and dynamically adjusts the live broadcast interaction to execute intelligent live broadcast interaction tasks.

7. The intelligent live streaming interaction method based on AI virtual human according to claim 1, characterized in that, The specific steps for making virtual human interaction decisions based on demand gap filling information and user emotional state map, and for dynamically adjusting live broadcast interaction to execute intelligent live broadcast interaction operations are as follows: Decision-making based on virtual human emotional resonance based on user emotional state profile, generating virtual human emotional resonance features; Based on the characteristics of virtual human emotional resonance, the real-time action frequency, tone speed and facial micro-expression parameters of virtual human are defined to obtain the virtual human emotional resonance mode. Based on the blank information to fill in the demand, make virtual human interaction decisions and generate real-time virtual human interaction information; Dynamic live-stream interaction control is implemented to regulate real-time information and emotional resonance patterns of virtual human interactions in order to execute intelligent live-stream interaction tasks.

8. An intelligent live streaming interactive system based on AI virtual humans, characterized in that, The method for performing intelligent live streaming interaction based on AI virtual human as described in claim 1 includes: The interaction perception module is used to identify real-time live interactive data streams, perform intelligent bullet screen filtering and multimodal user interaction perception, and construct a multimodal user interaction perception map. The demand prediction module is used to extract key bullet comments based on the multimodal interaction perception map and predict user demand, thereby generating user interaction demand features. The demand gap filling module is used to perform deep semantic space mapping on user interaction demand features, and then perform reverse live broadcast semantic gap analysis to obtain demand gap filling information. The emotion perception module is used to perform dynamic analysis of user behavior based on a multimodal user interaction perception map, and to fit the global evolution of real-time interactive emotions to construct a real-time interactive emotion map. The interactive control module is used to make virtual human interaction decisions based on the blank information to fill in the gaps and the real-time interactive emotion map, and to perform dynamic live broadcast interaction control in order to execute intelligent live broadcast interaction operations.

Citation Information

Patent Citations

  • Live broadcast interaction method and system based on AI digital human

    CN119071521A

  • Live broadcast system based on AI interaction

    CN120091164A