Knowledge interaction platform and method based on multi-modal emotion and message system
By adopting NATS message middleware and multi-modal emotion recognition technology in the knowledge interaction platform, the problem of multi-modal data collaborative processing in high-concurrency scenarios is solved, real-time emotional feedback and multi-person collaborative learning are realized, the personalization and stability of user interaction is improved, and it is suitable for a variety of high-concurrency scenarios.
Patent Information
- Application Number
- CN202510315962.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing knowledge interaction platform lacks the ability to collaborate multimodal data in high concurrency scenarios, and is difficult to support real-time emotional feedback and multi-person collaborative learning. Traditional message middleware such as RabbitMQ has high latency and high loss rate in large-scale user scenarios, which cannot meet users' needs for real-time interaction and personalized learning.
NATS message middleware is used as the communication framework, combined with multi-modal emotion recognition technology, a multi-dimensional sentiment analysis model is built, high concurrent connection is achieved through NATS publish/subscribe mode, emotional scheduling algorithms and virtual role dynamic scheduling strategies are designed, microservice modules are built to support rapid iteration and function expansion, and a dual mechanism of automatic generation and manual takeover is adopted when the knowledge base is insufficient.
It realizes the coherence and real-time nature of multi-modal emotional interaction, reduces message delay and loss, enhances the personalization and stability of user interaction, supports millions of concurrent connections, and ensures the continuity and efficiency of services.
Smart Images

Figure CN120255691A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of knowledge interaction platform implementation, and particularly to a knowledge interaction platform and method based on a multimodal emotion and message system. Background Art
[0002] With the in-depth global digital transformation, the way of knowledge acquisition is undergoing a structural change, accelerating from traditional offline education to an online and fragmented learning mode. However, there are still significant deficiencies in the interactive experience of knowledge payment platforms. The fixed mode of one-way video on demand and text comments is difficult to meet the deep-seated demands of users for real-time interaction and personalized learning scenarios. The lack of a two-way communication channel for immediate questions generated during the learning process directly leads to a reduction in the knowledge absorption efficiency; at the same time, platforms generally ignore the emotional transmission of non-verbal elements such as intonation and micro-expressions, making the emotional value of knowledge unable to effectively reach users, further weakening the learning stickiness. This lack of interaction dimension is becoming the key bottleneck restricting the sustainable development of the industry.
[0003] Multimodal interaction technology provides a new path for the experience upgrade of knowledge payment platforms. By integrating multi-dimensional information such as voice, text, and vision, this technology simulates the emotional expression and intention understanding mechanisms in human natural interaction, thereby enhancing the immersive interaction between users and the platform. For example, voice emotion recognition technology can real-time analyze the user's intonation, speech rate, and pause characteristics, accurately identify emotional states such as confusion or excitement, and dynamically adjust the knowledge explanation rhythm or supplement relevant cases, making the learning process more emotionally resonant and personalized. However, the implementation of the technology still faces two challenges: on the one hand, existing research mainly focuses on single-modal optimization in laboratory environments, lacking a systematic solution for the collaborative processing of multimodal data in high-concurrency scenarios; on the other hand, the surge in the user scale has made hundred-thousand-level concurrent requests normal. Due to architectural limitations, traditional message middleware such as RabbitMQ has a message delay of more than 800 milliseconds and a loss rate of more than 1.5%, making it difficult to support strong interaction functions such as real-time emotional feedback and multi-person collaborative learning.
[0004] Therefore, it is necessary to propose an emotional knowledge interaction platform and method that integrates multimodal interaction technology and an efficient message system, which can establish a collaborative model for multimodal emotion recognition and a high-concurrency system, break through the application limitations of emotional computing theory in distributed scenarios, provide a dynamic decision-making method for human-computer interaction research, achieve the efficient collaboration of multimodal data streams, and provide a reusable technical framework and practical path for the industry to build the next-generation knowledge service system with emotional temperature and real-time response capabilities. Summary of the Invention
[0005] In view of this, the present invention provides a knowledge interaction platform and method based on a multi-modal emotion and message system, which is used to solve the technical problems that most current knowledge interaction platforms focus on single-modal optimization in a laboratory environment, lack the collaborative processing of multi-modal data in high-concurrency scenarios, and are difficult to support strong interaction functions such as real-time emotion feedback and multi-person collaborative learning.
[0006] To achieve the above technical objectives, the present invention adopts the following technical solutions:
[0007] In the first aspect, the present invention provides a knowledge interaction platform based on a multi-modal emotion and message system. The platform uses the NATS message middleware as the communication framework, including an interaction module, and a user terminal module and a data module connected to the interaction module:
[0008] The user terminal module is used to receive multi-channel interaction requests from users. After authenticating the users' identities, it distributes the users to the knowledge interaction rooms based on the load balancing mechanism, updates the room status and user list in real time, and interacts with the users through the NATS message system;
[0009] The interaction module is used to create and manage virtual characters based on the cloud platform, and dynamically adjust the behaviors and interaction strategies of the virtual characters to achieve real-time interaction with users in high-concurrency scenarios;
[0010] The data module is used to analyze the users' emotional states in real time using multi-modal emotion recognition technology, construct and store knowledge base data, call the content of the knowledge base according to the users' emotional states, and provide basic data for the interaction strategies of the interaction module and the behaviors of the virtual characters.
[0011] Further, the user terminal module includes a front-end layer and an access layer;
[0012] The front-end layer is used to connect front-end customers and back-end customers connected through multiple channels, and receive multi-modal data inputs from each user;
[0013] The access layer is used to authenticate the users' identities, allocate user requests according to the connection numbers of each current room, obtain multi-modal recognition results and emotion analysis data information based on the NATS server, complete real-time interaction with the users, and ensure that each user obtains a complete interaction experience.
[0014] Further, the interaction module includes a business layer and a platform layer;
[0015] The business layer is used to classify user requests by topic based on the publish / subscribe mode of the NATS message system, send the classified user requests to the data module for processing, and route the processed interaction information to the corresponding users through virtual characters;
[0016] The platform layer is used to manage virtual human characters, interactive manual broadcast control, and interactive conversations, and is also used to configure the service menu, role permissions, security authentication, organization management, and operation logs of the interactive room, providing support for the operation of the platform.
[0017] Furthermore, the service layer includes a broadcast control console;
[0018] The broadcast control console includes a message distribution center module and a command center module;
[0019] Among them, the message distribution center module is used to classify user requests by topic through the publish / subscribe mode of NATS and dynamically route them to the data module for processing; the command center module is used to monitor the platform status in real time, dispatch virtual characters to go on and off stage through the role command library, and trigger the manual intervention mechanism; when the data module feedbacks complex requests not covered by the knowledge base, it automatically switches to the takeover mode and generates context-compliant answers through emotional interaction to ensure that user requests can be effectively responded to.
[0020] Furthermore, the data unit includes an engine layer and a storage layer;
[0021] The engine layer includes intelligent voice, emotion recognition, multimodal fusion, intelligent decision-making, and emotional interaction technology engines, which are used to perform emotion recognition on the received requests, perform multimodal fusion according to emotions, match and respond with the knowledge base, generate corresponding interactive content, and transmit it to the interaction module;
[0022] The storage layer is used to manage data storage using a hybrid storage solution that combines relational databases and non-relational databases.
[0023] Furthermore, the data unit includes an engine layer and a storage layer;
[0024] The engine layer includes intelligent voice, emotion recognition, multimodal fusion, intelligent decision-making, and emotional interaction technology engines, which are used to perform emotion recognition on the received requests, perform multimodal fusion according to emotions, match and respond with the knowledge base, generate corresponding interactive content, and transmit it to the interaction module;
[0025] The storage layer is used to manage data storage using a hybrid storage solution that combines relational databases and non-relational databases.
[0026] On the other hand, the present invention also provides a multimodal emotion knowledge interaction method, which is implemented by using the knowledge interaction platform based on multimodal emotion and message system described in the above technical solution, including:
[0027] Receiving the multi-channel interaction requests of users through the user terminal module and pushing them to the interaction module using the NATS message queue;
[0028] Classify user requests by theme through the interaction module, and send the theme requests and interaction data to the data module in real time;
[0029] Through the data module, combine voice emotion recognition with user historical behavior data, dynamically adjust the expression mode of the response content, retrieve the structured knowledge base, match the preset Q&A templates and standard answers, and convert the interaction content into emotional voice output;
[0030] Through the interaction module, use the NATS message queue to output voice data, monitor the system status, schedule the virtual character to go on and off stage, and transmit the interaction content to the user terminal module.
[0031] Furthermore, the method further includes:
[0032] For requests that fail to match after retrieving the structured knowledge base, use the ASR system to convert the interaction content into voice, and use the TTS system to output the voice of the virtual character through the emotion scheduling algorithm to ensure that the interaction is not interrupted.
[0033] Furthermore, the emotion scheduling algorithm is implemented based on principal component analysis and hidden Markov model, and is used for efficient recognition and dynamic adjustment of the user's emotional state.
[0034] Furthermore, the method further includes:
[0035] Verify the user's permissions according to the user's interaction request; if the verification is successful, assign the user to the corresponding interaction room, and send the basic information of the room to the user; the basic information of the room at least includes interaction basic information, course basic information, teacher information and NATS basic information;
[0036] When the user connects to NATS, determine whether the user is a lecturer or a teaching assistant. If the user is a lecturer or a teaching assistant, prompt the user to connect to Agora; otherwise, continue to push the basic information of the room;
[0037] Configure an interaction interface in the user's message display area to ensure that the user can obtain interaction information in real time;
[0038] Messages in the interaction room are transmitted through the NATS publish and subscribe mechanism; the chat message types include text, pictures, gift giving and muting messages.
[0039] Furthermore, the method further includes:
[0040] When the knowledge base query fails continuously for several times, perform an artificial interaction response through the ASR-TTS link.
[0041] Compared with the prior art, the knowledge interaction platform and method based on the multi-modal emotion and message system proposed by the present invention have the following advantages:
[0042] (1) The present invention breaks through the limitations of single-modal interaction in traditional knowledge interaction platforms and deeply integrates multi-modal emotion interaction technology. By constructing a multi-dimensional emotion analysis model, integrating voice signals (intonation, speech rate), text semantics, and interactive scenario features, cross-modal emotion recognition and dynamic feedback are achieved, enabling strict synchronization of virtual character voice output, facial expressions and movements, and interface feedback, solving the problem of fragmentation in traditional multi-modal interaction and enhancing the coherence of interaction.
[0043] (2) Aiming at the message delay and reliability bottlenecks in the large-scale user scenarios of knowledge payment platforms, the present invention achieves an innovative breakthrough in the message middleware architecture. The NATS publish / subscribe mode is adopted to replace traditional message queues (such as RabbitMQ), and a lightweight and high-throughput communication framework is constructed. A single server supports millions of concurrent connections. In addition, based on the NATS Stream mode, interactive data is persisted to distributed storage, supporting message backtracking and fault recovery, achieving a zero data loss rate in high-concurrency scenarios and verifying its reliability advantages.
[0044] (3) The present invention drives the interaction logic with emotional states and user behavior data, and designs a dynamic scheduling strategy for virtual characters based on emotional labels. For example, for users with negative emotions, encouraging characters (such as "affable senior") are preferentially matched, and their interaction styles (slower speech rate, softer intonation) are adjusted.
[0045] (4) The present invention realizes rapid iteration and function expansion through decoupled design and open interfaces. Modules such as multi-modal engines, message distribution, and knowledge bases are separated into microservices to build a microservice-based core module, supporting horizontal expansion on demand. At the same time, standardized interfaces are provided to access external AI models and data analysis tools, enabling open APIs to be integrated with third parties and expanding the system's ability boundaries.
[0046] (5) Through the "automatic generation + manual takeover" dual-mode guarantee mechanism, the problem of service continuity in complex scenarios is solved. When requests exceed the scope of the knowledge base, context-aware answers are generated immediately to achieve automated generalization responses and avoid interaction interruptions. At the same time, teaching assistants can take over abnormal requests in real time through the broadcast control console, and use ASR / TTS technology to convert manual input into virtual character responses, effectively ensuring the continuity and stability of services.
[0047] In summary, through the technical coupling of multi-modal emotion understanding and high-concurrency communication, the present invention reshapes the intelligent paradigm of knowledge interaction, providing a reusable technical framework and practical path for the industry to build the next-generation knowledge service system with emotional warmth and real-time response capabilities. Brief Description of the Drawings
[0048] Figure 1 It is an application schematic diagram of the knowledge interaction platform provided by the present invention;
[0049] Figure 2 Schematic diagram of the structure of the knowledge interaction platform based on the multi-modal emotion and message system provided by the present invention;
[0050] Figure 3 Schematic diagram of the knowledge interaction time sequence provided by the present invention;
[0051] Figure 4 Schematic diagram of the specific implementation process of the emotion fallback mechanism provided by the present invention;
[0052] Figure 5 Schematic diagram of the message subscription process provided by the present invention. Detailed implementation manners
[0053] The following will specifically describe the preferred embodiments of the present invention with reference to the accompanying drawings. The accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to explain the principle of the present invention, rather than to limit the scope of the present invention.
[0054] To better illustrate the application scenario of the present invention, first, the main business processes of the multi-modal emotion knowledge interaction of this application will be introduced. As Figure 1 shown, Figure 1 shows the business flow chart of the knowledge interaction platform. The knowledge interaction platform is mainly targeted at users who are keen on reading various types of information and knowledge and are willing to share their opinions. The design goal of the emotion knowledge interaction platform is to transmit emotions and knowledge through multiple modalities, ensuring that while knowledge is being transmitted, the emotional interaction needs of users are deeply met.
[0055] Users need to first register an account and complete the login authentication. After logging in, they can enter the corresponding knowledge interaction room according to their personal interests and hobbies, and comment, like, and give gifts to the content they are interested in. Users can initiate a request to go on air according to their own needs and participate in the interaction in the platform room. As the operator of the system's broadcast control console, the teaching assistant precisely controls the virtual character, calls the content in the knowledge base to interact with the user in terms of emotion and knowledge. When the content in the knowledge base is insufficient to meet the user's needs, the teaching assistant will act as the "person inside" role, using the automatic speech recognition (ASR) and text-to-speech (TTS) systems to convert into virtual human voice output to ensure the coherence and consistency of the interaction, and always ensure that the platform continuously outputs personalized content that meets the users.
[0056] Embodiment 1
[0057] Please refer to Figure 2 , this embodiment provides a knowledge interaction platform based on the multi-modal emotion and message system. The platform uses the NATS message middleware as the communication framework, including an interaction module, and a user terminal module and a data module connected to the interaction module:
[0058] The client module is used to receive multi-channel interaction requests from users. After authenticating the users, it allocates the users to the knowledge interaction room based on the load balancing mechanism, updates the room status and user list in real time, and interacts with the users through the NATS message system;
[0059] The interaction module is used to create and manage virtual characters based on the cloud platform, and dynamically adjust the behaviors and interaction strategies of the virtual characters to achieve real-time interaction with users in high-concurrency scenarios;
[0060] The data module is used to analyze the user's emotional state in real time using multi-modal emotion recognition technology, construct and store knowledge base data, and call the knowledge base content according to the user's emotional state to provide basic data for the interaction strategies of the interaction module and the behaviors of the virtual characters.
[0061] The knowledge interaction platform provided by this embodiment combines multi-modal emotion recognition technology, NATS message middleware, and load balancing mechanism to create an efficient, intelligent, and flexible knowledge interaction platform. It can meet the user's information needs while improving the user experience and participation through emotional interaction, thus realizing more personalized, real-time, and efficient services. This platform is applicable to various high-concurrency scenarios and has great application potential and development space.
[0062] As a preferred embodiment, the client module includes a front-end layer and an access layer;
[0063] The front-end layer is used to access front-end customers and back-end customers connected through multiple channels, and receive multi-modal data inputs from each user;
[0064] The access layer is used to authenticate the users and then allocate user requests according to the connection number of each current room. Based on the NATS server, it obtains multi-modal recognition results and emotion analysis data information, completes real-time interaction with the users, and ensures that each user obtains a complete interaction experience.
[0065] As a specific embodiment, the front-end layer interacts with users in various forms such as PC, WeChat mini-program, H5 / APP, and third-party systems, receives inputs from different users, and displays the outputs of the platform. Among them, front-end users can initiate operations such as "going on air", "going off air", "sending messages", and "entering the room" through the client, while back-end users can perform operations such as "managing users", "managing virtual human characters", "automatic broadcast control management", and "dialogue management" through the client.
[0066] The access layer ensures that user requests can be efficiently and securely transmitted to the back-end system through mechanisms such as load balancing, identity authentication, request routing, and security protection.
[0067] As a preferred embodiment, the interaction module includes a business layer and a platform layer;
[0068] The service layer is used to classify user requests by topic based on the publish / subscribe mode of the NATS messaging system, send the classified user requests to the data module for processing, and route the processed interaction information to the corresponding users through virtual characters;
[0069] The platform layer is used to manage virtual human characters, interactive manual broadcast control, and interactive conversations, and is also used to configure service menus, role permissions, security authentication, organization management, and operation logs for interactive rooms to provide support for platform operation.
[0070] As a preferred embodiment, the service layer includes a broadcast control console;
[0071] The broadcast control console includes a message distribution center module and a command center module;
[0072] Among them, the message distribution center module is used to classify user requests by topic through the publish / subscribe mode of NATS and dynamically route them to the data module for processing; the command center module is used to monitor the platform status in real time, dispatch virtual characters to go on and off stage through the role command library, and trigger the manual intervention mechanism; when the data module feedbacks complex requests not covered by the knowledge base, it automatically switches to the takeover mode and generates contextually appropriate answers through emotional interaction to ensure that user requests can be effectively responded to.
[0073] Furthermore, after receiving the classified user requests, the data module first performs knowledge base matching, preferentially retrieves the structured knowledge base, and matches the preset Q&A templates and standard answers. If the knowledge base cannot meet the requests, the system will dynamically adjust the expression way of the response content by combining voice emotion recognition and user historical behavior data to adapt to the emotional needs of users (such as using authoritative or humorous voice synthesis). For requests that cannot be matched, the AI large model will generate contextually relevant replies, which will be converted into emotional voice output through the TTS engine after manual review.
[0074] Such as Figure 3 shown, Figure 3 shows the knowledge interaction timing diagrams for both automatic and manual intervention states.
[0075] As a preferred embodiment, the data unit includes an engine layer and a storage layer;
[0076] The engine layer includes intelligent voice, emotion recognition, multi-modal fusion, intelligent decision-making, and emotional interaction technology engines, which are used to perform emotion recognition on the received requests, perform multi-modal fusion according to emotions, match and respond with the knowledge base, generate corresponding interaction content, and transmit it to the interaction module;
[0077] The storage layer is used to store and manage data using a hybrid storage solution that combines a relational database and a non-relational database.
[0078] With the above settings, the front-end layer and the access layer interact with the business layer through user requests. When processing user requests, the business layer calls the services and functions of the platform layer, and the platform layer depends on the technical support of the engine layer and the data storage of the data layer. Through the close connection and collaborative work among all parts, the efficient operation and function realization of the emotional knowledge interaction platform are ensured.
[0079] Embodiment 2
[0080] The present invention also provides a multi-modal emotional knowledge interaction method, which is implemented by using the knowledge interaction platform based on the multi-modal emotion and message system described in any one of the above technical solutions, and includes:
[0081] Receiving a multi-channel interaction request from a user through a user-end module, and pushing it to an interaction module by using a NATS message queue;
[0082] Classifying the user requests by theme through the interaction module, and sending the theme requests and interaction data to a data module in real time;
[0083] Combining voice emotion recognition and user historical behavior data through the data module, dynamically adjusting the expression mode of the response content, retrieving a structured knowledge base, matching a preset Q&A template and standard answer, and converting the interaction content into an emotional voice output;
[0084] Outputting voice data by using a NATS message queue through the interaction module, monitoring the system status, scheduling virtual characters to go on and off stage, and transmitting the interaction content to the user-end module.
[0085] As a preferred embodiment, the method further includes:
[0086] If a request that fails to match after retrieving the structured knowledge base is received, the interaction content is converted into voice by using an ASR system, and the voice of a virtual character is output through a TTS system to ensure that the interaction is not interrupted.
[0087] Specifically, interactive users access the platform through the client SDK and initiate voice or text interaction requests. After the request content is preliminarily processed by the cloud service platform where the interaction module is located, speech recognition (ASR) and emotion feature extraction are completed, providing structured data for subsequent processing. Based on the NATS publish / subscribe mode, the message distribution center classifies user requests by topic (such as knowledge Q&A, emotional interaction, role scheduling) and dynamically routes them to the instructor module, multi-modal engine, or knowledge base for processing. The data module performs knowledge base matching according to the received requests, preferentially retrieves the structured knowledge base, and matches the preset Q&A templates and standard answers. The platform combines voice emotion recognition and user historical behavior data to dynamically adjust the expression mode of the response content to meet the emotional needs of users (such as using authoritative or humorous speech synthesis). For requests that cannot be matched, the AI large model generates context-related responses, which are converted into emotional voice outputs through the TTS engine after manual review.
[0088] As Figure 4 shown, Figure 4 it shows the specific implementation process of the emotion fallback mechanism.
[0089] In the multi-modal interactive plot control framework, the emotion interaction tasks of the teaching assistant not only include regular teaching assistance, Q&A interaction, and emotional voice interaction, but also the function of maintaining a good interaction atmosphere. In addition, another important task of the teaching assistant is to ensure the stability of the system.
[0090] If the knowledge base cannot meet the request, this platform will implement an emotion fallback mechanism to ensure the stability of the platform and the coherence of the interaction. The teaching assistant, as the "person inside", quickly intervenes and acts as the "fallback" role. After the user inputs a question, the system first checks whether it matches the content of the knowledge base. If the match is successful, the script module calls the knowledge base to answer, and the role TTS system generates a voice output; if the match fails, the teaching assistant generates a situation-appropriate answer through emotional interaction, converts it into voice using the ASR system, and then outputs the voice of the virtual character through the TTS system to ensure that the interaction does not interrupt.
[0091] Furthermore, the emotion scheduling algorithm is implemented based on principal component analysis and hidden Markov model, and is used for the efficient recognition and dynamic adjustment of the user's emotional state.
[0092] The following details the specific implementation method of the emotion scheduling algorithm:
[0093] First, use PCA to perform dimensionality reduction processing on the user's multi-modal emotion data and extract key emotion features. Then, model these features through HMM and dynamically adjust the importance of each feature, thereby realizing the optimized scheduling of the virtual character's behavior and interaction strategy.
[0094] In the implementation logic of the emotion scheduling algorithm, it is first necessary to preprocess the multi-modal emotion data, including collecting data such as the user's voice, text, and behavior logs, and extracting emotion features from them, such as emotion intensity and emotion stability. Next, principal component analysis (PCA) is used to reduce the dimension of the extracted emotion features to extract key features. PCA projects the original data into a new coordinate system through a linear transformation to maximize the variance, thereby achieving dimension reduction and feature extraction. The specific mathematical expression is as follows:
[0095] Y = W T X (1)
[0096] where Y represents the data after dimension reduction, W is the projection matrix, and X is the original data.
[0097] Subsequently, the hidden Markov model (HMM) is used to model the emotion features after dimension reduction to dynamically adjust the importance of each feature. HMM is a statistical model suitable for time series data analysis. Its mathematical expression is:
[0098] λ = (A, B, π) (2)
[0099] where A is the state transition matrix, B is the observation matrix, and π is the initial state distribution.
[0100] Finally, based on the modeling results of HMM, dynamic scheduling instructions are generated to guide the behavior and interaction strategies of virtual characters. By dynamically adjusting the importance of each emotion feature, intelligent scheduling of virtual character behavior is achieved, thereby enhancing the personalization and intelligence level of interaction, improving the system resource utilization efficiency and user satisfaction.
[0101] It should be noted that according to the global management and local management of resources, emotion scheduling is divided into large scheduling and small scheduling.
[0102] (1) Design and implementation of large scheduling
[0103] Large scheduling is an important part of the emotion scheduling algorithm based on HMM and PCA. It is mainly oriented to global resource management and macro policy allocation, and focuses on system load balancing and long-term user experience. By dynamically analyzing the user's emotional state and interaction behavior, large scheduling can achieve intelligent scheduling of virtual characters and optimization of interaction strategies, thereby enhancing the overall performance of the system and user satisfaction.
[0104] In large-scale scheduling, the logic of character on-stage and off-stage is a crucial link. Its scheduling goal is to dynamically schedule virtual characters on-stage and off-stage based on the user's emotional intensity and knowledge matching degree, in order to balance resource occupancy and interaction quality. Specifically, when the user's emotional intensity is high (e.g., negative value > 0.7) or the knowledge base matching degree is high (> 0.8), the system will preferentially schedule expert characters (such as "authoritative lecturer") on-stage to provide more professional and targeted interactive services. On the contrary, when a character has no interaction for a continuous timeout (default 300 seconds) or the room load is high (> 80%), the platform will trigger the character off-stage command to release resources and ensure the stable operation of the system. This process can be described by the following mathematical expression:
[0105] S up = α·E neg +(1 - α)·K match (3)
[0106] Where, E neg is the negative emotion score, K match is the knowledge matching degree, and α is the weight coefficient with a value of 0.6.
[0107] The room expansion strategy is another important aspect of large-scale scheduling. When the number of users in a single room exceeds 500 or the emotional conflict (variance of user emotional differences) exceeds 0.4, the platform will use the consistent hashing algorithm
[36] to divert users to new rooms to achieve dynamic room partitioning. In addition, the system will reserve 10% of the computing resources (CPU / GPU) for sudden expansion to ensure that the expansion delay does not exceed 10 seconds. This resource pre-allocation strategy can effectively handle the rapid growth of the number of users and ensure the high availability of the platform and a good experience for users.
[0108] In the implementation process, first, PCA is used to reduce the dimensionality of the user's multi-modal emotion data and extract key emotion features. Then, HMM is used to model these features and dynamically adjust the importance of each feature, thereby achieving intelligent scheduling of virtual character behavior.
[0109] (2) Design and implementation of small-scale scheduling
[0110] Small-scale scheduling focuses on local interaction optimization and improves the immediate user experience through fine-grained strategies. In the emotional knowledge interaction platform, small-scale scheduling realizes the precise improvement of user participation through emotion-driven powder adding and dynamic priority strategies. Its design and verification methods integrate multi-modal analysis and reinforcement learning technologies, significantly exceeding the traditional threshold determination mode.
[0111] Emotion-driven follower addition means: constructing a reinforcement learning model based on Deep Deterministic Policy Gradient (DDPG), using the user's emotional state and historical interaction trajectories as the state space, the follower addition trigger action as the policy space, and the user retention rate and knowledge consumption as the reward function. Through off-policy optimization, dynamically adjust the emotional threshold for triggering follower addition.
[0112]
[0113] At the same time, introduce a multi-modal fusion verification mechanism, combine speech emotion recognition (based on HMM), text sentiment analysis (BERT), and behavioral logs (click-through rate, dwell time) to generate a comprehensive emotion confidence level. Only trigger follower addition when the confidence level > 0.85 and the conditions are met for 3 consecutive interactions to avoid misjudgment by a single modality.
[0114] The dynamic optimization of follower addition priority means:
[0115] Construct a multi-dimensional feature vector, including emotional stability (sliding window variance), interaction frequency (time-decay weighted), knowledge preference (TF-IDF topic distribution), and historical response rate. Use the XGBoost model to predict the user's response probability to follower addition, and the model output is the priority score:
[0116] Priority score = f(emotional stability, interaction frequency, knowledge preference, historical response rate)(5)
[0117] Analyze the feature contribution degree through SHAP values to ensure the interpretability of the policy.
[0118] Taking user satisfaction and knowledge consumption as the optimization objectives, construct a multi-objective optimization function:
[0119] Maximize w1·NPS + w2·knowledge consumption (w1 + w2 = 1)(6)
[0120] Conduct Bayesian hyperparameter search on the emotional stability threshold (variance range 0.05 - 0.2) and interaction frequency threshold (3 - 8 times / day), and use Gaussian process regression to fit the objective function surface to quickly approximate the Pareto optimal solution.
[0121] Furthermore, the output interaction results are distributed through the NATS cluster to ensure real-time and accuracy. The platform supports two modes: incremental synchronization and full synchronization. Incremental synchronization ensures millisecond-level latency by pushing real-time updates of interaction status, such as changes in the user list and virtual character actions. Full synchronization actively pulls the historical data of the room (such as chat records, gift logs) when a new user accesses, and fuses it with the incremental data to construct a complete interaction scenario.
[0122] The monitoring mechanism detects abnormal situations in real time and triggers corresponding processing procedures. For example, when the NATS cluster monitors that the node load exceeds the threshold, the system will automatically expand the message queue and divert requests to avoid service overload. If the emotion tags output by the multimodal engine do not match the context semantics, the command center will initiate a manual review process. For knowledge blind spots, after the knowledge base query fails continuously for a certain number of times (usually set to three times), the system will automatically transfer the conversation to the teaching assistant for manual response through the ASR-TTS link.
[0123] In some embodiments, the process for a user to enter an interactive room includes multiple steps to ensure that the user can smoothly participate in the interaction.
[0124] After the user initiates a request to enter the room, the system verifies the user's permissions. If the verification fails, the user is prompted that the entry into the room has failed; if the verification is successful, the user enters the room and obtains the basic information of the room, including interactive basic information, course basic information, teacher information, NATS basic information, Agora basic information, and Tencent Cloud basic information. The user connects to NATS and obtains the initial information of the room list. The system determines whether the user is a lecturer or a teaching assistant. If the user is a lecturer or a teaching assistant, the user connects to Agora; otherwise, the user continues to obtain the basic information of the room. The message display area includes a user list, a leaderboard, a whiteboard, a microphone order, a questionnaire, a vote, interactive room settings, course promotion, a question list, and on-the-wall views (on-the-wall questions) to ensure that the user can obtain interactive information in real time.
[0125] The transmission mechanism of voice live messages is as follows: The real-time voice messages between the users on stage are directly transmitted through the Agora channel and pushed to Tencent Cloud through the Agora background service; the users not on stage pull the voice live stream through Tencent Cloud to listen to the voice live.
[0126] Chat messages are transmitted through the publish (pub) and subscribe (sub) mechanisms of NATS. The types of chat messages include text, pictures, gift giving, and mute messages. Text messages contain pure text content and emojis; picture messages contain the picture address, and only 1 picture can be sent at a time; gift giving messages contain gift information, the gift giver's information, and the gift recipient's information; mute messages contain the information of the muted person.
[0127] The message system is the basis for system message transmission, responsible for efficient message distribution and synchronization. Its main functions include "publishing messages" and "subscribing to messages", and it also needs to stably handle incremental and full data synchronization to ensure that the interaction information of the client in the room is always consistent.
[0128] After completing the integration of incremental synchronization and full synchronization, the client enters the normal operation stage and dynamically updates the room status by continuously subscribing to NATS messages. Figure 5 Shows the specific process of message subscription. The specific process is as follows:
[0129] Step 1, the client connects to NATS: After the user enters the room, the client establishes a connection with the NATS server. This connection is the basis for message synchronization, ensuring that the client can receive incremental updates in real time.
[0130] Step 2, real-time message connection verification: The system verifies whether the real-time message connection is successful. If the connection fails, the system will retry or prompt an error to ensure the stability of the message connection.
[0131] Step 3, receive real-time messages: After the connection is successful, the client starts to receive real-time messages. These messages include information such as changes in the user list, updates of users on stage, and changes in interactive content. Through incremental synchronization, the client can instantly perceive the latest state of the room, ensuring that users always have the latest interactive dynamics at hand.
[0132] Step 4, request historical data: At the same time, the client requests full historical data through the backend interface, including user behavior records, interactive messages, chat records, and gift-giving information, etc. This full synchronization process provides users with a complete snapshot of the room, helping users comprehensively understand the historical dynamics of the room and providing the necessary context support for subsequent incremental updates to avoid information understanding biases and confusion.
[0133] Step 5, data integration and display: After receiving incremental and full data, the client will perform data integration and deduplication to ensure the uniqueness and accuracy of information. At the same time, the system will calculate the user's knowledge consumption data to accurately analyze the user's interactive and learning behaviors. This process not only improves the efficiency of data processing but also provides rich user behavior analysis data for the system, helping to further optimize the interactive experience.
[0134] The knowledge interaction platform and method based on the multi-modal emotion and message system provided by the present invention can, through multi-modal emotion recognition technology, analyze the user's emotional state in real time, adjust the behavior and interaction strategies of virtual characters, so as to achieve a more user-friendly interactive experience; combined with the NATS message system, in high-concurrency scenarios, virtual characters can reasonably allocate resources for interaction according to the load situation, thus ensuring that the system can continuously and stably operate and provide efficient services; through the dynamic adjustment of the virtual character behavior of the system, users can not only interact with the system in terms of knowledge but also experience more vivid and emotional conversations, enhancing the user's sense of participation. The present invention combines multi-modal emotion recognition technology, NATS message middleware, and load balancing mechanism to create an efficient, intelligent, and flexible knowledge interaction platform, which can, while meeting the user's information needs, improve the user experience and sense of participation through emotional interaction, so as to achieve more personalized, real-time, and efficient services. Such a platform is applicable to a variety of high-concurrency scenarios and has great application potential and development space.
[0135] As described above, it is only the preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A knowledge interaction platform based on a multi-modal emotion and message system, characterized in that, The platform adopts the NATS message middleware as the communication framework, including an interaction module, a client module connected to the interaction module, and a data module: The client module is used to receive multi-channel interaction requests from users. After authenticating the users, it allocates the users to the knowledge interaction room based on the load balancing mechanism, updates the room status and user list in real time, and interacts with the users through the NATS message system; The interaction module is used to create and manage virtual characters based on the cloud platform, and dynamically adjust the behaviors and interaction strategies of the virtual characters to achieve real-time interaction with users in high-concurrency scenarios; The data module is used to analyze the user's emotional state in real time using multi-modal emotion recognition technology, construct and store knowledge base data, call the knowledge base content according to the user's emotional state, and provide basic data for the interaction strategies of the interaction module and the behaviors of the virtual characters.
2. The knowledge interaction platform based on the multi-modal emotion and message system according to claim 1, characterized in that The client module includes a front-end layer and an access layer; The front-end layer is used to access front-end and back-end customers connected through multiple channels and receive multi-modal data inputs from each user; The access layer is used to authenticate the users and then allocate user requests according to the number of connections in each current room. Based on the NATS server, it obtains multi-modal recognition results and emotion analysis data information to complete real-time interaction with the users and ensure that each user obtains a complete interaction experience.
3. The knowledge interaction platform based on the multi-modal emotion and message system according to claim 1, characterized in that, The interaction module includes a business layer and a platform layer; The business layer is used to classify user requests by topic based on the publish / subscribe mode of the NATS message system, send the classified user requests to the data module for processing, and route the processed interaction information to the corresponding users through the virtual characters; The platform layer is used to manage virtual human characters, interactive manual control, and interactive conversations, and is also used to configure service menus, role permissions, security authentication, organizational management, and operation logs for the interaction room to provide support for the platform operation.
4. The knowledge interaction platform based on the multi-modal emotion and message system according to claim 3, characterized in that, The business layer includes a control console; The control console includes a message distribution center module and a command center module; Among them, the message distribution center module is used to classify user requests by topic through the publish / subscribe mode of NATS and dynamically route them to the data module for processing; the command center module is used to monitor the platform status in real time, dispatch virtual characters to go on and off stage through the role command library, and trigger the manual intervention mechanism; when the data module feedbacks complex requests not covered by the knowledge base, it automatically switches to the takeover mode and generates context-compliant answers through emotional interaction to ensure that user requests can be effectively responded to.
5. The knowledge interaction platform based on the multi-modal emotion and message system according to claim 1, characterized in that, The data unit includes an engine layer and a storage layer; The engine layer includes intelligent voice, emotion recognition, multi-modal fusion, intelligent decision-making, and emotional interaction technology engines, which are used to perform emotion recognition on the received requests, perform multi-modal fusion according to the emotions, match and respond with the knowledge base, generate corresponding interaction content, and transmit it to the interaction module; The storage layer is used to store and manage data using a hybrid storage solution that combines relational databases and non-relational databases.
6. A multimodal emotion knowledge interaction method, characterized in that, The implementation of the knowledge interaction platform based on multi-modal emotion and message system according to any one of claims 1-5 includes: Receive the multi-channel interaction requests of users through the client module and push them to the interaction module using the NATS message queue; Classify the user requests by topic through the interaction module and send the topic requests and interaction data to the data module in real time; Through the data module, combine voice emotion recognition and user historical behavior data, dynamically adjust the expression mode of the response content, retrieve the structured knowledge base, match the preset Q&A templates and standard answers, and convert the interaction content into an emotional voice output; Output the voice data through the interaction module using the NATS message queue, monitor the system status, schedule the virtual character to go on and off stage, and transmit the interaction content to the client module.
7. The multimodal emotion knowledge interaction method according to claim 6, characterized in that The method further includes: For requests that cannot be matched after retrieving the structured knowledge base, use the ASR system to convert the interaction content into voice, and use the TTS system to output the voice of the virtual character using the emotion scheduling algorithm to ensure that the interaction is not interrupted.
8. The knowledge interaction platform based on the multi-modal emotion and message system according to claim 7, characterized in that, The emotion scheduling algorithm is implemented based on principal component analysis and hidden Markov model, and is used for the recognition and dynamic adjustment of the user's emotional state.
9. The knowledge interaction platform based on the multi-modal emotion and message system according to claim 6, characterized in that, The method further includes: Verify the user's permissions according to the user interaction request; if the verification is successful, assign the user to the corresponding interaction room and send the basic information of the room to the user; the basic information of the room includes at least interaction basic information, course basic information, teacher information, and NATS basic information; When the user connects to NATS, determine whether the user is a lecturer or a teaching assistant. If the user is a lecturer or a teaching assistant, prompt the user to connect to Agora; otherwise, continue to push the room basic information; Configure an interaction interface in the message display area of the user to ensure that the user can obtain interaction information in real time; The messages in the interaction room are transmitted through the publish and subscribe mechanism of NATS; the chat message types include text, pictures, gift giving, and mute messages.
10. The knowledge interaction platform based on the multi-modal emotion and message system according to claim 6, characterized in that, The method further includes: When the knowledge base query fails for several consecutive times, perform an artificial interaction response through the ASR-TTS link.