Interaction control method and device based on user state perception, equipment and medium

By collecting and processing multimodal data, predicting changes in user emotions and adjusting interaction strategies in real time, the problem of lag in existing systems has been solved, enabling efficient and natural human-computer interaction in fintech and healthcare scenarios.

CN121808696APending Publication Date: 2026-04-07PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-07
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing voice interaction systems are slow to respond to real-time customer feedback and emotional changes, and cannot dynamically adjust the content of the interaction, resulting in mechanical and impersonal communication. This is especially true in scenarios such as financial consultation and medical diagnosis, where the interaction strategy is too simplistic, affecting efficiency and satisfaction.

Method used

Collect multimodal raw data, generate standardized user information, predict emotion change trends, generate preliminary interaction plans, and adjust strategies in real time during the interaction process to optimize emotion change trends and interaction plan generation process, thus constructing a closed-loop optimization system.

Benefits of technology

It achieves real-time adaptive control in complex interactive scenarios, improves response accuracy and interaction consistency, and enhances the naturalness and stability of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121808696A_ABST
    Figure CN121808696A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of financial science and technology, medical health and the like, and discloses an interaction control method, device, equipment and medium based on user state awareness, and the method comprises the steps: collecting multi-modal original data to generate standardized multi-modal user information; predicting the emotion change trend of the user according to the standardized multi-mode user information, and generating a preliminary interaction scheme in combination with the intention demand of the user and the current interaction state; generating a cooperative control strategy based on the preliminary interaction scheme; real-time multi-mode user information is obtained and compared in the process of executing the cooperative control strategy, and when an adjustment condition is triggered, a new interaction scheme is generated and the cooperative control strategy is updated; after interaction is finished, interaction performance is analyzed according to interaction data, and substandard results are optimized. According to the method, user state perception is realized through multi-modal information fusion, an adaptive closed-loop mechanism of emotion prediction and interaction control is constructed, and the response capability and interaction naturalness to a dynamic state are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to an interactive control method, apparatus, device, and medium based on user state perception. Background Technology

[0002] In the fintech business, existing voice interaction systems are typically used for automated outbound calls such as customer follow-ups, financial advice, and payment reminders. These systems largely rely on pre-set scripts and fixed response logic to complete business communication. However, in real-world applications, customer feedback is often immediate and diverse. Existing systems lag behind in handling real-time voice, emotional feedback, and semantic changes, failing to dynamically adjust interaction content based on changes in customer tone, emotional fluctuations, or temporary needs, resulting in a mechanical and impersonal communication process. Especially in scenarios involving financial advice or payment reminders, the customer's emotional state directly impacts communication effectiveness. Existing systems lack an effective coordination mechanism between emotion recognition and response strategies, leading to simplistic interaction strategies and low customer satisfaction. Furthermore, the current outbound call systems suffer from lengthy voice recognition and decision-making processes, with high latency in data transmission between multiple modules, making it difficult for voice robots to quickly generate appropriate scripts for real-time calls, impacting outbound call efficiency and call completion rates.

[0003] In the healthcare sector, voice interaction systems are commonly used in scenarios such as intelligent consultations, health follow-ups, and rehabilitation guidance. Existing technologies generally rely on rule-driven verbal logic, making it difficult to adjust in real-time to individual patient differences and sudden emotional reactions. For example, when patients experience anxiety, resistance, or confusion, traditional systems lack the ability to continuously perceive changes in their emotions, failing to adjust tone, speed, or communication style immediately, easily leading to decreased patient cooperation. Furthermore, in medical settings, patients' voice information, health consultation content, and behavioral feedback are often highly personalized. Existing systems lack sufficient multimodal data fusion and processing mechanisms, failing to create a dynamic linkage between emotional state, intentions, and interactive content, resulting in rigid communication processes, delayed responses, and impacting the medical communication experience and service quality. Summary of the Invention

[0004] The main objective of this invention is to provide an interactive control method, device, equipment, and storage medium based on user state perception, aiming to solve the technical problem that existing technologies cannot achieve dynamic emotion perception and interactive control based on the user's real-time multimodal state information, thus making it difficult to maintain the continuity and adaptability of response in complex and ever-changing interactive processes.

[0005] To achieve the above objectives, the present invention provides an interactive control method based on user state awareness, comprising: Collect and preprocess multimodal raw data generated during user interaction to generate standardized multimodal user information; Based on the standardized multimodal user information, predict the user's emotional change trend, and combine the user's emotional change trend, user's intention and needs with the current interaction state to generate a preliminary interaction plan; Based on the preliminary interaction scheme and the standardized multimodal user information, a collaborative control strategy is generated to guide the execution of the interaction. During the execution of the collaborative control strategy, real-time multimodal user information is continuously acquired and compared with the preliminary interaction scheme. When the adjustment condition is triggered, a new interaction scheme is re-planned and generated, and the collaborative control strategy is updated. After the interaction task is completed, the interaction performance is analyzed based on the interaction process data. When the analysis results do not meet the preset standards, the prediction process of the user's emotional change trend, the generation process of the collaborative control strategy, and the generation process of the interaction plan are optimized.

[0006] Furthermore, to achieve the above objectives, the present invention provides an interactive control device based on user state awareness, comprising: The multimodal perception and preprocessing module is used to collect and preprocess the multimodal raw data generated during user interaction to generate standardized multimodal user information. The emotion prediction and solution generation module is used to predict the user's emotion change trend based on the standardized multimodal user information, and generate a preliminary interaction solution by combining the user's emotion change trend, the user's intention and needs and the current interaction state. The collaborative control strategy generation module is used to generate a collaborative control strategy to guide the execution of the interaction based on the preliminary interaction scheme and the standardized multimodal user information. An adaptive interaction execution module is used to continuously acquire real-time multimodal user information during the execution of the collaborative control strategy, compare the real-time multimodal user information with the preliminary interaction scheme, and when the adjustment condition is triggered, re-plan and generate a new interaction scheme and update the collaborative control strategy. The interaction analysis and self-optimization module is used to analyze the interaction performance based on the interaction process data after the interaction task is completed. When the analysis results do not meet the preset standards, it optimizes the prediction process of user emotion change trends, the generation process of collaborative control strategies, and the generation process of interaction schemes.

[0007] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a user state-aware interactive control program stored in the memory and executable on the processor, wherein the user state-aware interactive control program, when executed by the processor, implements the steps of the user state-aware interactive control method as described above.

[0008] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a user-state-aware interactive control program, wherein the user-state-aware interactive control program, when executed by a processor, implements the steps of the user-state-aware interactive control method as described above.

[0009] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses an interactive control method, device, equipment, and medium based on user state perception, comprising: collecting and preprocessing multimodal raw data generated during user interaction to generate standardized multimodal user information; predicting the user's emotional change trend based on the standardized multimodal user information, and generating a preliminary interaction plan by combining the user's emotional change trend, intent needs, and current interaction state; generating a collaborative control strategy to guide interaction execution based on the preliminary interaction plan and standardized multimodal user information; continuously acquiring real-time multimodal user information during the execution of the collaborative control strategy, comparing the real-time multimodal user information with the preliminary interaction plan, and re-planning and generating a new interaction plan and updating the collaborative control strategy when adjustment conditions are triggered; analyzing the interaction performance based on the interaction process data after the interaction task ends, and optimizing the prediction process of emotional change trend, the generation process of collaborative control strategy, and the generation process of interaction plan when the analysis results do not meet preset standards. This invention achieves dynamic perception of user state by integrating multimodal information and constructs the emotional prediction, intent recognition, and collaborative control strategy generation processes into a closed-loop optimization system, thereby achieving real-time adaptive interactive control in complex interaction scenarios. By continuously monitoring and adjusting model parameters and strategy generation logic, the system's response accuracy and interaction consistency under different user states can be improved, significantly enhancing the naturalness and stability of human-computer interaction. Attached Figure Description

[0010] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a schematic diagram of an application environment for an interactive control method based on user state awareness, according to an embodiment of the present invention. Figure 2 This is a flowchart illustrating an embodiment of the user state-aware interactive control method of the present invention. Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the interactive control device based on user state awareness of the present invention. Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0011] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0012] The user-state-aware interactive control method provided in this embodiment of the invention can be applied to, for example... Figure 1 In this application environment, the client communicates with the server via a network. The server can collect and preprocess multimodal raw data generated during user interaction through the client, generating standardized multimodal user information; predicting the user's emotional change trend based on the standardized multimodal user information, and generating a preliminary interaction plan by combining the user's emotional change trend, intent needs, and current interaction state; generating a collaborative control strategy to guide the interaction execution based on the preliminary interaction plan and standardized multimodal user information; continuously acquiring real-time multimodal user information during the execution of the collaborative control strategy, comparing the real-time multimodal user information with the preliminary interaction plan, and re-planning and generating a new interaction plan and updating the collaborative control strategy when adjustment conditions are triggered; analyzing the interaction performance based on the interaction process data after the interaction task ends, and optimizing the prediction process of emotional change trend, the generation process of collaborative control strategy, and the generation process of interaction plan when the analysis results do not meet preset standards. This invention achieves dynamic perception of user state by integrating multimodal information and constructs the emotional prediction, intent recognition, and collaborative control strategy generation processes into a closed-loop optimization system, thereby achieving real-time adaptive interaction control in complex interaction scenarios. By continuously monitoring and adjusting model parameters and strategy generation logic, the system's response accuracy and interaction consistency under different user states can be improved, significantly enhancing the naturalness and stability of human-computer interaction. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.

[0013] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the user-state-aware interactive control method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0014] like Figure 2 As shown, the user state-aware interactive control method proposed in this invention includes the following steps: S10: Collect and preprocess the multimodal raw data generated during user interaction to generate standardized multimodal user information; In this embodiment, the acquisition and preprocessing of multimodal raw data generated during user interaction to generate standardized multimodal user information requires three consecutive operational steps: "acquisition," "preprocessing," and "standardized generation." The acquisition phase focuses on acquiring information from the user's multidimensional input signals during interaction, involving different modalities such as voice signals, semantic content, emotional expression, facial expressions, posture changes, and touch behavior. Voice signals can be captured at a high sampling rate using a microphone array to capture the raw audio stream. Semantic content can be converted from speech to text information using a speech recognition module. Emotional expression can be calculated by combining acoustic features such as timbre, speech rate, and energy envelope with semantic context. Visual input is extracted using a camera sensor to extract facial expression key points, eye movement trajectories, or posture features. Behavioral input is extracted from the terminal interaction log to extract temporal signals such as touch coordinates and response delays. All modal data is timestamped for subsequent temporal alignment and feature synchronization.

[0015] The goal of the preprocessing stage is to transform the collected multi-source heterogeneous data into a structured data format that can be uniformly modeled. Audio signals undergo denoising, speech activity detection, and channel equalization to eliminate environmental interference; text semantics are processed through word segmentation, stop word removal, and synonym grouping to reduce semantic redundancy; visual features are normalized and keypoint reprojection is performed to ensure scale and pose consistency; touch and behavioral data undergo temporal interpolation and missing data completion to maintain continuity. A data quality detection mechanism can be introduced during preprocessing to assess indicators such as signal-to-noise ratio, text integrity, and frame loss rate. Low-quality data can be repaired through weighted or resampling methods to ensure the reliability of subsequent model inputs.

[0016] The standardized generation process achieves uniformity in dimensionality, scale, and temporal alignment of cross-modal features. First, a timestamp alignment mechanism pairs features from different modalities according to interaction time windows. Second, feature normalization maps numerical features to a unified range; for example, audio energy features and visual brightness features can be processed using Z-score or min-max normalization. Third, an embedding mapping function converts discrete semantic information into a continuous vector space, maintaining dimensionality consistent with continuous features. Finally, a feature fusion algorithm weights and concatenates or attention-weights features from speech, semantics, vision, and behavior to form a standardized multimodal user information representation, which can then be directly called upon by subsequent sentiment prediction and interaction strategy generation modules.

[0017] During implementation, different data acquisition and preprocessing methods can be selected based on the application scenario and computing resources. Real-time acquisition and preprocessing can be achieved through edge terminals to reduce data transmission latency; alternatively, high-precision feature extraction and fusion can be performed through cloud computing nodes to improve computational stability. For speech signals, adaptive filtering algorithms can be used to dynamically suppress environmental noise; for text semantics, a contextual semantic model can be introduced to enhance semantic integrity; and for visual features, dynamic facial expression features can be extracted based on lightweight convolutional networks.

[0018] During the standardization process, a unified modality coding framework can be adopted to map features such as speech, text, vision, and behavior to a unified multimodal feature space through a shared embedding layer. Alternatively, a multimodal alignment network can be used to dynamically adjust the temporal offsets of different modal features through a sliding time window mechanism, ensuring temporal consistency of the fused features. For low-resource scenarios, feature compression algorithms can be used to reduce dimensional redundancy and improve processing efficiency.

[0019] This embodiment integrates multimodal data processing across the acquisition, preprocessing, and standardization stages, enabling the system to acquire complete, clean, and uniformly modelable user state information. This process effectively improves the input quality for subsequent emotion prediction and interactive control stages, thereby enhancing the overall system's accuracy in recognizing user states and the continuity of its response, achieving dynamic adaptation and real-time feedback to complex and changing user states.

[0020] S20, based on the standardized multimodal user information, predict the user's emotional change trend, and combine the user's emotional change trend, the user's intention and needs with the current interaction state to generate a preliminary interaction plan; In this embodiment, the process of predicting user emotion change trends based on standardized multimodal user information and generating preliminary interaction plans by combining user emotion change trends, user intent needs, and current interaction states aims to transform multimodal information into emotional dynamics and behavioral decisions. First, the standardized multimodal user information includes multi-source features such as speech, semantics, emotion, and behavior. After unified time alignment and feature normalization, it can be used as input for both the prediction model and the interaction decision module. The prediction of emotion change trends is based on temporal analysis, establishing emotional dependencies between continuous time steps in the multimodal feature space to estimate the user's future emotional state. Specifically, the standardized input features may include acoustic features (such as pitch, intensity, and spectral energy distribution), semantic features (such as sentence sentiment polarity and keyword frequency), visual features (such as facial expression encoding vectors and gaze direction vectors), and interaction behavior features (such as response latency and operation frequency). These features are vectorized and fused before being input into the emotion prediction module.

[0021] In the prediction phase, a temporal modeling unit can be used to achieve continuous state estimation. By learning the changing trend of the emotion vector on the time axis, the emotion trajectory for future time windows is output. The prediction result can be represented as a distribution or trend vector of emotion state changes over time, used to describe the stability, volatility, or abrupt change tendency of user emotions. This trend result is passed to the interaction strategy generation unit, working together with the user's intent and the current interaction state. The user's intent comes from the recognition results of the semantic understanding module and is used to determine the interaction goal, such as requesting information, expressing feedback, or refusing an action. The current interaction state is provided by a state cache maintained by the system, including the current session turn, completion rate, semantic context, and historical response records.

[0022] The comprehensive analysis phase establishes a mapping relationship between emotional trends, intentional needs, and interaction states in a multi-dimensional feature space, generating preliminary interaction schemes that align with user psychological expectations and task objectives. The generation process can determine the relative influence of different factors through weight fusion or attention allocation mechanisms. For example, when emotional fluctuations are significant, the system can increase the weight of emotional trends, prioritizing the generation of soothing or mitigating interaction content; when emotions are stable and the task orientation is clear, the weight of intentional needs can be increased to enhance task progress efficiency. The final generated preliminary interaction scheme includes interaction goals, tone control parameters, a semantic content framework, and a strategy logic sequence, providing the basic input for subsequent collaborative control strategy generation.

[0023] In the implementation process, various model structures and data fusion mechanisms can be used to adapt to different application environments. Multi-layer temporal networks can be employed to dynamically model standardized multimodal information, capturing the cross-modal temporal dependencies between speech, facial expressions, and behavior. Alternatively, an emotion state transition graph model can be used to express the changing trends of emotion at different stages in the form of latent states. For the fusion of intent requirements and interaction states, feature concatenation and nonlinear combination can be performed through a multi-input decision network to output a multi-dimensional weight vector to guide the generation of interaction schemes. In the emotion prediction part, a sliding time window mechanism can be used, where the system automatically updates the emotion trend prediction whenever a new batch of real-time data arrives, ensuring real-time scheme generation.

[0024] In different scenarios, the fusion strategy can be adjusted to optimize the interaction effect. In scenarios with rapid emotional changes, the state smoothing parameter can be reduced to make the model more sensitive to short-term emotional changes; in scenarios with gradual emotional changes, the weight of historical memory can be increased to enhance trend stability. For real-time monitoring of the interaction state, a low-latency caching synchronization mechanism can be adopted to ensure that the state update and prediction process is completed in milliseconds.

[0025] This embodiment combines emotional change trend prediction with intention needs and interaction states in a joint modeling process. This allows the system to generate preliminary solutions that align with the user's psychology and semantics before interaction, significantly improving the naturalness and relevance of the response. The combination of multimodal fusion and trend prediction enables the system to dynamically adapt, automatically adjusting dialogue strategies based on changes in user emotions. This avoids mechanical responses or delayed reactions, improving the overall interactive experience and task completion efficiency.

[0026] S30, Based on the preliminary interaction scheme and the standardized multimodal user information, generate a collaborative control strategy to guide the execution of the interaction; In this embodiment, the process of generating a collaborative control strategy to guide interaction execution based on a preliminary interaction plan and standardized multimodal user information is a crucial step in the fusion of emotion cognition and strategy decision-making. Its purpose is to enable the system to dynamically adjust tone, speech rate, content logic, and feedback mechanisms during interaction execution based on the user's real-time psychological state, semantic intent, and emotional trends, thereby achieving naturalness and relevance in human-computer interaction. This process first receives the preliminary interaction plan and standardized multimodal user information generated in the previous stage as input conditions. The preliminary interaction plan includes the interaction goal, semantic path, tone parameters, and content sequence, representing the system's predictive planning result for the current task; the standardized multimodal user information provides the current user's emotional state, semantic features, and behavioral signals to support the real-time and personalized nature of strategy generation.

[0027] In implementation, the system uses a strategy parsing module to perform feature alignment and context fusion on the input data. The semantic part extracts user intent keywords and interaction topics, determining the core content of the interaction through a keyword mapping table. The emotion part calculates tone control parameters based on the user's emotion state vector, including indicators such as emotion intensity, intonation curve, and energy range, to form a tone strength level; the speech rate control parameter is determined by statistically analyzing the mapping relationship between the rate of change of emotion trends and speech prosody features. The determination of the interaction guidance method depends on the directional judgment of the emotion change trend; for example, when the emotion changes in a positive direction, a progressive guidance can be used, while when the emotion tends towards negativity or neutrality, a soothing or interruptive guidance can be used.

[0028] During the control strategy generation phase, the system weighted and fused the aforementioned multi-dimensional control parameters with the preliminary interaction scheme, dynamically balancing the priority between emotional factors and task objectives through a weight allocation function. The strategy generation module constructs a set of collaborative control instructions based on the fused features, including voice output parameters, interaction logic sequences, and trigger condition mapping tables. Voice output parameters control the tone and speed of speech generated by the speech synthesis module; interaction logic sequences define the system's response paths under different emotional and semantic states; and the trigger condition mapping table describes the system's feedback mechanism under specific user states, enabling dynamic adjustment and personalized adaptation. After generation, the collaborative control strategy is cached in the control instruction library for real-time invocation or updating.

[0029] In different implementations, different strategy generation architectures can be selected based on the type of interaction task and the distribution of computing resources. A rule-based and model-driven strategy generation mechanism can be adopted, with logical rules dominating in scenarios with high semantic certainty and prediction models dominating in scenarios with significant emotional fluctuations. Alternatively, a multi-layered control architecture can be used, with lower-level control responsible for fine-tuning tone and speech rate, and higher-level control responsible for the dynamic planning of interaction logic, to achieve global control over complex multi-turn dialogues.

[0030] During the strategy generation process, attention mechanisms can be used to calculate emotion weights, thereby enhancing the contribution of high-impact features in strategy generation. Alternatively, a dynamic window mechanism can be employed to continuously update control parameters during interaction, enabling the strategy to respond instantly to changes in user state. For resource-constrained environments, compressed mapping networks can be used to reduce the number of control parameters, thereby reducing computational load without sacrificing the main control effects.

[0031] At the voice control level, the speech rate range and pitch range can be adaptively adjusted according to emotional state; at the logic control level, the strategy weights can be dynamically updated based on historical interaction results through reinforcement learning algorithms, enabling the system to have self-optimization capabilities in long-term operation.

[0032] This embodiment generates a collaborative control strategy by fusing a preliminary interaction scheme with standardized multimodal user information. The system can simultaneously adjust at the semantic, emotional, and behavioral levels, ensuring that the interaction response not only conforms to the task logic but also matches the user's psychological state. This process significantly improves the overall performance of the interaction system in terms of emotional sensitivity, naturalness of expression, and accuracy of response, reduces rigid tone and semantic mismatch issues, and achieves highly adaptive interactive control driven by emotion.

[0033] S40, during the execution of the collaborative control strategy, continuously acquire real-time multimodal user information, compare the real-time multimodal user information with the preliminary interaction scheme, and when the adjustment condition is triggered, re-plan and generate a new interaction scheme and update the collaborative control strategy. In this embodiment, during the execution of the collaborative control strategy, real-time multimodal user information is continuously acquired and compared with the initial interaction plan. When adjustment conditions are triggered, a new interaction plan is re-planned and generated, and the collaborative control strategy is updated. This is the core link in realizing dynamic closed-loop control of the interaction process. This process, based on continuous perception, comparative analysis, and adaptive updating, realizes the status monitoring and strategy optimization of the system during operation.

[0034] First, the system continuously collects real-time speech, semantic, and emotional information during the interactive process through a multimodal perception unit. Speech information is obtained in real-time through audio sensor sampling, yielding features such as pitch, energy, and speech rate. Semantic information is parsed into structured text and key semantic entities are extracted in real-time by the speech recognition and semantic understanding modules. Emotional information generates an emotion vector based on acoustic features and semantic context, representing the user's emotional state. All of the above multimodal inputs are timestamped and streamed in a time-synchronized manner to ensure the system has continuous perception capabilities of the user's state at every moment.

[0035] Secondly, the system performs a multi-dimensional comparison between real-time multimodal user information and predictive parameters in the initial interaction plan. Emotional level comparison calculates the distance or deviation between the current emotion vector and the emotion change trend vector to generate an emotion deviation score. Semantic level comparison calculates the semantic distance between the current semantic vector and the user's intended needs to generate a semantic matching score. Acoustic level comparison compares the real-time speech signal with the acoustic elements corresponding to the speech synthesis parameters in the collaborative control strategy to generate an acoustic similarity score. When any deviation or similarity index exceeds a preset threshold, the system determines that there is a state inconsistency and further calculates multi-dimensional deviation correlations to identify whether it belongs to a single modality anomaly or a cross-modal collaborative deviation.

[0036] Upon detecting a deviation event that meets the triggering conditions, the system enters the adaptive planning phase. The adaptive planning module selects the corresponding planning mechanism based on the deviation type. For example, when emotion deviates from the dominant trend, tone parameters and content guidance are adjusted first; when semantics deviates from the dominant trend, semantic path replanning is performed; and when multimodal collaborative deviation occurs, the interaction logic path is replanned. During the planning process, a new weight distribution can be calculated by combining real-time user status and task objectives, and structural constraints can be applied to the replanning results to ensure that the new interaction scheme responds to changes in user status while maintaining the coherence of task logic. After the new interaction scheme is generated, it will be mapped to updated execution instructions. Based on this, the system synchronously updates the speech synthesis parameters, interaction logic sequences, and feedback triggering conditions in the collaborative control strategy, achieving real-time adaptation at the strategy layer.

[0037] In implementation, multi-threaded data channels can be used to achieve parallel acquisition and fusion of information from different modalities, ensuring the system's real-time performance. Alternatively, an event-driven architecture can be adopted, triggering a replanning process immediately when modal deviations exceed limits, thereby reducing latency. The comparison process can be implemented by embedding spatial distance calculations, weighted similarity functions, or attention-based modal relevance networks. Triggering conditions can include various forms such as exceeding a single indicator threshold, exceeding the collaborative deviation ratio limit, or reversing the direction of emotion change. The system can dynamically adjust the judgment logic according to the task type.

[0038] In the replanning mechanism, a fast path planning algorithm can be used to reorder the interaction semantic tree and prioritize nodes with high matching degrees. Alternatively, a reinforcement learning mechanism can be employed to update policy weights based on historical adjustment records, allowing the system to gradually optimize replanning efficiency during continuous interactions. After the policy is updated, an asynchronous buffering mechanism can be used to smoothly switch policy parameters without interrupting the current dialogue, ensuring the continuity and naturalness of the interaction.

[0039] In different application environments, the sampling frequency and comparison period of real-time multimodal information can be dynamically adjusted according to system resources and task complexity. For example, in scenarios requiring high responsiveness, the sampling period can be controlled at the millisecond level; in scenarios with a slow task pace, a window averaging method can be used for batch updates to reduce computational load.

[0040] This embodiment achieves dynamic adaptive adjustment of the interaction process by continuously acquiring real-time multimodal user information and comparing it with the initial interaction plan. This enables the system to promptly identify changes in user state and automatically correct the interaction path and strategy parameters. This mechanism effectively avoids the response lag problem caused by static strategies, improves the system's emotional sensitivity and task execution robustness, and ensures that the interaction process remains natural, coherent, and meets user expectations.

[0041] S50, after the interaction task is completed, analyzes the interaction performance based on the interaction process data. When the analysis results do not meet the preset standards, it optimizes the prediction process of the user's emotional change trend, the generation process of the collaborative control strategy, and the generation process of the interaction plan.

[0042] In this embodiment, after the interaction task ends, the interaction performance is analyzed based on the interaction process data. When the analysis results do not meet the preset standards, the process of predicting the user's emotional change trend, generating the collaborative control strategy, and generating the interaction scheme are optimized. This is the core link in realizing the closed-loop self-learning and continuous optimization of the interaction system. This process uses the task end signal as the trigger point and performs structured analysis on the data recorded throughout the interaction to achieve performance backtracking and parameter correction, thereby enabling the system to gradually improve its adaptability and response accuracy in multiple interactions.

[0043] First, after the interaction ends, the system activates the effect analysis module. This module automatically reads the interaction records during the task, including the voice stream, semantic content, emotional feature trajectory, and strategy execution log. Interaction duration, voice energy distribution, speech rate change rate, semantic matching degree, and emotional vector sequence are extracted as the basic feature set to characterize the behavioral patterns and emotional dynamics of the interaction process. Through statistical analysis and time series modeling, the system calculates interaction goal completion indicators (e.g., task completion rate, information matching rate) and user cooperation indicators (e.g., response timeliness, semantic consistency). Furthermore, the system calculates an emotional fluctuation index based on the stability of the emotional trajectory to measure the accuracy of emotional trend prediction.

[0044] The system compares the aforementioned indicators with preset evaluation criteria to determine whether the interaction performance meets the expected level. Preset criteria may include dynamic thresholds based on historical average levels or static scoring limits defined by expert experience. When any indicator is detected to be below the standard threshold, the system automatically triggers the adaptive optimization module. This module adjusts three core processes through a multi-task optimization mechanism: first, it optimizes the parameter distribution of the time-series prediction model, correcting the emotional trend prediction weights through gradient backtracking to improve the stability of future predictions; second, it optimizes the parameters generated by the collaborative control strategy, focusing on adjusting the tone and speed control factors and guiding weight allocation to better match emotion-driven processes with semantic logic; and third, it updates the path planning parameters in the interaction scheme generation algorithm, redefining the interaction goals and content selection order to reduce the impact of sudden emotional changes on the task flow.

[0045] All optimization processes are based on joint learning of historical data and the latest interaction records. After optimization, the system automatically stores the new parameter set and updates the runtime environment so that subsequent tasks can directly use the improved strategy and model.

[0046] In practical implementation, a multi-channel data acquisition architecture can be adopted to process interaction logs asynchronously, thereby improving analysis efficiency. Alternatively, an incremental learning framework can be used to achieve local parameter updates, reducing interference with the system's global model. Optimization of sentiment change trends can be achieved by dynamically adjusting the prior distribution of the time-series prediction model based on a Bayesian update mechanism, making the prediction results more adaptable to new data.

[0047] In the optimization of collaborative control strategies, a policy weight update method based on reinforcement learning can be used to automatically adjust policy parameters by backpropagating the reward function of historical interaction replays. In the optimization of interaction schemes, a multi-objective optimization mechanism can be introduced to solve for the weighted solution of task success rate and emotional stability, generating path planning parameters with balanced performance.

[0048] In various scenarios, the system can flexibly set the optimization cycle according to the type of interactive task. In high-frequency interaction environments, the optimization cycle can be set to update immediately after each task ends; in resource-constrained environments, periodic batch optimization can be used to balance computing resources and learning speed.

[0049] This embodiment enables continuous performance evolution by performing interaction process analysis and adaptive optimization after the interactive task concludes. This mechanism effectively compensates for the inability of static models to cope with diverse user behaviors, allowing the system to continuously adjust its prediction and decision-making logic based on historical feedback, thus forming a self-learning closed loop in long-term operation. The optimized model exhibits more stable and accurate performance in emotion trend prediction, tone control, and task guidance, significantly improving interaction quality and task completion rate.

[0050] This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses an interactive control method, device, equipment, and medium based on user state perception, comprising: collecting and preprocessing multimodal raw data generated during user interaction to generate standardized multimodal user information; predicting the user's emotional change trend based on the standardized multimodal user information, and generating a preliminary interaction plan by combining the user's emotional change trend, intent needs, and current interaction state; generating a collaborative control strategy to guide interaction execution based on the preliminary interaction plan and standardized multimodal user information; continuously acquiring real-time multimodal user information during the execution of the collaborative control strategy, comparing the real-time multimodal user information with the preliminary interaction plan, and re-planning and generating a new interaction plan and updating the collaborative control strategy when adjustment conditions are triggered; analyzing the interaction performance based on the interaction process data after the interaction task ends, and optimizing the prediction process of emotional change trend, the generation process of collaborative control strategy, and the generation process of interaction plan when the analysis results do not meet preset standards. This invention achieves dynamic perception of user state by integrating multimodal information and constructs the emotional prediction, intent recognition, and collaborative control strategy generation processes into a closed-loop optimization system, thereby achieving real-time adaptive interactive control in complex interaction scenarios. By continuously monitoring and adjusting model parameters and strategy generation logic, the system's response accuracy and interaction consistency under different user states can be improved, significantly enhancing the naturalness and stability of human-computer interaction.

[0051] In one embodiment, step S10 above includes: S101 collects user voice information through a voice recognition module; S102, the semantic understanding module identifies the user's intention and needs from the user's voice information; S103, the emotion analysis module determines the user's emotional state based on the user's voice information; S104, construct user profile information by combining historical interaction data with the user profile module; S105, perform noise reduction processing on the user voice information to generate noise-reduced user voice information; S106, perform semantic cleaning on the user intent requirements to generate cleaned user intent requirements; S107, Standardize the user's emotional state to generate a standardized user emotional state; S108, extract features from the user profile information to generate user profile features; S109, the noise-reduced user voice information, the cleaned user intent needs, the standardized user emotional state, and the user profile features are fused to generate standardized multimodal user information.

[0052] In this embodiment, the goal is to transform the multi-source heterogeneous data generated during the interaction process into unified, standardized multimodal user information that can be directly used in subsequent stages.

[0053] The speech recognition module's acquisition of user voice information revolves around acoustic signal acquisition and ensuring transcribing capability. The acquisition unit uses a pickup unit to transfer the continuous audio stream into a buffer, while simultaneously adding time stamps and device identifiers to ensure subsequent timing synchronization and source traceability. To improve speech transcribing capability, pre-processing performs voice activity determination to eliminate silent segments, and performs echo suppression and channel equalization to limit amplitude and phase distortion caused by differences in device and environment. Speech segments are divided into short frames, generating frame-level metadata as input for subsequent textualization and acoustic feature construction. This module's source is a necessary entry point for speech-to-text conversion; examples include customer responses and follow-up questions in outbound call scenarios, and health statements and feedback statements in remote follow-ups.

[0054] The semantic understanding module identifies user intent from voice information by revolving around textualized sequences, context extraction, and intent summarization. Voice segments are transcribed into text sequences by a transcription engine, along with start and end times and confidence labels. Trigger words, object words, and constraint words are extracted from the text sequences and combined with a context window to form semantic fragments. Intent needs are summarized through the structured combination of semantic fragments and mapped to intent labels and slot sets, including action tendencies, object scope, and limitations. This module draws upon structured expressions of users' genuine needs, such as inquiries, confirmations, rejections, clarifications, and appointments.

[0055] The emotion analysis module, which determines a user's emotional state based on their voice information, revolves around the joint representation of acoustic prosody and semantic cues. It extracts prosodic elements such as pitch, intensity, speech rate, and formants from segmented speech frames, and combines these with sentiment-related words and negation / transition words in the text sequence to generate frame-level emotion cues. Through temporal window aggregation, it obtains segment-level emotion vectors, which contain a joint representation of emotion category and intensity level, and outputs the emotion trajectory that changes over time. The source is non-content-level information carried in human-computer interaction; examples include the consistent reflection of states such as calmness, tension, hesitation, and pleasure in prosody and word choice.

[0056] The user profiling module, which combines historical interaction data to construct user profiles, revolves around extracting long-term preferences and interaction habits. Historical records extract long-term attributes such as communication preferences, content preferences, and compliance levels, while also extracting behavioral statistics such as communication frequency, response to prompts, and acceptance of explanations. Data from different sources is uniformly encoded and written into the profile vector, forming a searchable and incrementally updated user-side context. Sources are past interaction facts, such as a preference for brief answers, a tendency to accept detailed explanations, and sensitivity to reminders.

[0057] The implementation of noise reduction processing for user voice information to generate denoised user voice information revolves around environmental interference suppression and speech intelligibility preservation. Noise spectrum estimation and residual noise suppression are performed at the frame level, with transient noise suppression superimposed to address keystroke and cough-like impulse interference. Finally, energy normalization is performed to stabilize the amplitude distribution. The output maintains the same time stamp as the original frame to ensure alignment with subsequent text and emotional trajectories. This is driven by robustness requirements of the acoustic front-end, such as noisy office areas, outdoor wind noise, and echoing rooms.

[0058] The semantic cleansing of user intent requests to generate cleaned user intent requests revolves around ambiguity resolution and structural unification. This involves merging synonyms, restoring the original references to pronouns, filling in missing slots with fillable fields, and ensuring consistent expression of tense and negation. The output retains the reference relationships and temporal positions of the original trigger words and object words, guaranteeing its ability to integrate with acoustic and emotional trajectories. The source is the need for consistent expression across time periods and wording; an example is mapping "want to know about benefits" and "what are the benefits?" to the same intent framework.

[0059] The standardization of user emotional states revolves around category alignment and strength alignment. First, emotional categories from different sources are mapped to a unified label space. Then, the strength levels are mapped to a unified interval, making emotional representations comparable across different conversations, devices, and speaking rates. The output includes both the label and level of the current segment and retains a short-term smoothed time trajectory for subsequent coupling with intent, personas, and text content. The source is the unification of dimensions before cross-modal fusion; for example, mapping elation and excitement to a unified "happy" category and unifying the strength level measurement.

[0060] The implementation of feature extraction from user profile information to generate user profile features revolves around the vectorization and reusability of long-term features. Category fields in the profile information are grouped and encoded, behavioral statistics fields are scaled and standardized, and timeliness fields are weighted with decaying weights to generate fixed-length vectors that are decoupled from user identifiers, allowing preference invocation without revealing the user's identity. The vectors retain the source labels for each dimension, supporting weighted referencing during the fusion stage. The source is cross-conversational transfer and personalization control; an example is mapping long-term attributes of slow speech preference to one-dimensional weights in a fixed dimension.

[0061] The process of fusing denoised user voice information, cleaned user intent needs, standardized user emotional states, and user profile features to generate standardized multimodal user information revolves around temporal alignment, dimensional alignment, and weighted fusion. Temporal alignment uses time tags as anchors to align voice frames, text segments, and emotional trajectories to a unified sliding window; dimensional alignment maps discrete labels and continuous vectors to a unified embedding space; weighted fusion assigns weights based on signal quality and task relevance, and performs robust smoothing on abrupt changes within the short window, resulting in two types of products: a window-level unified representation and a session-level summary representation. The unified representation includes indexes of four types of source fields, ensuring source transparency and traceability. This fusion product directly serves as input for subsequent emotion trend prediction, intent-driven planning, and execution control.

[0062] This embodiment unifies cross-modal, cross-temporal, and cross-device data into a fusion representation with temporal and dimensional constraints through fine-grained acquisition, cleaning, and standardization of data from four sources: voice, semantics, emotion, and user profile. Quality control and temporal alignment mechanisms suppress noise, missing data, and abrupt changes that could interfere with subsequent processing. Retaining source indexes and traceable tags allows subsequent stages to reference data as needed and apply precise weights. The resulting standardized multimodal user information is more stable in terms of completeness and consistency, and its coupling with emotion trend prediction and interactive control is smoother. This reduces propagation bias caused by input uncertainty and achieves higher recognition reliability and better strategy controllability under the same resource conditions.

[0063] In one embodiment, step S20 above includes: S201, Extract user emotional features and user intent needs from the standardized multimodal user information; S202, Use a time-series prediction model to process the user's emotional characteristics and predict the user's emotional change trend; S203, Obtain the current interaction state maintained by the state management unit; S204, Based on the emotional change trend and the user's intention needs, determine the constraints of the interaction path; S205, Based on the constraints, the current interaction state, and the user's communication preferences, generate a preliminary interaction plan.

[0064] In this embodiment, the goal is to form a preliminary interaction scheme for subsequent execution based on standardized multimodal user information, thereby achieving a closed-loop connection from state awareness to executable planning.

[0065] Extracting user emotional features and intent requests from standardized multimodal user information employs parallel temporal and semantic alignment. Temporal alignment, based on a unified timestamp, aligns acoustic frame-level features with text fragments to a sliding time window, ensuring a one-to-one correspondence between acoustic intensity, pitch curves, speech rate changes, and textual sentiment cues within the same time window. Emotional features are jointly constructed from prosodic elements and emotional lexical indicators, forming a numerical vector and its temporal trajectory, while retaining strength levels and confidence intervals to support subsequent weighting. Intent requests are processed through text normalization and slot extraction to obtain intent labels and parameter sets, preserving trigger word positions and cross-sentence citation relationships for integration with historical context. To avoid mutual interference, the emotional and intent parsing pathways remain decoupled at the feature layer, and are fused using gated weighting at the time window output layer, outputting two products: an emotional vector and a structured representation of intent.

[0066] The implementation of a time-series prediction model to process user emotional features and predict trends in user emotions employs a two-stage process: sequence modeling and robust post-processing. Sequence modeling takes a sequence of emotional vectors with equal time windows as input, learns the correlation between short-term fluctuations and medium-term tendencies, and outputs trend vectors and trend confidence scores for several future time windows. Robust post-processing smooths the trend vectors and removes outliers to avoid directional misjudgments caused by transient noise. Simultaneously, it attenuates or amplifies the trend amplitude based on the confidence score, forming a trend expression that can be directly used to constrain generation. The trend expression maintains the same time reference as the original emotional trajectory, facilitating window-by-window fusion with intention and state information.

[0067] The implementation of obtaining the current interaction state maintained by the state management unit relies on a unified state cache and a consistent reading mechanism. The state management unit maintains fields such as session round, completed goals, current topic node, historical response results, and the most recent exception and rollback flags, writing them incrementally over time and providing atomic reads. A snapshot is read at the start of the planning cycle to ensure consistency with the time window of sentiment trends and intent requirements. When cross-thread updates occur, version number verification is used; if a version lag is detected, a lightweight reread is triggered to ensure consistency between the planning basis and the actual execution progress.

[0068] The constraints for determining the interaction path based on user sentiment trends and intent needs are achieved through a combination of rule mapping and continuous constraints. Rule mapping maps trend direction and intent type to path-level restrictions or tendencies, such as increasing explanation density or limiting follow-up question depth. Continuous constraints calculate weights from trend magnitude, trend confidence, and intent certainty to generate a differentiable cost weight function, which applies penalties or rewards to different candidate branches during the path search process. Discrete constraints and continuous weights are merged in a unified constraint table, which includes the permissibility, priority, cost bias, and interruptibility of each candidate node, supporting rapid reference in subsequent search processes.

[0069] The initial interaction plan, generated based on constraints, the current interaction state, and user communication preferences, employs a dual-path collaborative approach of constrained semantic path search and expression parameter synthesis. Semantic path search starts with the current topic node and, guided by a constraint table, expands candidate nodes in a bounded manner, prioritizing branches that satisfy permissibility and minimize total cost, generating a finite-length content sequence and jump logic. The expression parameter synthesis path reads communication preferences and trend expressions, mapping tone strength levels and speech rate ranges to executable expression parameters, and assigning expression parameter fragments to each node in the content sequence. The outputs of the two paths are aligned in the planning synthesizer to form an initial interaction plan containing the content sequence, node jump conditions, and expression parameter timeline, while retaining constraint references and state version numbers to support execution-time tracking and rapid correction.

[0070] This embodiment analyzes emotional features and intent requirements from standardized multimodal user information in parallel while maintaining temporal consistency. It then uses temporal prediction to provide credible emotional change trends and generates path-level constraints using the current interaction state of the state management unit as an anchor point. This enables the collaborative synthesis of content paths and expression parameters within the same planning cycle. This mechanism unifies the user's current state, short-term trends, and task progress into the same planning space, reducing path jitter and expression distortion caused by information asynchrony. It improves the matching degree and executability of content selection and tone / speed settings, provides stable input for subsequent collaborative control and adaptive updates during the execution period, and reduces the frequency of invalid searches and repeated corrections under equal resource conditions, resulting in more coherent interaction responses and faster planning convergence.

[0071] In one embodiment, step S202 includes: S2021, Perform time-series standardization processing on the user's emotional features to generate a standardized emotional feature sequence; S2022, The standardized emotion feature sequence is input into the feature extraction layer of the time-series prediction model for local feature extraction to obtain local emotion features; S2023, The local emotion features are input into the memory network layer of the time-series prediction model to model the long-short-term dependency relationship and generate a time-series feature sequence; S2024, Based on the temporal feature sequence, determine the probability distribution of the emotional state at future time steps through the output layer of the temporal prediction model; S2025, Generate the user's emotional change trend based on the probability distribution of emotional state at the future time step.

[0072] In this embodiment, the goal is to transform user emotional characteristics into a predictable time trajectory, thereby estimating the emotional trend over several future time periods. The input is a vector sequence of user emotional characteristics that changes over time, and the output is a time-series trajectory of emotional change trends that contains directional and strength information, along with a corresponding confidence level.

[0073] Temporal standardization aims to unify the timeline, scale, and distribution. Original emotion features are resampled at uniform sampling intervals to align timestamps from different sources and devices. Scale normalization and centering are performed on each feature dimension to suppress amplitude drift caused by individual and device differences. Robust interpolation and anomaly suppression are applied to abrupt changes and missing segments, employing window smoothing and boundary protection strategies to improve sequence stability without erasing genuine transitions. The processed result forms a standardized emotion feature sequence, containing continuous vectors and quality markers within equal-length time windows, facilitating unified integration into subsequent models.

[0074] The feature extraction layer is responsible for capturing local patterns within short time windows. Standardized emotion feature sequences are sliced ​​into fixed-length windows with a fixed step size, constructing a neighborhood context for each window. Within each window, local elements such as rhythm, fluctuation amplitude, local trends, and change rates are extracted. These multidimensional elements are mapped into a compressed representation through learnable linear or nonlinear transformations, yielding local emotion features. This representation emphasizes short-term dynamics and micro-morphology, preserving temporal sequence information and local boundary information, and is then output to subsequent modeling units.

[0075] The memory network layer is used to establish long-term and short-term dependencies. Continuous local emotional feature sequences are input into temporal units with memory and gating mechanisms. Short-term pathways remain highly sensitive to rapid changes in the nearest window, while long-term pathways accumulate memory of slow drifts across multiple windows. The gating units automatically adjust the information passage ratio based on input strength and historical state to avoid gradient vanishing and bursting. Multi-layer stacking and residual connections are used to simultaneously retain information at different time scales, outputting a temporal feature sequence containing a context-aware representation and internal state summary for each time position, carrying the global and local evidence needed for trend judgment.

[0076] The output layer maps the temporal feature sequence to a probability distribution of emotional states at future time steps. Several future time positions are defined according to the prediction range, and a probability vector for discrete emotional categories or intervalized continuous emotional dimensions is calculated for each position. Temperature scaling or edge calibration is used to improve the reliability of the probability scale, and uncertainty assessment is introduced to quantify the model's confidence differences across different future positions. To suppress the misleading influence of occasional noise, lightweight smoothing and monotonicity constraints are introduced after the output layer to reduce the probability of sharp reversals when historical evidence points to a stable trend.

[0077] The sentiment change trend is generated from a complete probability distribution. Based on the probability vectors of each future position, the expected trajectory, directional derivative, and strength of change index are calculated to form a time-stamped trend curve. Uncertainty is used as a weighting factor to apply confidence weights to the trajectory, resulting in the main trend and alternative branches, while retaining the probability of occurrence and switching conditions for each branch. The final generated sentiment change trend includes not only direction and amplitude but also confidence structure and time alignment information that can be used by upstream decision-making modules, ensuring seamless integration with intent parsing and interaction status.

[0078] This embodiment unifies the timeline and scale, jointly models on both local and global time scales, and probabilistically characterizes future uncertainties, transforming them into actionable trend trajectories. This provides stable and reliable emotional trajectory input for subsequent planning, thereby reducing the impact of input noise and individual differences on prediction, minimizing path jitter caused by misjudgments during dialogue, and making content selection and expression parameter settings closer to the user's upcoming psychological state, thus improving response consistency and decision hit rate.

[0079] In one embodiment, step S30 above includes: S301, The user's emotional state and intentional needs are parsed from the standardized multimodal user information; S302, determine the core part of the interactive content based on the user's intent requirements; S303, determine the strength level of the interactive tone and the speed level of the interactive speech based on the user's emotional state. S304, determine the guidance method for interactive content based on the changing trend of the user's emotional state; S305, combining the preliminary interaction scheme, the core part of the interaction content, the strength level of the interaction tone, the speed level of the interaction speed, and the guidance method of the interaction content, a collaborative control strategy containing execution instructions is generated.

[0080] In this embodiment, the input is a preliminary interaction scheme and standardized multimodal user information, and the output is a collaborative control strategy to guide the execution of the interaction.

[0081] First, user emotional states and intent needs are analyzed from standardized multimodal user information. Emotional states are derived from acoustic prosody, semantic sentiment cues, and recent emotional trajectories, expressed as vectors representing strength levels and stability. Intent needs are derived from trigger words, object words, and constraints in textual fragments, forming intent tags and parameter sets. The analysis process maintains time window consistency, mapping emotional vectors and intent structures to the same time base, and includes confidence weights and source markers to facilitate subsequent weighted fusion and source tracing.

[0082] The core components of the interactive content are determined based on the user's intent. To this end, a content element library and a topic index are introduced. Intent tags and parameter sets are bidirectionally matched with semantic units in the element library, outputting a set of core units and their sequential relationships. If the intent contains multiple objectives or is ambiguous, conflict resolution is performed according to confidence weights and contextual coherence, preserving the main sequence and candidate branches, and recording entry and exit conditions to ensure that subsequent execution stages can switch according to state.

[0083] The strength and speed of the interactive tone are determined based on the user's emotional state. The tone level is estimated by both emotional intensity and fluctuation amplitude, while the speed level is estimated by both prosody and attention load indicators. Both are generated as discrete levels using a segmented mapping relationship, with continuously adjustable intervals for fine-grained adjustments during the synthesis and execution phases. To avoid abrupt changes in level that could cause jarring shifts, a time smoothing gate is introduced, which weights and merges the output from the previous period with the current estimated result to output a smoother level sequence.

[0084] The guidance method for interactive content is determined based on the changing trends of the user's emotional state. A trend vector describes the direction and magnitude of the interaction. Combined with the completion level of the core components and the most recent user feedback, the guidance strategy is presumed to be one or a combination of progressive, explanatory, reassuring, or convergent approaches, and a switching threshold and observation window length are provided. When trend uncertainty increases, a more fault-tolerant guidance method is prioritized, and the observation window is extended to reduce interaction jitter caused by frequent switching.

[0085] Finally, the initial interaction plan, the core parts of the interaction content, the strength levels of the interaction tone, the speed levels of the interaction speed, and the guidance methods are synthesized to generate a collaborative control strategy containing execution instructions. The synthesis process proceeds in parallel along two pathways: semantic and expressive. The semantic pathway, based on the topic nodes and jump conditions in the initial interaction plan, inserts core units and guiding nodes to form a content sequence and branching structure, along with entry and exit criteria. The expressive pathway generates a speech output parameter table based on the tone and speed levels, and aligns the parameter fragments with the content sequence one by one to form a timetable. The products of the two pathways are aligned and validated within the synthesizer, outputting a unified strategy object containing the content sequence, jump logic, trigger conditions, speech output parameters, and fallback anchors, while retaining source markers and version numbers for easy updates and rollbacks during execution.

[0086] This embodiment, by combining emotional state, intent, and trend information on the same time base, first determines the content focus, then sets the tone and speed level, and simultaneously selects a matching guidance method. Finally, it synthesizes this with the initial interaction plan to form a unified execution instruction, achieving consistency in both content and expression. This reduces expression deviations and unnecessary content wandering caused by emotional mismatch, lowers the frequency of strategy switching and ineffective back-and-forth, improves response consistency and controllability, and provides a structured foundation with version and source tags for adaptive updates during execution.

[0087] In one embodiment, step S40 above includes: S401, during the execution of the collaborative control strategy, real-time voice information, real-time semantic information and real-time emotion information are collected in parallel by the multimodal perception unit; S402, compare the real-time emotion information with the user's emotion change trend by emotion vector comparison, compare the real-time semantic information with the user's intention needs by semantic distance comparison, and compare the real-time voice information with the acoustic elements corresponding to the strength level of the interactive tone and the speed level of the interactive speech. S403, when any comparison result exceeds the corresponding threshold range or multiple comparison results show coordinated deviation, it is determined that the multidimensional adjustment condition is triggered. S404, select the corresponding adaptive planning module to replan the interaction path based on the specific adjustment dimension triggered; S405, Based on the replanned interaction path and the real-time emotion information, generate a new interaction scheme that includes the strength level of the updated interaction tone and the content sequence. S406, the new interaction scheme is mapped into an execution instruction, and the speech synthesis parameters, interaction logic sequence and feedback triggering conditions in the collaborative control strategy are updated based on the execution instruction.

[0088] In this embodiment, during the execution of the collaborative control strategy, the parallel acquisition of real-time speech, semantic, and emotional information by the multimodal perception unit relies on low-latency data channels and a unified time reference. The multimodal perception unit establishes separate buffer queues for the audio, text, and emotional signals, and adds a timestamp and source marker to each segment during sampling. Real-time speech information is processed at the short-time frame level, preserving acoustic elements such as pitch contours, energy envelopes, speech rate estimation, and formant trajectories. Real-time semantic information consists of segment-level text output from speech transcription and a context window, including structured representations of trigger words, object words, and qualifier words. Real-time emotional information is composed of a vector expression obtained by fusing acoustic prosody and semantic sentiment cues, including emotion category, intensity level, and confidence weight. The three signals are aligned within a sliding time window, ensuring a one-to-one correspondence between speech, semantics, and emotion within the same window, facilitating synchronous calculations in subsequent comparison stages.

[0089] The process of comparing real-time emotion information with the user's emotion change trend using emotion vectors is based on a vector space of the same dimension and a calibrated time axis. In each window, the system calculates the distance metric and directional consistency index between the current emotion vector and the trend vector, and weights the distance based on trend confidence, outputting emotion deviation and directional deviation signs to quantify the difference from the expected trend. The process of comparing real-time semantic information with the user's intent needs using semantic embedding space is based on a shared semantic embedding space. The system calculates the matching degree between fragment-level semantic representations and intent tags and their slot sets, while simultaneously evaluating the satisfaction of named entities, consistency constraints, and negation transitions, outputting semantic matching degree and slot coverage. The process of comparing real-time voice information with the acoustic elements corresponding to the strength and speed levels of interactive intonation and speech rate uses a policy-side expected parameter table as a reference. By mapping intonation levels to energy and pitch intervals, and speech rate levels to beat and pause intervals, interval consistency and offset calculations are performed with the statistics of real-time acoustic elements to obtain acoustic consistency and offset direction. The three comparison results retain the time index and source marker, which are used for multidimensional combination in subsequent trigger determination.

[0090] When any comparison result exceeds the corresponding threshold range or multiple comparison results show coordinated deviation, the process of determining the trigger for multi-dimensional adjustment conditions is based on a threshold table and a set of coordinated rules. The threshold table defines the allowable and alarm intervals for sentiment deviation, semantic matching, and acoustic consistency, respectively. The set of coordinated rules is used to identify cross-modal consistency imbalances, such as the simultaneous occurrence of negative sentiment deviation and decreased semantic matching, or high semantic matching with persistently low acoustic consistency. The system first performs threshold determination on single-modal indicators; if any indicator enters the alarm interval, a single-dimensional trigger marker is generated. Subsequently, a consistency check is performed on cross-modal combinations; if a rule-matched combination event occurs, a coordinated trigger marker is generated. Multi-dimensional adjustment conditions consist of a set of trigger dimensions, deviation intensity, duration, and uncertainty weights, used to guide the subsequent selection of appropriate planning channels.

[0091] The process of replanning the interaction path based on the specific adjustment dimension triggered uses a dimension-to-module mapping table as the entry point. When the trigger is primarily based on the emotion dimension, expression adjustment planning is prioritized, resetting the tone and speed levels and adjusting the guidance method. When the trigger is primarily based on the semantic dimension, semantic path planning is prioritized, reordering or replacing the content sequence and node jump conditions. When there is a co-deviation between emotion and semantics, the joint planning module is activated, simultaneously adjusting the expression parameters and semantic path in a coupled manner. Path replanning is performed under constraints and the current session state snapshot. A cost function is used to evaluate candidate branches, with the cost term consisting of emotion deviation, semantic matching gap, and policy stability. Backtracking anchors are introduced during the search process to ensure smooth rollback when necessary.

[0092] The process of generating a new interaction scheme based on the replanned interaction path and real-time emotion information, including the updated strength levels of the interactive tone and the content sequence, is characterized by the parallel synthesis of semantic and expressive pathways. The semantic pathway generates the content sequence and node transition conditions on the new path and performs temporal positioning of sequence segments by window. The expressive pathway generates a speech output parameter timetable based on the real-time emotion vector and the updated tone and speed levels, and smooths transitions at boundary segments to avoid abrupt changes. The two pathways complete alignment and consistency checks in the scheme synthesizer, outputting a new interaction scheme that includes the content sequence, expressive parameters, and switching conditions, while retaining the trigger source, threshold hit record, and version number to support traceable updates to subsequent strategies.

[0093] The process of mapping new interaction schemes to execution commands and updating the speech synthesis parameters, interaction logic sequences, and feedback triggering conditions in the collaborative control strategy based on these commands requires atomic updates and seamless switching. The mapping phase converts the content sequence into executable logical nodes and edges, generating jump condition tables and fallback anchor point tables; it converts the expression parameter timetable into parameter blocks that the speech synthesis engine can load and marks their effective time; it binds feedback triggering conditions to comparison indicators, clarifying the thresholds and observation windows used in the next round of monitoring. The strategy update employs a versioned loading and dual-caching mechanism. New parameter blocks take effect at predetermined boundary times, and old versions are released after the transition is complete, ensuring continuity and stability during the dialogue.

[0094] This embodiment continuously collects multimodal real-time signals during execution and compares the emotion, semantic, and acoustic pathways in parallel. It combines thresholds and collaborative rules for multidimensional trigger determination, and then selects an adaptive planning channel based on the trigger dimension to synchronously reconstruct the path and expression. This enables the interaction to simultaneously align with the user's current state in terms of content selection and expression parameters. This closed-loop mechanism reduces the lag and mismatch caused by static strategies, decreases invalid branches and frequent round trips, improves the controllability and smoothness of strategy switching, and achieves uninterrupted replacement through versioned instructions and dual-caching updates. Thus, under conditions of equal computing resources, it achieves higher response continuity and execution stability, providing high-quality, traceable input for subsequent rounds of monitoring and readjustment.

[0095] In one embodiment, step S50 above includes: S501, after the interactive task is completed, start the effect analysis module; S502, The effect analysis module collects interaction duration data and user emotion change data during the interaction process; S503, determine the interaction duration index based on the interaction duration data, and determine the emotion stability index, the completion index of the interaction goal, and the user cooperation index based on the user emotion change data. S504, compare the interaction duration index, the emotional stability index, the completion index, and the cooperation index with preset analysis standards; S505: When the comparison results show that there are indicators that do not meet the standards, the adaptive optimization module is activated. S506, the adaptive optimization module adjusts the model parameters of the time-series prediction model, optimizes the strategy parameters of tone and speech rate control during the collaborative control strategy generation process, and updates the path planning parameters for generating the interaction scheme.

[0096] In this embodiment, an offline review and parameter correction process is triggered after the interaction task ends. First, the effect analysis module takes over the full record of the current session, synchronously retrieving the voice stream, text transcription, emotion vector trajectory, strategy execution log, and trigger event list according to the session identifier and timeline, generating replayable data slices and summary metadata. Interaction duration data extracts basic quantities from the session start and end times, the proportion of silent segments, the number of effective turns, and the response interval. After excluding waiting access and non-interaction periods, the net interaction duration and turn density are calculated and aligned to the session timeline using a unified time base. User emotion change data is aggregated from frame-level emotion vectors into segment-level trajectories, preserving the emotion category sequence, strength level sequence, and confidence weights, and performing robust interpolation on missing segments to ensure trajectory continuity.

[0097] The metrics are calculated using a unified standard. The interaction duration metric is obtained by weighting net interaction duration with turn-taking density, and retaining quantile summaries to reflect the duration distribution. The emotion stability metric is jointly defined based on the fluctuation amplitude of the emotion vector trajectory, the number of direction reversals, and the confidence weighted variance, outputting a stability score on a zero-mean scale. The completion rate metric for interaction goals is derived from parsing the arrival and exit markers of task nodes from the strategy execution log, statistically analyzing the achievement rate and key node coverage based on a preset goal set. The user cooperation metric combines response interval, interruption frequency, the proportion of negative tones, and the frequency of clarification requests to calculate a comprehensive cooperation score, providing segmented results at different dialogue stages to pinpoint the stage at which problems occur. Metrics and original segments are mutually referenced, and each metric is bound to its calculation source and time window, facilitating subsequent tracing and diagnosis.

[0098] The comparison phase compares interaction duration, emotional stability, completion rate, and cooperation rate indicators against the analysis standards one by one. The analysis standards consist of three categories: historical distribution threshold, business expectation line, and safety threshold, which can be further subdivided by group, scenario, and time period. A ternary judgment is used for comparison, outputting "pass," "warning," and "failure," respectively, and recording the deviation magnitude and confidence interval. When at least one indicator fails to meet the standard, or multiple indicators simultaneously enter the warning interval and the collaboration rule is triggered, the process transitions to the adaptive optimization module. The collaboration rule is used to identify structural anomalies between indicators, such as combinations of low completion rate and poor emotional stability, or low cooperation rate and excessively long duration, to avoid over-correction caused by occasional occurrences of a single indicator.

[0099] The adaptive optimization module performs parameter updates for three object pathways. For the time-series prediction model, an incremental sample set is first constructed on the interaction records, keeping the weights of the original training distribution unchanged. The current session's emotion vector trajectory and trend error are introduced, the gradient is calculated, and only the sensitive sublayer and normalization scale are updated to limit large drifts. At the same time, the regularization strength and time smoothing coefficient are adjusted according to the deviation direction of the emotion stability index to make future predictions more robust in noisy regions. For the policy parameters of tone and speed control in the collaborative control strategy generation process, the affected parameter dimensions are extracted from the table mapping the unmet indicators to the parameter set, such as the interval boundaries of the tone level mapping table, the target beat and pause ratio of the speed level, and the threshold for switching the guidance method. Small-step updates are performed according to the deviation magnitude. After the update, a fast simulation is performed on the session playback engine to verify the expression smoothness and interval consistency. After passing the consistency check, it is written into the new version. For the path planning parameters used to generate interaction solutions, the weights of the emotion cost term, semantic matching gap term, and stability penalty term in the cost function are recalibrated. When the completion rate is low, the weights of key nodes are increased and the exploration upper limit of unnecessary branches is reduced. When the cooperation rate is low, the priority of explanatory and reassuring branches is increased and the observation window is extended to reduce frequent switching. After updating the path planning parameters, a bundle search replay is performed on typical dialogue segments to evaluate branch jitter and convergence depth. The solution is released after meeting the threshold.

[0100] All parameter updates employ a versioning and atomic replacement strategy. The adaptive optimization module creates new versions in the parameter repository, recording the version number, impact scope, triggering indicators, and deviation summary. The deployment process uses double-buffering loading, seamlessly switching between old and new versions at safety boundaries. If subsequent monitoring detects regression risks, the previous version is immediately restored via rollback anchors. The performance analysis module and the adaptive optimization module exchange comparison results and update receipts via an event bus, forming a closed-loop link. When the next session starts, the new time-series prediction model, strategy parameters, and path planning parameters take effect simultaneously, with monitoring metrics updated synchronously, facilitating continuous evaluation and further fine-tuning.

[0101] Example Description: In a fintech outbound calling system, this interactive control method is applied to an intelligent voice outbound calling service platform to achieve real-time emotion perception, adaptive strategy adjustment, and interaction process optimization during customer communication. When initiating an outbound call task, the system simultaneously receives customer voice signals, semantic content, and historical interaction data through a multimodal acquisition module. The speech recognition module transcribes the voice signal into structured text and extracts acoustic elements such as pitch, rhythm, intensity, and speech rate; the semantic understanding module identifies the customer's intent from the text, such as inquiring about loan interest rates, verifying transaction information, or modifying account settings; the emotion analysis module judges the customer's emotional state, such as indifference, confusion, or anxiety, based on voice energy distribution and semantic polarity features; and the user profiling module combines historical interaction records and business tags to generate profile information such as customer preferences, communication habits, and risk levels. After noise reduction, semantic cleaning, and standardization, the system integrates these multimodal features to form standardized multimodal user information with a unified structure.

[0102] Based on standardized user information, the system extracts customer emotional characteristics and intent needs. Emotional characteristics are represented by vectors based on dimensions such as tone, speech rate, and semantic polarity, while intent needs are expressed as a set of semantic slots. The system inputs these emotional characteristics into a time-series prediction model for trend prediction. The model captures local features of tone and emotional fluctuations in the feature extraction layer, identifies patterns of change over time in the memory network layer, outputs the probability distribution of emotional states for future time steps, and generates a trend of customer emotional changes. This trend is combined with the customer's current intent needs and the outbound call task status to determine the constraints of the interaction path. For example, when the model predicts that the customer's emotions are becoming tense, the system reduces the frequency of marketing statements in the path planning and switches the response strategy to explanation and reassurance-oriented to maintain dialogue stability.

[0103] Combining emotional trends, intent needs, and the current interaction state, the system generates a preliminary interaction plan, clarifying the priority, tone level, and strategy triggering conditions of the response content. The collaborative control strategy generation module further refines the execution parameters based on the preliminary interaction plan and standardized multimodal information. First, it analyzes the customer's current emotional state and business intent to determine the core part of the interaction content, such as focusing on explaining the repayment plan or fee details. Then, it determines the tone strength and speech rate levels based on the emotional state; when the customer's emotions tend to be negative, the speech rate is reduced and pauses are increased; conversely, the pace can be appropriately increased to maintain attention. The system then determines the guidance method based on emotional trends; for example, direct guidance is used when the customer's emotions are predicted to be stable, while indirect guidance is used when emotions are fluctuating. Finally, the tone level, speech rate level, core content, and guidance method are integrated to generate a collaborative control strategy containing executable instructions, covering speech synthesis parameters, interaction logic nodes, and feedback triggering conditions.

[0104] During the execution of the collaborative control strategy, the multimodal perception unit collects real-time speech, semantic, and emotion information in parallel and compares it with the initial interaction plan. The system performs vector comparison between real-time emotion information and predicted emotion trends to calculate the deviation degree and direction; it compares real-time semantic information with the target intent demand to determine whether the customer has deviated from the predetermined topic; and it matches real-time speech features with the tone and speed parameters defined in the strategy using acoustic elements. If any result exceeds the threshold range or multiple collaborative deviations occur, the system determines that multi-dimensional adjustment conditions are triggered and selects the adaptive planning module to replan the interaction path. When the trigger dimension is emotion, the system adjusts the tone parameters to reduce the sales pitch; when the semantic deviation is significant, the system regenerates the content sequence to guide the customer back to the target topic. The synthesized new plan is mapped to execution instructions, and the speech synthesis parameters, logical sequences, and feedback conditions of the collaborative control strategy are atomically updated to ensure that the task is not interrupted.

[0105] After a complete outbound call task is completed, the performance analysis module is activated to analyze the data from the entire call process. The system collects data on interaction duration, emotional change trajectory, goal achievement rate, and customer cooperation level. Interaction duration reflects communication efficiency, emotional stability indicates the amplitude of customer emotional fluctuations, goal achievement rate measures the proportion of business completed, and cooperation level reflects the customer's responsiveness. The system compares these indicators with preset standards, and if any items fail to meet the standards, the adaptive optimization module is triggered. The optimization module refines and adjusts the time weight of the time-series prediction model, collaborative control strategy, and path planning parameters. When emotional prediction is unstable, the system adjusts the time weight of the memory layer; when the tone strategy deviates from customer expectations, the interval boundaries of the tone level mapping table are optimized; when the path planning complexity is too high, leading to an increase in task interruption rate, the branch cost weights are reallocated to reduce cognitive burden. All parameter updates use versioned loading and dual-caching switching to ensure smooth system upgrades and traceability.

[0106] In fintech scenarios, this mechanism enables the intelligent outbound calling system to perceive and adapt its strategies to the real-time status of customers. The system can dynamically adjust dialogue across various business processes, including outbound marketing, account verification, and repayment reminders. When it detects hesitation or negative emotions in a customer's tone, it automatically switches to explanatory statements to ease the atmosphere; when the customer shows positive cooperation, it promptly advances the business confirmation process. Through continuous data collection and parameter optimization, the system gradually develops dynamic interaction patterns tailored to different customer groups, making subsequent outbound calling tasks more efficient and more aligned with individual characteristics, significantly improving call completion rates and customer experience consistency.

[0107] In the healthcare field, this interactive control method can be applied to intelligent terminal systems for chronic disease management, remote rehabilitation guidance, or elderly health companionship, establishing a dynamic, emotion-aware interaction mechanism between the device and the user. The system embeds multimodal perception components to collect multi-source information from user speech, facial expressions, tone changes, and behavioral actions. In the speech recognition module, the user's speech signal is analyzed frame by frame to extract elements such as pitch, energy, and speech rate. Simultaneously, the semantic understanding module identifies the user's intended needs within the language content, such as medication reminders, rehabilitation training guidance, or health status inquiries. The emotion analysis module determines emotional states, such as anxiety, fatigue, or positivity, based on acoustic and semantic features. This is combined with long-term accumulated lifestyle habits, emotional patterns, and health records from the user profiling module to construct an individualized user profile. Through data fusion, standardized multimodal user information is formed, ensuring consistency in input feature dimensions and providing a unified expression for subsequent predictions.

[0108] After acquiring standardized multimodal user information, the system extracts key dimensions reflecting tone intensity, pitch fluctuations, and semantic polarity through an emotion feature extraction module. These dimensions are then combined with the user's intent and needs input into a time-series prediction model for trend inference. The model uses a local feature extraction layer to identify short-term emotion changes, a memory network layer to establish long-term and short-term dependencies, and outputs a time-series feature sequence. The output layer calculates the probability distribution of emotional states over future time periods to infer potential trends in user emotions. For example, in long-term rehabilitation support, if a declining trend in the user's emotional state is predicted, the system can adjust the interaction method in advance, reducing the intensity of instructions or extending the feedback interval to prevent emotional fluctuations from affecting rehabilitation execution.

[0109] By combining user emotional trends, intent needs, and current interaction status, the system determines the constraints of the interaction path. If user emotional fluctuations are predicted, the redundancy of the interaction content and the speed of feedback are adjusted; if the user's emotions are stable, the task complexity is appropriately increased to enhance their motivation. Based on these conditions and user communication preferences, a preliminary interaction plan is generated, which includes a sequence of content instructions, tone parameters, and behavioral feedback strategies, providing input for subsequent collaborative control.

[0110] The collaborative control strategy generation process takes a preliminary interaction plan and standardized user information as input. First, it analyzes the current emotional state and intended needs, extracts the core parts of the interaction content, and determines the levels of tone strength and speech rate through parameter mapping. The system generates guidance methods based on emotional change trends; for example, it guides users through rehabilitation training with a gentle tone or maintains stable interaction with a neutral tone during health reminders. The generated collaborative control strategy includes not only speech synthesis parameters but also logical control sequences and feedback trigger conditions, ensuring adaptive stability of interaction execution under different states.

[0111] During the execution of the collaborative control strategy, the system's multimodal perception unit collects speech, semantic, and emotion information in real time and compares the real-time data with the parameters of the initial interaction plan. Real-time emotion information is compared with predicted trends using vector comparison; semantic information is calculated based on semantic distance to the user's intent; and speech features are matched with acoustic elements at the current tone and speed level. When any result exceeds a threshold or exhibits multidimensional deviation, the system triggers the adaptive planning module to replan the interaction path. The new interaction plan is updated synchronously with the guidance strategy and tone parameters, such as automatically reducing the speech speed and adding encouraging statements when user emotional fluctuations are detected, thereby stabilizing the interactive experience. The replanned plan is mapped to new execution instructions, updating the speech synthesis parameters and logical sequence to ensure the control strategy remains consistent with the user's state.

[0112] After a complete health interaction task is completed, the system enters the analysis phase. The effect analysis module collects data from the entire interaction process, calculates indicators such as interaction duration, emotional stability, task completion rate, and user cooperation, and compares them with preset analysis standards. When any indicators fail to meet the standards, the adaptive optimization module is activated, fine-tuning the parameters of the time-series prediction model, the tone and speed control parameters of the collaborative control strategy, and the path planning parameters. For example, when user cooperation is low, the system reduces instruction density and increases semantic redundancy to improve comprehension; when emotional stability indicators are insufficient, it adds emotional smoothing terms or increases the weight of positive feedback. Through this optimization mechanism, the system has better emotional prediction capabilities and interaction matching in subsequent interactions.

[0113] In practical applications in the healthcare field, this mechanism manifests as an intelligent health assistant automatically adjusting its communication style based on the voice state of elderly or rehabilitated patients. When the system detects signs of fatigue in a user's tone, it automatically switches to a soothing mode and extends the response time to reduce cognitive stress. After detecting an improvement in mood, it resumes normal speech rate and task rhythm, achieving continuous adaptive health companionship and behavioral intervention. Simultaneously, the system's analysis and optimization of interaction data ensures that subsequent interactions better align with the user's psychological state and rhythmic characteristics, forming a closed-loop emotional perception and dynamic adjustment capability. This significantly enhances the naturalness and continuity of health management interactions in non-hospital environments.

[0114] This embodiment calculates interaction duration, emotional stability, completion, and cooperation metrics using a unified time benchmark and traceability after the interaction ends. These metrics are then compared in a structured manner using analytical standards. Deviations are mapped to the model parameters of the time-series prediction model, the policy parameters for tone and speed control during the collaborative control strategy generation process, and the path planning parameters for generating the interaction plan, achieving targeted and incremental corrections with small steps. This closed loop makes predictions more robust, expressions more relevant, and paths more convergent, reducing trend misjudgments and invalid branch explorations in subsequent conversations, decreasing the frequency of abrupt changes in expression and strategy jitter, and improving the achievement rate and continuity of the next round of interaction without increasing computational and bandwidth overhead. Furthermore, versioning and rollback mechanisms ensure the security, controllability, and reproducibility of the update process.

[0115] In one embodiment, a user-state-aware interactive control device is provided, which corresponds one-to-one with the user-state-aware interactive control method described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the user state-aware interactive control device of the present invention. The modules include a multimodal perception and preprocessing module 10, an emotion prediction and solution generation module 20, a collaborative control strategy generation module 30, an adaptive interaction execution module 40, and an interaction analysis and self-optimization module 50. Detailed descriptions of each functional module are as follows: The multimodal perception and preprocessing module 10 is used to collect and preprocess the multimodal raw data generated during user interaction to generate standardized multimodal user information. The emotion prediction and solution generation module 20 is used to predict the user's emotion change trend based on the standardized multimodal user information, and generate a preliminary interaction solution by combining the user's emotion change trend, the user's intention and needs and the current interaction state. The collaborative control strategy generation module 30 is used to generate a collaborative control strategy to guide the execution of the interaction based on the preliminary interaction scheme and the standardized multimodal user information. The adaptive interaction execution module 40 is used to continuously acquire real-time multimodal user information during the execution of the collaborative control strategy, compare the real-time multimodal user information with the preliminary interaction scheme, and when the adjustment condition is triggered, re-plan and generate a new interaction scheme and update the collaborative control strategy. The interaction analysis and self-optimization module 50 is used to analyze the interaction performance based on the interaction process data after the interaction task is completed. When the analysis results do not meet the preset standards, it optimizes the prediction process of the user's emotional change trend, the generation process of the collaborative control strategy, and the generation process of the interaction scheme.

[0116] In one embodiment, the multimodal sensing and preprocessing module 10 is specifically used for: The user's voice information is collected through a speech recognition module; The semantic understanding module identifies the user's intent and needs from the user's voice information; The emotion analysis module determines the user's emotional state based on the user's voice information. User profile information is constructed by combining historical interaction data with the user profile module; The user's voice information is subjected to noise reduction processing to generate noise-reduced user voice information; The user intent requirements are semantically cleaned to generate cleaned user intent requirements; The user's emotional state is standardized to generate a standardized user emotional state. Feature extraction is performed on the user profile information to generate user profile features; The noise-reduced user voice information, the cleaned user intent and needs, the standardized user emotional state, and the user profile features are fused together to generate standardized multimodal user information.

[0117] In one embodiment, the emotion prediction and solution generation module 20 is specifically used for: Extract user emotional features and user intent needs from the standardized multimodal user information; The user's emotional characteristics are processed using a time-series prediction model to predict the trend of user emotional changes. Obtain the current interaction state maintained by the state management unit; Based on the described emotional change trends and the user's intentions and needs, determine the constraints of the interaction path; Based on the constraints, the current interaction state, and the user's communication preferences, a preliminary interaction plan is generated.

[0118] In one embodiment, the emotion prediction and solution generation module 20 is specifically used for: The user's emotional features are subjected to time-series standardization processing to generate a standardized emotional feature sequence; The standardized emotion feature sequence is input into the feature extraction layer of the time-series prediction model for local feature extraction to obtain local emotion features; The local emotion features are input into the memory network layer of the time-series prediction model to model long-short-term dependencies and generate a time-series feature sequence. Based on the temporal feature sequence, the probability distribution of emotional state at future time steps is determined through the output layer of the temporal prediction model; The user's emotional change trend is generated based on the probability distribution of emotional state at the future time step.

[0119] In one embodiment, the cooperative control strategy generation module 30 is specifically used for: The user's emotional state and intentional needs are extracted from the standardized multimodal user information. Determine the core part of the interactive content based on the user's intent and needs; The strength level of the interactive tone and the speed level of the interactive speech are determined based on the user's emotional state. The guidance method for interactive content is determined based on the changing trends of the user's emotional state; By combining the preliminary interaction scheme, the core part of the interaction content, the strength level of the interaction tone, the speed level of the interaction speed, and the guidance method of the interaction content, a collaborative control strategy containing execution instructions is generated.

[0120] In one embodiment, the adaptive interactive execution module 40 is specifically used for: During the execution of the collaborative control strategy, real-time speech information, real-time semantic information, and real-time emotion information are collected in parallel by a multimodal perception unit. The real-time emotion information is compared with the user's emotion change trend to form an emotion vector; the real-time semantic information is compared with the user's intention needs to form a semantic distance; and the real-time voice information is compared with the acoustic elements corresponding to the strength level of the interactive tone and the speed level of the interactive speech. When any comparison result exceeds the corresponding threshold range or multiple comparison results show coordinated deviation, the multidimensional adjustment condition is determined to be triggered. Based on the specific adjustment dimension triggered, select the corresponding adaptive planning module to replan the interaction path; Based on the replanned interaction path and the real-time emotion information, a new interaction scheme is generated, which includes the strength level of the updated interaction tone and the content sequence. The new interaction scheme is mapped to execution instructions, and the speech synthesis parameters, interaction logic sequences, and feedback triggering conditions in the collaborative control strategy are updated based on the execution instructions.

[0121] In one embodiment, the interactive analysis and self-optimization module 50 is specifically used for: After the interactive task is completed, start the effect analysis module; The effect analysis module collects interaction duration data and user emotion change data during the interaction process. Based on the interaction duration data, an interaction duration indicator is determined; based on the user emotion change data, an emotion stability indicator, an interaction goal completion indicator, and a user cooperation indicator are determined. The interaction duration index, the emotional stability index, the completion index, and the cooperation index are compared with preset analysis standards. When the comparison results show that there are indicators that have not met the standards, the adaptive optimization module is activated; The adaptive optimization module adjusts the model parameters of the time-series prediction model, optimizes the strategy parameters for tone and speed control during the collaborative control strategy generation process, and updates the path planning parameters for generating the interaction scheme.

[0122] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides deterministic and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a user-state-aware interactive control method on the server side.

[0123] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements client-side functions or steps of a user-state-aware interactive control method.

[0124] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Collect and preprocess multimodal raw data generated during user interaction to generate standardized multimodal user information; Based on the standardized multimodal user information, predict the user's emotional change trend, and combine the user's emotional change trend, user's intention and needs with the current interaction state to generate a preliminary interaction plan; Based on the preliminary interaction scheme and the standardized multimodal user information, a collaborative control strategy is generated to guide the execution of the interaction. During the execution of the collaborative control strategy, real-time multimodal user information is continuously acquired and compared with the preliminary interaction scheme. When the adjustment condition is triggered, a new interaction scheme is re-planned and generated, and the collaborative control strategy is updated. After the interaction task is completed, the interaction performance is analyzed based on the interaction process data. When the analysis results do not meet the preset standards, the prediction process of the user's emotional change trend, the generation process of the collaborative control strategy, and the generation process of the interaction plan are optimized.

[0125] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Collect and preprocess multimodal raw data generated during user interaction to generate standardized multimodal user information; Based on the standardized multimodal user information, predict the user's emotional change trend, and combine the user's emotional change trend, user's intention and needs with the current interaction state to generate a preliminary interaction plan; Based on the preliminary interaction scheme and the standardized multimodal user information, a collaborative control strategy is generated to guide the execution of the interaction. During the execution of the collaborative control strategy, real-time multimodal user information is continuously acquired and compared with the preliminary interaction scheme. When the adjustment condition is triggered, a new interaction scheme is re-planned and generated, and the collaborative control strategy is updated. After the interaction task is completed, the interaction performance is analyzed based on the interaction process data. When the analysis results do not meet the preset standards, the prediction process of the user's emotional change trend, the generation process of the collaborative control strategy, and the generation process of the interaction plan are optimized.

[0126] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0127] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0128] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0129] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

[0130] The user personal information involved in this application embodiment is all authorized (knowing and consenting) by the relevant parties or fully authorized by all parties, and the executing entity can obtain it through various open, legal and compliant means. The collection, storage, use, processing, transmission, provision and disclosure of the information, data and signals involved all comply with the relevant laws and regulations of the relevant countries and regions, and do not violate public order and good morals.

Claims

1. An interactive control method based on user state awareness, characterized in that, Includes the following steps: Collect and preprocess multimodal raw data generated during user interaction to generate standardized multimodal user information; Based on the standardized multimodal user information, predict the user's emotional change trend, and combine the user's emotional change trend, user's intention and needs with the current interaction state to generate a preliminary interaction plan; Based on the preliminary interaction scheme and the standardized multimodal user information, a collaborative control strategy is generated to guide the execution of the interaction. During the execution of the collaborative control strategy, real-time multimodal user information is continuously acquired and compared with the preliminary interaction scheme. When the adjustment condition is triggered, a new interaction scheme is re-planned and generated, and the collaborative control strategy is updated. After the interaction task is completed, the interaction performance is analyzed based on the interaction process data. When the analysis results do not meet the preset standards, the prediction process of the user's emotional change trend, the generation process of the collaborative control strategy, and the generation process of the interaction plan are optimized.

2. The user-state-aware interactive control method as described in claim 1, characterized in that, Collect and preprocess multimodal raw data generated during user interaction to generate standardized multimodal user information, including: The user's voice information is collected through a speech recognition module; The semantic understanding module identifies the user's intent and needs from the user's voice information; The emotion analysis module determines the user's emotional state based on the user's voice information. User profile information is constructed by combining historical interaction data with the user profile module; The user's voice information is subjected to noise reduction processing to generate noise-reduced user voice information; The user intent requirements are semantically cleaned to generate cleaned user intent requirements; The user's emotional state is standardized to generate a standardized user emotional state. Feature extraction is performed on the user profile information to generate user profile features; The noise-reduced user voice information, the cleaned user intent and needs, the standardized user emotional state, and the user profile features are fused together to generate standardized multimodal user information.

3. The user-state-aware interactive control method as described in claim 1, characterized in that, Based on the standardized multimodal user information, predict the user's emotional change trend, and combine the user's emotional change trend, user intention needs, and current interaction state to generate a preliminary interaction plan, including: Extract user emotional features and user intent needs from the standardized multimodal user information; The user's emotional characteristics are processed using a time-series prediction model to predict the trend of user emotional changes. Obtain the current interaction state maintained by the state management unit; Based on the described emotional change trends and the user's intentions and needs, determine the constraints of the interaction path; Based on the constraints, the current interaction state, and the user's communication preferences, a preliminary interaction plan is generated.

4. The user-state-aware interactive control method as described in claim 3, characterized in that, The user's emotional characteristics are processed using a time-series prediction model to predict the trend of changes in the user's emotions, including: The user's emotional features are subjected to time-series standardization processing to generate a standardized emotional feature sequence; The standardized emotion feature sequence is input into the feature extraction layer of the time-series prediction model for local feature extraction to obtain local emotion features; The local emotion features are input into the memory network layer of the time-series prediction model to model long-short-term dependencies and generate a time-series feature sequence. Based on the temporal feature sequence, the probability distribution of emotional state at future time steps is determined through the output layer of the temporal prediction model; The user's emotional change trend is generated based on the probability distribution of emotional state at the future time step.

5. The user-state-aware interactive control method as described in claim 1, characterized in that, Based on the preliminary interaction scheme and the standardized multimodal user information, a collaborative control strategy is generated to guide the execution of the interaction, including: The user's emotional state and intentional needs are extracted from the standardized multimodal user information. Determine the core part of the interactive content based on the user's intent and needs; The strength level of the interactive tone and the speed level of the interactive speech are determined based on the user's emotional state. The guidance method for interactive content is determined based on the changing trends of the user's emotional state; By combining the preliminary interaction scheme, the core part of the interaction content, the strength level of the interaction tone, the speed level of the interaction speed, and the guidance method of the interaction content, a collaborative control strategy containing execution instructions is generated.

6. The user-state-aware interactive control method as described in claim 1, characterized in that, During the execution of the collaborative control strategy, real-time multimodal user information is continuously acquired and compared with the initial interaction scheme. When adjustment conditions are triggered, a new interaction scheme is re-planned and generated, and the collaborative control strategy is updated, including: During the execution of the collaborative control strategy, real-time speech information, real-time semantic information, and real-time emotion information are collected in parallel by a multimodal perception unit. The real-time emotion information is compared with the user's emotion change trend to form an emotion vector; the real-time semantic information is compared with the user's intention needs to form a semantic distance; and the real-time voice information is compared with the acoustic elements corresponding to the strength level of the interactive tone and the speed level of the interactive speech. When any comparison result exceeds the corresponding threshold range or multiple comparison results show coordinated deviation, the multidimensional adjustment condition is determined to be triggered. Based on the specific adjustment dimension triggered, select the corresponding adaptive planning module to replan the interaction path; Based on the replanned interaction path and the real-time emotion information, a new interaction scheme is generated, which includes the strength level of the updated interaction tone and the content sequence. The new interaction scheme is mapped to execution instructions, and the speech synthesis parameters, interaction logic sequences, and feedback triggering conditions in the collaborative control strategy are updated based on the execution instructions.

7. The user-state-aware interactive control method as described in claim 1, characterized in that, After the interaction task is completed, the interaction performance is analyzed based on the interaction process data. When the analysis results do not meet the preset standards, the process of predicting user emotion change trends, generating collaborative control strategies, and generating interaction solutions are optimized, including: After the interactive task is completed, start the effect analysis module; The effect analysis module collects interaction duration data and user emotion change data during the interaction process. Based on the interaction duration data, an interaction duration indicator is determined; based on the user emotion change data, an emotion stability indicator, an interaction goal completion indicator, and a user cooperation indicator are determined. The interaction duration index, the emotional stability index, the completion index, and the cooperation index are compared with preset analysis standards. When the comparison results show that there are indicators that have not met the standards, the adaptive optimization module is activated; The adaptive optimization module adjusts the model parameters of the time-series prediction model, optimizes the strategy parameters for tone and speed control during the collaborative control strategy generation process, and updates the path planning parameters for generating the interaction scheme.

8. An interactive control device based on user state awareness, characterized in that, The user-state-aware interactive control device includes: The multimodal perception and preprocessing module is used to collect and preprocess the multimodal raw data generated during user interaction to generate standardized multimodal user information. The emotion prediction and solution generation module is used to predict the user's emotion change trend based on the standardized multimodal user information, and generate a preliminary interaction solution by combining the user's emotion change trend, the user's intention and needs and the current interaction state. The collaborative control strategy generation module is used to generate a collaborative control strategy to guide the execution of the interaction based on the preliminary interaction scheme and the standardized multimodal user information. An adaptive interaction execution module is used to continuously acquire real-time multimodal user information during the execution of the collaborative control strategy, compare the real-time multimodal user information with the preliminary interaction scheme, and when the adjustment condition is triggered, re-plan and generate a new interaction scheme and update the collaborative control strategy. The interaction analysis and self-optimization module is used to analyze the interaction performance based on the interaction process data after the interaction task is completed. When the analysis results do not meet the preset standards, it optimizes the prediction process of user emotion change trends, the generation process of collaborative control strategies, and the generation process of interaction schemes.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and a user-state-aware interactive control program stored in the memory and executable on the processor, wherein the user-state-aware interactive control program, when executed by the processor, implements the steps of the user-state-aware interactive control method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a user-state-aware interactive control program, which, when executed by a processor, implements the steps of the user-state-aware interactive control method as described in any one of claims 1-7.

Citation Information

Cited By

  • Interactive guiding method and system of intelligent question-answering system and storage medium

    CN122064795A

  • An interactive guiding method and system of an intelligent question-answering system and a storage medium

    CN122064795B