Emotion adaptive multi-modal dialogue generation method and system based on biological signals

By simultaneously processing biosignals, video and audio, and text-based multimodal interaction information, extracting features, and performing cross-modal fusion, the problem of insufficient emotion perception and dynamic adaptability in existing technologies is solved, thereby improving the naturalness and emotional fit of dialogue interaction.

CN121303369BActive Publication Date: 2026-04-21XIAMEN UNIV OF TECH +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIAMEN UNIV OF TECH
Filing Date
2025-12-15
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing multimodal dialogue generation technologies have limitations in terms of emotion perception and dynamic adaptability. In particular, when emotion expression is not significant or cross-modal emotions are inconsistent, it is difficult to perceive changes in user emotions in real time and adaptively adjust dialogue strategies, resulting in insufficient naturalness of dialogue interaction.

Method used

By simultaneously processing biosignals, video and audio, and text-based multimodal interaction information, extracting features and performing cross-modal fusion, and combining emotional consistency assessment, the system monitors user emotional changes in real time, generates dialogue content that matches the user's current emotional state, and achieves adaptive emotional response.

Benefits of technology

It improves the naturalness of dialogue interaction, ensures that the emotional state recognition matches the user's true psychological state, and enhances the semantic coherence and emotional fit of dialogue responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121303369B_ABST
    Figure CN121303369B_ABST
Patent Text Reader

Abstract

The application provides an emotion adaptive multi-modal dialogue generation method and system based on biological signals, and relates to the technical field of data processing.The method comprises the following steps: step 1, processing the collected biological signals and multi-modal interaction information, extracting video space-time features, audio time sequence features and text semantic features; step 2, cross-modal fusion of the video space-time features, the audio time sequence features and the text semantic features, obtaining deep semantic connections between multi-modal data through fusion semantic analysis, and establishing a cross-modal semantic correspondence relationship; using a sentiment consistency evaluation link to analyze the consistency degree of different modalities in emotional expression, and obtaining a comprehensive emotional state representation.The application can realize real-time perception of user emotional changes and adaptive adjustment of dialogue strategies, and improve the naturalness of dialogue interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method and system for generating emotion-adaptive multimodal dialogues based on biosignals. Background Technology

[0002] Dialogue generation technologies based on multimodal information can combine multiple input modalities such as text, speech, and vision to generate responses. Most of them utilize natural language processing models to analyze the semantic content of user input and refer to the dialogue context to maintain topic coherence. However, existing methods still have some limitations in multimodal emotion perception and dynamic adaptability. For example, in dialogue scenarios involving emotional support, existing systems may rely heavily on explicit emotional cues in text and speech, while making relatively limited use of biological signals that reflect the user's internal emotional state. This may lead to an incomplete perception of changes in the user's emotions, especially when emotional expression is not significant or there is cross-modal emotional inconsistency. In addition, although existing natural language generation models have a certain ability to adapt to context, they mostly lack dynamic tracking and adjustment mechanisms for continuous fluctuations in the user's emotional state during long-term interactions. Summary of the Invention

[0003] The technical problem to be solved by this invention is to provide a method and system for generating emotion-adaptive multimodal dialogue based on biosignals, which can perceive changes in user emotions in real time and adaptively adjust dialogue strategies to improve the naturalness of dialogue interaction.

[0004] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0005] Firstly, a method for generating emotion-adaptive multimodal dialogue based on biological signals, the method comprising:

[0006] Step 1: Process the collected biological signals and multimodal interaction information to extract video spatiotemporal features, audio temporal features, and text semantic features;

[0007] Step 2 involves cross-modal fusion of video spatiotemporal features, audio temporal features, and text semantic features. Through fusion semantic analysis, deep semantic connections between multimodal data are obtained, and cross-modal semantic correspondences are established. The emotional consistency assessment step is used to analyze the degree of consistency in emotional expression across different modalities, resulting in a comprehensive emotional state representation.

[0008] Step 3: Based on the comprehensive emotional state representation and combined with the contextual information of the multi-turn dialogue, use a natural language generation model to obtain dialogue text content that matches the user's current emotional state, and convert it into speech output through a speech conversion process.

[0009] Step 4: During the voice output process, monitor the dynamic changes of the user's biosignals and voice signals in real time, organize the newly collected multimodal data into a time-series data stream, and construct an emotional state analysis benchmark; by detecting key turning points and extreme points in the time-series data stream, extract key event points representing the boundaries of emotional states, and identify two emotional change patterns.

[0010] Step 5: Analyze the two emotion change patterns, determine the fluctuation range of the emotion state by combining the results of the fusion semantic analysis, and construct the emotion state change trajectory by selecting reference instances of the emotion state; generate dynamic adjustment parameters of the emotion state based on the emotion state change trajectory, and adjust the dialogue strategy and generated content in real time to achieve adaptive emotion response.

[0011] Secondly, an emotion-adaptive multimodal dialogue generation system based on biological signals includes:

[0012] The extraction module is used to process the collected biological signals and multimodal interaction information to extract video spatiotemporal features, audio temporal features, and text semantic features.

[0013] The fusion module is used to perform cross-modal fusion of video spatiotemporal features, audio temporal features, and text semantic features. Through fusion semantic analysis, it obtains deep semantic connections between multimodal data and establishes cross-modal semantic correspondences. The sentiment consistency assessment step is used to analyze the degree of consistency in sentiment expression across different modalities to obtain a comprehensive emotional state representation.

[0014] The generation module is used to obtain dialogue text content that matches the user's current emotional state based on the comprehensive emotional state representation and the contextual information of the multi-turn dialogue, and then convert it into speech output through a speech conversion process.

[0015] The recognition module is used to monitor the dynamic changes of the user's biosignals and speech signals in real time during the speech output process, organize the newly acquired multimodal data into a time-series data stream, and construct a benchmark for emotion state analysis; by detecting key turning points and extreme points in the time-series data stream, it extracts key event points that represent the boundaries of the emotion state and identifies two emotion change patterns.

[0016] The adjustment module analyzes two emotion change patterns, determines the fluctuation range of the emotion state by combining the results of fusion semantic analysis, and constructs the emotion state change trajectory by selecting reference instances of the emotion state. Based on the emotion state change trajectory, it generates dynamic adjustment parameters for the emotion state, and adjusts the dialogue strategy and generated content in real time to achieve adaptive emotion response.

[0017] The above-described solution of the present invention has at least the following beneficial effects:

[0018] By simultaneously processing biosignals and multimodal interaction information from video, audio, and text, spatiotemporal features of video, temporal features of audio, and semantic features of text are extracted. Combined with the physiological objectivity of biosignals, this reduces emotional misjudgments caused by relying solely on language or behavioral features, making emotional state recognition more closely aligned with the user's true psychological state. Through cross-modal fusion and semantic analysis, a deep semantic correspondence between multimodal data is established. At the same time, an emotional consistency assessment step is added to quantify the consistency of different modalities in emotional expression, ensuring the reliability of the comprehensive emotional state representation. Based on the comprehensive emotional state representation and multi-turn dialogue context information, a natural language generation model is used to generate dialogue text that fits the user's current emotion, and the text is output through a speech conversion process, so that the dialogue response not only meets semantic coherence but also matches the user's emotional needs. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating the emotion-adaptive multimodal dialogue generation method based on biosignals provided in an embodiment of the present invention.

[0020] Figure 2 This is a schematic diagram of an emotion-adaptive multimodal dialogue generation system based on biosignals provided in an embodiment of the present invention. Detailed Implementation

[0021] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0022] like Figure 1 As shown, embodiments of the present invention propose an emotion-adaptive multimodal dialogue generation method based on biosignals, the method comprising the following steps:

[0023] Step 1: Process the collected biological signals and multimodal interaction information to extract video spatiotemporal features, audio temporal features, and text semantic features;

[0024] Step 2 involves cross-modal fusion of video spatiotemporal features, audio temporal features, and text semantic features. Through fusion semantic analysis, deep semantic connections between multimodal data are obtained, and cross-modal semantic correspondences are established. The emotional consistency assessment step is used to analyze the degree of consistency in emotional expression across different modalities, resulting in a comprehensive emotional state representation.

[0025] Step 3: Based on the comprehensive emotional state representation and combined with the contextual information of the multi-turn dialogue, use a natural language generation model to obtain dialogue text content that matches the user's current emotional state, and convert it into speech output through a speech conversion process.

[0026] Step 4: During the voice output process, monitor the dynamic changes of the user's biosignals and voice signals in real time, organize the newly collected multimodal data into a time-series data stream, and construct an emotional state analysis benchmark; by detecting key turning points and extreme points in the time-series data stream, extract key event points representing the boundaries of emotional states, and identify two emotional change patterns.

[0027] Step 5: Analyze the two emotion change patterns, determine the fluctuation range of the emotion state by combining the results of the fusion semantic analysis, and construct the emotion state change trajectory by selecting reference instances of the emotion state; generate dynamic adjustment parameters of the emotion state based on the emotion state change trajectory, and adjust the dialogue strategy and generated content in real time to achieve adaptive emotion response.

[0028] In this embodiment of the invention, by simultaneously processing biosignals and multimodal interaction information of video, audio, and text, spatiotemporal features of video, temporal features of audio, and semantic features of text are extracted. Combined with the physiological objectivity of biosignals, the misjudgment of emotions caused by relying solely on language or behavioral features is reduced, making the emotional state recognition more consistent with the user's true psychological state. A deep semantic correspondence relationship of multimodal data is established through cross-modal fusion and semantic analysis. At the same time, an emotional consistency assessment step is added to quantify the consistency of different modalities in emotional expression, ensuring the reliability of the comprehensive emotional state representation. Based on the comprehensive emotional state representation and multi-turn dialogue context information, a natural language generation model is used to generate dialogue text that fits the user's current emotion, and the text is output through a speech conversion process, so that the dialogue response not only meets semantic coherence but also matches the user's emotional needs.

[0029] In a preferred embodiment of the present invention, step 1 above, which processes the collected biosignals and multimodal interaction information to extract video spatiotemporal features, audio temporal features, and text semantic features, may include:

[0030] In this embodiment of the invention, a non-invasive sensing device is used to simultaneously acquire the user's EEG, ECG, and EEG response signals. Timestamps are recorded to ensure time alignment with multimodal information. Signal processing and feature extraction are as follows: The EEG signal is bandpass filtered from 5Hz to 30Hz and denoised using a 50-sampling-point moving average. It is then segmented into 1-second segments, and the mean, variance, and energy percentages of alpha waves (8Hz-13Hz) and beta waves (14Hz-30Hz) are extracted from each segment. The ECG signal is thresholded to identify R-wave peak values ​​and remove interfering waveforms. After normalization to the 0-1 range, the heart rate, the number of R-waves per minute, and the mean and standard deviation of the R-wave interval are calculated. The EEG response signal is denoised using a median filter of 20 consecutive sampling points and then segmented into 1-second segments. The process involves extracting peak values, rise rates, and sums of values ​​from video segments; extracting spatiotemporal features from the video, standardizing the video to 1920×1080 resolution and 30fps; processing faces by detecting and cropping 224×224 pixel face regions and enhancing contrast through histogram equalization; removing blurry frames (clarity < 60% of the mean) and motionless frames (sum of feature point displacements < threshold), retaining valid frames; for static features, locating 68 key facial feature points per frame, calculating 30 key distance parameters, 15 feature vector angle parameters, and 100 10×10 pixel block texture parameters (mean + variance), integrating them into a 145-dimensional single-frame spatial feature vector; and for dynamic features, based on the spatial features of 10 consecutive frames, calculating 9 inter-frame difference vectors and extracting the parameters. Average rate of change, amplitude of change, and cumulative change; spatiotemporal fusion, splicing single-frame spatial and dynamic features to form a 580-dimensional video spatiotemporal feature vector, containing static facial features and dynamic change information; audio temporal feature extraction, the speech signal is standardized to a 16kHz sampling rate and 16-bit deep mono, the high-frequency signal is enhanced by pre-emphasis, and the frame is divided into 20ms frames and 10ms frames, with 320 sampling points per frame, a Hanning window is applied to reduce spectral leakage, and silent frames with short-time energy <10% of the mean are removed; feature extraction, extracting temporal features, short-time energy, zero-crossing rate, fundamental frequency and frequency domain features, spectral amplitude and spectral centroid of 256 frequency points; temporal integration, arranging the effective frame features in chronological order to form a 260-dimensional / frame audio temporal feature vector. The text sequence is analyzed; semantic features are extracted by encoding the text uniformly in UTF-8, removing special symbols, numbers, and punctuation, performing part-of-speech tagging after word segmentation, removing stop words and correcting typos, and retaining core semantic words; sentiment and semantic features are analyzed based on a Chinese sentiment dictionary, counting the number of positive / negative sentiment words, and calculating the total sentiment intensity and average sentiment intensity; a context window is constructed, and the overlap rate of core words is calculated; the above indicators are integrated with the semantic encoding of core words to form a feature vector; biosignal features, video spatiotemporal features, audio temporal features, and text semantic features are aligned by timestamp to form a unified multimodal feature set, ensuring that the features of each modality are synchronized in time, and comprehensively covering the user's internal physiological state, external visual expression, acoustic emotional cues, and semantic sentiment tendency.

[0031] In a preferred embodiment of the present invention, step 2 above involves cross-modal fusion of video spatiotemporal features, audio temporal features, and text semantic features. Through fused semantic analysis, deep semantic connections between multimodal data are obtained, and cross-modal semantic correspondences are established. The emotional consistency assessment step analyzes the degree of consistency in emotional expression across different modalities to obtain a comprehensive emotional state representation, which may include:

[0032] In this embodiment of the invention, step 220 involves cross-modal semantic alignment of video spatiotemporal features, audio temporal features, and text semantic features, and analyzing the interaction strength between different modal features through a multimodal attention process. Specifically, this includes: firstly, basic semantic labeling is performed on video spatiotemporal features, audio temporal features, and text semantic features. The semantic labels for each modality are divided into 5 core semantic dimensions. The semantic label dimensions for video spatiotemporal features include facial expressions, head posture, body movements, eye state, and facial muscle movements. Each dimension is further subdivided into specific labels, such as facial expressions including smiling, frowning, calm, pouting, and surprise. The semantic label dimensions for audio temporal features include pitch changes, speech rate rhythm, volume strength, intonation, and audio prosody. Subdivided labels, such as pitch changes, include rising, falling, steady, and fluctuating. The semantic label dimensions for text semantic features include lexical sentiment, sentence structure and tone, semantic tendency, expressive intent, and emotional intensity. Subdivided labels, such as lexical sentiment, include positive, neutral, negative, commendatory, and derogatory.

[0033] A cross-modal semantic alignment mapping table is established, using emotional expression as the core anchor point. Labels with the same or highly similar semantic meanings across the three modalities are paired and grouped to form semantic groups. For example, the first semantic group includes video facial expressions and smiling, head posture and nodding, audio pitch changes and rising, speech rate and rhythm, text vocabulary emotion and positive, and semantic tendency and positive; the second semantic group includes video facial expressions and frowning, eye state and tension, audio pitch changes and lowing, intonation and harshness, text vocabulary emotion and negative, and emotional intensity and strength, and so on. This ensures that each semantic group is built around the same core emotion or semantic meaning, achieving preliminary semantic alignment. For each semantic group, feature fragments corresponding to the three modalities within that group are extracted, and the semantic details contained in each feature fragment are analyzed one by one, statistically analyzing the features of each modality. The number of matching details between the fragment and the core meaning of the semantic group is calculated. For example, the core meaning of the first semantic group is positive emotional expression. The video feature fragment contains 2 matching details of smiling and nodding, the audio feature fragment contains 2 matching details of rising pitch and brisk speech, and the text feature fragment contains 2 matching details of positive words and positive semantics. The total number of details in this semantic group is 6, which is the sum of the number of matching details of the three modalities. Then, the local correlation coefficient between each modality within a single semantic group is calculated. The local correlation coefficient between the video modality and the audio modality is (number of matching details in the video + number of matching details in the audio) ÷ total number of details in the semantic group, i.e., (2+2) ÷ 6 ≈ 0.67; the local correlation coefficient between the video modality and the text modality is (2+2) ÷ 6 ≈ 0.67; the local correlation coefficient between the audio modality and the text modality is (2+2) ÷ 6 ≈ 0.67.

[0034] The local association coefficients of all semantic groups are summarized, and the average of the local association coefficients between each modality and the other two modalities is calculated. This average represents the interaction association strength between features of different modalities. Assuming a total of 10 semantic groups are constructed, the local association coefficients of video and audio in the 10 semantic groups are 0.67, 0.7, 0.65, 0.72, 0.68, 0.71, 0.66, 0.69, 0.73, and 0.64, respectively. These 10 values ​​are added together to obtain a total of 6.85. Dividing this by the total number of semantic groups (10), we obtain the interaction association strength between video and audio as 0.685. Similarly, the sum of the 10 local association coefficients between video and text is 6.78, with an interaction association strength of 0.678; the sum of the 10 local association coefficients between audio and text is 6.82, with an interaction association strength of 0.682.

[0035] Step 221: Based on the interaction association strength, fuse the features of each modality to generate a semantically aligned cross-modal joint representation. Specifically, this includes: first, determining the weight coefficient of each modality feature. The weight coefficient is directly taken as the average value of the interaction association strength obtained in step 220, that is, the weight coefficient of the video modality = (interaction association strength between video and audio + interaction association strength between video and text) ÷ 2, which is calculated as (0.685 + 0.678) ÷ 2 ≈ 0.6815; the weight coefficient of the audio modality = (interaction association strength between video and audio + interaction association strength between audio and text) ÷ 2, that is, (0.685 + 0.682) ÷ 2 ≈ 0.6835; the weight coefficient of the text modality = (interaction association strength between video and text + interaction association strength between audio and text) ÷ 2, that is, (0.678 + 0.682) ÷ 2 = 0.68.

[0036] The video spatiotemporal features, audio temporal features, and text semantic features are processed to unify their dimensions. Assuming the original dimensions of the video spatiotemporal features are 128, the audio temporal features are 64, and the text semantic features are 256, the largest dimension of 256 is selected as the unified dimension. Dimensional completion is performed on the video spatiotemporal features and audio temporal features. The video spatiotemporal features are expanded from 128 to 256 dimensions by repeatedly filling in feature dimensions with high semantic contribution. The audio temporal features are expanded from 64 to 256 dimensions by interpolating to fill in missing dimension features, ensuring that all three modalities have 256 dimensions, meeting the dimensionality requirements for fusion computation. For each feature dimension, feature values ​​for each of the three modalities are extracted. For example, in dimension 1, the feature value for the video modality is 0.35, the feature value for the audio modality is 0.42, and the feature value for the text modality is 0.38; in dimension 2, the feature value for the video modality is 0.29, the feature value for the audio modality is 0.36, and the feature value for the text modality is 0.36. The eigenvalue of the modality is 0.41, and so on. Then, the eigenvalues ​​of each dimension are weighted: the video modality eigenvalue is multiplied by its weight coefficient of 0.6815, the audio modality eigenvalue by 0.6835, and the text modality eigenvalue by 0.68, resulting in three weighted eigenvalues. Taking the first dimension as an example, the weighted eigenvalue for video is 0.35 × 0.6815 ≈ 0.2385, and the weighted eigenvalue for audio is 0.42 × 0.68. 35≈0.2871, the weighted feature value of the text is 0.38×0.68≈0.2584; then add these three weighted feature values ​​together to get the joint feature value of this dimension = 0.2385+0.2871+0.2584≈0.784; calculate the joint feature values ​​of 256 feature dimensions in the above manner, and arrange all the joint feature values ​​in order from the 1st dimension to the 256th dimension to form a 256-dimensional semantically aligned cross-modal joint representation.

[0037] Step 222, based on semantic alignment, cross-modal joint representation, constructs a semantic association matrix between modalities by calculating the semantic similarity between feature segments of different modalities; specifically, it includes: firstly, dividing video spatiotemporal features, audio temporal features, and text semantic features into segments of fixed time length, setting the time length of each segment to 2 seconds. Assuming the total duration of the original multimodal data is 20 seconds, the features of each modality are evenly divided into 10 feature segments, each segment corresponding to a 2-second time interval, and the time intervals of the segments of the three modalities are completely consistent, such as segment 1 corresponding to 0-2 seconds, segment 2 corresponding to 2-4 seconds, ..., segment 10 corresponding to 18-20 seconds, achieving segment-level time alignment; in the cross-modal joint representation... Using semantic groups as a reference, core semantic keywords are determined for each feature segment. Specifically, the number of labels from each semantic group contained in each feature segment is counted, and the semantic group with the most labels is selected as the core semantic group of that segment. Then, 2-3 key labels are extracted from the core semantic group as core semantic keywords. For example, in video segment 3, the positive emotion expression semantic group has the most labels, so its core semantic keywords are smiling, nodding, and positive; in the corresponding audio segment 3, the positive emotion expression semantic group has the most labels, and its core semantic keywords are rising pitch and brisk speech; in the corresponding text segment 3, the positive emotion expression semantic group has the most labels, and its core semantic keywords are positive vocabulary and positive semantics.

[0038] To calculate the semantic similarity between any two feature segments of different modalities, the calculation rule is as follows: First, count the number of identical core semantic keywords in the two feature segments. Then, count the number of semantically similar keywords. Add the number of identical keywords to the number of semantically similar keywords multiplied by 0.5, and then divide by the sum of the total number of keywords in the two feature segments to obtain the semantic similarity of a single segment pair. Taking video segment 3 and audio segment 3 as an example, the number of identical keywords is 0; the number of semantically similar keywords is 3; the total number of keywords in video segment 3 is 3, and the total number of keywords in audio segment 3 is 2, with a total of 5 keywords. Therefore, the semantic similarity between the two is (0 + 3 × 0.5) ÷ 5 = 1.5 ÷ 5 = 0.3. Taking video segment 3 and text segment 3 as another example: the number of identical keywords is 1; the number of semantically similar keywords is 2; the total number of keywords in video segment 3 is 3, and the total number of keywords in text segment 3 is 2, with a total of 5 keywords. 5; Semantic similarity = (1 + 2 × 0.5) ÷ 5 = (1 + 1) ÷ 5 = 0.4; Construct the intermodal semantic association matrix. The matrix is ​​divided into three groups: video and audio semantic association matrix, audio and text semantic association matrix, and video and text semantic association matrix. The number of rows and columns in each group of matrices is equal to the number of feature segments (10 rows and 10 columns). The row index of the matrix represents the feature segment number of the first modality (1-10), and the column index represents the feature segment number of the second modality (1-10). Each element in the matrix is ​​the semantic similarity between the two corresponding segments. For example, in the video and audio semantic association matrix, the element in the 3rd row and 3rd column is the semantic similarity of video segment 3 and audio segment 3 (0.3), and the element in the 5th row and 5th column is the semantic similarity of video segment 5 and audio segment 5. Fill all matrix elements in sequence to complete the construction of the three groups of semantic association matrices.

[0039] Step 223 involves analyzing the semantic association matrix to identify cross-modal semantic association patterns and establish structured cross-modal semantic correspondences. Specifically, this includes: first, setting a semantic similarity threshold; calculating the average value of all elements in all semantic association matrices; using this average value as the threshold; assuming three sets of matrices, each with 10 × 10 = 100 elements, for a total of 300 elements, the sum of their semantic similarity values ​​is 120, then the average value = 120 ÷ 300 = 0.4, i.e., setting the semantic similarity threshold to 0.4; traversing the semantic association matrix between each modality, filtering out elements with values ​​greater than the threshold of 0.4; the feature segment pairs corresponding to these elements are semantically highly related segment pairs. For example, in the video and text semantic association matrix, the element in the 2nd row and 2nd column is 0.5, greater than 0.4, so video segment 2 and text segment 2 are highly related segment pairs; the element in the 7th row and 7th column is 0.6, greater than 0.4, so video segment 7 and text segment 7 are highly related segment pairs, and so on, summarizing all highly related segments in the three sets of matrices. Yes; we analyze these highly related segment pairs to identify cross-modal semantic association patterns. Two core association patterns are defined: First, a comprehensive matching association pattern, judged by the fact that the semantic similarity of a segment pair across all semantic groups is greater than a threshold of 0.4, and the number of semantic groups involved accounts for no less than 80% of the total number of semantic groups, meaning the segment pair is highly matched across multiple semantic dimensions. Second, a partial matching association pattern, judged by the fact that the semantic similarity of a segment pair is greater than a threshold of 0.4 only in some semantic groups, and the number of semantic groups involved accounts for between 30% and 80% of the total number of semantic groups, meaning the segment pair is highly matched only in a specific semantic dimension. For example, a segment pair consisting of video segment 5, audio segment 5, and text segment 5 has a similarity greater than 0.4 in 9 out of 10 semantic groups, accounting for 90%, thus it is judged as a comprehensive matching pattern; a segment pair consisting of video segment 8 and audio segment 8 has a similarity greater than 0.4 in 4 out of 10 semantic groups, accounting for 40%, thus it is judged as a partial matching pattern.

[0040] For each association pattern, all highly related fragment pairs under that pattern are summarized, and detailed information for each fragment pair is recorded, including the type of the first modality, fragment number, core semantic keywords, association pattern type, the type of the corresponding associated modality, and fragment number. For example, a record under the full matching pattern is: Modality type - video, fragment number - 5, core semantic keywords - calm, relaxed, neutral, association pattern - full matching, corresponding associated modality - audio fragment 5, text fragment 5; a record under the partial matching pattern is: Modality type - audio, fragment number - 8, core semantic keywords - slow speech rate, even tone, association pattern - partial matching, corresponding associated modality - video fragment 8. All records are organized into a structured cross-modal semantic correspondence table, which contains fields such as modality type, fragment number, core semantic keywords, association pattern, corresponding modality 1 type, corresponding modality 1 fragment number, corresponding modality 2 type, and corresponding modality 2 fragment number. This table clearly presents the semantic correspondence between all feature fragments of the three modalities.

[0041] Step 224: Based on cross-modal semantic correspondence, extract corresponding visual sentiment features, acoustic sentiment features, and semantic sentiment features from video modality, audio modality, and text modality. Specifically, this includes: firstly, selecting sentiment-related semantic groups from the cross-modal semantic correspondence table, dividing them into three categories according to sentiment type: positive sentiment group, containing core semantic keywords such as pleasure, excitement, satisfaction, and relief; neutral sentiment group, containing core semantic keywords such as calm, indifferent, and no obvious emotion; and negative sentiment group, containing core semantic keywords such as sadness, anger, anxiety, and dissatisfaction. Then, determine the feature fragments of the three modalities corresponding to each sentiment semantic group, i.e., the association patterns in the correspondence table. For segment pairs that are either fully or partially matched, visual emotional features are extracted. For corresponding feature segments of the video modality, specific features are extracted from four dimensions: facial expression, head posture, eye state, and facial muscle movement. The facial expression dimension extracts the degree of upward turn of the corners of the mouth, divided into 0-5 levels (0 for no upward turn, 5 for maximum upward turn), and eyebrow state, categorized as frowning, relaxed, and tense. The head posture dimension extracts the head tilt angle, categorized as horizontal, tilting left, tilting right, looking down, and looking up, and the nodding frequency, calculated per second. The eye state dimension extracts eye brightness, divided into 0-5 levels, and blinking frequency, calculated per second. The facial muscle movement dimension extracts cheek... Muscle tension is graded from 0 to 3, and the number of forehead wrinkles is categorized as none, few, and many. For example, in a video clip from a positive emotion group, the extracted visual emotion features are: mouth corner upturn level 4, eyebrows relaxed, head level, nodding frequency 1.5 times / second, eye brightness level 4, blinking frequency 0.8 times / second, cheek muscle tension level 1, and no forehead wrinkles. Acoustic emotion features are extracted for corresponding feature segments of the audio modality, from four dimensions: pitch variation, speech rate and rhythm, volume, and intonation. The pitch variation dimension extracts the average pitch, calculated in Hertz, ranging from 100-500 Hertz; the pitch fluctuation range is categorized as maximum pitch to minimum pitch. The pitch is calculated; speech rate and pause time are extracted in seconds; the volume is extracted in decibels, ranging from 30 to 80 decibels, and the volume fluctuation is calculated as the maximum volume minus the minimum volume; the intonation is extracted in frequency of pauses, calculated in seconds, and the intonation intensity is divided into levels 0 to 3. For example, in the audio clip of the positive emotion group mentioned above, the extracted acoustic emotion features are: average pitch 250 Hz, pitch fluctuation range 50 Hz, speech rate 3.2 words / second, pause time 0.2 seconds, average volume 60 decibels, volume fluctuation range 10 decibels, intonation frequency 0.5 times / second, and intonation intensity level 1.

[0042] Semantic sentiment features are extracted. For corresponding feature segments of the text modality, specific information is extracted from four dimensions: lexical sentiment, sentence structure and tone, sentiment intensity, and semantic tendency. The lexical sentiment dimension counts the number of positive, neutral, and negative words, and calculates the percentage of sentimental words (number of sentimental words / total number of words). The sentence structure and tone dimension determines the sentence type and the number of modal particles. The sentiment intensity dimension is divided into levels 0-1, with level 0 representing no obvious sentiment and level 1 representing strong sentiment, adjusted based on the intensity of the sentiment words. The semantic tendency dimension determines positive, neutral, and negative sentiment; for example, the positive sentiment group mentioned above... The extracted semantic sentiment features from the text fragments are as follows: 2 positive words, 3 neutral words, and 0 negative words, with sentiment words accounting for 40%; the sentence structure is exclamatory, with 1 modal particle; the sentiment intensity is 0.9; and the semantic tendency is positive. The extracted visual sentiment features, acoustic sentiment features, and semantic sentiment features are classified and stored according to sentiment semantic groups. Three modal sentiment feature subfolders are created under each sentiment group. The files in the folders are named in the format of fragment number-feature type-feature details to ensure that the three modal sentiment features under each sentiment category can be quickly linked and queried.

[0043] Step 225 involves judging visual, acoustic, and semantic emotional features to form emotional tendency descriptions for video, audio, and text modalities. Specifically, this includes: First, establishing a detailed emotional judgment standard library, classifying emotions into three categories: positive, neutral, and negative. For each category, clear multi-dimensional judgment criteria are set, with each criterion containing a specific numerical range or classification standard. For positive emotions, the judgment criteria are: visual criteria, requiring at least 3 items to be met: mouth corners raised ≥3 levels, eyebrows relaxed, head nodding frequency ≥1. Speech rate ≥ 2.5 words / second, eye brightness ≥ 4, cheek muscle tension ≤ 1; Acoustic criteria, requiring ≥ 3: average pitch 200-300 Hz, speech rate ≥ 2.5 words / second, average volume 50-70 dB, pause frequency ≤ 0.8 times / second, pause intensity ≤ 1; Textual criteria, requiring ≥ 2: positive vocabulary ≥ 30%, sentence structure exclamatory / declarative, emotional intensity ≥ 0.7, semantic tendency positive; Neutral emotion judgment criteria: Visual criteria, requiring ≥ 3: mouth corner upturn 2-3 levels, eyebrows flat, nodding frequency 0.5-1 times / second, Eye brightness level 2-3, facial muscle tension level 1-2; Acoustic criteria, meeting ≥3 criteria: average pitch 150-200 Hz or 300-350 Hz, speech rate 1.5-2.5 words / second, average volume 40-50 dB or 70-75 dB, pause frequency 0.8-1.2 times / second, pause intensity level 1-2; Textual criteria, meeting ≥2 criteria: positive vocabulary ratio 10%-30% and negative vocabulary ratio 10%-30%, sentence structure declarative, emotional intensity level 0.3-0.7, semantic tendency neutral; Criteria for negative emotion judgment: visual Based on the following criteria, at least three items must be met: the degree of upward turn of the mouth is ≤2, the eyebrows are furrowed / tight, the nodding frequency is ≤0.5 times / second, the eye brightness is ≤2, and the degree of facial muscle tension is ≥2; based on the following criteria, at least three items must be met: average pitch is ≤150 Hz or ≥350 Hz, speech rate is ≤1.5 words / second, average volume is ≤40 dB or ≥75 dB, the frequency of pauses is ≥1.2 times / second, and the intensity of pauses is ≥2; based on the following criteria, at least two items must be met: the proportion of negative words is ≥30%, the sentence structure is declarative / interrogative, the emotional intensity is ≥0.7 (negative direction), and the semantic tendency is negative.

[0044] To determine the emotional tendency of a video modality, the emotional features of the video modality under a specific emotional group are extracted. Each feature is then compared against the visual criteria for positive, neutral, and negative emotions. The number of criteria matching each emotional category is counted. For example, for a certain video feature, the following criteria match: 4 levels of upward curve of the mouth, relaxed eyebrows, horizontal head position, 4 levels of eye brightness, and 1 level of facial muscle tension. These match 4 criteria for positive emotions, 1 for neutral emotions, and 0 for negative emotions. Since the number of criteria matching positive emotions is the highest and meets the requirement of ≥3, the emotional tendency of the video modality is determined. To assess the positive sentiment, a description of the emotional tendency in the video modality is formed, leaning towards positive emotions. This is based on visual features such as a level 4 increase in the corners of the mouth, relaxed eyebrows, level 4 brightness in the eyes, and level 1 tension in the cheek muscles. The same method is used to assess the emotional tendency in the audio modality. The corresponding emotional features of the audio modality are extracted and compared against the acoustic criteria for the three emotion categories. The number of matches is counted. For example, the audio features mentioned above, such as an average pitch of 250 Hz, a speech rate of 3.2 words / second, an average volume of 60 dB, a pause frequency of 0.5 times / second, and a pause intensity of level 1, have a total of [number missing] matches. Five criteria were used to determine the emotional tendency of the audio modality: positive (5 criteria), neutral (1 criterion), and negative (0 criteria). A description of the audio modality's emotional tendency was formed, indicating a positive emotional tendency. This was based on acoustic features such as an average pitch of 250 Hz, a speech rate of 3.2 words / second, an average volume of 60 dB, a pause frequency of 0.5 times / second, and a pause intensity of level 1. To determine the emotional tendency of the text modality, the corresponding emotional features were extracted and compared against the three emotional criteria, and the number of matches was counted. For example, the above text features included positive vocabulary. The text modality was determined to be positive in sentiment based on the following characteristics: 40% positive vocabulary, exclamatory sentence structure, emotional intensity level of 0.9, and positive semantic tendency. It met four criteria for positive sentiment, one for neutral sentiment, and zero for negative sentiment. A sentiment tendency description for the text modality was then generated, indicating a positive emotional tendency. Based on these text features, the sentiment tendency description for each modality was compiled to ensure that the description clearly included the sentiment tendency category and core criteria for judgment.

[0045] Step 226 involves cross-comparing the sentiment tendency descriptions of the video, audio, and text modalities to assess the inherent consistency between the sentiment tendency descriptions of different modalities. Based on this inherent consistency, a corresponding fusion contribution is determined for the sentiment tendency description of each modality. Specifically, this includes: first, determining the three sets of objects for cross-comparison: video modality and audio modality, audio modality and text modality, and text modality and video modality. Each comparison focuses on the degree of fit of the sentiment tendency descriptions. Referring to the calculation logic of the sum of interior angles of a polygon, the three modalities are regarded as the three vertices of a triangle, and the three comparison results correspond to the three interior angles of the triangle. The sum of the interior angles of the triangle is fixed at 180°. Therefore, the total consistency score of the three comparisons is fixed at 3 points, with each comparison receiving a maximum score of 1 point. The sum of the scores of the three interior angles is 3 points, ensuring the rigor and fairness of the evaluation logic. At the same time, consistency scoring rules are established: if the sentiment tendency descriptions of two modalities are completely identical (i.e., the sentiment categories are consistent and the intensity deviation is ≤0.1), and if both are sentiment... For a positive emotion with an intensity of 0.9, the group scores 1 point, corresponding to a 60° interior angle of the triangle, indicating the highest degree of fit due to balanced angles. If the emotion tendencies of the two modalities are similar (i.e., the emotion categories are the same but the intensity difference is 0.1-0.3, or the emotion categories are adjacent and the intensity difference is ≤0.2), such as one group being inclined towards positive emotion with an intensity of 0.9 and the other being inclined towards positive emotion with an intensity of 0.7; or one group being inclined towards neutral to slightly positive emotion with an intensity of 0.6 and the other being inclined towards positive emotion with an intensity of 0.6, the score is 1 point. If the degree is 0.7, the comparison score for this group is 0.5, corresponding to an interior angle of 45° in the triangle. The angles are relatively balanced, and the degree of fit is moderate. If the emotional tendency descriptions of the two modalities are completely different, that is, the emotional categories are not adjacent or the intensity deviation is >0.3, such as one group tending to positive emotions and the other group tending to negative emotions; or one group tending to positive emotions with an intensity of 0.9 and the other group tending to positive emotions with an intensity of 0.5, the comparison score for this group is 0, corresponding to an interior angle of 30° in the triangle. The angles are unbalanced, and the degree of fit is the lowest.

[0046] The consistency scores for the three comparison groups were calculated sequentially according to the above rules. Assuming the sentiment tendency descriptions of the three modalities in this comparison are as follows: Video modality, tending towards positive emotion, intensity level 0.9; Audio modality, tending towards positive emotion, intensity level 0.8; Text modality, tending towards positive emotion, intensity level 0.7; In the first group, video and audio have the same sentiment category and an intensity deviation of 0.1, meeting the similarity standard and scoring 0.5 points; In the second group, audio and text have the same sentiment category and an intensity deviation of 0.1, meeting the similarity standard and scoring 0.5 points; In the third group, text and video have the same sentiment category and an intensity deviation of... 0.2, meeting the similarity standard, receives 0.5 points; the total score of the three comparisons = 0.5 + 0.5 + 0.5 = 1.5 points, corresponding to the sum of the three interior angles of a triangle being 180°, and the total score conforms to the allocation logic of a fixed value of 3 points; calculate the fusion contribution of each modality. The core logic of the fusion contribution is that the contribution of each modality is positively correlated with the sum of the scores of the two comparisons in which that modality participates. Each comparison score in the total score involves two modalities, therefore it needs to be divided by twice the total score to ensure that the sum of the contributions of the three modalities is 1. The specific calculation steps are as follows: the two comparisons in which the video modality participates... Video and audio score 0.5 each, text and video score 0.5 each, summing the scores for both groups = 0.5 + 0.5 = 1 point; For audio modality comparisons, video and audio score 0.5 each, audio and text score 0.5 each, summing the scores for both groups = 0.5 + 0.5 = 1 point; For text modality comparisons, audio and text score 0.5 each, text and video score 0.5 each, summing the scores for both groups = 0.5 + 0.5 = 1 point; Double the total score = 1.5 × 2 = 3 points; The fusion contribution of the video modality = sum of scores for the two video modalities ÷ double the total score = 1 ÷ 3 ≈ 0.333 The fusion contribution of the audio modality = the sum of the scores of the two audio-related groups ÷ twice the total score = 1 ÷ 3 ≈ 0.333; the fusion contribution of the text modality = the sum of the scores of the two text-related groups ÷ twice the total score = 1 ÷ 3 ≈ 0.334; finally, the sum of the fusion contributions of the three modalities is verified to be 0.333 + 0.333 + 0.334 = 1, which meets the basic requirements of contribution allocation. If there is a slight deviation, the contribution of one modality is fine-tuned to ensure that the total contribution is strictly 1, which can objectively reflect the reliability and fit of the emotional description of each modality.

[0047] Step 227: Based on the fusion contribution, the sentiment tendency descriptions of each modality are proportionally fused to form a preliminary comprehensive sentiment representation. This preliminary comprehensive sentiment representation is then normalized to obtain a final comprehensive emotional state representation. Specifically, this includes: first, converting the sentiment tendency description of each modality into quantifiable sentiment values. The conversion rules are based on sentiment category and intensity level, with a base value of 1 for positive sentiment, 0.5 for neutral sentiment, and 0 for negative sentiment; intensity levels are converted from 0 to 1. The intensity level is directly used as the correction factor; for example, an intensity level of 0.9 corresponds to a correction factor of 0.9. The final sentiment value = base value × correction factor. For example, in the video modality, with a positive sentiment category, a base value of 1, and an intensity level of 0.9, the sentiment value = 1 × 0.9 = 0.9; in the audio modality, with a positive sentiment category, a base value of 1, and an intensity level of 0.8, the sentiment value = 1 × 0.8 = 0.8; in the text modality, with a positive sentiment category, a base value of 1, and an intensity level of 0.7, the sentiment value = 1 × 0.7 = 0.7. For neutral to slightly positive and negative... For mixed sentiment tendencies such as neutral or slightly positive, the base value is the average of the base values ​​of the two corresponding sentiment types, multiplied by an intensity correction coefficient. For example, the base value for a neutral-slightly positive sentiment is (1 + 0.5) ÷ 2 = 0.75, and the sentiment value for an intensity level of 0.6 is 0.75 × 0.6 = 0.45. Based on the fusion contribution, a preliminary comprehensive sentiment representation is calculated. The calculation logic is as follows: the sentiment value of each modality is multiplied by its corresponding fusion contribution to obtain the weighted sentiment value for that modality. The weighted sentiment values ​​of the three modalities are then added together, and the sum is the preliminary comprehensive sentiment representation. Substituting the values ​​into the current sentiment score and fusion contribution calculation, the video modality weighted sentiment score = video sentiment score × video fusion contribution = 0.9 × 0.333 ≈ 0.2997; the audio modality weighted sentiment score = audio sentiment score × audio fusion contribution = 0.8 × 0.333 ≈ 0.2664; the text modality weighted sentiment score = text sentiment score × text fusion contribution = 0.7 × 0.334 ≈ 0.2338; the preliminary comprehensive sentiment representation = 0.2997 + 0.2664 + 0.2338 ≈ 0.8.

[0048] Normalization is performed by first determining the effective range of sentiment values ​​as 0 to 1, where 0 corresponds to extreme negativity, 0.5 to absolute neutrality, and 1 to extreme positivity. The purpose of normalization is to ensure that the initial comprehensive sentiment representation falls within this effective range, avoiding deviations due to calculation errors or special circumstances. The processing rule is as follows: if the initial comprehensive sentiment representation is within the range of 0 to 1, then the normalized result = initial comprehensive sentiment representation; if the initial comprehensive sentiment representation > 1, then the normalized result = 1; if the initial comprehensive sentiment representation < 0, then the normalized result = 0. In this case, the initial comprehensive sentiment representation is 0.8, which falls within the range of 0 to 1. Therefore, the normalized result... If the result is still 0.8, a structured comprehensive emotional state representation is generated by combining the emotional category corresponding to the emotional value and the core basis. The representation format must include three parts: emotional category, emotional intensity, and core basis modality. For example, the comprehensive emotional state representation in this case is: comprehensive emotional state is positive, emotional intensity is 0.8, and the core basis is that the emotional tendencies of the three modalities of video, audio, and text are highly consistent. Among them, the positive emotional features of the video modality are the most significant, and the positive emotional features of the audio and text modalities corroborate each other. This representation not only clarifies the emotional category and intensity, but also explains the core basis, ensuring the credibility and traceability of the comprehensive emotional state.

[0049] By using multi-dimensional semantic tagging, cross-modal semantic alignment, and interaction correlation strength analysis, explicit emotional cues from video, audio, and text modalities were comprehensively integrated. With the help of strict scoring rules and contribution allocation logic in emotional consistency assessment, the reliability of emotional expressions in different modalities was effectively distinguished.

[0050] In a preferred embodiment of the present invention, step 3 above, which involves obtaining dialogue text content that matches the user's current emotional state based on the comprehensive emotional state representation and the contextual information of the multi-turn dialogue using a natural language generation model, and converting it into speech output through a speech conversion process, may include:

[0051] In this embodiment of the invention, step 330 involves fusing the comprehensive emotional state representation with the vectorized multi-turn dialogue context to generate a contextual semantic representation. Specifically, this includes: determining the specific forms of two core inputs; the comprehensive emotional state representation is a fixed-dimensional numerical vector containing quantified values ​​for three core dimensions—valence, arousal, and dominance—as well as emotional contribution values ​​corresponding to video, audio, and text modalities, with each dimension's value falling within a preset uniform range; the multi-turn dialogue context refers to the historical dialogue content prior to the current interaction, including user-sent statements and previously generated system response statements, arranged chronologically to form a dialogue sequence; vectorizing the multi-turn dialogue context by first splitting each sentence in the dialogue sequence into independent lexical units, removing stop words without actual semantic meaning, and retaining core nouns, verbs, adjectives, and other semantic words; assigning corresponding semantic weights to each core word, based on the word's importance in the dialogue, such as a weight of 3 for core appeal words, 1 for descriptive words, and 0.5 for connective words; mapping each word to a fixed-length base vector, and then calculating the semantic weight of each word by multiplying the base vector by the semantic weight. The process begins by weighting the words into a single sentence vector. This initial vector is obtained by summing the weighted vectors of all words in the sentence. Following the dialogue sequence, these initial vectors are concatenated sequentially. Then, a weighted recursive method of multiplying the vector of the previous sentence by 0.7 and the vector of the next sentence by 0.3 is used to integrate all sentence vectors, resulting in a vectorized encoding of the multi-turn dialogue context. This encoding result maintains the same vector dimension as the comprehensive emotional state representation. A fusion calculation is then performed to generate a contextual semantic representation. First, fusion weights are assigned to the comprehensive emotional state representation and the multi-turn dialogue context encoding vector, based on the interaction scenario. For example, in an emotional support scenario, the weight of the comprehensive emotional state representation is 0.6, and the weight of the multi-turn dialogue context encoding vector is 0.4; in a regular consultation scenario, both weights are 0.5. Then, the sum of the comprehensive emotional state representation vector multiplied by its corresponding fusion weight and the multi-turn dialogue context encoding vector multiplied by its corresponding fusion weight is calculated to obtain the initial fusion vector. Finally, the values ​​of each dimension in the initial fusion vector are normalized by dividing each dimension's value by the sum of all dimension values, ensuring that the sum of all dimension values ​​equals 1. This results in a contextual semantic representation that simultaneously contains emotional and contextual semantic information.

[0052] Step 331 involves inputting the contextual semantic representation into a pre-trained natural language generation model to generate dialogue text that matches the current emotional state in terms of semantic content and expression. Specifically, this includes: constructing a pre-trained natural language generation model by selecting a large-scale general dialogue corpus containing dialogue samples from various scenarios such as casual conversation, consulting services, and emotional guidance, as well as an emotion-labeled corpus of dialogues. Each dialogue sample is labeled with its corresponding emotion type, such as happy, angry, depressed, or anxious, as the data source for training the natural language generation model; and inputting samples from the general dialogue corpus into the basic language to allow the natural language generation model to learn semantic coherence. The training process involves several steps: first, identifying the subject matter, topic relevance, and standardization of language expression; then, classifying samples from an emotion-labeled corpus of dialogues by emotion type and inputting each type into a pre-learned general semantic base model; adjusting the parameters within the natural language generation model to prioritize activating the vocabulary and sentence templates related to the corresponding emotion when generating text; during training, comparing the semantic similarity and emotion matching between the text generated by the natural language generation model and the original samples in the corpus, continuously refining the parameters of the natural language generation model until the semantic accuracy and emotion matching of the generated text reach the preset standards, thus completing the construction of the pre-trained natural language generation model.

[0053] The contextual semantic representation is input into the pre-trained natural language generation model. The model first parses the multi-turn dialogue semantic information in the contextual semantic representation to clarify the core topic of the current dialogue, the user's expressed needs, and the historical interaction logic, ensuring that the generated text is consistent with the topic. Then, it extracts the emotional information in the contextual semantic representation to determine the user's current emotional type and intensity. It matches the corresponding emotional vocabulary according to the emotional type and adjusts the usage of emotional vocabulary according to the emotional intensity. At the same time, it determines the core function of the generated text by combining the contextual semantics and selects sentence patterns that match the function and emotion. Finally, through the semantic organization logic inside the natural language generation model, the selected vocabulary and the determined sentence patterns are combined to generate dialogue text that fits the dialogue context and is highly consistent with the user's current emotional state.

[0054] Step 332 involves inputting the dialogue text and the comprehensive emotional state representation into the emotional speech synthesizer, mapping the corresponding prosody, pitch, and rhythm parameters based on the comprehensive emotional state. Specifically, this includes: constructing the emotional speech synthesizer by collecting a database of real-person speech samples containing various emotional types. Each speech sample corresponds to specific text content and an emotion label, while also annotating the sample's prosody, pitch, and rhythm parameters; breaking down the text content in the speech sample database into syllable units and establishing a correspondence database of text syllables, speech parameters, and emotion labels; building a speech parameter mapping rule system to associate the core dimensions of the comprehensive emotional state representation with the speech parameters; and optimizing and improving this mapping rule system using data from the sample database to ensure that the system understands the correspondence between text syllables and speech parameters under different emotional states. By continuously optimizing the associated parameters in the rule system, it can accurately output the corresponding speech parameters based on the input emotional state and text content, thus completing the construction of the emotional speech synthesizer.

[0055] Parameter mapping is performed by inputting the dialogue text into the emotional speech synthesizer. The text is first segmented into words, sentences, and syllables. Then, a comprehensive emotional state representation is input into the synthesizer. The synthesizer first extracts a valence value to determine the positive or negative tendency of the emotion; a valence value higher than a preset median indicates a positive emotion, while a value lower indicates a negative emotion. It then extracts an arousal value to determine the intensity of the emotion; an arousal value higher than a preset median indicates a high-intensity emotion, while a value lower indicates a low-intensity emotion. Based on the positive or negative tendency and intensity of the emotion, basic parameters are matched from a text syllable-speech parameter-emotion label mapping database. For example, a positive, high-intensity emotion corresponds to a higher basic pitch and a higher basic rhythm. A faster pace and lower intensity negative emotions correspond to a lower base pitch and slower base rhythm. This is then combined with the modal emotional contribution in the overall emotional state representation. For example, a high audio modal emotional contribution indicates significant user emotional expression, requiring enhancement of the pitch fluctuation parameter; a high video modal emotional contribution indicates significant user facial emotions, requiring enhancement of the prosodic pause parameter. The base parameters are then corrected, such as audio modal emotional contribution × 0.3 + base pitch = corrected pitch, and video modal emotional contribution × 0.2 + base pause duration = corrected pause duration. Finally, prosodic, pitch, and rhythm parameters that accurately match the overall emotional state are generated.

[0056] Step 333 involves performing speech synthesis on the dialogue text based on prosody, pitch, and rhythm parameters to obtain a natural speech output that carries corresponding emotional expression. Specifically, this includes: preprocessing the dialogue text by breaking it down into the smallest pronunciation units based on the labeled syllables and labeling the standard pronunciation duration of each unit; adjusting the actual pronunciation duration of each unit in conjunction with the pronunciation speed in the rhythm parameters; adjusting the pronunciation features according to the parameters; setting the baseline pitch of the speech based on the fundamental frequency in the pitch parameters; adjusting the pitch changes of the corresponding pronunciation units based on the frequency fluctuation values ​​of each syllable to ensure that the pitch changes meet the needs of emotional expression; inserting pauses at corresponding positions in the dialogue text according to the pause positions and durations in the prosody parameters to avoid abrupt speech; performing speech synthesis and optimization by sequentially splicing the adjusted pronunciation units in the text order to form a continuous speech stream; smoothing the speech stream; and fine-tuning the volume of the speech stream by referring to the emotional expression patterns of real human speech; ultimately obtaining a speech output that conforms to the semantics of the dialogue text, carries emotional expression consistent with the user's current emotional state, and is natural and fluent.

[0057] By deeply integrating the overall emotional state with the context of multi-turn dialogue, the generated dialogue text is ensured to closely adhere to the semantic logic of the preceding text, stay on topic of discussion, and accurately match the user's current true emotional state.

[0058] In a preferred embodiment of the present invention, step 4 above involves real-time monitoring of the dynamic changes in the user's biosignals and speech signals during the speech output process, organizing the newly acquired multimodal data into a time-series data stream, and constructing an emotional state analysis benchmark. By detecting key turning points and extreme points in the time-series data stream, key event points representing the boundaries of emotional states are extracted, and two emotional change patterns are identified, which may include:

[0059] In this embodiment of the invention, step 440 involves using a preset duration sliding time window to extract continuously acquired raw biological signals and speech signals, obtaining biological signal segments and speech signal segments within the current time window. Specifically, this includes: determining the setting standards for the preset duration and sliding step size; combining the emotional response timeliness of biological signals and speech signals, setting the preset duration of the sliding time window to 2 seconds and the sliding step size to 1 second; determining the type and format standards of the acquired raw signals; the raw biological signals include electroencephalogram (EEG) signals, electrocardiogram (ECG) signals, and electrodermal response signals, with the acquisition format of these three signals consistent with step 1; the acquisition format of the raw speech signals is also consistent with step 1, being a 16kHz sampling rate, 16-bit depth mono signal; performing signal extraction operations; and extracting the speech output... Real-time acquisition begins instantly, using the current system time as a reference. A window capture is triggered every second. During capture, the current trigger time is used as the end point of the window, and two seconds prior is used as the starting point to determine the time range of the window. Segments are captured and stored separately according to signal type. For EEG, ECG, and TENS signals, all sampling point data within the corresponding time range are captured to form three independent biosignal segments. At the same time, speech signal sampling point data within the corresponding time range are captured to form speech signal segments. Each segment is labeled with a corresponding timestamp and named and stored in the format of signal type-window sequence number-timestamp to ensure the traceability of the segments. For example, EEG-1-0-2s represents the EEG signal segment within the first window and the time range of 0-2 seconds.

[0060] Step 441: Extract features from the biosignal and speech signal segments within the current time window to obtain real-time feature vectors representing the user's current physiological and speech states. Specifically, this includes: extracting features from the biosignal segments, ensuring the extraction dimensions remain consistent with those in Step 1 to guarantee feature uniformity; for EEG signal segments, the 2-second EEG segment is divided into two smaller segments of 1 second each; the mean of each segment is calculated; the sum of all sampled values ​​within a segment is divided by the number of sampled values ​​in that segment; the variance is calculated; the squares of the differences between each sampled value and the segment mean are summed and divided by the number of sampled values; and the energy percentages of alpha waves (8-13Hz) and beta waves (14-30Hz) are calculated (energy percentage of a certain frequency band = the sum of the squares of all sampled values ​​within that frequency band divided by the sum of the squares of all sampled values ​​in the entire smaller segment); then, the average of the features corresponding to the two smaller segments is calculated (the average of the two smaller segments). The four characteristic parameters of EEG are obtained by summing the mean values ​​of the segments and dividing by 2, and by similarly for variance and energy percentage. For ECG segment characteristics, the number of R-wave peaks within a 2-second ECG segment is counted, multiplied by 30 to obtain the heart rate per minute (since 2 seconds is 1 / 30 of a minute, the number of peaks multiplied by 30 equals the heart rate per minute). The time interval between consecutive R-wave peaks is calculated to obtain several RR interval values. The mean (sum of all RR interval values ​​divided by the number of intervals) and standard deviation (sum of the squares of the differences between each RR interval value and the mean divided by the number of intervals) of these values ​​are then calculated, resulting in the three characteristic parameters of ECG. For GSR segment characteristics, the peak value of a 2-second GSR segment is calculated, along with the maximum value of all sampling points within the segment, the rate of rise (the difference between the peak value and the value at the beginning of the segment, divided by the difference between the peak occurrence time and the segment start time), and the sum of values ​​(the sum of all sampling point values ​​within the segment), ultimately yielding the three characteristic parameters of GSR.

[0061] The four features of EEG, three features of ECG, and three features of GSR are arranged in the order of EEG-ECG-GSR to form a 10-dimensional biosignal feature sub-vector. Next, feature extraction is performed on the speech signal segment, maintaining the same dimensionality as the audio features in step 1. For temporal features, the 2-second speech segment is divided into several frames with a 20ms frame length and a 10ms frame shift. Silent frames with short-time energy below 10% of the average of all frames are removed. For the remaining valid frames, the short-time energy, zero-crossing rate, and fundamental frequency are calculated. The average value of the corresponding features for all valid frames is then calculated. The short-time energy of all valid frames is summed and divided by the number of valid frames. The zero-crossing rate and fundamental frequency are calculated similarly, resulting in three temporal feature parameters. For frequency domain features, Fourier transform correlation processing is performed on each valid frame to obtain the spectral amplitude at 256 frequency points. The average spectral amplitude of all valid frames at corresponding frequency points is calculated. The amplitude values ​​of all valid frames at each frequency point are summed and divided by the number of valid frames to form a 256-dimensional spectral amplitude feature. Then, the spectral centroid is calculated by multiplying the frequency value of each frequency point by the average spectral amplitude of that point, summing all products, and dividing by the sum of the average spectral amplitudes of all frequency points to obtain a spectral centroid parameter. The three time-domain features, the 256-dimensional spectral amplitude feature, and the 1 spectral centroid feature are arranged in order to form a 260-dimensional speech signal feature sub-vector. The 10-dimensional biosignal feature sub-vector and the 260-dimensional speech signal feature sub-vector are concatenated in the order of biosignal and speech signals to form a vector with a dimension of 10 + 260 = 270. This vector is the real-time feature vector representing the user's current physiological and speech states.

[0062] Step 442 involves concatenating and aligning the real-time feature vectors with a preset reference feature sequence to form a unified and continuous temporal data stream. Specifically, this includes determining the source and composition of the preset reference feature sequence. The reference feature sequence is a sequence of feature vectors collected from the user in a calm state before the start of speech output. The collection process involves continuously collecting biosignals and speech signals for five consecutive windows with a 2-second window and a 1-second step size. Five real-time feature vectors are extracted using the method in step 441. These five vectors are arranged in chronological order of collection time to form the preset reference feature sequence. The sequence length is 5, each vector has a dimension of 270, and each vector is labeled with a corresponding collection timestamp. After the start of speech output, each real-time feature vector is labeled with its corresponding window time. The timestamp is concatenated with the timestamp of the last vector in the reference feature sequence to ensure that the timestamps of all vectors are continuous without any breaks. Each vector in the reference feature sequence and the newly generated real-time feature vector are both 270-dimensional, and the feature types corresponding to each dimension are completely consistent. No dimension adjustment is required. The vectors are directly concatenated according to the order of their timestamps. The reference feature sequence is used as the initial part of the data stream. After that, each time a new real-time feature vector is generated, it is added to the end of the data stream in the order of its timestamps, so that the length of the data stream gradually increases over time. Each element in this data stream is a 270-dimensional feature vector, and they are arranged continuously according to their timestamps to form a unified and continuous time-series data stream that fully records the changes in physiological and voice features of the user from a calm state to voice interaction.

[0063] Step 443 involves performing aggregation analysis on the reference feature vectors corresponding to the user's calm state intervals in the time-series data stream, calculating the central trend of the reference feature vector numerical distribution, and constructing an emotional state baseline, i.e., an emotional state analysis baseline. Specifically, this includes: filtering reference feature vectors for the calm state intervals; extracting an initial reference feature sequence from the time-series data stream; these vectors correspond to the feature data of the user in a calm state before speech output, free from emotional fluctuation interference, and are determined as the set of reference feature vectors for the calm state intervals; performing aggregation analysis on each feature dimension separately, calculating the central trend, and using the mean as the central trend indicator, balancing computational simplicity and representativeness; for each dimension of the 270-dimensional features, extracting the values ​​of 5 reference feature vectors in that dimension, adding these 5 values, and then dividing by the number of reference vectors (5) to obtain the mean of that dimension. This mean is the central trend value of that feature dimension in the calm state. For example, the 5 reference values ​​for the first dimension (EEG mean) are 0.3, 0.32, 0.28, 0.31, and 0.29, which are added together to get 1.5, and then divided by 5. The value 0.3 is obtained, which is the central tendency value of the first dimension. An emotional state baseline is constructed by arranging the central tendency values ​​of the 270 feature dimensions in the original feature dimension order, forming a 270-dimensional baseline vector. This baseline vector is the emotional state baseline. The value of each dimension in the baseline represents the typical level of that feature when the user is calm. For example, the baseline value of the 11th dimension is 0.4, indicating that the typical value of the user's short-term vocal energy when calm is 0.4. The degree of deviation between the real-time feature vector and this baseline can be compared to determine whether the emotion has changed. The baseline is then verified and calibrated by calculating the Euclidean distance between each of the five reference feature vectors and the baseline. The squares of the differences between each dimension value and the corresponding dimension value of the baseline are summed, and the square root of the sum is taken. If the Euclidean distance of all vectors is less than a preset deviation threshold, the baseline is valid. If there is a vector exceeding the threshold, the abnormal reference vector is removed, and the mean values ​​of each dimension of the remaining four vectors are recalculated to update the baseline. This process continues until the distances between all reference vectors and the baseline meet the requirements, ensuring the reliability of the baseline.

[0064] Step 444 involves dividing the time-series data stream into fixed-duration windows, calculating the temporal and frequency domain features of the biosignals within each window, and generating a high-dimensional feature sequence. Specifically, this includes setting parameters for the fixed-duration window: based on the update frequency of the time-series data stream, the duration of the fixed-duration window is set to 1 second, and the step size is set to 0.5 seconds. This means each window contains two consecutive data stream vectors. The fixed-duration window is divided starting from the beginning timestamp of the first vector of the time-series data stream, with windows divided sequentially according to a duration of 1 second and a step size of 0.5 seconds. For example, the first window's time range is 0-1 second. The second window is 0.5-1.5 seconds, the third is 1-2 seconds, and so on, ensuring that each fixed window corresponds to a portion of the vectors in the data stream and covers the entire time range of the data stream. Biosignal features are extracted within each fixed window. For temporal feature extraction, for each fixed window, biosignal feature sub-vectors are extracted from all data stream vectors covered by that window. For each dimension, the mean of all values ​​in that dimension within the window is calculated. All values ​​in that dimension within the window are summed and divided by the number of values ​​and the variance. The squares of the differences between each value and the mean are summed and divided by the number of values ​​and the maximum value. The maximum and minimum values ​​of each dimension, and the minimum value of that dimension within the window, yield 4 time-domain features for each dimension, resulting in 40 time-domain feature parameters across 10 dimensions. For frequency-domain feature extraction, the biological signal feature sub-vectors within each fixed window are subjected to Fourier transform processing according to each dimension. The spectral energy of each dimension is calculated, along with the sum of the squares of all frequency points in the frequency domain, the spectral centroid, and the product of the frequency value of each frequency point and its corresponding frequency domain value. All these products are then summed and divided by the sum of the frequency domain values ​​of all frequency points, resulting in 2 frequency-domain features for each dimension, totaling 20 frequency-domain features across 10 dimensions. Feature parameters; finally, a high-dimensional feature sequence is generated by fusing the 40 time-domain feature parameters and 20 frequency-domain feature parameters of each fixed window in the order of time domain and frequency domain to form a 60-dimensional window feature vector; then, the 60-dimensional feature vector of each window is concatenated with the speech signal feature sub-vector in the data stream vector covered by the window to form a 60+260=320-dimensional high-dimensional feature vector; according to the division order of the fixed windows, all high-dimensional feature vectors are arranged to form a high-dimensional feature sequence with the sequence length consistent with the number of fixed windows, and each vector is labeled with the corresponding window timestamp.

[0065] Step 445 involves performing first-order difference calculations on the high-dimensional feature sequence to obtain a feature change rate sequence, detecting zero-crossing points and local extrema in the feature change rate sequence, and identifying candidate key event points. Specifically, this includes: calculating the first-order difference to obtain the feature change rate sequence; performing first-order difference calculations on each of the 320 feature dimensions in the high-dimensional feature sequence; for each dimension, subtracting the feature value of the previous window from the feature value of the next window according to the window order of the high-dimensional feature sequence to obtain the difference value for that dimension, i.e., the feature change rate; and then applying the difference values ​​of all windows... The 320 subsequences of change rates are sequentially arranged to form a subsequence of change rate for each dimension. These subsequences are then concatenated in their original dimensional order to form a 320-dimensional feature change rate sequence. The sequence length is one less than the length of the higher-dimensional feature sequence. For each subsequence of change rate, the signs of adjacent difference values ​​are checked one by one. If the preceding difference value is positive and the following is negative, or vice versa, a zero-crossing point is determined to exist between the windows corresponding to these two difference values. The window position and its corresponding feature dimension are recorded. For example, the third difference value in a certain dimension... The difference value of the first window is 0.2, and the fourth window is -0.1. Therefore, there is a zero-crossing point between the third and fourth windows. The window position is recorded as 4, and the dimension is the feature dimension. For each subsequence of the rate of change in each dimension, the relationship between each difference value and its two adjacent difference values ​​is checked one by one. The first and last difference values ​​are compared only with their adjacent values. If a difference value is greater than its two adjacent difference values, it is determined to be a local maximum; if a difference value is less than its two adjacent difference values, it is determined to be a local minimum. Each local extremum is recorded. The window position and feature dimension corresponding to the point are recorded. For example, if the difference value of the 5th window in a certain dimension is 0.3, the 4th is 0.2, and the 6th is 0.1, then the difference value of the 5th window is a local maximum point. The window position is recorded as 5, and the dimension is the feature dimension. All records of zero-crossing points and local extreme points are summarized, and duplicate window positions are removed to obtain a list of candidate key event points. Each candidate point in the list contains two core pieces of information: window position and the set of feature dimensions involved. These candidate points represent potential time nodes where emotional state changes may occur.

[0066] Step 446: Candidate key event points are filtered using preset amplitude and duration thresholds, merging closely adjacent candidate key event points to form a key event point sequence. Specifically, the amplitude threshold is used to filter candidate points with significant feature changes. It is calculated by extracting all difference values ​​in the calm state interval before speech output for each dimension of the feature change rate sequence, calculating the mean of the absolute values ​​of these values, summing all absolute values, dividing by the number of values, and then multiplying the mean by 2 to obtain the amplitude threshold for that dimension. The amplitude thresholds for 320 dimensions are arranged in dimensional order to form a 320-dimensional amplitude threshold vector. For each candidate key event point, the absolute value of the difference values ​​in all feature dimensions involved is checked. If at least 50% of the dimensions have absolute difference values ​​greater than the amplitude threshold of the corresponding dimension, the candidate point is retained; if less than 50%, it is determined to be a false candidate point caused by noise interference and is removed. The duration threshold is used to merge closely adjacent candidate points to avoid duplicate counting. Based on the temporal characteristics of emotional changes, it is set to 0.5 seconds, i.e., the time interval between two candidate points. When the interval is less than 0.5 seconds, it is determined to be too close. The filtered candidate key event points are sorted in ascending order according to the timestamp corresponding to the window position. The time interval between two adjacent candidate points is calculated. If the time interval is less than 0.5 seconds, the two candidate points are merged into one key event point. The timestamp of the merged point is the average of the timestamps of the two points. The set of feature dimensions involved is the union of the dimension sets of the two points. If the time interval is greater than or equal to 0.5 seconds, the two independent candidate points are retained and not merged. For example, if the timestamp of candidate point A is 1.2 seconds and the timestamp of candidate point B is 1.5 seconds, the interval is 0.3 seconds < 0.5 seconds, and they are merged into a key event point with a timestamp of 1.35 seconds. The set of dimensions is the result of merging the dimension sets of A and B. All the merged key event points are sorted in ascending order according to the timestamp. Each point contains the timestamp, the set of feature dimensions involved, and the maximum absolute value of the corresponding difference value. They are arranged in order to form a key event point sequence. Each point in this sequence represents a clear time node where the emotional state may change, and the feature changes are significant and there is no duplication or redundancy.

[0067] Step 447: Based on the key event point sequence, calculate the time interval and feature amplitude difference between adjacent key event points in the sequence. Specifically, this includes: calculating the time interval between adjacent key event points; arranging the key event point sequence in ascending order of timestamps; assuming the sequence contains N key event points; sequentially pairing the 1st with the 2nd, 2nd with the 3rd, ..., N-1th with the Nth point to form N-1 pairs; for each pair, subtracting the timestamp of the preceding point from the timestamp of the following point to obtain the time interval. For example, if the timestamp of the 1st point is 2.0 seconds and the timestamp of the 2nd point is 3.5 seconds, the time interval is 3.5 seconds - 2.0 seconds = 1.5 seconds; if the timestamps of the 2nd and 3rd points are 3.5 seconds and 5.0 seconds respectively, the time interval is 1.5 seconds. Calculate the time interval for all adjacent pairs sequentially to obtain N-1 time interval values; calculate the feature amplitude difference between adjacent key event points; for each adjacent pair, extract the previous... The high-dimensional feature vectors corresponding to each point are used to calculate the feature difference for each feature dimension. The absolute value of this difference is then taken. The absolute values ​​of the feature differences across the 320 dimensions are summed and divided by the total number of feature dimensions (320) to obtain the average feature amplitude difference for that pair. This value reflects the overall intensity of feature changes between two key event points. For example, if the sum of the absolute values ​​of the feature differences across the 320 dimensions for a pair is 64, the average feature amplitude difference is 64 ÷ 320 = 0.2. If the sum for another pair is 96, the average feature amplitude difference is 96 ÷ 320 = 0.3. The average feature amplitude difference for all adjacent pairs is calculated sequentially to obtain N-1 amplitude difference values. The time interval value of each adjacent pair is then bound to its corresponding average feature amplitude difference value to form N-1 time interval and amplitude difference data pairs. Each data pair corresponds to a continuous emotional change process in the sequence of key event points.

[0068] Step 448 compares the time interval and the difference in characteristic amplitude with preset thresholds to distinguish between abrupt and gradual emotional states. Specifically, this includes: determining the time interval threshold and the characteristic amplitude difference threshold, both based on the characteristic change patterns in a calm state. The time interval threshold is calculated by setting half of the average time interval between key event points corresponding to the calm state interval. For example, if the average time interval in a calm state is 2 seconds, the time interval threshold = 2 seconds ÷ 2 = 1 second. The characteristic amplitude difference threshold is calculated by setting three times the average characteristic amplitude difference between all time interval and amplitude difference data pairs in the calm state interval. For example, if the average amplitude difference in a calm state is 0.1, the characteristic amplitude difference threshold = 0.1 × 3 = 0.3. A pattern differentiation rule is established based on the characteristics of emotional changes, clarifying the criteria for judging the two patterns: abrupt change corresponds to a rapid change in emotion within a short period, while gradual change corresponds to a slow change in emotion over a longer period. Specifically, if in a certain time interval and amplitude difference data pair, the time interval is less than the time interval threshold and the characteristic amplitude difference is greater than the characteristic amplitude difference threshold, then... If a data pair shows a sudden change in emotional state, and the time interval threshold is 1 second and the amplitude difference threshold is 0.3, then a data pair with a time interval of 0.8 seconds (<1 second) and an amplitude difference of 0.4 (>0.3) is considered to have a sudden change. If a data pair shows a time interval greater than or equal to the time interval threshold and an amplitude difference less than or equal to the amplitude difference threshold, then the data pair is considered to have a gradual change in emotional state. For example, a data pair with a time interval of 1.2 seconds (≥1 second) and an amplitude difference of 0.2 (≤0.3) is considered to have a gradual change. If neither of the above two rules is met, the emotional state baseline is used to assist in the judgment. The Euclidean distance between the high-dimensional feature vectors of the two key event points corresponding to the data pair and the baseline is calculated. The squares of the differences between each dimension value and the corresponding dimension value of the baseline are added together, and the square root is taken. If the difference between the distances of the two points and the baseline is greater than twice the average distance of the baseline, it is determined to be a sudden change pattern; if the difference is less than or equal to twice the average distance of the baseline, it is determined to be a gradual change pattern. According to the order of the time interval and amplitude difference data pairs, the patterns corresponding to each data pair are arranged in sequence to form an emotional change pattern sequence.

[0069] It enables dynamic capture of users' physiological and voice signals, ensuring timely perception of subtle emotional fluctuations, avoiding delayed responses to emotional changes, reducing interference from individual differences and environmental noise on emotion recognition, and improving the accuracy of emotion judgment.

[0070] In a preferred embodiment of the present invention, step 5 above, which analyzes two emotion change patterns, determines the fluctuation range of the emotion state by combining the results of fusion semantic analysis, and selects reference instances of the emotion state to construct an emotion state change trajectory; and generates dynamic adjustment parameters of the emotion state based on the emotion state change trajectory to adjust the dialogue strategy and generated content in real time to achieve an adaptive emotion response, may include:

[0071] In this embodiment of the invention, step 550 involves calculating the time span and characteristic amplitude range of emotional state fluctuations corresponding to each of the mutation and gradual change modes, respectively, to form preliminary fluctuation range parameters. Specifically, this includes: determining the key event point sequences corresponding to each of the mutation and gradual change modes; for the mutation mode, firstly selecting all adjacent key event points under this mode, calculating the time difference between each pair of adjacent key event points (i.e., the timestamp of the later key event point minus the timestamp of the earlier key event point), and then summing all these time differences to obtain the sum, which is the emotional state fluctuation time span corresponding to the mutation mode; simultaneously, extracting the real-time feature vector values ​​corresponding to each key event point under the mutation mode, and finding... The maximum and minimum values ​​of these feature vectors are extracted, and the minimum value is subtracted from the extracted maximum value. The result is the range of emotional state feature amplitude corresponding to the mutation mode. For the gradual change mode, the same calculation method is used as for the mutation mode. First, all time differences between adjacent key event points in this mode are calculated and summed to obtain the time span of emotional state fluctuation in the gradual change mode. Then, the real-time feature vector values ​​corresponding to all key event points in this mode are extracted, the maximum and minimum values ​​are found, and the difference is calculated to obtain the range of emotional state feature amplitude in the gradual change mode. The fluctuation time span and feature amplitude range of the mutation mode and the fluctuation time span and feature amplitude range of the gradual change mode are used as the initial fluctuation interval parameters for the two modes, respectively.

[0072] Step 551: Align the preliminary fluctuation range parameters with the emotional semantic information obtained from cross-modal fusion semantic analysis, and use the emotional semantic information to correct and calibrate the boundaries of the fluctuation range to determine the final emotional state fluctuation range. Specifically, this includes: acquiring the emotional semantic information obtained from cross-modal fusion semantic analysis, which includes emotional semantic labels corresponding to multimodal data, time segments corresponding to emotional semantics, and emotional semantic intensity descriptions; mapping the time range corresponding to the fluctuation time span in the preliminary fluctuation range parameters to the time segments in the emotional semantic information to determine all emotional semantic labels and their corresponding intensity descriptions within the preliminary fluctuation time range; simultaneously, associating the feature amplitude range in the preliminary fluctuation range parameters with the emotional semantic intensity descriptions, and then performing boundary correction. If the preliminary fluctuation time span... If the start time is earlier than the start time of the corresponding core emotional semantic in the emotional semantic information, then the start time of the initial fluctuation time span is adjusted to the start time of that core emotional semantic; if the end time of the initial fluctuation time span is later than the end time of the corresponding core emotional semantic in the emotional semantic information, then the end time of the initial fluctuation time span is adjusted to the end time of that core emotional semantic; if the upper limit of the characteristic amplitude range of the initial fluctuation is lower than the upper limit of the conventional characteristic amplitude corresponding to the emotional semantic intensity description, then the upper limit of the characteristic amplitude range is corrected to that conventional amplitude upper limit; if the lower limit of the characteristic amplitude range of the initial fluctuation is higher than the lower limit of the conventional characteristic amplitude corresponding to the emotional semantic intensity description, then the lower limit of the characteristic amplitude range is corrected to that conventional amplitude lower limit. After the above calibration of the time boundary and amplitude boundary, the final emotional state fluctuation range is obtained.

[0073] Step 552: Based on the emotional state fluctuation range, retrieve samples from the emotional state instance library that match the emotional features and emotional semantic labels within the current fluctuation range, and select emotional state reference instances. Specifically, this includes: determining the core information contained in the sample data stored in the emotional state instance library, including the emotional fluctuation time span, emotional feature amplitude range, emotional feature vector set, emotional semantic label, and description of emotional change trend for each sample; then, using the finally determined emotional state fluctuation range as the search condition, firstly, filter out samples in the instance library whose difference between the emotional fluctuation time span and the current fluctuation time span is within a preset reasonable range and whose overlap ratio between the emotional feature amplitude range and the current feature amplitude range reaches a preset threshold; secondly, further compare the emotional semantic labels from the samples selected in the first step, retaining samples that are completely consistent with or highly similar to the emotional semantic labels obtained from the current cross-modal fusion semantic analysis; thirdly, extract the emotional feature vector set within the current fluctuation range, calculate the similarity between these feature vectors and the emotional feature vector set of the samples selected in the second step, and select several samples with the highest similarity by comparing the matching degree of corresponding dimension feature values; finally, determine these selected samples as emotional state reference instances.

[0074] Step 553 involves organizing the emotional state reference instances in chronological order and, in conjunction with the determined emotional state fluctuation range, constructing an emotional state change trajectory reflecting the continuous change process of emotional states. Specifically, this includes: extracting the time series information corresponding to each emotional state reference instance and the emotional feature values ​​and emotional semantic labels corresponding to each time node; then, using the timeline of the current human-computer interaction process as a benchmark, uniformly mapping the time series of all emotional state reference instances to ensure that the time order of the reference instances is consistent with the time order of the current user's emotional changes; next, combining the final emotional state fluctuation range, marking the emotional feature values ​​and emotional semantic labels corresponding to each mapped time node at the corresponding positions on the current timeline; for blank time periods between adjacent time nodes, supplementing transitional emotional feature values ​​based on the changing trends of the emotional feature values ​​between the two nodes, so that the changes in emotional feature values ​​form a continuous curve; finally, integrating the continuous curves marked with time nodes, emotional feature values, and emotional semantic labels to form an emotional state change trajectory that fully reflects the continuous change process of the user from the starting point of emotional fluctuation to the current time point, and includes potential future changing trends.

[0075] Step 554 involves analyzing the trajectory of emotional state changes, extracting the rate of change in emotional intensity and the pattern of emotional type transition, and forming emotional evolution trend characteristics. Specifically, this includes: extracting the emotional intensity value corresponding to each time node from the emotional state change trajectory. This value is obtained by integrating emotional features after multimodal fusion, reflecting the intensity of the emotion. Then, the rate of change in emotional intensity is calculated. Two consecutive adjacent time nodes in the trajectory are selected, and the emotional intensity value of the latter time node is subtracted from the emotional intensity value of the former time node to obtain the emotional intensity difference. This difference is then divided by the time interval between the two time nodes, i.e., the timestamp of the latter time node minus the timestamp of the former time node, to obtain the difference between these two time nodes. The rate of change of emotional intensity between points is calculated. Following the above method, the rate of change of emotional intensity between all adjacent time points in the trajectory is calculated to form a complete sequence of the rate of change of emotional intensity. Then, the emotional type transfer pattern is analyzed. The emotional semantic label corresponding to each time point in the trajectory is observed, and it is recorded whether the label changes, the start time and end time of the change, and the emotional type before and after the change. If the emotional semantic label remains unchanged in multiple consecutive time points, it is determined to be a stable emotional type pattern. If the label changes, it is determined to be an emotional type transfer pattern, and the direction and speed of the transfer are identified. Finally, the obtained sequence of the rate of change of emotional intensity and the emotional type transfer pattern are integrated to form the emotional evolution trend characteristics.

[0076] Step 555: Based on the characteristics of emotional evolution trends and combined with the contextual information of multi-turn dialogues, generate dialogue strategy parameters and content generation parameters through a preset mapping relationship. Dialogue strategy parameters include response timing and the degree of dialogue initiative; content generation parameters include emotional vocabulary density and sentence complexity. Specifically, this includes: acquiring contextual information from multi-turn dialogues, including the topic content of historical dialogues, the core needs previously expressed by the user, the generated dialogue response content, and the user's feedback on historical responses; then, performing correlation analysis between the characteristics of emotional evolution trends and this contextual information to clarify the correlation between the current emotional evolution and historical dialogues; next, invoking the preset mapping relationship; if the emotional evolution trend is a rapid increase in emotional intensity, a shift in negative emotional type, or a change in context... If a user is dissatisfied because their needs haven't been met, the mapping will result in a parameter combination of early response timing, low dialogue initiative, low emotional vocabulary density, and simple sentence complexity. If the emotional evolution trend is a slow increase in emotional intensity, a stable positive emotional type, and the context indicates that the user is interested in the current topic, the mapping will result in a parameter combination of normal response timing, high dialogue initiative, high emotional vocabulary density, and medium sentence complexity. If the emotional evolution trend is a slow decrease in emotional intensity, a stable negative emotional type, and the context indicates that the user is still expressing their feelings, the mapping will result in a parameter combination of delayed response timing, medium dialogue initiative, medium emotional vocabulary density, and simple sentence complexity. Through the above mapping process, the corresponding dialogue strategy parameters and content generation parameters are finally generated.

[0077] Step 556: Apply the dialogue strategy parameters to the dialogue management process to determine the timing of responses and the guidance method for the dialogue. Specifically, this includes: analyzing the response timing parameter in the dialogue strategy parameters. If the parameter is marked as "early," the response preparation process will begin immediately after the user finishes their current expression, without executing the usual response waiting interval. If the parameter is marked as "normal," the response preparation will begin strictly according to the preset usual waiting interval. If the parameter is marked as "delayed," the response preparation will be started after extending the preset time based on the usual waiting interval to ensure the user has sufficient space to express themselves. Then, analyze the dialogue initiative parameter. If the parameter is marked as "high," extended guidance related to the current user's emotions and the dialogue topic will be added to the response content. If the parameter is marked as "medium," the response content will only echo the core content of the current user's expression, without adding additional extended topics, only maintaining the current dialogue direction. If the parameter is marked as "low," the response content will only directly address the specific problem or request expressed by the user, without containing any guiding statements to avoid interfering with the user's subsequent expression. By implementing these two parameters, the response timing and dialogue guidance method for this interaction will be ultimately determined.

[0078] Step 557 involves inputting content generation parameters into the natural language generation process to control the intensity of emotional expression and linguistic complexity of the generated text. Specifically, this includes parsing the emotional lexical density parameter in the content generation parameters. If the parameter is marked as high, emotional lexical words matching the current user's emotional state are prioritized during natural language generation, and the proportion of emotional lexical words in the generated text is not less than a preset ratio. For example, when the user's emotion is pleasant, words like "happy," "fantastic," "reassuring," and "exhilarating" are used. If the parameter is marked as medium, emotional lexical words are used moderately, with the proportion controlled at a preset middle ratio, balancing the accuracy of semantic expression and the appropriateness of emotional transmission. If the parameter is marked as low, subjective emotional lexical words are avoided as much as possible in the generated text. The system primarily uses objective and straightforward descriptive vocabulary to convey only core semantic information. Then, it analyzes sentence complexity parameters. If the parameter is marked as "simple," the generated text mainly consists of short sentences with simple subject-verb-object structures, avoiding complex clauses and multiple layers of modifiers, with sentence length controlled within a preset short range. If the parameter is marked as "medium," it combines short sentences with structurally simple long sentences, allowing a small number of simple clauses, with sentence length controlled within a preset medium range. If the parameter is marked as "complex," it can appropriately use multiple layers of modifiers and complex clauses to enrich the sentence structure, and sentence length can be appropriately extended according to semantic expression needs. Through these parameter controls, the emotional intensity and linguistic complexity of the generated text are precisely matched to the user's current emotional state.

[0079] Step 558: Based on the dialogue management process and natural language generation process, a voice response consistent with the emotional state is formed to achieve an emotion-adaptive response. Specifically, this includes: integrating response timing and dialogue guidance methods, as well as text content adapted to the intensity of emotional expression and linguistic complexity; clarifying the core elements of the voice response; inputting the integrated text content into the emotional speech synthesis process; simultaneously calling the prosody, pitch, and rhythm parameters corresponding to the comprehensive emotional state representation obtained through previous cross-modal fusion; and initiating speech synthesis at a specified time point according to the response timing requirements: if the response timing is early, synthesis is performed immediately after text generation; if it is normal or delayed, synthesis is performed according to the corresponding time point. During the synthesis process, the emotional concentration of the voice is adjusted based on the emotional expression intensity of the text, and the pause position and duration of the voice are adjusted based on the linguistic complexity. Finally, the synthesized voice is transmitted to the user through a voice output device. This voice not only conforms to the response timing and guidance logic required by the dialogue strategy, but also highly matches the user's current emotional state in terms of text content, emotional intensity, and linguistic rhythm, thereby achieving an emotion-adaptive response.

[0080] By calculating the fluctuation parameters of the two emotion change patterns separately and combining them with semantic information for calibration, the determination of the emotional state fluctuation range is more comprehensive and accurate, thereby improving the ability to capture the user's true emotional state.

[0081] like Figure 2 As shown, embodiments of the present invention also provide an emotion-adaptive multimodal dialogue generation system based on biosignals, comprising:

[0082] The extraction module is used to process the collected biological signals and multimodal interaction information to extract video spatiotemporal features, audio temporal features, and text semantic features.

[0083] The fusion module is used to perform cross-modal fusion of video spatiotemporal features, audio temporal features, and text semantic features. Through fusion semantic analysis, it obtains deep semantic connections between multimodal data and establishes cross-modal semantic correspondences. The sentiment consistency assessment step is used to analyze the degree of consistency in sentiment expression across different modalities to obtain a comprehensive emotional state representation.

[0084] The generation module is used to obtain dialogue text content that matches the user's current emotional state based on the comprehensive emotional state representation and the contextual information of the multi-turn dialogue, and then convert it into speech output through a speech conversion process.

[0085] The recognition module is used to monitor the dynamic changes of the user's biosignals and speech signals in real time during the speech output process, organize the newly acquired multimodal data into a time-series data stream, and construct a benchmark for emotion state analysis; by detecting key turning points and extreme points in the time-series data stream, it extracts key event points that represent the boundaries of the emotion state and identifies two emotion change patterns.

[0086] The adjustment module analyzes two emotion change patterns, determines the fluctuation range of the emotion state by combining the results of fusion semantic analysis, and constructs the emotion state change trajectory by selecting reference instances of the emotion state. Based on the emotion state change trajectory, it generates dynamic adjustment parameters for the emotion state, and adjusts the dialogue strategy and generated content in real time to achieve adaptive emotion response.

[0087] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.

[0088] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. An emotion adaptive multi-modal dialogue generation method based on biological signals, characterized in that, The method includes: Step 1: Process the collected biological signals and multimodal interaction information to extract video spatiotemporal features, audio temporal features, and text semantic features; Step 2 involves cross-modal fusion of video spatiotemporal features, audio temporal features, and text semantic features. Through fusion semantic analysis, deep semantic connections between multimodal data are obtained, and cross-modal semantic correspondences are established. The emotional consistency assessment step is used to analyze the degree of consistency in emotional expression across different modalities, resulting in a comprehensive emotional state representation. Step 3: Based on the comprehensive emotional state representation and combined with the contextual information of the multi-turn dialogue, use a natural language generation model to obtain dialogue text content that matches the user's current emotional state, and convert it into speech output through a speech conversion process. Step 4: During the speech output process, monitor the dynamic changes of the user's biosignals and speech signals in real time, organize the newly acquired multimodal data into a time-series data stream, and construct a baseline for emotion state analysis. By detecting key turning points and extreme points in the time-series data stream, extract key event points representing the boundaries of the emotion state and identify two emotion change patterns. Specifically, this includes: dividing the time-series data stream into fixed-duration windows, calculating the time-domain and frequency-domain features of the biosignals within each window, and generating a high-dimensional feature sequence; performing first-order difference calculation on the high-dimensional feature sequence to obtain a feature change rate sequence, and detecting feature changes. Zero-crossing points and local extrema in the rate sequence are used to identify candidate key event points. These candidate key event points are then filtered using preset amplitude and duration thresholds, merging closely adjacent candidate key event points to form a key event point sequence. Based on this sequence, the time interval and characteristic amplitude difference between adjacent key event points are calculated. The time interval and characteristic amplitude difference are compared with preset thresholds to distinguish between abrupt and gradual change patterns in emotional states. Abrupt change pattern is defined as a short time interval with a large amplitude change, while a gradual change pattern is defined as a uniform time interval with a small amplitude change. Step 5: Based on the mutation and gradual change modes, calculate the time span and characteristic amplitude range of emotional state fluctuations for each mode to form preliminary fluctuation interval parameters; align the preliminary fluctuation interval parameters with the emotional semantic information obtained from cross-modal fusion semantic analysis, and use the emotional semantic information to correct and calibrate the boundaries of the fluctuation intervals to determine the final emotional state fluctuation intervals; based on the emotional state fluctuation intervals, retrieve samples from the emotional state instance library that match the emotional features and emotional semantic labels within the current fluctuation interval, and select emotional state reference instances; organize the emotional state reference instances in chronological order, and combine them with the determined emotional state fluctuation intervals to construct an emotional state change trajectory reflecting the continuous change process of emotional states; [Further details on the emotional state fluctuation intervals are needed for accurate translation.] The system analyzes the trajectory of emotional state changes, extracts the rate of change in emotional intensity and the pattern of emotional type transition, and forms emotional evolution trend characteristics. Based on these characteristics and contextual information from multiple rounds of dialogue, dialogue strategy parameters and content generation parameters are generated through a pre-defined mapping relationship. The dialogue strategy parameters include response timing and the degree of dialogue initiative, while the content generation parameters include emotional vocabulary density and sentence complexity. The dialogue strategy parameters are applied to the dialogue management process to determine the timing of responses and the guidance method for dialogue. The content generation parameters are input into the natural language generation process to control the emotional expression intensity and linguistic complexity of the generated text. Based on the dialogue management process and the natural language generation process, a voice response consistent with the emotional state is generated, achieving an adaptive emotional response. 2.The emotion-adaptive multi-modal dialogue generation method based on biosignals according to claim 1, wherein, This approach involves cross-modal fusion of video spatiotemporal features, audio temporal features, and text semantic features. Through fused semantic analysis, deep semantic connections between multimodal data are obtained, and cross-modal semantic correspondences are established, including: Cross-modal semantic alignment is performed on video spatiotemporal features, audio temporal features, and text semantic features. The interaction and correlation strength between different modal features is analyzed through a multimodal attention process. Based on the interaction correlation strength, the features of each modality are fused to generate a semantically aligned cross-modal joint representation; The cross-modal joint representation based on semantic alignment constructs a semantic association matrix between modalities by calculating the semantic similarity between feature fragments of different modalities. The semantic association matrix is ​​analyzed to identify cross-modal semantic association patterns and establish structured cross-modal semantic correspondences. 3.The emotion-adaptive multi-modal dialogue generation method based on biosignals according to claim 2, characterized in that, The consistency of emotion expression across different modalities is analyzed using an affective consistency assessment process to obtain a comprehensive representation of emotional state, including: Based on cross-modal semantic correspondence, corresponding visual sentiment features, acoustic sentiment features, and semantic sentiment features are extracted from video modality, audio modality, and text modality; Visual emotional features, acoustic emotional features, and semantic emotional features are discriminated to form emotional tendency descriptions for video modality, audio modality, and text modality. The sentiment tendency descriptions of video, audio, and text modalities are cross-compared to evaluate the intrinsic consistency among the sentiment tendency descriptions of different modalities. Based on the intrinsic consistency, a corresponding fusion contribution is determined for the sentiment tendency description of each modality. Based on the fusion contribution, the emotional tendency descriptions of each modality are proportionally fused to form a preliminary comprehensive emotional representation. The preliminary comprehensive emotional representation is then normalized to obtain the final comprehensive emotional state representation. 4.The emotion-adaptive multi-modal dialogue generation method based on biosignals according to claim 3, characterized in that, Based on the comprehensive emotional state representation and combined with the contextual information from multiple rounds of dialogue, a natural language generation model is used to obtain dialogue text content that matches the user's current emotional state. This text is then converted into speech output through a speech-to-speech process, including: By fusing a comprehensive emotional state representation with a vectorized, multi-turn dialogue context, a contextual semantic representation is generated. Inputting the contextual semantic representation into a pre-trained natural language generation model generates dialogue text that matches the current emotional state in terms of semantic content and expression. The dialogue text and the comprehensive emotional state representation are input into the emotional speech synthesizer, and the corresponding prosody, pitch and rhythm parameters are mapped out according to the comprehensive emotional state. Based on prosody, pitch, and rhythm parameters, the dialogue text is processed by speech synthesis to obtain natural speech output that carries corresponding emotional expression.

5. The emotion-adaptive multimodal dialogue generation method based on biosignals according to claim 4, characterized in that, During voice output, the dynamic changes of the user's biosignals and voice signals are monitored in real time. The newly acquired multimodal data is organized into a time-series data stream to construct a baseline for emotion state analysis, including: The raw biological signals and speech signals collected continuously are extracted using a sliding time window of preset duration to obtain biological signal segments and speech signal segments within the current time window; Feature extraction is performed on biosignal and speech signal segments within the current time window to obtain real-time feature vectors representing the user's current physiological and speech states; The real-time feature vectors are concatenated and aligned with the preset reference feature sequence to form a unified and continuous time-series data stream; Aggregate analysis is performed on the reference feature vectors corresponding to the user's calm state interval in the time series data stream, calculate the central trend of the numerical distribution of the reference feature vectors, and construct the emotional state baseline, i.e., the emotional state analysis baseline.

6. An emotion-adaptive multi-modal dialogue generation system based on bio-signals, the system implementing the method of any one of claims 1 to 5, characterized in that, include: The extraction module is used to process the collected biological signals and multimodal interaction information to extract video spatiotemporal features, audio temporal features, and text semantic features. The fusion module is used to perform cross-modal fusion of video spatiotemporal features, audio temporal features, and text semantic features. Through fusion semantic analysis, it obtains deep semantic connections between multimodal data and establishes cross-modal semantic correspondences. The sentiment consistency assessment step is used to analyze the degree of consistency in sentiment expression across different modalities to obtain a comprehensive emotional state representation. The generation module is used to obtain dialogue text content that matches the user's current emotional state based on the comprehensive emotional state representation and the contextual information of the multi-turn dialogue, and then convert it into speech output through a speech conversion process. The recognition module is used to monitor the dynamic changes of the user's biosignals and speech signals in real time during the speech output process, organize the newly acquired multimodal data into a time-series data stream, and construct a benchmark for emotion state analysis; by detecting key turning points and extreme points in the time-series data stream, it extracts key event points that represent the boundaries of the emotion state and identifies two emotion change patterns. The adjusting module is used for analyzing two emotional change modes, determining a fluctuation interval of the emotional state in combination with a result of the semantic analysis, and selecting an emotional state reference instance to construct an emotional state change track; and generates an emotional state dynamic adjustment parameter based on the emotional state change track to adjust the dialogue strategy and the generated content in real time, so as to realize the emotional adaptive response.

Citation Information

Patent Citations

  • Real-time emotion perception and voice interaction system for intelligent cockpit

    CN121009400A