A method and system for intelligent interaction of a virtual mental health scenario for adolescents
By collecting multimodal data in virtual mental health interaction scenarios, forming a baseline feature library of cooperative interactions and identifying adversarial interaction features, the problem of difficulty in monitoring dynamic mental states in existing technologies is solved, and stable assessment and adaptive feedback of mental health risks in adolescents are achieved, thereby improving the effectiveness of intervention.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU ZHUODUN INFORMATION TECH CO LTD
- Filing Date
- 2026-05-26
- Publication Date
- 2026-06-26
Smart Images

Figure CN122290898A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent interaction technology, and more specifically, to an intelligent interaction method and system for a virtual mental health scenario for teenagers. Background Technology
[0002] Current interventions for adolescent mental health typically employ face-to-face counseling or standardized online interactive Q&A for psychological state assessment and intervention guidance. These methods rely heavily on subjective human judgment and static questionnaire data, lacking continuous monitoring of dynamic psychological changes in adolescent users during interaction. They also struggle to effectively identify uncooperative or confrontational behavioral patterns that emerge during interaction, making it difficult to adjust and optimize intervention strategies in a timely manner to meet individual psychological needs, thus impacting the effectiveness of the intervention. Summary of the Invention
[0003] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide an intelligent interaction method and system for a virtual mental health scenario for adolescents to solve the problems mentioned in the background art.
[0004] To achieve the above objectives, the present invention provides the following technical solution: A method for intelligent interaction in a virtual mental health scenario for teenagers includes the following steps: S1: Collect multimodal interaction data generated in virtual mental health interaction scenarios to obtain multimodal interaction data sequences; S2: Perform cluster analysis and statistical behavior profiling based on multimodal interaction data sequences to extract cooperative interaction patterns and form a cooperative interaction benchmark feature library; S3: Based on the deviation of the multimodal interaction data sequence from the cooperative interaction benchmark feature library, extract adversarial interaction candidate features to form an adversarial interaction feature set; S4: Based on the cooperative interaction benchmark feature library and the adversarial interaction feature set, evaluate the time series consistency and cross-modal consistency of multimodal psychological cues, and calculate the multimodal interaction credibility weight distribution; S5: Based on the multimodal interaction data sequence and the multimodal interaction credibility weight distribution, an adversarial robust feature fusion method is used to generate psychological risk assessment features and obtain the psychological risk assessment results for adolescents; S6: Construct an adaptive virtual dialogue strategy decision-making model based on the results of adolescent psychological risk assessment, and output virtual mental health interactive feedback information that addresses the real psychological needs of adolescents.
[0005] In a preferred embodiment, S1 specifically refers to: Collect data on text dialogues, voice and audio signals, facial expression images, and interaction timing rhythm generated by adolescent users in virtual mental health interactive scenarios; Extract semantic features from text dialogue; Extracting speech features such as pitch, speech rate, and emotion from speech audio signals; Facial motion unit features are extracted from facial expression images; Extract features of interaction pause duration and response interval from interaction time rhythm data; By summarizing semantic features, speech features, facial expression features, and rhythmic features, a multimodal interaction data sequence is obtained.
[0006] In a preferred embodiment, S2 specifically refers to: Based on multimodal interaction data sequences, cluster analysis was used to classify semantic features, speech features, facial expression features, and rhythm features into patterns, and to determine the category attribution of adolescent user interaction behaviors in virtual mental health interaction scenarios. Based on the category classification results, the distribution characteristics of semantic features, speech features, facial expression features and rhythm features are processed to create behavioral profiles of typical interactive behavior patterns of adolescent users in virtual mental health interaction scenarios. The category classification results are correlated with behavioral profiles of typical interaction patterns to form a baseline feature library for cooperative interaction.
[0007] In a preferred embodiment, S3 specifically refers to: Based on the multimodal interaction data sequence and the cooperative interaction benchmark feature library, semantic deviation, speech deviation, facial expression deviation and rhythm deviation are calculated; Extract candidate features for language inverse representation based on semantic deviation; Extract emotional speech inverse candidate features from speech features based on speech deviation; Based on the degree of facial expression deviation, candidate features for opposite emotional expressions are extracted from facial expression features; Extract abnormal rhythm candidate features from rhythm features based on rhythm deviation; Candidate features for reverse language expression, reverse emotional speech, reverse emotional facial expression, and abnormal rhythm are combined in chronological order to form candidate features for adversarial interaction, which are then categorized and summarized to form a set of adversarial interaction features.
[0008] In a preferred embodiment, S4 specifically refers to: Based on the cooperative interaction benchmark feature library and the adversarial interaction feature set, the temporal stability of semantic features, speech features, facial expression features and rhythm features in the continuous interaction process is evaluated, and the temporal series consistency index of each feature is calculated. Correlation analysis was performed on the cross-modal feature consistency among semantic features, speech features, facial expression features, and rhythm features, and a consistency index was calculated between each group of cross-modal features. Feature credibility weights are calculated based on time series consistency index and cross-modal feature consistency index to obtain the multimodal interaction credibility weight distribution corresponding to semantic features, speech features, facial expression features and rhythm features.
[0009] In a preferred embodiment, S5 specifically refers to: The semantic features, speech features, facial expression features, and rhythm features are weighted according to the multimodal interaction credibility weight distribution to obtain weighted semantic features, weighted speech features, weighted facial expression features, and weighted rhythm features. The weighted semantic features, weighted speech features, weighted facial expression features, and weighted rhythm features are aligned, and feature denoising, feature concatenation, and feature mapping are performed sequentially to generate psychological risk assessment features. Input psychological risk assessment features into a pre-trained risk assessment model and output psychological risk assessment results for adolescents.
[0010] In a preferred embodiment, S6 specifically refers to: Based on the results of adolescent psychological risk assessment, parameters are configured for semantic expression strategies, voice expression strategies, facial expression strategies, and interaction rhythm control strategies in virtual psychological health interaction scenarios, generating a set of strategy parameters. Based on the set of strategy parameters, the text output content, voice output method, facial expression presentation method and interaction time rhythm in the virtual mental health interaction scenario are adjusted to construct an adaptive virtual dialogue strategy decision model; The text output content, voice output method, facial expression presentation method and interaction time rhythm of the adaptive virtual dialogue strategy decision model are combined to generate virtual mental health interactive feedback information.
[0011] On the other hand, the present invention provides an intelligent interactive system for a virtual mental health scenario for teenagers, comprising: Data acquisition module: Collects multimodal interaction data generated in virtual mental health interaction scenarios to obtain multimodal interaction data sequences; Benchmark database construction module: Based on multimodal interaction data sequences, cluster analysis and statistical behavior profiling are performed to extract cooperative interaction patterns and form a cooperative interaction benchmark feature library; Adversarial identification module: Based on the deviation of the multimodal interaction data sequence from the cooperative interaction benchmark feature library, it extracts adversarial interaction candidate features to form an adversarial interaction feature set; Credibility Assessment Module: Based on the cooperative interaction benchmark feature library and the adversarial interaction feature set, the module assesses the temporal and cross-modal consistency of multimodal psychological cues and calculates the multimodal interaction credibility weight distribution. Risk fusion module: Based on the multimodal interaction data sequence and the multimodal interaction credibility weight distribution, an adversarial robust feature fusion method is used to generate psychological risk assessment features and obtain the psychological risk assessment results for adolescents; Strategy generation module: Constructs an adaptive virtual dialogue strategy decision-making model based on the results of adolescent psychological risk assessment, and outputs virtual mental health interactive feedback information that addresses the real psychological needs of adolescents.
[0012] The technical effects and advantages of the intelligent interactive method and system for virtual mental health scenarios for teenagers proposed in this invention are as follows: By collecting multimodal interaction data in virtual mental health interaction scenarios and forming a baseline feature library of cooperative interactions based on cluster analysis and statistical behavioral profiling, the cooperative behavior patterns of adolescents in natural interaction processes can be stably represented. Based on the deviation of the multimodal interaction data sequence from the baseline feature library of cooperative interactions, an adversarial interaction feature set is extracted, which can identify adversarial and deceptive interaction behaviors of adolescent users in virtual scenarios from multiple dimensions such as language, voice, facial expressions, and rhythm. The credibility weight distribution of multimodal interactions is calculated using time series consistency and cross-modal consistency assessment, allowing for differentiated quantitative expression of the credibility of different modalities in psychological risk assessment. Combining the multimodal interaction data sequence and the credibility weight distribution of multimodal interactions, an adversarial robust feature fusion method is used to generate psychological risk assessment features, which is beneficial for obtaining stable and reliable adolescent psychological risk assessment results under the condition of adversarial interaction interference. Finally, based on the adolescent psychological risk assessment results, an adaptive virtual dialogue strategy decision model is constructed, which can dynamically adjust semantic expression, voice expression, facial expression presentation, and interaction rhythm, making the virtual mental health interaction feedback information more closely match the real psychological needs of adolescents and improving the pertinence and sustainability of virtual mental health intervention. Attached Figure Description
[0013] Figure 1 This is a schematic diagram of an intelligent interaction method for a virtual mental health scenario for teenagers according to the present invention; Figure 2 This is a schematic diagram of the structure of an intelligent interactive system for a virtual mental health scenario for teenagers, according to the present invention. Detailed Implementation
[0014] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0015] Example 1: Figure 1 This invention provides an intelligent interaction method for a virtual mental health scenario for teenagers, which includes the following steps: S1: Collect multimodal interaction data generated in virtual mental health interaction scenarios to obtain multimodal interaction data sequences; S2: Perform cluster analysis and statistical behavior profiling based on multimodal interaction data sequences to extract cooperative interaction patterns and form a cooperative interaction benchmark feature library; S3: Based on the deviation of the multimodal interaction data sequence from the cooperative interaction benchmark feature library, extract adversarial interaction candidate features to form an adversarial interaction feature set; S4: Based on the cooperative interaction benchmark feature library and the adversarial interaction feature set, evaluate the time series consistency and cross-modal consistency of multimodal psychological cues, and calculate the multimodal interaction credibility weight distribution; S5: Based on the multimodal interaction data sequence and the multimodal interaction credibility weight distribution, an adversarial robust feature fusion method is used to generate psychological risk assessment features and obtain the psychological risk assessment results for adolescents; S6: Construct an adaptive virtual dialogue strategy decision-making model based on the results of adolescent psychological risk assessment, and output virtual mental health interactive feedback information that addresses the real psychological needs of adolescents.
[0016] S1: Collect multimodal interaction data generated in virtual mental health interaction scenarios to obtain multimodal interaction data sequences, including: Collect data on text dialogues, voice and audio signals, facial expression images, and interaction timing and rhythm generated by adolescent users in virtual mental health interactive scenarios; The virtual mental health interaction scenario continuously interacts with adolescent users through a psychological intervention dialogue program. During the interaction, each round of the virtual mental health interaction scenario presents the adolescent users with problem prompts or guidance, and receives text input, voice input, and visible facial images from the adolescent users. Text dialogue is collected by logging the dialogue window used for text input in the virtual mental health interaction scenario. Each time an adolescent user submits a text message, the text string content and corresponding timestamp information are recorded. Voice audio signals are collected by continuously sampling the microphone input channel bound to the virtual mental health interaction scenario. For example, the voice sampling frequency can be set to 16000 Hz or 44100 Hz, and the voice sampling bit depth can be set to 16 bits. The voice audio signals are buffered in chronological order and bound to corresponding timestamps. Facial expression image collection is achieved by enabling the camera video capture function in the virtual mental health interaction scenario. The video stream captured by the camera extracts keyframe images at fixed time intervals, such as 5 frames per second or 10 frames per second. Each frame of facial expression image is bound to the corresponding frame's timestamp. Interaction time rhythm data collection is obtained by recording the time sequence of various interaction events in the virtual mental health interaction scenario. Interaction events include the time point when the virtual mental health interaction scenario outputs prompt information, the time point when the adolescent user starts inputting text dialogue, the time point when the adolescent user ends inputting text dialogue and sends a message, the time point when the adolescent user starts and ends voice input, and the time point when the camera is turned on or off. All interaction events are recorded as timestamps with a unified time reference.
[0017] Extract semantic features from text dialogue; Each text dialogue undergoes sentence segmentation and word segmentation. Sentence segmentation is based on periods, question marks, exclamation marks, and line breaks. Word segmentation employs dictionary-based or statistical-based algorithms to break the text string into word sequences. After obtaining the word sequences, lists of emotion-related words, psychological state-related words, and keywords related to self-evaluation and peer evaluation are constructed. By statistically analyzing the frequency of emotion-related words, psychological state-related words, and self-evaluation / peer evaluation words in the word sequences, frequency statistical features reflecting the text's emotional tendency and psychological focus are generated. Simultaneously, based on a pre-trained word vector model or sentence vector model, the text dialogue is mapped to fixed-dimensional vectors, with dimensions such as 128, 256, or higher. The semantic features corresponding to the text dialogue are then combined with these frequency statistical features and vectors.
[0018] Extracting speech features such as pitch, speech rate, and emotion from speech audio signals; The speech audio signal is denoised, for example, by using frequency domain filtering to remove environmental noise. The denoised speech audio signal is then divided into frames according to a fixed duration window, for example, the frame length can be set to 20 milliseconds to 40 milliseconds. The pitch, speech rate, and emotional speech features of each frame of the speech audio signal are then calculated. The pitch features are obtained using a fundamental frequency extraction algorithm, such as the YIN algorithm. The speech rate features are determined by calculating the average speech rate per second, i.e., the number of syllables per second. The emotional speech features are obtained by extracting Mel-frequency cepstral coefficients, for example, setting the dimension of the Mel-frequency cepstral coefficients to 13 dimensions. Finally, the pitch features, speech rate features, and emotional speech features corresponding to each frame of the speech audio signal are concatenated to form the speech features.
[0019] Facial motion unit features are extracted from facial expression images; In facial expression images, a cascaded classifier-based face detection method is used to crop and geometrically normalize the face regions, aligning them to a uniform scale and angle. Facial key points (FKs) are then detected within the aligned face regions, including the corners of the eyes, brow peaks, nostrils, and corners of the mouth. The relative distances, angles, and displacements between these FKs are calculated to obtain the intensity of facial action units (FAUs) related to facial expressions. FKs include actions such as raising and lowering eyebrows, closing eyelids, raising and lowering the corners of the mouth. For each frame of the facial expression image, the intensity of each FK is calculated, and the values are normalized to the 0-1 range using training samples. The combined intensity of all FKs forms the facial action unit features corresponding to the facial expression image, which are then bound to the timestamp of the facial expression image to constitute the expression features.
[0020] Extract features of interaction pause duration and response interval from interaction time rhythm data; For interaction timing data, the pause duration and response interval need to be calculated based on the timestamps of interaction events recorded in the virtual mental health interaction scenario. The pause duration is defined as the time difference between the output of a prompt or question in the virtual mental health interaction scenario and the start of corresponding text dialogue or voice input by the adolescent user. The response interval is defined as the time interval between two consecutive text dialogues or between two voice inputs by the adolescent user. These calculations form the rhythm characteristics.
[0021] By summarizing semantic features, speech features, facial expression features, and rhythmic features, a multimodal interaction data sequence is obtained; Semantic features, speech features, facial expression features, and rhythm features are aligned according to a unified time dimension, that is, they are synchronized on the time axis by interpolation methods, such as linear interpolation. Semantic features, speech features, facial expression features, and rhythm features are then concatenated to form a multimodal interactive data sequence.
[0022] S2: Based on the multimodal interaction data sequences, cluster analysis and statistical behavioral profiling are performed to extract cooperative interaction patterns, forming a cooperative interaction baseline feature library, including: Based on multimodal interaction data sequences, cluster analysis was used to classify semantic features, speech features, facial expression features, and rhythm features into patterns, and to determine the category attribution of adolescent user interaction behaviors in virtual mental health interaction scenarios. The clustering analysis method employs the K-means clustering algorithm to extract semantic, speech, facial expression, and rhythmic features from multimodal interaction data sequences, forming a high-dimensional feature vector. This high-dimensional feature vector is then normalized, for example, using a min-max normalization method to normalize the value range of each feature dimension to the interval between 0 and 1. The silhouette coefficient method is used to determine the number of clusters, K. Specifically, for K within multiple integer ranges (e.g., 2 to 10), cluster analysis is performed, the corresponding silhouette coefficients are calculated, and the cluster with the largest silhouette coefficient is selected as the number of clusters, K. For example, let's set it to 5. Then, randomly initialize the normalized high-dimensional feature vectors into K cluster centers, and calculate the Euclidean distance between the high-dimensional feature vectors and the cluster centers. Assign each high-dimensional feature vector to the category corresponding to the nearest cluster center. Then, based on all high-dimensional feature vectors within each category, recalculate new cluster centers, i.e., calculate the mean of the high-dimensional feature vectors within each category as the new cluster centers. Repeat the cluster division and cluster center update until the change in cluster center position is less than a preset threshold. The preset threshold for the change in cluster center position is determined through multiple experiments, for example, it can be set to less than or equal to 0.001. When the cluster division is completed, the final K cluster centers are obtained, and the category assignment corresponding to each high-dimensional feature vector is determined, i.e., the category assignment of adolescent user interaction behavior in the virtual mental health interaction scenario is determined.
[0023] Based on the category classification results, the distribution characteristics of semantic features, speech features, facial expression features and rhythm features are processed to create behavioral profiles of typical interactive behavior patterns of adolescent users in virtual mental health interaction scenarios. Statistical distribution parameters, including mean, variance, median, and quartiles, are calculated for all semantic, speech, facial expression, and rhythmic features within a category to describe the typical feature distribution of each category. For example, for semantic features within a category, the statistical mean and variance of the frequency of emotion-related words and the frequency of self-evaluation or others' evaluation words are calculated; for speech features within a category, the statistical mean and standard deviation of pitch, speech rate, and emotional speech features are calculated; for facial expression features within a category, the mean and variance of the facial motor unit intensity distribution are calculated; for rhythmic features within a category, the mean and variance of the interaction pause duration and response interval are calculated. Through statistical calculations, a behavioral profile is formed for each category. For example, the behavioral profile of a typical interaction behavior pattern for a category is characterized by a combination of typical features: high frequency of emotion-related words, low pitch and slow speech rate, facial motor unit intensity manifested as lowered eyebrows and downturned corners of the mouth, relatively long interaction pause duration, and long response interval.
[0024] The category classification results are correlated with the behavioral profiles of typical interaction behavior patterns to form a baseline feature library for cooperative interaction. Based on the categorization results, category labels are determined. Each clustering result category is assigned a unique identifier. A mapping relationship is established between each category label and the behavioral profiles of typical interaction patterns within that category. Using the category labels as indexes and the behavioral profiles of typical interaction patterns as data content, a baseline feature library for cooperative interaction is constructed. For example, when the category label is Category 1, the corresponding behavioral profile includes semantic features such as an emotional word frequency of 0.7, a self-evaluation word frequency of 0.5, a mean speech pitch of 150 Hz with a standard deviation of 10 Hz, a mean speech rate of 3 syllables per second, a mean Mel frequency cepstral coefficient of a specific 13-dimensional vector, a mean facial action unit intensity of 0.3 for eyebrow raising and 0.6 for mouth corner raising, an average interaction pause duration of 3 seconds, and an average response interval of 2 seconds. Through this method, behavioral profile data for typical interaction patterns corresponding to each category are generated and aggregated to form the baseline feature library for cooperative interaction.
[0025] S3: Based on the deviation of the multimodal interaction data sequence from the cooperative interaction baseline feature library, extract adversarial interaction candidate features to form an adversarial interaction feature set, including: Based on the multimodal interaction data sequence and the cooperative interaction benchmark feature library, semantic deviation, speech deviation, facial expression deviation and rhythm deviation are calculated; The semantic deviation is calculated by comparing the semantic features in the multimodal interaction data sequence with the corresponding semantic features in the cooperative interaction benchmark feature library. The semantic features in the multimodal interaction data sequence are represented as vectors of fixed dimensions, for example, a vector dimension of 256. The semantic features of each category in the cooperative interaction benchmark feature library are also represented as vectors of the same dimension; for example, the semantic feature vectors in the behavioral profiles of typical interaction behavior patterns also have a dimension of 256. Then, the cosine similarity between the semantic feature vector of the current interaction data and the semantic feature vectors of each category of typical interaction behavior patterns is calculated. Cosine similarity is used as a quantitative indicator to measure the difference between semantic features. The semantic deviation is obtained by inverting the cosine similarity or subtracting 1 from the cosine similarity. When the cosine similarity is close to 1, the difference is small; when the cosine similarity is close to 0, the difference is large.
[0026] The speech deviation is calculated by comparing the speech features in the multimodal interaction data sequence with the corresponding speech features in the cooperative interaction benchmark feature library. The absolute difference between the pitch features, speech rate features, and emotional speech features of each frame of speech signal and the corresponding features in the cooperative interaction benchmark feature library is calculated, and the average of the absolute differences of each frame is used to form a unified speech deviation index. For example, if the average pitch feature of a typical interaction behavior pattern of a certain category in the cooperative interaction benchmark feature library is 150 Hz, the average speech rate feature is 3 syllables per second, and the average Mel-frequency cepstral coefficient of the emotional speech feature is a specific 13-dimensional vector, then the absolute difference between the pitch of the current speech feature and 150 Hz, the absolute difference between the current speech rate and 3 syllables per second, and the Euclidean distance between the Mel-frequency cepstral coefficient of the current emotional speech feature and the specific 13-dimensional vector are calculated. The data is then standardized and a weighted average is taken to obtain the speech deviation. The weight coefficients of each feature difference value are determined based on the statistical analysis of adversarial interaction behavior in historical data. For example, the weight coefficient of tone is set to 0.3, the weight coefficient of speech rate is set to 0.3, and the weight coefficient of emotional speech features is set to 0.4.
[0027] The difference between facial expression features in multimodal interaction data sequences and corresponding facial expression features in the cooperative interaction benchmark feature library is calculated to obtain the expression deviation: Euclidean distance is used to measure the difference between the current facial action unit intensity vector and the facial action unit intensity vectors of typical interaction behavior patterns in the cooperative interaction benchmark feature library. For example, facial action unit vectors include the intensity of action units such as raising eyebrows, lowering eyebrows, closing eyelids, raising corners of the mouth, and lowering corners of the mouth. The Euclidean distance is calculated as the expression deviation.
[0028] The rhythm deviation is calculated by comparing the rhythm features in the multimodal interaction data sequence with the corresponding rhythm features in the cooperative interaction benchmark feature library. The absolute differences between the current interaction pause duration and response interval and the mean values of the interaction pause duration and response interval for the corresponding typical interaction behavior patterns in the cooperative interaction benchmark feature library are calculated respectively. The two absolute differences are then normalized and weighted to form the rhythm deviation. For example, the weight coefficient for the interaction pause duration difference can be set to 0.5 and the weight coefficient for the response interval difference can be set to 0.5 to obtain the rhythm deviation.
[0029] Extract candidate features for language inverse representation based on semantic deviation; A pre-trained text sentiment classifier is used to identify the sentiment polarity of each text segment, that is, to determine whether the sentiment tendency of the text segment is positive, neutral, or negative. The text sentiment classifier can adopt a convolutional neural network model, with network parameters such as 3 convolutional layers, 1 pooling layer, and 2 fully connected layers, the activation function is the ReLU function, and the output layer uses the Softmax function to perform a three-class classification task, outputting the probability distributions of positive, neutral, and negative, respectively. The input of the convolutional neural network model is the semantic feature vector corresponding to the text segment, and the output is the sentiment polarity category with the highest probability. The sentiment polarity classification results are compared with the sentiment polarity in the behavioral profiles of typical interaction behavior patterns in the cooperative interaction benchmark feature library. Text segments with sentiment classification results that are different from the sentiment polarity categories in the behavioral profiles, that is, text segments that show the opposite sentiment polarity, are selected as candidate features for language reverse expression.
[0030] Extract emotional speech inverse candidate features from speech features based on speech deviation; Based on the pitch features, speech rate features, and emotional speech features corresponding to each speech segment, the statistical mean of the corresponding speech features in the cooperative interaction benchmark feature library is compared with the statistical mean of the pitch features. For the pitch features of all speech segments in the cooperative interaction benchmark feature library, the statistical mean and standard deviation of the pitch features are calculated. For the speech rate features of all speech segments in the cooperative interaction benchmark feature library, the statistical mean and standard deviation of the speech rate features are calculated. For the emotional speech feature vectors of all speech segments in the cooperative interaction benchmark feature library, the statistical mean and standard deviation of each feature dimension of the emotional speech feature vector are calculated. For the current speech segment to be detected, calculate the difference between the pitch feature of the speech segment and the statistical mean of the pitch feature in the cooperative interaction benchmark feature library, and record the absolute value of the difference as the pitch deviation; calculate the difference between the speech rate feature of the speech segment and the statistical mean of the speech rate feature in the cooperative interaction benchmark feature library, and record the absolute value of the difference as the speech rate deviation; calculate the difference between the feature value of the emotional speech feature vector of the speech segment in each feature dimension and the statistical mean of the corresponding feature dimension in the cooperative interaction benchmark feature library, and record the absolute value of the difference in each feature dimension as the deviation of the emotional speech in each dimension. The deviation thresholds are determined based on the standard deviations in the cooperative interaction benchmark feature library. The pitch deviation threshold is set to the range outside the interval boundary formed by the sum of the statistical mean of the pitch features and 1.5 times the standard deviation of the pitch features. The speech rate deviation threshold is set to the range outside the interval boundary formed by the sum of the statistical mean of the speech rate features and 1.5 times the standard deviation of the speech rate features. The deviation thresholds for each dimension of the emotional speech features are set to the range outside the interval boundary formed by the sum of the statistical mean of the corresponding feature dimension and 1.5 times the standard deviation of the feature dimension. When the pitch deviation is greater than the pitch deviation threshold and the speech rate deviation is greater than the speech rate deviation threshold, or when the deviation of each dimension of the emotional speech feature vector is greater than the corresponding feature dimension deviation threshold, the speech segment is marked as an emotional speech reverse candidate feature.
[0031] Based on the degree of facial expression deviation, candidate features for opposite emotional expressions are extracted from facial expression features; Each frame of facial expression features in the multimodal interaction data sequence is analyzed, and the facial action unit intensity corresponding to the facial expression features of that frame is extracted. Then, it is compared with the facial action unit intensity contained in the facial expression features of the behavioral profile corresponding to the typical interaction behavior pattern in the cooperative interaction benchmark feature library. If the intensity of one or more facial action units of the facial expression features of that frame differs from the intensity of the corresponding feature in the cooperative interaction benchmark feature library, for example, if the difference in facial action unit intensity exceeds the set action unit intensity difference threshold, then the corresponding expression frame is determined as a candidate feature of reverse emotion expression. The action unit intensity difference threshold is determined based on the normalization range of action unit intensity. For example, when the action unit intensity is normalized to the range of 0 to 1, the action unit intensity difference threshold can be set to 0.3.
[0032] Extract abnormal rhythm candidate features from rhythm features based on rhythm deviation; The pause duration and response interval of each interaction event in the multimodal interaction data sequence are calculated and compared with the statistical mean and standard deviation of the corresponding rhythm features in the cooperative interaction benchmark feature library. If the pause duration or response interval of a certain interaction event exceeds the range of ±2 times the statistical mean of the corresponding feature in the cooperative interaction benchmark feature library, the interaction segment corresponding to the interaction event is identified as an abnormal rhythm candidate feature.
[0033] The candidate features of reverse language expression, reverse emotional speech, reverse emotional expression, and abnormal rhythm are combined in chronological order to form adversarial interaction candidate features, which are then classified and summarized to form an adversarial interaction feature set. The adversarial interaction candidate feature sequence is formed by combining the timestamps bound to the candidate features of language inverse expression, emotion speech inverse candidate features, emotion inverse facial expression candidate features, and abnormal rhythm candidate features. Then, each candidate feature in the adversarial interaction candidate feature sequence is categorized and summarized according to multimodal feature categories, namely semantics, speech, facial expression, and rhythm, thus forming the adversarial interaction feature set.
[0034] S4: Based on the cooperative interaction benchmark feature library and the adversarial interaction feature set, assess the time-series consistency and cross-modal consistency of multimodal psychological cues, and calculate the multimodal interaction credibility weight distribution, including: Based on the cooperative interaction benchmark feature library and the adversarial interaction feature set, the temporal stability of semantic features, speech features, facial expression features and rhythm features in the continuous interaction process is evaluated, and the temporal series consistency index of each feature is calculated. The temporal series consistency index of semantic features is obtained by calculating the cosine similarity between the semantic feature vectors of multiple consecutive interaction events: the cosine similarity is calculated for the semantic feature vectors corresponding to each pair of consecutive interaction events; then the cosine similarity of multiple consecutive interaction events is arithmetically averaged to obtain the temporal series consistency index of semantic features.
[0035] A dynamic time warping algorithm is employed to calculate the temporal consistency index of speech features. The dynamic time warping algorithm calculates the dynamic time warping distance between the speech feature vectors in multiple consecutive interaction events and the speech feature vectors in the cooperative interaction benchmark feature library. Two speech feature vectors to be compared are denoted as Sequence 1 and Sequence 2, with lengths M and N respectively. The Euclidean distance between any two feature frames of the two sequences is used as the local distance to construct a dynamic time warping distance matrix of size M×N. Then, the cumulative distance of the optimal path from the start frame to the end frame of the sequence is calculated using dynamic programming. The cumulative distance is divided by the optimal path length to obtain the dynamic time warping distance. The dynamic time warping distance is then inverted or normalized by subtracting 1 from the dynamic time warping distance, and used as the temporal consistency index of the speech features.
[0036] The autocorrelation function analysis method is used to calculate the time series consistency index of facial expression features: the facial action unit intensity of each facial expression feature in multiple consecutive interaction events is used as input, and the sum of the facial action unit intensities of the facial expression feature vectors in each frame is used as a univariate sequence for autocorrelation function calculation: the mean of the sequence is calculated, and the sequence value at each time step is subtracted from the sequence mean to form a centered sequence; then, for each delay order, the sum of the products of the centered sequence and itself at the corresponding delay order is calculated, and divided by the sequence variance and the sequence length to obtain the autocorrelation coefficient of the corresponding delay order; the autocorrelation coefficient with a delay order of 1 is used as the time series consistency index of facial expression features.
[0037] The time series consistency index of rhythm features is calculated by taking the reciprocal of the standard deviation: the standard deviation of the pause duration and response interval of multiple consecutive interactive events is calculated separately, and the two standard deviations are weighted and averaged, and the reciprocal is taken as the time series consistency index of rhythm features; for example, if the standard deviations of the pause duration and response interval of 20 consecutive interactive events are T1 and T2 respectively, and the weight coefficients are each set to 0.5, then the time series consistency index of rhythm features is 1 / (0.5×T1+0.5×T2).
[0038] Correlation analysis was performed on the cross-modal feature consistency among semantic features, speech features, facial expression features, and rhythm features, and a consistency index was calculated between each group of cross-modal features. The Pearson correlation coefficient is calculated between semantic features and speech features, speech features and facial expression features, semantic features and facial expression features, rhythmic features and speech features, rhythmic features and facial expression features, and semantic features and rhythm. For multiple consecutive interactive events, each feature of the two cross-modal features to be analyzed is represented as a scalar sequence, forming two scalar sequences of equal length. Each scalar sequence is mean-centered. Then, the two mean-centered scalar sequences are multiplied one by one and summed, and divided by the product of the standard deviation and length of each scalar sequence to obtain the Pearson correlation coefficient. Pearson correlation coefficient; the Pearson correlation coefficient ranges from -1 to +1, with values close to ±1 indicating high correlation and values close to 0 indicating low correlation. For example, to calculate the consistency index between semantic features and speech features, the Pearson correlation coefficient is calculated using the frequency statistical feature sequence of emotion-related words in semantic features and the cepstral coefficient sequence of the first Vimel frequency in speech features of emotion. Similar methods are used to calculate the consistency index between speech features and facial expression features, between semantic features and facial expression features, between rhythm features and speech features, between rhythm features and facial expression features, and between semantic features and rhythm.
[0039] The feature credibility weights are calculated based on the time series consistency index and the cross-modal feature consistency index to obtain the multimodal interaction credibility weight distribution corresponding to semantic features, speech features, facial expression features and rhythm features; The time-series consistency index of each feature is normalized to a uniform range of 0 to 1. Then, the consistency index of each group of cross-modal features is also normalized to a range of 0 to 1. The average of the normalized time-series consistency index of each feature and the normalized cross-modal consistency index of the corresponding feature is used as the initial credibility weight score for each feature. All initial credibility weight scores are normalized so that the sum of the credibility weight distributions corresponding to semantic features, speech features, facial expression features, and rhythm features is 1, thus obtaining the multimodal interaction credibility weight distribution. For example, if the normalized time-series consistency index of the semantic feature is 0.8 and the average value of the normalized cross-modal consistency index is 0.7, then the initial credibility weight score of the semantic feature is (0.8 + 0.7) / 2 = 0.75. The credibility weight scores of speech features, facial expression features, and rhythm features are calculated similarly to obtain the multimodal interaction credibility weight distribution.
[0040] S5: Based on the multimodal interaction data sequence and the multimodal interaction credibility weight distribution, an adversarial robust feature fusion method is used to generate psychological risk assessment features, resulting in psychological risk assessment results for adolescents, including: The semantic features, speech features, facial expression features, and rhythm features are weighted according to the multimodal interaction credibility weight distribution to obtain weighted semantic features, weighted speech features, weighted facial expression features, and weighted rhythm features. The credibility weight distribution is derived from the multimodal interaction credibility weight distribution, where each feature corresponds to a unique weight, representing the importance of each feature in the comprehensive assessment of psychological risk. For example, if the credibility weight corresponding to semantic features is 0.4, the credibility weight corresponding to speech features is 0.3, the credibility weight corresponding to facial expression features is 0.2, and the credibility weight corresponding to rhythm features is 0.1, then the value of each dimension in the semantic feature vector is multiplied by 0.4, the pitch feature, speech rate feature, and emotional speech feature corresponding to each frame in the speech feature vector are multiplied by 0.3, the facial action unit intensity of each frame in the facial expression feature is multiplied by 0.2, and the interaction pause duration and response interval corresponding to each interaction event in the rhythm feature are multiplied by 0.1, ultimately yielding weighted semantic features, weighted speech features, weighted facial expression features, and weighted rhythm features, respectively.
[0041] The weighted semantic features, weighted speech features, weighted facial expression features, and weighted rhythm features are aligned, and feature denoising, feature concatenation, and feature mapping are performed sequentially to generate psychological risk assessment features. A unified time reference sequence is established, for example, using the timestamp sequence of speech feature vectors as a unified reference. Synchronization along the timeline is then performed using this unified time reference sequence as the target. A linear interpolation method is used; that is, when the timestamp of a sampling point for a certain feature does not completely coincide with the timestamp of the unified reference sequence, the feature value corresponding to the missing timestamp is calculated by linear interpolation between adjacent timestamps. For example, if two consecutive timestamps of a weighted semantic feature are t1 and t2, with corresponding feature values f1 and f2 respectively, and the unified reference sequence includes an intermediate timestamp t1.5, then the feature value at t1.5 is calculated using linear interpolation as f1 + (f2 - f1) * (t1.5 - t1) / (t2 - t1). After interpolating the weighted semantic features, weighted speech features, weighted facial expression features, and weighted rhythm features respectively, the time-aligned weighted semantic features, weighted speech features, weighted facial expression features, and weighted rhythm features are finally obtained.
[0042] After feature alignment, low-pass filtering is used to perform feature denoising on each weighted feature: a Butterworth low-pass filter with an adjustable cutoff frequency is used to filter each weighted feature dimension by dimension. The order and cutoff frequency of the Butterworth filter are determined by analyzing historical psychological risk assessment feature data and experiments on the filtering effect in different scenarios. For example, the filter order can be set to 2nd order, and the cutoff frequency is set by analyzing the spectral characteristics of the interactive data sequence and taking the frequency at which the cumulative energy distribution of the sequence spectrum reaches 95% as the cutoff frequency, for example, the cutoff frequency is set between 0.1 Hz and 0.5 Hz. After filtering, weighted semantic features, weighted speech features, weighted facial expression features, and weighted rhythm features are obtained after removing high-frequency fluctuation noise.
[0043] After noise reduction, the weighted features are concatenated to form a unified multimodal fusion feature sequence: the weighted semantic features, weighted speech features, weighted facial expression features, and weighted rhythm features corresponding to the same timestamp are concatenated to form a multimodal fusion feature sequence.
[0044] Feature mapping is performed on the multimodal fusion feature sequence after feature concatenation to generate psychological risk assessment features: A pre-trained autoencoder neural network model is constructed. The multimodal fusion feature sequence is input into the encoding part of the autoencoder for dimensionality reduction mapping, and the low-dimensional encoding vector output by the encoder is extracted as the psychological risk assessment features. The autoencoder neural network structure includes an input layer, multiple hidden layers, and an encoding layer. The number of nodes in the hidden layers decreases layer by layer. For example, three hidden layers are set with 128, 64, and 32 nodes respectively, and the number of nodes in the encoding layer is 16. The autoencoder is pre-trained using the mean squared error loss function and the Adam optimizer. During the training process, historical multimodal fusion feature sequences are input, and the reconstruction error is kept below a set threshold, for example, the error threshold is set to 0.01, to ensure that the encoder can extract key psychological risk assessment information. After training, the real-time generated multimodal fusion feature sequence is input, and the psychological risk assessment features are output after being mapped by the encoder.
[0045] Input psychological risk assessment features into a pre-trained risk assessment model and output psychological risk assessment results for adolescents; The risk assessment model is a supervised learning neural network model, such as a three-layer feedforward neural network, including an input layer, two hidden layers, and an output layer. The number of nodes in the hidden layers is set to 32 and 16 respectively. The output layer uses the Softmax activation function to output the probability distribution of psychological risk levels, such as low risk, medium risk, and high risk. Supervised training specifically involves collecting historical psychological risk level data annotated by experts, inputting the corresponding psychological risk assessment features, and training with a cross-entropy loss function. The learning rate is set to, for example, 0.001, and the Adam optimization algorithm is used as the optimizer. The training process iterates until the change in the loss function is less than a set threshold, for example, an error threshold of 0.005. After training is complete, real-time psychological risk assessment features are input into the risk assessment model, and the risk level with the highest probability output is the psychological risk assessment result for the adolescent.
[0046] S6: Construct an adaptive virtual dialogue strategy decision-making model based on the results of adolescent psychological risk assessment, and output virtual mental health interactive feedback information that addresses the real psychological needs of adolescents, including: Based on the results of adolescent psychological risk assessment, parameters are configured for semantic expression strategies, voice expression strategies, facial expression strategies, and interaction rhythm control strategies in virtual psychological health interaction scenarios, generating a set of strategy parameters. When the psychological risk assessment result is low risk, the parameter configuration method for the semantic expression strategy is to increase the proportion of positive feedback and encouraging words, for example, the proportion of positive feedback words is set to 80% to 90%, and the proportion of encouraging words is set to 10% to 20% of the total semantic expression words; the parameter configuration for the speech expression strategy is to increase the pitch features of the speech output, for example, the pitch features are set to 1.1 to 1.2 times the statistical mean of the corresponding pitch in the cooperative interaction baseline feature library, the speech rate features are increased to 5 to 6 syllables per second, and the Mel frequency cepstral coefficient of the emotional speech features is set close to the mean of the positive emotional speech features; the parameter configuration for the facial expression strategy is to enhance the intensity of positive facial action units, for example, the intensity of eyebrow raising and mouth corner raising action units is increased to the range of 0.6 to 0.8 respectively; the parameter configuration for the interaction rhythm control strategy is to reduce the duration of interaction pauses and response intervals, for example, the duration of interaction pauses is set to 1 to 1.5 seconds, and the response interval is set to less than 1 second.
[0047] When the psychological risk assessment result is medium risk, the parameter configuration method for the semantic expression strategy is to maintain the proportion of neutral expression content, for example, the proportion of neutral expression vocabulary is set to 60% to 70%, and the proportion of positive and guiding expression is set to 30% to 40%; the parameter configuration for the speech expression strategy is to use medium pitch features and speech rate features, for example, the pitch feature is set to 0.9 to 1.0 times the statistical mean of the pitch corresponding to the cooperative interaction baseline feature library, and the speech rate feature is set to about 4 syllables per second; the Mel frequency cepstral coefficient of the emotional speech feature is set to be near the mean of the neutral emotional speech feature; the parameter configuration for the facial expression presentation strategy is to use moderate intensity facial action units, for example, the intensity of the eyebrow raising action unit is set to 0.4 to 0.6, and the intensity of the corner of the mouth raising action unit is set to 0.3 to 0.5; the parameter configuration for the interaction rhythm control strategy is to set the interaction pause duration and response interval to a medium range, for example, the interaction pause duration is set to 2 to 3 seconds, and the response interval is set to 1.5 to 2.5 seconds.
[0048] When the psychological risk assessment result is high risk, the semantic expression strategy parameters are configured to enhance reassuring and empathetic expressions, for example, reassuring and empathetic vocabulary accounts for 70% to 80% of the overall semantic expression, while reducing the proportion of guiding or positive expressions to 20% to 30%; the speech expression strategy parameters are configured to reduce pitch and speech rate features, for example, the pitch feature is set to 0.7 to 0.8 times the statistical mean of the pitch corresponding to the cooperative interaction baseline feature library, and the speech rate feature is set to 2 to 3 syllables per second; the Mel frequency cepstral coefficient of the emotional speech feature is set to be close to the mean of the reassuring emotional speech feature; the facial expression presentation strategy parameters are configured to enhance empathetic facial action units, for example, the intensity of eyebrow drooping and eyelid closing action units is increased to 0.7 to 0.9; the interaction rhythm control strategy parameters are configured to appropriately extend the interaction pause duration and response interval, for example, the interaction pause duration is set to 3 to 4 seconds, and the response interval is set to 2 to 3 seconds.
[0049] By setting strategy parameters based on the psychological risk assessment results, and summarizing them to form a set of strategy parameters, including a subset of semantic expression strategy parameters, a subset of voice expression strategy parameters, a subset of facial expression strategy parameters, and a subset of interaction rhythm control strategy parameters.
[0050] Based on the set of strategy parameters, the text output content, voice output method, facial expression presentation method and interaction time rhythm in the virtual mental health interaction scenario are adjusted to construct an adaptive virtual dialogue strategy decision model; The adaptive virtual dialogue strategy decision-making model employs a supervised learning Long Short-Term Memory (LSTM) neural network. The LSM neural network structure includes an input layer, two LSM hidden layers, and an output layer. The number of nodes in the hidden layers is adjusted based on historical interaction data and model prediction performance; for example, the first hidden layer is set to 64 nodes, and the second hidden layer to 32 nodes. The output layer outputs text content feature vectors, speech output feature vectors, facial expression and action unit feature vectors, and interaction rhythm feature values. The model training process involves inputting a large amount of historical psychological interaction data and a set of policy parameters labeled with corresponding psychological risk assessment results. A hybrid loss function of cross-entropy and mean squared error is used, and the Adam optimizer is employed for model optimization training. The learning rate is adjusted incrementally; for example, it can be set to 0.003. The training iteration terminates when the model loss function converges to a preset threshold, such as a loss function value less than or equal to 0.01.
[0051] After model training, a set of policy parameters is input in real time, and the adaptive virtual dialogue policy decision model outputs text content feature vectors, speech output feature vectors, facial expression and action unit feature vectors, and interaction rhythm feature values. The text content feature vectors are then matched and combined with predefined semantic templates, for example, selecting the template with the closest feature vector similarity from a predefined semantic template library to form the final text output content. The speech output feature vectors are fed into a speech synthesis model, such as the Tacotron2 model, with parameters such as a 3-layer convolutional network plus a 1-layer recurrent neural network, to output synthesized speech signals. The facial expression and action unit feature vectors drive facial expression animations within the virtual interaction scene, adjusting the animation presentation in real time. The interaction rhythm feature values directly determine the interaction output intervals and pause durations in the virtual interaction scene, completing the comprehensive dynamic adjustment of multimodal feedback in the virtual mental health interaction scene.
[0052] The text output content, voice output method, facial expression presentation method and interaction time rhythm output by the adaptive virtual dialogue strategy decision model are combined to generate virtual mental health interactive feedback information. Feedback information is presented in a multimodal form through a virtual mental health interactive scenario interface, including a text window displaying adaptively adjusted text content, a speaker playing synthesized speech signals, a facial model presenting facial animations with adjusted facial expression units, and a predetermined interaction time rhythm controlling the presentation rhythm between each modality, thereby achieving accurate responses to changes in the psychological state of adolescent users.
[0053] Example 2: The difference between Example 2 and Example 1 is that this example introduces an intelligent interactive system for a virtual mental health scenario for teenagers.
[0054] Figure 2A schematic diagram of the structure of an intelligent interactive system for a virtual mental health scenario for teenagers is provided. The intelligent interactive system for a virtual mental health scenario for teenagers includes: Data acquisition module: Collects multimodal interaction data generated in virtual mental health interaction scenarios to obtain multimodal interaction data sequences; Benchmark database construction module: Based on multimodal interaction data sequences, cluster analysis and statistical behavior profiling are performed to extract cooperative interaction patterns and form a cooperative interaction benchmark feature library; Adversarial identification module: Based on the deviation of the multimodal interaction data sequence from the cooperative interaction benchmark feature library, it extracts adversarial interaction candidate features to form an adversarial interaction feature set; Credibility Assessment Module: Based on the cooperative interaction benchmark feature library and the adversarial interaction feature set, the module assesses the temporal and cross-modal consistency of multimodal psychological cues and calculates the multimodal interaction credibility weight distribution. Risk fusion module: Based on the multimodal interaction data sequence and the multimodal interaction credibility weight distribution, an adversarial robust feature fusion method is used to generate psychological risk assessment features and obtain the psychological risk assessment results for adolescents; Strategy generation module: Constructs an adaptive virtual dialogue strategy decision-making model based on the results of adolescent psychological risk assessment, and outputs virtual mental health interactive feedback information that addresses the real psychological needs of adolescents.
[0055] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0056] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0057] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0058] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0059] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0060] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.
[0061] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0062] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0063] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for intelligent interaction in a virtual mental health scenario for teenagers, characterized in that, Includes the following steps: S1: Collect multimodal interaction data generated in virtual mental health interaction scenarios to obtain multimodal interaction data sequences; S2: Perform cluster analysis and statistical behavior profiling based on multimodal interaction data sequences to extract cooperative interaction patterns and form a cooperative interaction benchmark feature library; S3: Based on the deviation of the multimodal interaction data sequence from the cooperative interaction benchmark feature library, extract adversarial interaction candidate features to form an adversarial interaction feature set; S4: Based on the cooperative interaction benchmark feature library and the adversarial interaction feature set, evaluate the time series consistency and cross-modal consistency of multimodal psychological cues, and calculate the multimodal interaction credibility weight distribution; S5: Based on the multimodal interaction data sequence and the multimodal interaction credibility weight distribution, an adversarial robust feature fusion method is used to generate psychological risk assessment features and obtain the psychological risk assessment results for adolescents; S6: Construct an adaptive virtual dialogue strategy decision-making model based on the results of adolescent psychological risk assessment, and output virtual mental health interactive feedback information that addresses the real psychological needs of adolescents.
2. The intelligent interaction method for a virtual mental health scenario for teenagers according to claim 1, characterized in that, S1, specifically: Collect data on text dialogues, voice and audio signals, facial expression images, and interaction timing rhythm generated by adolescent users in virtual mental health interactive scenarios; Extract semantic features from text dialogue; Extracting speech features such as pitch, speech rate, and emotion from speech audio signals; Facial motion unit features are extracted from facial expression images; Extract features of interaction pause duration and response interval from interaction time rhythm data; By summarizing semantic features, speech features, facial expression features, and rhythmic features, a multimodal interaction data sequence is obtained.
3. The intelligent interaction method for a virtual mental health scenario for teenagers according to claim 2, characterized in that, S2, specifically: Based on multimodal interaction data sequences, cluster analysis was used to classify semantic features, speech features, facial expression features, and rhythm features into patterns, and to determine the category attribution of adolescent user interaction behaviors in virtual mental health interaction scenarios. Based on the category classification results, the distribution characteristics of semantic features, speech features, facial expression features and rhythm features are processed to create behavioral profiles of typical interactive behavior patterns of adolescent users in virtual mental health interaction scenarios. The category classification results are correlated with behavioral profiles of typical interaction patterns to form a baseline feature library for cooperative interaction.
4. The intelligent interaction method for a virtual mental health scenario for teenagers according to claim 3, characterized in that, S3, specifically: Based on the multimodal interaction data sequence and the cooperative interaction benchmark feature library, semantic deviation, speech deviation, facial expression deviation and rhythm deviation are calculated; Extract candidate features for language inverse representation based on semantic deviation; Extract emotional speech inverse candidate features from speech features based on speech deviation; Based on the degree of facial expression deviation, candidate features for opposite emotional expressions are extracted from facial expression features; Extract abnormal rhythm candidate features from rhythm features based on rhythm deviation; Candidate features for reverse language expression, reverse emotional speech, reverse emotional facial expression, and abnormal rhythm are combined in chronological order to form candidate features for adversarial interaction, which are then categorized and summarized to form a set of adversarial interaction features.
5. The intelligent interaction method for a virtual mental health scenario for teenagers according to claim 4, characterized in that, S4, specifically: Based on the cooperative interaction benchmark feature library and the adversarial interaction feature set, the temporal stability of semantic features, speech features, facial expression features and rhythm features in the continuous interaction process is evaluated, and the temporal series consistency index of each feature is calculated. Correlation analysis was performed on the cross-modal feature consistency among semantic features, speech features, facial expression features, and rhythm features, and a consistency index was calculated between each group of cross-modal features. Feature credibility weights are calculated based on time series consistency index and cross-modal feature consistency index to obtain the multimodal interaction credibility weight distribution corresponding to semantic features, speech features, facial expression features and rhythm features.
6. The intelligent interaction method for a virtual mental health scenario for teenagers according to claim 5, characterized in that, S5, specifically: The semantic features, speech features, facial expression features, and rhythm features are weighted according to the multimodal interaction credibility weight distribution to obtain weighted semantic features, weighted speech features, weighted facial expression features, and weighted rhythm features. The weighted semantic features, weighted speech features, weighted facial expression features, and weighted rhythm features are aligned, and feature denoising, feature concatenation, and feature mapping are performed sequentially to generate psychological risk assessment features. Input psychological risk assessment features into a pre-trained risk assessment model and output psychological risk assessment results for adolescents.
7. The intelligent interaction method for a virtual mental health scenario for teenagers according to claim 6, characterized in that, S6, specifically: Based on the results of adolescent psychological risk assessment, parameters are configured for semantic expression strategies, voice expression strategies, facial expression strategies, and interaction rhythm control strategies in virtual psychological health interaction scenarios, generating a set of strategy parameters. Based on the set of strategy parameters, the text output content, voice output method, facial expression presentation method and interaction time rhythm in the virtual mental health interaction scenario are adjusted to construct an adaptive virtual dialogue strategy decision model; The text output content, voice output method, facial expression presentation method and interaction time rhythm of the adaptive virtual dialogue strategy decision model are combined to generate virtual mental health interactive feedback information.
8. A smart interactive system for a virtual mental health scenario for adolescents, used to implement the smart interactive method for a virtual mental health scenario for adolescents as described in any one of claims 1-7, characterized in that, include: Data acquisition module: Collects multimodal interaction data generated in virtual mental health interaction scenarios to obtain multimodal interaction data sequences; Benchmark database construction module: Based on multimodal interaction data sequences, cluster analysis and statistical behavior profiling are performed to extract cooperative interaction patterns and form a cooperative interaction benchmark feature library; Adversarial identification module: Based on the deviation of the multimodal interaction data sequence from the cooperative interaction benchmark feature library, it extracts adversarial interaction candidate features to form an adversarial interaction feature set; Credibility Assessment Module: Based on the cooperative interaction benchmark feature library and the adversarial interaction feature set, the module assesses the temporal and cross-modal consistency of multimodal psychological cues and calculates the multimodal interaction credibility weight distribution. Risk fusion module: Based on the multimodal interaction data sequence and the multimodal interaction credibility weight distribution, an adversarial robust feature fusion method is used to generate psychological risk assessment features and obtain the psychological risk assessment results for adolescents; Strategy generation module: Constructs an adaptive virtual dialogue strategy decision-making model based on the results of adolescent psychological risk assessment, and outputs virtual mental health interactive feedback information that addresses the real psychological needs of adolescents.