User psychological state monitoring method and system based on voice and semantic recognition
Through cross-modal joint coding network and dynamic weight adjustment technology, the problems of insufficient single modal analysis and lack of dynamic adjustment capabilities in the existing technology are solved, the comprehensive integration of multimodal information and accurate capture of emotional changes are achieved, and the accuracy and real-time nature of psychological state monitoring are improved.
Patent Information
- Application Number
- CN202510301463.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-01
AI Technical Summary
The existing technology has problems such as insufficient single mode analysis, inability to dynamically adjust the weights and model parameters in user psychological state monitoring, insufficient modeling of mood swing trends, and multimodal feature alignment and fusion problems, resulting in insufficient accuracy and real-timeness of monitoring results.
The cross-modal joint coding network model is used to align and fusion the speech spectrum and text features, combine Bayesian network and Gaussian hybrid model for dynamic weight adjustment and mood fluctuation trend modeling, and extract local information and fusion of features through time sequence alignment algorithm and convolutional neural network, and finally classify psychological states through softmax function.
It has achieved a comprehensive integration of multimodal information, dynamically adjusted analysis strategies, accurately captured the laws of emotional changes, and improved the accuracy and real-time nature of psychological state monitoring.
Smart Images

Figure CN120236609A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal information processing, and particularly to a method and system for monitoring user mental state based on speech and semantic recognition. Background Art
[0002] With the continuous development of artificial intelligence technology, speech recognition and natural language processing technologies have been widely applied in multiple fields. However, traditional speech and semantic analysis methods mainly focus on information extraction and understanding, and have relatively limited capabilities in monitoring and analyzing user mental states. Mental state monitoring is of great significance in fields such as mental health assessment, human-computer interaction, and intelligent customer service. However, the current methods have the following deficiencies:
[0003] 1. Most traditional mental state monitoring methods rely on single-modal data, such as speech or text. Speech analysis usually focuses on acoustic features such as pitch and speech rate, while text analysis focuses more on emotional vocabulary and semantic content. However, single-modal analysis is difficult to comprehensively reflect the user's mental state because the expression of mental state is often multimodal, that is, it is reflected through the combination of speech and language.
[0004] 2. In the process of feature extraction and analysis in the prior art, fixed weights and model parameters are often used, and it is impossible to dynamically adjust the analysis strategy according to the user's emotional fluctuations and changes in mental state. This results in insufficient accuracy and real-time performance of the monitoring results when facing complex and changeable mental states.
[0005] 3. The change of mental state is a dynamic process, and the modeling of emotional fluctuations is crucial for accurately monitoring mental state. However, the prior art has deficiencies in the modeling of emotional fluctuation trends, and it is impossible to effectively capture the change rules of emotions, thus affecting the accurate classification of mental states.
[0006] 4. Dynamic mental state modeling introduces time series modeling technology to reflect the change trend of mental state by capturing the time series changes of speech and text. However, the time series features of speech and text have different dynamic characteristics, and it is difficult to synchronously process short-term features such as speech pause time and fundamental frequency fluctuation and long-term features such as semantic coherence and emotional tendency of text in time series modeling. In addition, the fusion of psychological scales and AI analysis results requires dynamically adjusting the confidence weights through a Bayesian network, but there is a matching problem between the discreteness of scale data and the continuity of AI analysis results, which may lead to deviations or lags in the dynamic adjustment process of the model, affecting the accuracy and real-time performance of mental state monitoring.
[0007] 5. Although multi-modal fusion technology has been applied in other fields, in the monitoring of mental states, how to effectively align and fuse speech spectral features and text semantic features remains a technical challenge. Existing methods have deficiencies in the joint modeling of multi-modal features and semantic space alignment, resulting in the fused features being unable to fully reflect the true mental state of the user. Summary of the Invention
[0008] To solve the problems existing in the above-mentioned prior art, the object of the present invention is to propose a method and system for monitoring the mental state of users based on speech and semantic recognition, which solves the problems existing in the prior art through technical means such as cross-modal joint coding, dynamic weight adjustment, and emotional fluctuation trend modeling.
[0009] To achieve the above object, the present invention provides the following solutions:
[0010] A method for monitoring the mental state of users based on speech and semantic recognition, including:
[0011] Obtain the speech data of the user and its corresponding text data, and respectively extract the speech spectral feature vector and text feature vector in the speech data;
[0012] Input the speech spectral feature vector and the text feature vector into a cross-modal joint coding network model, and use the multi-head attention mechanism to align the speech spectrum and the text embedding space to obtain joint features;
[0013] Adopt a Bayesian network model to extract the confidence weight of the joint feature to obtain a weighted feature, use a Gaussian mixture model to fit the emotional fluctuation trend in the weighted feature to obtain an emotional distribution probability;
[0014] Dynamically adjust the confidence weight of the joint feature according to the emotional distribution probability to further obtain a new weighted feature, perform alignment processing on the new weighted feature representation through a time series alignment algorithm, use a convolutional neural network model to extract local information from the aligned feature sequence, and perform feature fusion. For the fused feature vector, calculate the classification probability through the softmax function to determine the final classification result of the user's mental state.
[0015] Optionally, extracting the speech spectral feature vector in the speech data includes:
[0016] Perform spectral analysis on the speech data to obtain a spectrogram, extract Mel frequency cepstral coefficients from the spectrogram to generate Mel system data, calculate the fundamental frequency value according to the spectrogram to obtain fundamental frequency data, extract the energy spectrum from the spectrogram to generate energy spectrum data, and further analyze the duration and syllable number of the speech data to determine the speech rate value;
[0017] Integrate the Mel - based data, the fundamental frequency data, the speech rate value, and the energy spectrum data into an acoustic parameter set, and input the acoustic parameter set into a feature extractor to generate a speech feature vector.
[0018] Optionally, extracting the text feature vector from the speech data includes:
[0019] Perform semantic parsing on the text data. If the text contains specific emotional words, extract the emotional tendency information, further obtain keywords from the text, and analyze the context relationship of the keywords;
[0020] Construct a text feature vector according to the emotional tendency information and the context relationship of the keywords.
[0021] Optionally, obtaining the joint feature includes:
[0022] Perform temporal modeling on the speech spectrum feature vector and the text feature vector to obtain the temporal feature;
[0023] Input the speech spectrum feature vector and the text feature vector into a cross - modal joint encoding network model. According to the temporal feature, use the multi - head attention mechanism to analyze the correlation between the speech spectrum feature vector and the text feature vector, map different modal features to a unified semantic space, and if the speech spectrum is aligned with the text embedding space, obtain the joint feature.
[0024] Optionally, obtaining the weighted feature includes:
[0025] Obtain psychological scale data, use a Bayesian network model to model the psychological scale data, characterize the conditional dependence relationship between the dimensional indicators of the psychological scale, and obtain the psychological feature distribution;
[0026] Calculate the confidence weight of the joint feature based on the psychological feature distribution, and obtain the weighted value of the joint feature according to the confidence weight of the joint feature as the weighted feature.
[0027] To achieve the above object, the present invention also provides a user psychological state monitoring system based on speech and semantic recognition, including:
[0028] An information acquisition module, configured to acquire the speech data of the user and its corresponding text data, and respectively extract the speech spectrum feature vector and the text feature vector from the speech data;
[0029] A feature extraction module, configured to input the speech spectrum feature vector and the text feature vector into a cross - modal joint encoding network model, use the multi - head attention mechanism to align the speech spectrum with the text embedding space, and obtain the joint feature;
[0030] An emotion distribution probability extraction module, which is used to extract the confidence weights of the joint features by using a Bayesian network model, obtain weighted features, fit the emotional fluctuation trend in the weighted features by using a Gaussian mixture model, and obtain the emotion distribution probability;
[0031] A user mental state monitoring module, which is used to dynamically adjust the confidence weights of the joint features according to the emotion distribution probability, further obtain new weighted features, perform alignment processing on the representations of the new weighted features through a time series alignment algorithm, extract local information from the aligned feature sequence by using a convolutional neural network model, and perform feature fusion. For the fused feature vectors, calculate the classification probability through a softmax function to determine the final classification result of the user mental state.
[0032] Optionally, the information acquisition module includes:
[0033] A voice feature acquisition unit, which is used to perform spectrum analysis on the voice data to obtain a spectrogram, extract Mel frequency cepstral coefficients from the spectrogram to generate Mel series data, calculate the fundamental frequency value according to the spectrogram to obtain fundamental frequency data, extract the energy spectrum from the spectrogram to generate energy spectrum data, and further analyze the duration and syllable number of the voice data to determine the speech rate value; integrate the Mel series data, the fundamental frequency data, the speech rate value, and the energy spectrum data into an acoustic parameter set, and input the acoustic parameter set into a feature extractor to generate a voice feature vector.
[0034] Optionally, the information acquisition module further includes:
[0035] A text feature acquisition unit, which is used to perform semantic parsing on the text data. If the text contains specific emotional words, extract the emotional tendency information, further obtain keywords from the text, analyze the context relationship of the keywords, and construct a text feature vector according to the emotional tendency information and the context relationship of the keywords.
[0036] Optionally, the feature extraction module includes:
[0037] A feature extraction unit, which is used to perform time series modeling on the voice spectrum feature vector and the text feature vector to obtain the time series features, input the voice spectrum feature vector and the text feature vector into a cross-modal joint encoding network model, analyze the correlation between the voice spectrum feature vector and the text feature vector by using a multi-head attention mechanism according to the time series features, map different modal features to a unified semantic space, and if the voice spectrum is aligned with the text embedding space, obtain the joint features.
[0038] Optionally, the emotion distribution probability extraction module includes:
[0039] A psychological feature distribution extraction unit, which is used to obtain psychological scale data, build a Bayesian network model for the psychological scale data, characterize the conditional dependence relationship between the index of each dimension of the psychological scale, and obtain the psychological feature distribution; calculate the confidence weight of the joint feature based on the psychological feature distribution, and obtain the weighted value of the joint feature according to the confidence weight of the joint feature as the weighted feature;
[0040] An emotional distribution probability extraction unit, which is used to fit the emotional fluctuation trend in the weighted feature by using a Gaussian mixture model to obtain the emotional distribution probability.
[0041] The beneficial effects of the present invention are as follows:
[0042] Comprehensiveness of multimodal fusion: Through the cross-modal joint coding network model, the present invention aligns and fuses the speech spectrum features and text features, makes full use of the complementarity of speech and semantic information, and comprehensively reflects the user's psychological state. This multimodal fusion method can more accurately capture the user's emotional and psychological changes, and improve the accuracy and reliability of psychological state monitoring.
[0043] Flexibility of the dynamic adjustment mechanism: The present invention adopts a Bayesian network model and a Gaussian mixture model to dynamically adjust the confidence weight of the joint feature according to the emotional distribution probability. This dynamic adjustment mechanism can respond to the changes of the user's psychological state in real time, adapt to the psychological state monitoring requirements in different scenarios, and improve the flexibility and adaptability of the system.
[0044] Effective modeling of the emotional fluctuation trend: The present invention fits the emotional fluctuation trend in the weighted feature by using a Gaussian mixture model. The present invention can effectively capture the change law of emotions, and processes the feature sequence through a time series alignment algorithm to further improve the accuracy of the emotional fluctuation trend modeling. This provides strong support for the dynamic monitoring of the psychological state, enabling the system to more accurately identify the user's emotional changes.
[0045] Accuracy of feature extraction: In the process of speech and text feature extraction, the present invention introduces a variety of advanced technical means, such as Mel Frequency Cepstral Coefficients, sentiment analysis, keyword context relationship analysis, etc. These technologies can accurately extract the key information in speech and text, and provide high-quality feature input for subsequent psychological state monitoring.
[0046] Efficiency of the system architecture: The user psychological state monitoring system proposed by the present invention modularizes functions such as information acquisition, feature extraction, emotional distribution probability extraction, and psychological state classification through modular design, improving the scalability and maintainability of the system. At the same time, based on the classification probability calculation of a convolutional neural network and a softmax function, it can efficiently output the classification result of the user's psychological state, meeting the requirements of real-time monitoring.
[0047] In summary, through innovative technologies such as multi-modal fusion, dynamic adjustment mechanism, and emotional fluctuation trend modeling, the present invention significantly improves the accuracy and real-time performance of user mental state monitoring, providing a new technical solution for fields such as mental health assessment and human-computer interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0049] Figure 1 It is a flowchart of a method for monitoring user mental state based on speech and semantic recognition according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0051] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0052] As Figure 1 shown, this embodiment discloses a method for monitoring user mental state based on speech and semantic recognition, including: obtaining the speech data of the user and its corresponding text data, and respectively extracting the speech spectrum feature vector and text feature vector in the speech data; inputting the speech spectrum feature vector and text feature vector into a cross-modal joint encoding network model, using the multi-head attention mechanism to align the speech spectrum and text embedding spaces, and obtaining joint features; using a Bayesian network model to extract the confidence weights of the joint features, obtaining weighted features, using a Gaussian mixture model to fit the emotional fluctuation trend in the weighted features, and obtaining the emotional distribution probability; dynamically adjusting the confidence weights of the joint features according to the emotional distribution probability, further obtaining new weighted features, performing alignment processing on the new weighted feature representation through a time series alignment algorithm, using a convolutional neural network model to extract local information from the aligned feature sequence, and performing feature fusion. For the fused feature vector, calculating the classification probability through the softmax function to determine the final user mental state classification result.
[0053] Furthermore, extracting the speech spectral feature vectors from the speech data includes: performing spectral analysis on the speech data to obtain a spectrogram, extracting Mel-frequency cepstral coefficients from the spectrogram to generate Mel-based data, calculating the fundamental frequency value according to the spectrogram to obtain fundamental frequency data, extracting the energy spectrum from the spectrogram to generate energy spectrum data, further analyzing the duration and syllable count of the speech data to determine the speech rate value; integrating the Mel-based data, fundamental frequency data, speech rate value, and energy spectrum data into an acoustic parameter set, and inputting the acoustic parameter set into a feature extractor to generate speech feature vectors.
[0054] Specifically, a pre-trained model is used to perform spectral analysis on the speech signal to obtain a spectrogram; Mel-frequency cepstral coefficients are extracted from the spectrogram to generate Mel-based data, the fundamental frequency value is calculated according to the spectrogram to obtain fundamental frequency data, the duration and syllable count of the speech signal are analyzed to determine the speech rate value, the energy spectrum is extracted from the spectrogram to generate energy spectrum data, and the Mel-based data, fundamental frequency data, speech rate value, and energy spectrum data are integrated into an acoustic parameter set. The acoustic parameter set is input into a feature extractor to generate speech feature vectors:
[0055] By performing frequency-domain conversion on the speech waveform using a pre-trained model, a spectrogram reflecting the characteristics of sound frequency and time variation can be obtained. Taking a segment of Mandarin speech as an example, the spectrogram obtained by using the short-time Fourier transform shows that the speech signal lasts for about two seconds on the time axis, and the frequency range covers zero to eight thousand Hertz. Based on this spectrogram, extracting Mel-frequency cepstral coefficients can better simulate the perceptual characteristics of the human ear to sound. The Mel-frequency cepstral coefficients usually take twelve to sixteen dimensions, and each dimension reflects the energy distribution characteristics of different frequency bands. When calculating the fundamental frequency from the spectrogram, it is necessary to analyze the periodic characteristics of the sound signal. The fundamental frequency range of Mandarin tones is generally between eighty and four hundred Hertz. For example, the fundamental frequency curve of the first tone is flat, while the fourth tone shows a characteristic of rising first and then falling. By analyzing the fundamental frequency variation, different tones can be effectively distinguished, providing an important reference for speech recognition. Speech rate analysis involves the relationship between the number of syllables and the duration. For example, if a two-second speech contains six syllables, the speech rate is three syllables per second. By comparing with the standard speech rate range, the speech rate of the speaker can be judged, which has guiding significance for adjusting the naturalness of speech synthesis. The energy spectrum reflects the intensity variation of the speech signal at different time points. For example, the energy of the vowel segment is usually significantly higher than that of the consonant segment, and energy jumps often occur at the junction of the initial consonant and the final vowel. Taking the Chinese syllable "ma" as an example, its energy curve shows a characteristic of being weak first and then strong, which is consistent with the pronunciation characteristics of its initial consonant and final vowel. After integrating the above acoustic parameters, a multi-dimensional feature vector can be formed. For example, the feature vector of a certain syllable may include twelve-dimensional Mel coefficients, one-dimensional fundamental frequency value, one-dimensional speech rate value, and one-dimensional energy value, totaling fifteen dimensions. These feature vectors can be used for subsequent tasks such as speech recognition and sentiment analysis. The entire process of feature extraction reflects the conversion from the original speech signal to high-level speech features, enabling the machine to better understand and process human speech. This multi-level acoustic feature extraction method can comprehensively capture various characteristics of the speech signal, and can be used not only for speech recognition, but also in fields such as speaker recognition and speech synthesis. By reasonably combining these features, the representation ability of the system for the speech signal can be improved, thereby enhancing the performance of related applications.
[0056] Furthermore, extracting the text feature vector from the speech data includes: performing semantic parsing on the text data. If the text contains specific sentiment words, extracting the sentiment tendency information, further obtaining keywords from the text, and analyzing the context relationship of the keywords; constructing the text feature vector according to the sentiment tendency information and the context relationship of the keywords.
[0057] Specifically, a pre-trained language model is used to load a text dataset and perform semantic parsing on the text content. If the text contains specific sentiment words, sentiment tendency information is extracted, keywords are obtained from the text, the relevance of the keywords in the context is analyzed, and according to the sentiment tendency and the context relationship of the keywords, text semantic features are generated, and the semantic features and the keyword relevance are integrated to construct a text feature vector:
[0058] The pre-trained language model serves as the basic framework for text analysis and usually adopts a deep neural network structure. After being pre-trained on a large-scale corpus, it can understand the semantic information of the text. Taking news comment analysis as an example, after the model inputs a text such as "This movie is extremely wonderful and the acting skills of the actors are very excellent", semantic understanding is performed through multiple layers of neural networks. During the semantic parsing process, the model first identifies sentiment words. For example, when analyzing a restaurant review "The service attitude of this store is very poor, but the taste of the dishes is good", the system will extract sentiment words such as "poor" and "good", and determine that the first half is a negative sentiment and the second half is a positive sentiment. For reviews in a specific field, such as "The doctor has a kind attitude and the treatment effect is remarkable" in medical service evaluation, the system can extract sentiment words such as "kind" and "remarkable". The keyword extraction link focuses on the core words in the text. Taking product evaluation as an example, in "The camera function of this mobile phone is excellent, but the battery life is weak", the system will extract "camera" and "battery life" as keywords. When analyzing the context relevance of these keywords, the system can understand that "camera" and "excellent" form a positive evaluation, and "battery life" and "weak" form a negative evaluation. The generation of text semantic features comprehensively considers the sentiment tendency and the keyword relationship. For example, when analyzing a recruitment information "This position requires the applicant to have strong communication skills and teamwork spirit", the system extracts "communication skills" and "teamwork" as key features and identifies the positive nature of these requirements. By integrating these features, a feature vector describing the text semantics is formed.
[0059] Furthermore, obtaining the joint features includes: performing temporal modeling on the speech spectrum feature vector and the text feature vector to obtain temporal features; inputting the speech spectrum feature vector and the text feature vector into a cross-modal joint encoding network model, and according to the temporal features, using the multi-head attention mechanism to analyze the relevance between the speech spectrum feature vector and the text feature vector, mapping different modal features to a unified semantic space, and if the speech spectrum is aligned with the text embedding space, the joint features are obtained.
[0060] Specifically, a joint encoding network is used to obtain speech feature vectors and text feature vectors. The speech spectrum and text embedding space are processed through a multi-head attention mechanism. If the speech spectrum and text embedding space are aligned, a joint feature representation is obtained. The consistency of the semantic space is judged based on the joint feature representation. An attention layer is used to analyze the correlation between the speech vector and the text vector. The cross-modal semantic alignment result is determined through the feature representation, and the final output of the joint encoding is generated according to the alignment result:
[0061] The joint encoding network adopts a two-stream architecture to process speech and text inputs respectively, and maps features of different modalities to a unified semantic space through a shared encoder. Taking speech emotion analysis as an example, when a user says "The weather is really nice today", the speech spectrum will show a specific frequency distribution and energy change pattern, and the text features will contain positive semantic information. The multi-head attention mechanism can simultaneously focus on multiple feature dimensions of speech and text, such as the pitch, intensity, speech rate of speech, etc. and the keywords, syntactic structures, etc. in the text. In the speech education application scenario, the system needs to judge whether the pronunciation of the learner is accurate. When the learner reads the sentence "Spring is coming", the joint encoding network will analyze the acoustic features of the speech and the standard pronunciation features of the text. By calculating the similarity between the speech spectrum and the standard pronunciation template through the attention mechanism, the system can accurately locate pronunciation deviations, such as inaccurate tones or mispronounced phonemes. In the intelligent customer service scenario, the joint feature representation can accurately understand the user's speech instructions and text information. For example, when the user asks "When will this product be shipped", the system will comprehensively analyze the urgency shown by the tone, intonation, etc. in the speech, as well as the time query intention in the text. By judging the actual demand urgency of the user through feature alignment, a more accurate service response can be provided. During the cross-modal semantic alignment process, the attention layer will dynamically adjust the weights of different features. In the speech news summary task, for a news like "The stock market rose sharply yesterday", the system will focus on the stressed part expressing "sharply" in the speech, as well as the key data information in the text. By calculating the attention scores of different modality features, important information can be highlighted to generate an accurate news summary. The final output of the joint encoding needs to ensure the semantic consistency of the speech and text features. In the speech navigation application, when the user says "Navigate to the nearest supermarket", the system will convert the speech instruction into a standardized text command and ensure the accuracy of understanding through feature alignment. If inconsistencies between the speech and text features are detected, the system will ask the user to confirm to avoid incorrect navigation. This multi-modal feature fusion method can improve the robustness and reliability of the system and provide a more natural and accurate human-computer interaction experience for users.
[0062] A bidirectional long short-term memory network is used to perform temporal modeling on the speech feature vector and the text feature vector to obtain a temporal feature representation. For the temporal feature representation, the speech spectrum and the text embedding space are aligned through a multi-head attention mechanism to obtain a joint feature representation:
[0063] The bidirectional long short-term memory network captures the temporal dependence relationships of speech and text sequences by establishing two temporal channels, a forward channel and a backward channel. For example, in speech emotion recognition, for a piece of speech expressing anger, the trend of its pitch and the change in speech rate both have obvious temporal characteristics. The network can obtain the process of rising intonation from the forward direction and the characteristic of accelerating speech rate from the backward direction, and comprehensively form a judgment of the emotion. The multi-head attention mechanism learns the correspondence between speech and text from different perspectives by setting multiple attention heads. Taking speech navigation as an example, when the user says "turn forward", the attention mechanism will focus on the keywords "forward" and "turn" in the speech, and at the same time correspond to the direction and action information in the text instruction to achieve accurate semantic alignment. Each attention head can focus on different feature dimensions, such as intonation, rhythm, etc. When the convolutional neural network extracts local features, it uses a multi-layer convolutional structure to gradually extract high-level semantic information. In speech command recognition, the first-layer convolution may focus on the fundamental frequency feature, the second layer extracts the phoneme combination feature, and the third layer identifies the complete speech command pattern. By stacking multiple convolutional layers, the network can effectively extract feature hierarchies from simple to complex.
[0064] Furthermore, obtaining the weighted features includes: obtaining psychometric scale data, using a Bayesian network model to model the psychometric scale data, depicting the conditional dependence relationships between the various dimensional indicators of the psychometric scale, and obtaining the psychological feature distribution; calculating the confidence weights of the joint features based on the psychological feature distribution, and obtaining the weighted values of the joint features according to the confidence weights of the joint features as the weighted features.
[0065] Specifically, according to the psychometric scale data, use a Bayesian network to dynamically adjust the confidence weights of the joint feature representation to obtain the weighted feature representation. Use a Bayesian network to model the psychometric scale data to obtain the psychological feature distribution. For the psychological feature distribution, calculate the confidence weights of the joint feature representation, and dynamically adjust the weighted values of the joint feature representation according to the confidence weights:
[0066] Bayesian networks depict the conditional dependence relationships among the dimensional indicators of psychological scales through probabilistic graphical models. For example, in a depression scale, dimensional indicators such as sleep quality, appetite status, and mood swings influence each other. By building a Bayesian network, the probability distributions of these psychological characteristics can be obtained. For instance, the score distribution of a certain subject in the depression dimension is 0.6 for severe depression, 0.3 for moderate depression, and 0.1 for mild depression. Calculate the confidence weights of the joint feature representation based on the psychological characteristic distribution, with a focus on the reliability and representativeness of the features. Taking depressive symptoms as an example, if the speech feature shows a low mood, the text feature also reflects negative thinking, and the psychological scale indicates a tendency towards severe depression, then the confidence weights of these features will be relatively high. On the contrary, if there are significant differences among the features of each modality, the weights will be correspondingly reduced. When dynamically adjusting the weighting values, the temporal changes of the features need to be considered. For example, if the speech emotion features of a certain patient show a fluctuating trend within a week, and the text expression also changes repeatedly from negative to positive, the weights should be dynamically adjusted according to the consistency of the temporal features at this time, enabling the model to capture this emotional fluctuation pattern.
[0067] Use a Gaussian mixture model to fit the emotional fluctuation trend in the weighted feature representation to obtain the emotion distribution probability. According to the emotion distribution probability, use a dynamic adjustment algorithm to update the confidence weights of the weighted feature representation. For the updated confidence weights, use the joint feature representation algorithm to calculate the new weighted feature representation. Align the new weighted feature representation through the temporal alignment algorithm to obtain the aligned feature sequence. Use a convolutional neural network to extract local information from the aligned feature sequence to obtain the local feature representation. According to the local feature representation, use a fully connected layer for feature fusion to obtain the fused feature vector. For the fused feature vector, calculate the classification probability through the softmax function to determine the final classification result:
[0068] The Gaussian mixture model models the distribution characteristics of different emotional fluctuations and fits complex emotional change trends through the superposition of multiple Gaussian distributions. For example, in a day, someone's mood may be calm in the morning, excited at noon, and slightly tired in the afternoon. The emotional state in each time period can be described by a Gaussian distribution, and the weights and parameters of these distributions reflect the probability distribution of emotional changes. The dynamic adjustment algorithm updates the weights of feature representations based on the emotional distribution probability. For instance, when it is detected that someone shows similar emotional fluctuations at the same time period for several consecutive days, the system will increase the weight of the features in this time period. For example, during the morning commute on weekdays, if anxiety is continuously observed, the system will increase the confidence weight of the features in this time period. The joint feature representation algorithm comprehensively considers multi-dimensional information. For example, data from multiple dimensions such as speech features, facial expressions, and heart rate changes are fused. When phenomena such as increased speech rate, tense facial expressions, and elevated heart rate occur simultaneously, the system can more accurately identify a state of tension or anxiety. The time series alignment algorithm processes feature sequences at different time scales. For example, by comparing the emotional states at fixed time periods every day, it is found that someone shows low work enthusiasm on Monday mornings every week, while showing high enthusiasm on Friday afternoons. Through alignment processing, this periodic emotional change pattern can be accurately captured. When the convolutional neural network extracts local features, it focuses on the emotional change features within a specific time window. For example, by analyzing the emotional fluctuations within an hour, it may be found that nervousness appears during the preparation stage before a meeting, gradually calms down during the meeting, and shows a relaxed state after the meeting. This fine-grained local feature can reflect the impact of specific events on emotions. The feature fusion layer integrates feature information at different levels to form a more comprehensive feature representation. For example, by combining short-term emotional fluctuation features with long-term emotional trend features, it can not only reflect the current emotional state but also the emotional development trend over a period of time. For example, someone has high work pressure recently and shows an overall negative emotional trend, but their short-term emotions will improve significantly when participating in team activities. The classification probability calculation finally maps the fused features to specific emotional categories. For example, the system may identify that there is a 60% probability that the current state is mild anxiety, a 30% probability of being in a normal state, and a 10% probability of other emotional states. This probabilistic output can better reflect the uncertainty and complexity of emotional states.
[0069] If the deviation between the emotional classification result and the psychological scale data exceeds the preset threshold, the confidence weights of the Bayesian network are readjusted, and the weighted feature representation is updated.
[0070] Calculate the deviation value between the emotion classification result and the psychological scale data. If the deviation value exceeds the preset threshold, obtain the current confidence weight of the Bayesian network. For the obtained confidence weight, use an optimization algorithm to adjust it. Based on the adjusted confidence weight, recalculate the weighted feature representation. According to the updated weighted feature representation, retrain the emotion classification model. Use the trained model to reclassify the psychological scale data. According to the reclassification result, determine whether the preset deviation threshold requirement is met.
[0071] Specifically, the deviation analysis between the emotion classification result and the psychological scale data usually involves multi-dimensional evaluations. Taking a certain mental health assessment system as an example, after a patient fills out the Beck Depression Inventory and gets a total score of 30 points, the system determines it as mild depression based on the emotion classification model, while the scale assessment result shows moderate depression. At this time, it is necessary to calculate the deviation value between the two assessment results. Suppose the preset acceptable deviation threshold is 15%. When the current deviation exceeds the threshold, it is necessary to adjust the model parameters. The confidence weight of the Bayesian network reflects the influence degree of different features on emotion judgment. For example, in the recognition of depressive emotions, features such as sleep quality, appetite status, and social activities have different weights. The initial weight may set the sleep quality as 0.3, the appetite status as 0.2, and the social activity as 0.25. By analyzing a large amount of clinical data, it is found that the current weight setting may underestimate the importance of social activities. When using the gradient descent method to adjust the weights by the optimization algorithm, gradually adjust the weights of each feature according to the deviation value. For example, increase the weight of social activities to 0.35, and at the same time appropriately reduce the weights of other features to keep the total weight as 1. The correlation between features needs to be considered during the adjustment process to avoid over-reliance on a single feature. The recalculation of the weighted feature representation involves the update of the feature vector. Suppose the feature data of a certain patient shows: the sleep quality score is 8 points, the appetite status score is 7 points, and the social activity score is 5 points. Applying the adjusted weights, a new feature representation is obtained, which more accurately reflects the comprehensive influence of each index on the emotional state. The retraining of the emotion classification model needs to consider the distribution characteristics of the samples. In the actual application in medical institutions, the situation of unbalanced sample distribution may be encountered, such as more mild depression samples and fewer severe depression samples. At this time, techniques such as oversampling or undersampling need to be used to balance the sample distribution and improve the model's recognition ability for various emotional states. During the reclassification process of the psychological scale data, the clinical significance of the classification result needs to be concerned. For example, a certain patient is determined to be moderately depressed after reclassification, which is consistent with the scale assessment result, and the deviation value drops to 5%, meeting the preset threshold requirement. At the same time, the time continuity of the classification result needs to be considered to avoid drastic fluctuations in the short term. Through the dynamic adjustment process, the classification result of the model gradually approaches the clinical judgment of professional doctors, improving the accuracy and reliability of mental state assessment.
[0072] This embodiment also provides a user mental state monitoring system based on speech and semantic recognition, including: an information acquisition module for acquiring the speech data of a user and its corresponding text data, and respectively extracting the speech spectrum feature vector and the text feature vector from the speech data; a feature extraction module for inputting the speech spectrum feature vector and the text feature vector into a cross-modal joint encoding network model, and using a multi-head attention mechanism to align the speech spectrum and the text embedding space to obtain joint features; an emotion distribution probability extraction module for using a Bayesian network model to extract the confidence weights of the joint features to obtain weighted features, and using a Gaussian mixture model to fit the emotion fluctuation trend in the weighted features to obtain the emotion distribution probability; a user mental state monitoring module for dynamically adjusting the confidence weights of the joint features according to the emotion distribution probability, further obtaining new weighted features, performing alignment processing on the new weighted feature representation through a time series alignment algorithm, using a convolutional neural network model to extract local information from the aligned feature sequence, and performing feature fusion. For the fused feature vector, classification probability calculation is performed through a softmax function to determine the final user mental state classification result.
[0073] Further, the information acquisition module includes: a speech feature acquisition unit for performing spectrum analysis on the speech data to obtain a spectrogram, extracting Mel frequency cepstral coefficients from the spectrogram to generate Mel series data, calculating the fundamental frequency value according to the spectrogram to obtain fundamental frequency data, extracting the energy spectrum from the spectrogram to generate energy spectrum data, and further analyzing the duration and syllable number of the speech data to determine the speech rate value; integrating the Mel series data, the fundamental frequency data, the speech rate value, and the energy spectrum data into an acoustic parameter set, and inputting the acoustic parameter set into a feature extractor to generate a speech feature vector.
[0074] Further, the information acquisition module also includes: a text feature acquisition unit for performing semantic parsing on the text data. If the text contains specific emotional words, then extract the emotional tendency information, further obtain keywords from the text, and analyze the context relationship of the keywords. According to the emotional tendency information and the context relationship of the keywords, construct a text feature vector.
[0075] Further, the feature extraction module includes: a feature extraction unit for performing time series modeling on the speech spectrum feature vector and the text feature vector to obtain time series features, inputting the speech spectrum feature vector and the text feature vector into a cross-modal joint encoding network model, and according to the time series features, using a multi-head attention mechanism to analyze the correlation between the speech spectrum feature vector and the text feature vector, mapping different modal features to a unified semantic space, and if the speech spectrum and the text embedding space are aligned, then obtain joint features.
[0076] Further, the emotion distribution probability extraction module includes: a psychological feature distribution extraction unit, configured to obtain psychological scale data, model the psychological scale data using a Bayesian network model, depict the conditional dependence relationship between the dimensional indexes of the psychological scale, and obtain the psychological feature distribution; calculate the confidence weight of the joint feature based on the psychological feature distribution, and obtain the weighted value of the joint feature according to the confidence weight of the joint feature as the weighted feature; an emotion distribution probability extraction unit, configured to fit the emotion fluctuation trend in the weighted feature using a Gaussian mixture model to obtain the emotion distribution probability.
[0077] The embodiments described above are only descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A method for monitoring user psychological state based on speech and semantic recognition, characterized in that: include: Acquire the user's voice data and its corresponding text data, and respectively extract the voice spectrum feature vector and the text feature vector from the voice data; Inputting the speech spectrum feature vector and the text feature vector into a cross-modal joint encoding network model, aligning the speech spectrum and the text embedding space using a multi-head attention mechanism, and obtaining joint features; A Bayesian network model is used to extract the confidence weight of the joint feature to obtain a weighted feature, and a Gaussian mixture model is used to fit the emotion fluctuation trend in the weighted feature to obtain the emotion distribution probability; The confidence weight of the joint feature is dynamically adjusted according to the emotion distribution probability, and new weighted features are further obtained. The new weighted feature representation is aligned through a temporal alignment algorithm, and a convolutional neural network model is used to extract local information from the aligned feature sequence, and feature fusion is performed. For the fused feature vector, the classification probability is calculated through a softmax function to determine the final user psychological state classification result.
2. The user psychological state monitoring method based on speech and semantic recognition according to claim 1 is characterized in that: Extracting the speech spectrum feature vector from the speech data includes: Performing spectral analysis on the speech data to obtain a spectrogram, extracting Mel-frequency cepstrum coefficients from the spectrogram to generate Mel-frequency data, calculating a fundamental frequency value according to the spectrogram to obtain fundamental frequency data, extracting an energy spectrum from the spectrogram to generate energy spectrum data, and further analyzing the duration and number of syllables of the speech data to determine a speech rate value; The Mel-series data, the fundamental frequency data, the speech rate value and the energy spectrum data are integrated into an acoustic parameter set, and the acoustic parameter set is input into a feature extractor to generate a speech feature vector.
3. The user psychological state monitoring method based on speech and semantic recognition according to claim 1 is characterized in that: Extracting the text feature vector from the speech data includes: Perform semantic analysis on the text data, if the text contains specific sentiment words, extract sentiment tendency information, further obtain keywords from the text, and analyze the contextual relationship of the keywords; A text feature vector is constructed according to the contextual relationship between the sentiment tendency information and the keywords.
4. The user psychological state monitoring method based on speech and semantic recognition according to claim 1 is characterized in that: Acquiring the joint feature includes: Performing time series modeling on the speech spectrum feature vector and the text feature vector to obtain the time series feature; The speech spectrum feature vector and the text feature vector are input into a cross-modal joint encoding network model. According to the temporal features, a multi-head attention mechanism is used to analyze the correlation between the speech spectrum feature vector and the text feature vector, and different modal features are mapped to a unified semantic space. If the speech spectrum is aligned with the text embedding space, the joint feature is obtained.
5. The user psychological state monitoring method based on speech and semantic recognition according to claim 1 is characterized in that: Acquiring the weighted feature includes: Obtaining psychological scale data, modeling the psychological scale data using a Bayesian network model, describing the conditional dependency relationship between indicators of each dimension of the psychological scale, and obtaining psychological characteristic distribution; The confidence weight of the joint feature is calculated based on the psychological feature distribution, and the weighted value of the joint feature is obtained according to the confidence weight of the joint feature as the weighted feature.
6. A user psychological state monitoring system based on speech and semantic recognition, characterized in that: include: An information acquisition module, used to acquire the user's voice data and its corresponding text data, and to extract the voice spectrum feature vector and the text feature vector from the voice data respectively; A feature extraction module, used to input the speech spectrum feature vector and the text feature vector into a cross-modal joint encoding network model, align the speech spectrum and the text embedding space using a multi-head attention mechanism, and obtain joint features; The emotion distribution probability extraction module is used to extract the confidence weight of the joint feature using a Bayesian network model to obtain weighted features, and to fit the emotion fluctuation trend in the weighted features using a Gaussian mixture model to obtain the emotion distribution probability; The user psychological monitoring module is used to dynamically adjust the confidence weight of the joint feature according to the emotion distribution probability, further obtain new weighted features, align the new weighted feature representations through the time series alignment algorithm, use the convolutional neural network model to extract local information from the aligned feature sequence, and perform feature fusion. For the fused feature vector, the classification probability is calculated through the softmax function to determine the final user psychological state classification result.
7. The user psychological state monitoring method based on speech and semantic recognition according to claim 6 is characterized in that: The information acquisition module includes: A speech feature acquisition unit is used to perform spectral analysis on the speech data to obtain a spectrogram, extract Mel-frequency cepstral coefficients from the spectrogram to generate Mel-frequency data, calculate the fundamental frequency value according to the spectrogram to obtain the fundamental frequency data, extract the energy spectrum from the spectrogram to generate energy spectrum data, further analyze the duration and number of syllables of the speech data to determine the speech rate value; integrate the Mel-frequency data, the fundamental frequency data, the speech rate value and the energy spectrum data into an acoustic parameter set, input the acoustic parameter set into a feature extractor, and generate a speech feature vector.
8. The method for monitoring user psychological state based on speech and semantic recognition according to claim 6, characterized in that: The information acquisition module also includes: The text feature acquisition unit is used to perform semantic analysis on the text data. If the text contains specific emotional words, the emotional tendency information is extracted, and keywords are further acquired from the text. The contextual relationship of the keywords is analyzed, and a text feature vector is constructed based on the emotional tendency information and the contextual relationship of the keywords.
9. The method for monitoring user psychological state based on speech and semantic recognition according to claim 6, characterized in that: The feature extraction module comprises: A feature extraction unit is used to perform time series modeling on the speech spectrum feature vector and the text feature vector to obtain the time series features, input the speech spectrum feature vector and the text feature vector into a cross-modal joint encoding network model, analyze the correlation between the speech spectrum feature vector and the text feature vector according to the time series features using a multi-head attention mechanism, map different modal features to a unified semantic space, and obtain the joint features if the speech spectrum is aligned with the text embedding space.
10. The method for monitoring user psychological state based on speech and semantic recognition according to claim 6, characterized in that: The emotion distribution probability extraction module includes: A psychological characteristic distribution extraction unit is used to obtain psychological scale data, model the psychological scale data using a Bayesian network model, characterize the conditional dependency relationship between indicators of each dimension of the psychological scale, and obtain psychological characteristic distribution; calculate the confidence weight of the joint feature based on the psychological characteristic distribution, and obtain the weighted value of the joint feature according to the confidence weight of the joint feature as the weighted feature; The emotion distribution probability extraction unit is used to use the Gaussian mixture model to fit the emotion fluctuation trend in the weighted features and obtain the emotion distribution probability.
Citation Information
Cited By
Psychological risk early warning method and system based on multi-dimensional emotion data of user
CN120413055A
Postoperative psychological state evaluation system and method based on MPNFS theory
CN120674085A
A postoperative psychological state evaluation system and method based on MPNFS theory
CN120674085B
Emotion evaluation method and device based on depth time sequence modeling
CN120959742A
Speech processing method
CN121054044A