Psychological state assessment method and device, electronic equipment and computer program product

By collecting user voice signals and physiological detection signals, performing multi-dimensional feature extraction and consistency verification, generating a masking penalty factor, and optimizing the feature weight matrix, the problem of low accuracy in psychological state assessment in existing technologies is solved, and a more accurate psychological state assessment is achieved.

CN122376100APending Publication Date: 2026-07-14
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Filing Date
2026-04-17
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing technologies cannot accurately distinguish between deliberate concealment and natural emotional expression when faced with complex psychological state expressions, especially in scenarios of emotional faking, resulting in poor accuracy in psychological state assessment.

Method used

By collecting users' voice signals and physiological detection signals, multi-dimensional feature extraction and consistency verification are performed to generate a concealment penalty factor. Combined with machine learning algorithms, the feature weight matrix is ​​optimized to conduct psychological state assessment.

Benefits of technology

It improves the accuracy of psychological state assessment, solves the problem of low assessment accuracy caused by the "smiling depression" phenomenon, and achieves accurate assessment of multi-dimensional psychological state.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122376100A_ABST
    Figure CN122376100A_ABST
Patent Text Reader

Abstract

The application discloses a psychological state evaluation method and device, electronic equipment and computer program product, and relates to the technical field of psychological state evaluation. A psychological evaluation guide task is displayed; a voice signal during user voice interaction and a physiological detection signal of the user are collected; multi-dimensional feature extraction is performed on the voice signal; a physiological emergency index is extracted from the physiological detection signal; consistency verification is performed on a semantic emotional valence feature, an acoustic dynamics feature and a psychological state indicated by the physiological emergency index; a concealment penalty factor is generated based on the physiological emergency index when the consistency verification fails; a corresponding feature weight matrix is called to weight and fuse each dimensional feature to obtain an initial emotion score; the initial emotion score is corrected based on the concealment penalty factor to obtain a corrected emotion score, and a psychological state evaluation result is generated according to the corrected emotion score. The application can improve the accuracy of user psychological state evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of psychological state assessment technology, specifically relating to a psychological state assessment method, device, electronic equipment, and computer program product. Background Technology

[0002] Psychological state assessment refers to the accurate identification of a user's emotional state, cognitive load, and psychological characteristics through multi-dimensional data collection and analysis. This provides crucial support for intelligent systems to offer personalized interaction strategies and improve user experience quality. Current mainstream assessment technologies primarily rely on physiological signal monitoring (such as EEG and heart rate), facial expression recognition, voice feature analysis, and behavioral log mining, combined with machine learning algorithms to construct state prediction models. However, existing technologies have significant limitations in dealing with complex psychological state expressions, particularly in their inability to identify scenarios of feigned emotions, leading to substantial discrepancies between assessment results and the true psychological state.

[0003] Specifically, in interactive scenarios, users may actively conceal their true emotions to protect their privacy, resulting in outward behaviors that contradict their inner experiences. For example, users may use a fake smile or tone of voice to mask their true emotional expression. This phenomenon of "smiling depression" often makes it difficult for traditional psychological state assessment techniques to distinguish between active concealment and natural emotional expression, leading to poor accuracy in user psychological state assessment and a significantly increased false alarm rate.

[0004] Therefore, how to provide an effective solution to improve the accuracy of user psychological state assessment has become an urgent problem to be solved in existing technologies. Summary of the Invention

[0005] The purpose of this invention is to provide a method, device, electronic device, and computer program product for assessing psychological state, in order to solve the aforementioned problems existing in the prior art.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a method for assessing psychological states, comprising: The psychological assessment guidance task is displayed based on the subjective state evaluation value input by the user and the preset trigger strategy. Collect the user's voice signals and physiological detection signals when the user interacts with the user in response to the psychological assessment guidance task; Multi-dimensional speech features are obtained by performing multi-dimensional feature extraction on the speech signal. The multi-dimensional speech features include semantic sentiment valence features and acoustic dynamic features. Physiological stress indicators characterizing the state of the nervous system are extracted from the physiological detection signals; The consistency of the semantic emotional valence features, the acoustic dynamic features, and the psychological states indicated by the physiological stress indicators is verified. If the consistency check fails, it is determined that there is emotional concealment behavior, and a concealment punishment factor is generated based on the physiological emergency indicators. Retrieve the feature weight matrix corresponding to the psychological assessment dimension of the psychological assessment guidance task, and perform weighted fusion on the subjective state assessment value input by the user, the physiological emergency index and the features of each dimension in the multi-dimensional voice features to obtain the initial emotion score; The initial emotion score is corrected based on the concealment penalty factor to obtain a corrected emotion score, and the psychological state assessment result of the user in the psychological assessment dimension to which the psychological assessment guidance task belongs is generated based on the corrected emotion score. Different psychological assessment dimensions correspond to different feature weight matrices. The feature weight matrix corresponding to any psychological assessment dimension is obtained by iteratively optimizing the subjective state assessment value input by the sample user during voice interaction during the sample psychological assessment guidance task of that psychological assessment dimension, the multi-dimensional voice features of the corresponding sample voice signal, the physiological emergency index of the sample physiological detection signal corresponding to the sample user, and the initial emotion score of the sample corresponding to the sample user as training samples through machine learning algorithms.

[0007] In one possible design, multi-dimensional feature extraction is performed on the speech signal to obtain multi-dimensional speech features, including: The semantic sentiment valence features of the speech signal are extracted through natural speech processing; High-dimensional acoustic representation vectors of the speech signal are extracted using a pre-trained speech deep neural network model. The speech signal is separated into a foreground human voice segment and a background silent segment; Extract the acoustic dynamics features of the foreground human voice segment; Based on the duration ratio of the speech signal to the background silence segment, the psychomotor retardation characteristics are determined. The high-dimensional acoustic representation vector, the semantic emotional valence feature, the acoustic dynamics feature, and the psychomotor lag feature are used as the multi-dimensional speech features.

[0008] In one possible design, the psychomotor retardation characteristic is Rpause=(Ttotal) Tspeech / Ttotal, where Ttotal represents the total duration of the speech signal, Tspeech represents the duration of the foreground human voice segment, and Ttotal... Tspeech represents the duration of the background silence segment.

[0009] In one possible design, consistency verification is performed on the semantic emotional valence features, the acoustic dynamic features, and the psychological states indicated by the physiological stress indicators, including: Determine whether the psychological state indicated by the semantic emotional valence features, the acoustic dynamic features, and the physiological stress indicators is consistent; If the psychological states indicated by the semantic emotional valence features, the acoustic dynamic features, and the physiological emergency indicators are all consistent, then the consistency check is deemed to have passed. Otherwise, the consistency check is deemed to have failed; The generation of the masking punishment factor based on the physiological stress indicators includes: Based on the aforementioned physiological emergency indicators, a masking penalty factor is generated through a preset continuous mapping function or neural network model.

[0010] In one possible design, a coordinated assessment method combining a single-question high-frequency assessment mode (daily capsule) and a multi-dimensional correlation assessment mode (deep scan) can be used for psychological state assessment. This involves presenting guided psychological assessment tasks based on pre-set trigger strategies, including: When the set daily task trigger time is reached or a user-initiated task trigger request is received, determine whether a new psychological state assessment cycle has been reached and whether the psychological state assessment results of the previous consecutive days are all abnormal. If a new psychological state assessment cycle has not yet been reached and at least one day in the previous consecutive days of psychological state assessment results shows that the psychological state is normal, then the subjective state assessment value entered by the user will trigger and display a psychological assessment guidance task for one psychological assessment dimension. If a new psychological state assessment cycle is reached or the psychological state assessment results for several consecutive days in the past are all abnormal, the subjective state assessment value input by the user will trigger psychological assessment guidance tasks for multiple psychological assessment dimensions. These tasks will be displayed sequentially to collect the user's voice signals and physiological detection signals when responding to the multiple psychological assessment guidance tasks and engaging in voice interaction, in order to conduct psychological state assessments for multiple psychological assessment dimensions.

[0011] In one possible design, after assessing mental state across multiple psychological dimensions, the method further includes: Based on the dispersion of the corrected emotion scores corresponding to multiple psychological assessment dimensions, the user's emotional stability and psychological energy indicators are determined. A circular closed curve is generated, and the radial perturbation amplitude of the edge of the circular closed curve is modulated using the emotional stability index to obtain the modulated closed curve. The psychological energy index is mapped to warm and cool color temperature values, and the modulated closed curve is color-rendered to obtain a user's psychological state profile.

[0012] In one possible design, after generating the user's psychological state assessment result for the psychological assessment dimension to which the psychological assessment guidance task belongs based on the modified emotion score, the method further includes: Based on the modified emotion score, a corresponding delivery message is generated.

[0013] In a second aspect, the present invention provides a psychological state assessment device, comprising: The triggering unit is used to display the psychological assessment guidance task based on the subjective state evaluation value input by the user and the preset triggering strategy. The acquisition unit is used to acquire the voice signals of the user when responding to the psychological assessment guidance task and the user's physiological detection signals. The first extraction unit is used to extract multi-dimensional features from the speech signal to obtain multi-dimensional speech features, the multi-dimensional speech features including semantic sentiment valence features and acoustic dynamic features; The second extraction unit is used to extract physiological emergency indicators characterizing the state of the nervous system from the physiological detection signals. A consistency verification unit is used to verify the consistency of the semantic emotional valence features, the acoustic dynamic features, and the psychological state indicated by the physiological emergency indicators. The judgment and generation unit is used to determine the presence of emotional concealment behavior when the consistency check fails and to generate a concealment punishment factor based on the physiological emergency indicators. The computation unit is used to retrieve the feature weight matrix corresponding to the psychological assessment dimension of the psychological assessment guidance task, and to perform weighted fusion of the subjective state assessment value input by the user, the physiological emergency index and the features of each dimension in the multi-dimensional voice features to obtain an initial emotion score. The correction and generation unit is used to correct the initial emotion score based on the concealment penalty factor to obtain a corrected emotion score, and generate the psychological state assessment result of the user in the psychological assessment dimension to which the psychological assessment guidance task belongs based on the corrected emotion score. Different psychological assessment dimensions correspond to different feature weight matrices. The feature weight matrix corresponding to any psychological assessment dimension is obtained by iteratively optimizing the subjective state assessment value input by the sample user during voice interaction during the sample psychological assessment guidance task of that psychological assessment dimension, the multi-dimensional voice features of the corresponding sample voice signal, the physiological emergency index of the sample physiological detection signal corresponding to the sample user, and the initial emotion score of the sample corresponding to the sample user as training samples through machine learning algorithms.

[0014] Thirdly, the present invention provides an electronic device comprising a memory, a processor, and a transceiver sequentially and communicatively connected, wherein the memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the psychological state assessment method as described in the first aspect or any possible design of the first aspect.

[0015] Fourthly, the present invention provides a computer-readable storage medium storing instructions that, when executed on a computer, perform the psychological state assessment method described in the first aspect or any possible design of the first aspect.

[0016] Fifthly, the present invention provides a computer program product containing instructions that, when executed on a computer, cause the computer to perform the mental state assessment method as described in the first aspect or any possible design of the first aspect.

[0017] Beneficial effects: This invention discloses a psychological state assessment scheme that addresses the problem of poor accuracy in user psychological state assessment caused by the inability to distinguish between active concealment and natural emotional expression due to the "smiling depression" phenomenon in existing technologies, thereby improving the accuracy of user psychological state assessment. Specifically, the scheme first presents a psychological assessment guidance task based on the user's input subjective state assessment value and a preset triggering strategy; it then collects the user's voice signals during voice interaction in response to the psychological assessment guidance task, as well as the user's physiological detection signals; it performs multi-dimensional feature extraction on the voice signals to obtain multi-dimensional voice features, including semantic-emotional valence features and acoustic-dynamic features; it extracts physiological emergency indicators representing the state of the nervous system from the physiological detection signals; it performs a consistency check on the psychological state indicated by the semantic-emotional valence features, the acoustic-dynamic features, and the physiological emergency indicators; if the consistency check fails, it determines that there is emotional concealment behavior and generates a concealment penalty factor based on the physiological emergency indicators; it retrieves the feature weight matrix corresponding to the psychological assessment dimension of the psychological assessment guidance task and applies it to the user's input subjective state assessment value. The initial emotion score is obtained by weighted fusion of features from physiological emergency indicators and multi-dimensional speech features. A modified emotion score is then obtained by correcting the initial emotion score based on a masking penalty factor. Based on the modified emotion score, a psychological state assessment result for the user in the psychological assessment guidance task is generated. Different psychological assessment dimensions correspond to different feature weight matrices. The feature weight matrix for any psychological assessment dimension is obtained by iteratively optimizing the training samples obtained during the sample psychological assessment guidance task for that dimension, using the subjective state assessment value input by the sample user during voice interaction, the multi-dimensional speech features of the corresponding sample speech signal, the physiological emergency indicators of the sample physiological detection signal corresponding to the sample user, and the initial emotion score corresponding to the sample user as training samples through machine learning algorithms. Thus, a cross-modal consistency verification mechanism is introduced into the psychological state assessment process. When the psychological state indicated by semantic emotional valence features, acoustic dynamic features, and physiological emergency indicators is inconsistent, a masking penalty factor is generated to correct the assessment result. This solves the problem of poor accuracy in user psychological state assessment caused by the active masking of emotional expression in "smiling depression," thereby improving the accuracy of user psychological state assessment. Meanwhile, it supports dynamic weight allocation based on different psychological assessment dimensions. Through machine learning algorithms, the weight matrix of each feature of different psychological assessment dimensions is iterated and optimized, so that it can be used to assess the psychological state of different psychological assessment dimensions, realize the accurate assessment of users' multi-dimensional psychological state, and facilitate practical application and promotion. Attached Figure Description

[0018] Figure 1 A flowchart of the psychological state assessment method provided in the embodiments of this application; Figure 2 A block diagram of the psychological state assessment device provided in the embodiments of this application; Figure 3 This is a block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the present invention will be briefly introduced below in conjunction with the accompanying drawings and descriptions of the embodiments or the prior art. Obviously, the following description of the structure of the accompanying drawings is only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. It should be noted that the description of these embodiments is for the purpose of helping to understand the present invention, but does not constitute a limitation of the present invention.

[0020] It should be understood that although the terms first, second, etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit, without departing from the scope of the exemplary embodiments of the invention.

[0021] It should be understood that the term "and / or" that may appear in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, B exists alone, and A and B exist simultaneously. The term " / and" that may appear in this document describes another relationship between related objects, indicating that two relationships can exist. For example, A / and B can mean: A exists alone, and A and B exist alone. In addition, the character " / " that may appear in this document generally indicates that the related objects before and after it are in an "or" relationship.

[0022] It should be understood that specific details are provided in the following description to facilitate a complete understanding of the exemplary embodiments. However, those skilled in the art will understand that the exemplary embodiments can be implemented without these specific details. For example, the system may be shown in block diagrams to avoid obscuring the example with unnecessary details. In other instances, well-known processes, structures, and techniques may be shown without unnecessary details to avoid obscuring the exemplary embodiments.

[0023] To improve the accuracy of user psychological state assessment, embodiments of this application provide a psychological state assessment method, device, electronic device, and computer program product. When the psychological states indicated by multi-dimensional features are inconsistent, the psychological state assessment method, device, electronic device, and computer program product can generate a masking penalty factor to correct the assessment results, thereby improving the accuracy of user psychological state assessment.

[0024] The psychological state assessment method provided in this application can be applied to electronic devices, such as smartwatches, smart bracelets, smartphones, personal computers, or tablets. It is understood that the execution entity described herein does not constitute a limitation on the embodiments of this application.

[0025] The psychological state assessment method provided in the embodiments of this application will be described in detail below.

[0026] like Figure 1 The diagram shown is a flowchart of a psychological state assessment method provided in the first aspect of the present application. The psychological state assessment method may include, but is not limited to, the following steps S101-S108.

[0027] Step S101. Based on the subjective state evaluation value input by the user and the preset triggering strategy, display the psychological assessment guidance task.

[0028] The preset triggering strategies can be set according to actual conditions. For example, a psychological assessment guidance task can be triggered and displayed daily at a set time, or psychological assessment guidance tasks of multiple dimensions can be triggered and displayed sequentially at certain intervals.

[0029] In one or more embodiments, a psychological assessment guidance task is displayed based on the subjective state evaluation value input by the user and a preset triggering strategy, which may include, but is not limited to, the following steps S1011-S1013.

[0030] Step S1011. When the set daily task trigger time is reached at the current time or a task trigger request is received from the user, determine whether a new psychological state assessment cycle has been reached and whether the psychological state assessment results of the previous consecutive days are all abnormal.

[0031] The psychological state assessment period can be one week, half a month, or one month, and the preceding consecutive days can be three consecutive days or five consecutive days, etc., but no specific limitation is made in this application embodiment.

[0032] Step S1012. If a new psychological state assessment cycle has not yet been reached and at least one day in the previous consecutive days of psychological state assessment results shows that the psychological state is normal, then a psychological assessment guidance task based on the subjective state assessment value input by the user is triggered and displayed for a psychological assessment dimension.

[0033] The content of the psychological assessment guidance task triggered and displayed can also vary depending on the subjective state evaluation value input by the user.

[0034] Step S1013. If a new psychological state assessment cycle has been reached or the psychological state assessment results for several consecutive days in the past have all been abnormal, then the subjective state assessment value input by the user triggers psychological assessment guidance tasks for multiple psychological assessment dimensions, and displays psychological assessment guidance tasks for multiple psychological assessment dimensions in sequence, so as to collect the voice signals and physiological detection signals of the user when responding to multiple psychological assessment guidance tasks and conducting voice interaction to conduct psychological state assessment for multiple dimensions.

[0035] Among them, psychological assessment guidance tasks with multiple psychological assessment dimensions may include, but are not limited to, psychological assessment guidance tasks for user interest deficiency dimension, user physical and mental energy dimension, user future outlook dimension, and user psychological resilience dimension.

[0036] Step S102. Collect the user's voice signals and physiological detection signals when the user responds to the psychological assessment guidance task and engages in voice interaction.

[0037] Each time a psychological assessment guidance task is presented to assess the user's psychological state, a guidance screen and / or voice can be generated to guide the user to interact via voice.

[0038] More specifically, each time a psychological assessment guidance task is presented to assess the user's mental state, a guiding screen and / or voice can be generated first to guide the user to breathe in accordance with the frequency. After the user follows the breathing frequency, voice questions are generated (such as "What do you like to do in the past? Do you feel the same when you do these things recently?", "Have you been sleeping well at night recently? Do you feel heavy or don't want to move during the day?"). The user is guided to interact with the voice questions. During this process, the user's voice signals during the voice interaction can be collected, as well as the user's physiological detection signals.

[0039] Among them, the user's physiological detection signals may include, but are not limited to, the user's heart rate, electrodermal activity (EDA), etc.

[0040] Step S103. Perform multi-dimensional feature extraction on the speech signal to obtain multi-dimensional speech features.

[0041] The multidimensional speech features include semantic emotional valence features and acoustic dynamic features, and may also include some speech features that can reflect the user's psychological state.

[0042] In one or more embodiments, multi-dimensional speech features are extracted from the speech signal, which may include, but is not limited to, the following steps S1031-S1036.

[0043] Step S1031. Extract the semantic sentiment valence features of the speech signal through natural speech processing.

[0044] Step S1032. Extract the high-dimensional acoustic representation vector of the speech signal using a pre-trained speech deep neural network model.

[0045] The speech deep neural network model can be, but is not limited to, the Wav2Vec model (a self-supervised learning model proposed by the Facebook AIResearch team, mainly used for speech recognition and speech understanding tasks) or the HuBERT (Hidden-unit BERT, a self-supervised speech representation learning model) model. The high-dimensional acoustic representation vector refers to the latent space feature representation extracted directly from the original speech waveform through training the speech deep neural network model. Unlike traditional shallow acoustic features that rely on manual design (such as fundamental frequency F0 and Mel-frequency cepstral coefficients), this high-dimensional acoustic representation vector is a dense vector that integrates time-frequency information. In terms of physical and pathological dimensions, it can, but is not limited to, implicitly containing the following four types of deep features: 1. Microscopic rhythmic characteristics and emotional tension: It includes subtle intonation shifts, non-linear changes in fundamental frequency perturbations, and dynamic distribution of emotional stress during pronunciation. This feature can bypass the masking of semantic content and directly represent the emotional tension and true emotional coloring at the user's subconscious level.

[0046] 2. Vocal dynamics and muscle control characteristics: It includes information on the smoothness of formant trajectories and the clarity of articulation boundaries. Under psychological stress or depression, the nervous system's control over the muscles of the articulatory organs (such as the vocal cords, lips, and tongue) decreases. This vector can implicitly capture early physiological acoustic representations of "psychomotor retardation," such as unclear pronunciation, prolonged vowels, or difficulty in articulation.

[0047] 3. Implicit sound source and sound quality characteristics: It contains microscopic information reflecting the vocal cord closure state and airflow dynamics. It is used to map changes in vocal cord tension caused by psychological abnormalities, such as hoarseness, breathiness, and deepness in the voice, which are micromechanical alterations in sound quality.

[0048] 4. Temporal context and nonverbal event characteristics: It contains long-term temporal speech dependencies captured by multi-layered self-attention mechanisms. It can encode and contextualize non-verbal sound events such as breathing sounds, non-lexical sighs, unconscious lip smacking, or hesitant pauses on the timeline, thus providing richer information about mental load than simply counting the "total duration of pauses".

[0049] Step S1033. Separate the speech signal into a foreground human voice segment and a background silence segment.

[0050] In one or more embodiments, a Voice Activity Detection (VAD) algorithm can be used to separate a speech signal into a foreground human voice segment and a background silent segment.

[0051] Step S1034. Extract the acoustic dynamic features of the foreground human voice segment.

[0052] The acoustic dynamics features mentioned therein may be, but are not limited to, acoustic energy, fundamental frequency jitter, etc.

[0053] Step S1035. Based on the duration ratio of the speech signal to the background silence segment, determine the psychomotor retardation characteristics.

[0054] Psychomotor retardation can be referred to as the pause ratio, which is a key indicator for measuring silence or pauses in speech.

[0055] The psychomotor retardation can be represented as Rpause=(Ttotal) Tspeech / Ttotal, where Ttotal represents the total duration of the speech signal, Tspeech represents the duration of the foreground human voice segment, and Ttotal... Tspeech represents the duration of the background silence segment.

[0056] Step S1036. The high-dimensional acoustic representation vector, the semantic emotional valence feature, the acoustic dynamics feature, and the psychomotor lag feature are used as the multi-dimensional speech features.

[0057] Step S104. Extract physiological stress indicators that characterize the state of the nervous system from physiological detection signals.

[0058] The physiological emergency indicators may be, but are not limited to, time-domain indicators (SDNN) or frequency-domain indicators of heart rate variability (HRV), and are not specifically limited in the embodiments of this application.

[0059] Step S105. Perform consistency verification on the psychological state indicated by semantic emotional valence features, acoustic dynamic features, and physiological emergency indicators.

[0060] In one or more embodiments, the consistency verification of the semantic emotional valence features, the acoustic dynamic features, and the psychological state indicated by the physiological emergency indicators may include, but is not limited to, the following steps S1051-S1053.

[0061] Step S1051. Determine whether the psychological state indicated by the semantic emotional valence feature, the acoustic dynamic feature, and the physiological emergency index is consistent.

[0062] Step S1052. If the psychological states indicated by the semantic emotional valence features, the acoustic dynamic features, and the physiological emergency indicators are all consistent, then the consistency check is deemed to have passed.

[0063] Step S1053. If the psychological state indicated by the semantic emotional valence feature, the acoustic dynamics feature, and the physiological emergency index is inconsistent, then the consistency check is determined to have failed.

[0064] For example, if the semantic affective valence feature indicates a "positive" or "neutral" psychological state, while the physiological stress indicator is in the high-pressure abnormal range (i.e., a high physiological stress state, indicating an anxious or fatigued psychological state), then the consistency check is deemed to have failed.

[0065] Step S106. If the consistency check fails, determine that there is emotional concealment behavior and generate a concealment punishment factor based on physiological stress indicators.

[0066] If the consistency check fails, it is determined that the user has engaged in emotional concealment behavior. In this case, a concealment penalty factor can be generated based on the physiological emergency indicators through a preset continuous mapping function or neural network model.

[0067] In one or more embodiments, when the physiological emergency indicator is a time-domain indicator of heart rate variability, generating a masking penalty factor may include, but is not limited to, the following steps S1061-S1062.

[0068] Step S1061. When the physiological emergency index is less than the first preset threshold, set the concealment penalty factor to P1.

[0069] Step S1062. When the physiological emergency index is greater than or equal to the first preset threshold and less than the second preset threshold, let the concealment penalty factor be P2, where the first preset threshold is less than the second preset threshold and P1 is greater than P2.

[0070] For example, the first preset threshold can be 30ms, the second preset threshold can be 50ms, the value of P1 can be 1.5, and the value of P2 can be 0.5.

[0071] Step S107. Retrieve the feature weight matrix corresponding to the psychological assessment dimension of the psychological assessment guidance task, and perform weighted fusion of the subjective state assessment value, physiological emergency index and multi-dimensional voice features input by the user to obtain the initial emotion score.

[0072] Different psychological assessment dimensions correspond to different feature weight matrices. The feature weight matrix records the weights of each dimension's features (including multi-dimensional speech features and physiological emergency indicators). The feature weight matrix corresponding to any psychological assessment dimension is obtained by iteratively optimizing the sample user's subjective state assessment value, multi-dimensional speech features of the corresponding sample speech signal, physiological emergency indicators of the sample physiological detection signal, and the initial emotion score of the sample user during the sample psychological assessment guidance task of that psychological assessment dimension through machine learning algorithms.

[0073] Step S108. Based on the concealment penalty factor, the initial emotion score is corrected to obtain the corrected emotion score, and the psychological state assessment result of the user in the psychological assessment dimension to which the psychological assessment guidance task belongs is generated based on the corrected emotion score.

[0074] In one or more embodiments, the correction formula for adjusting the initial sentiment score can be expressed as Scorefinal=max(0,Scorebase-Penalty), where Scorefinal represents the adjusted sentiment score, Scorebase represents the initial sentiment score, Penalty represents the masking penalty factor, and max() represents taking the maximum value.

[0075] The larger the value corresponding to the modified emotion score, the more positive the psychological state assessment result of the user in the psychological assessment dimension to which the psychological assessment guidance task belongs.

[0076] In one or more embodiments, after generating the user's psychological state assessment result for the psychological assessment dimension to which the psychological assessment guidance task belongs based on the modified emotion score, a corresponding script can also be generated based on the modified emotion score.

[0077] In one or more embodiments, after assessing the psychological state across multiple psychological evaluation dimensions, a psychological state profile of the user can be generated, which may include, but is not limited to, the following steps S201-S203.

[0078] Step S201. Based on the dispersion of the corrected emotion scores corresponding to multiple psychological assessment dimensions, determine the user's emotional stability and psychological energy indicators.

[0079] For example, a user's emotional stability index can be determined based on the standard deviation of the modified emotion score corresponding to the most recent consecutive psychological state assessments, and a user's psychological energy index can be determined based on the mean of the multi-dimensional voice features corresponding to the most recent consecutive psychological state assessments.

[0080] Step S202. Generate a circular closed curve and use the emotion stability index to modulate the radial perturbation amplitude of the edge of the circular closed curve to obtain the modulated closed curve.

[0081] In one or more embodiments, the shape of the closed curve can represent the stability of a psychological state (emotion). The rounder the closed curve, the more stable the emotion. The greater the radial perturbation amplitude at the edge of the closed curve, the more unstable the surface emotion.

[0082] Step S203. Map the psychological energy index to warm and cool color temperature values ​​and perform color rendering on the modulated closed curve to obtain a user's psychological state profile.

[0083] In one or more embodiments, blue may be used to represent low psychological energy indicators and orange to represent high psychological energy indicators.

[0084] The psychological state assessment method provided by this invention displays a psychological assessment guidance task based on the subjective state assessment value input by the user and a preset trigger strategy; it collects the user's voice signals and physiological detection signals when the user interacts with the psychological assessment guidance task; it extracts multi-dimensional voice features from the voice signals, including semantic-emotional valence features and acoustic-dynamic features; it extracts physiological emergency indicators representing the state of the nervous system from the physiological detection signals; it performs a consistency check on the psychological state indicated by the semantic-emotional valence features, the acoustic-dynamic features, and the physiological emergency indicators; if the consistency check fails, it determines that there is emotional masking behavior and generates a masking penalty factor based on the physiological emergency indicators; it retrieves the feature weight matrix corresponding to the psychological assessment dimension to which the psychological assessment guidance task belongs, and applies it to the subjective state input by the user. The initial emotion score is obtained by weighted fusion of the subjective state evaluation value, physiological emergency indicators, and multi-dimensional speech features. A modified emotion score is then obtained by correcting the initial emotion score based on a masking penalty factor. Based on the modified emotion score, a psychological state evaluation result for the user in the psychological evaluation dimension of the psychological evaluation guidance task is generated. Different psychological evaluation dimensions correspond to different feature weight matrices. The feature weight matrix for any psychological evaluation dimension is obtained by iteratively optimizing the sample user's subjective state evaluation value input during voice interaction, the multi-dimensional speech features of the corresponding sample speech signal, the physiological emergency indicators of the sample physiological detection signal, and the initial emotion score of the sample user during the psychological evaluation guidance task for that psychological evaluation dimension, using these as training samples. This is achieved through machine learning algorithms. Thus, a cross-modal consistency verification mechanism is introduced into the psychological state evaluation process. When the psychological states indicated by semantic emotional valence features, acoustic dynamic features, and physiological emergency indicators are inconsistent, a masking penalty factor is generated to correct the evaluation results. This solves the problem of poor accuracy in user psychological state evaluation caused by the active masking of emotional expression in "smiling depression," thereby improving the accuracy of user psychological state evaluation. Simultaneously, it supports dynamic weight allocation based on different psychological assessment dimensions. Machine learning algorithms iterate and optimize the weight matrices of features across different psychological assessment dimensions, enabling the assessment of psychological states across various dimensions and achieving accurate multi-dimensional assessment of user psychological states. Furthermore, in terms of data processing, it simultaneously collects user voice signals and physiological detection signals, and uses a voice activity detection algorithm to extract features from both foreground voice segments and background silent segments, achieving multi-dimensional feature extraction and further improving the accuracy of psychological state assessment, facilitating practical application and promotion.

[0085] Please see Figure 2 The second aspect of this application provides a psychological state assessment device, comprising: The triggering unit is used to display the psychological assessment guidance task based on the subjective state evaluation value input by the user and the preset triggering strategy. The acquisition unit is used to acquire the voice signals of the user when responding to the psychological assessment guidance task and the user's physiological detection signals. The first extraction unit is used to extract multi-dimensional features from the speech signal to obtain multi-dimensional speech features, the multi-dimensional speech features including semantic sentiment valence features and acoustic dynamic features; The second extraction unit is used to extract physiological emergency indicators characterizing the state of the nervous system from the physiological detection signals. A consistency verification unit is used to verify the consistency of the semantic emotional valence features, the acoustic dynamic features, and the psychological state indicated by the physiological emergency indicators. The judgment and generation unit is used to determine the presence of emotional concealment behavior when the consistency check fails and to generate a concealment punishment factor based on the physiological emergency indicators. The computation unit is used to retrieve the feature weight matrix corresponding to the psychological assessment dimension of the psychological assessment guidance task, and to perform weighted fusion of the subjective state assessment value input by the user, the physiological emergency index and the features of each dimension in the multi-dimensional voice features to obtain an initial emotion score. The correction and generation unit is used to correct the initial emotion score based on the concealment penalty factor to obtain a corrected emotion score, and generate the psychological state assessment result of the user in the psychological assessment dimension to which the psychological assessment guidance task belongs based on the corrected emotion score. Different psychological assessment dimensions correspond to different feature weight matrices. The feature weight matrix corresponding to any psychological assessment dimension is obtained by iteratively optimizing the subjective state assessment value input by the sample user during voice interaction during the sample psychological assessment guidance task of that psychological assessment dimension, the multi-dimensional voice features of the corresponding sample voice signal, the physiological emergency index of the sample physiological detection signal corresponding to the sample user, and the initial emotion score of the sample corresponding to the sample user as training samples through machine learning algorithms.

[0086] The working process, working details and technical effects of the psychological state assessment device provided in the second aspect of this embodiment can be found in the first aspect of the embodiment, and will not be repeated here.

[0087] like Figure 3 As shown, a third aspect of this application provides an electronic device, including a memory, a processor, and a transceiver that are sequentially and communicatively connected, wherein the memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the psychological state assessment method as described in the first aspect of the embodiment.

[0088] Specifically, the memory may include, but is not limited to, random access memory (RAM), read-only memory (ROM), flash memory, first-in-first-out (FIFO) memory, and / or last-in-first-out (FILO) memory, etc.; the processor may not be limited to microprocessors of the STM32F105 series, ARM (Advanced RISC Machines), x86 architecture processors, or processors with integrated NPU (neural-network processing units); the transceiver may be, but is not limited to, WiFi (Wireless Fidelity) wireless transceivers, Bluetooth wireless transceivers, General Packet Radio Service (GPRS) wireless transceivers, ZigBee (a low-power LAN protocol based on the IEEE 802.15.4 standard), 3G transceivers, 4G transceivers, and / or 5G transceivers, etc.

[0089] This fourth aspect of the embodiment provides a computer-readable storage medium storing instructions containing the psychological state assessment method described in the first aspect of the embodiment. Specifically, the computer-readable storage medium stores instructions that, when executed on a computer, perform the psychological state assessment method as described in the first aspect. The computer-readable storage medium refers to a data storage medium, which may include, but is not limited to, floppy disks, optical disks, hard disks, flash memory, USB flash drives, and / or memory sticks. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.

[0090] The fifth aspect of this embodiment provides a computer program product containing instructions that, when executed on a computer, cause the computer to perform the psychological state assessment method as described in the first aspect of the embodiment, wherein the computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device.

[0091] Finally, it should be noted that the above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for assessing psychological state, characterized in that, include: The psychological assessment guidance task is displayed based on the subjective state evaluation value input by the user and the preset trigger strategy. Collect the user's voice signals and physiological detection signals when the user interacts with the user in response to the psychological assessment guidance task; Multi-dimensional speech features are obtained by performing multi-dimensional feature extraction on the speech signal. The multi-dimensional speech features include semantic sentiment valence features and acoustic dynamic features. Physiological stress indicators characterizing the state of the nervous system are extracted from the physiological detection signals; The consistency of the semantic emotional valence features, the acoustic dynamic features, and the psychological states indicated by the physiological stress indicators is verified. If the consistency check fails, it is determined that there is emotional concealment behavior, and a concealment punishment factor is generated based on the physiological emergency indicators. Retrieve the feature weight matrix corresponding to the psychological assessment dimension of the psychological assessment guidance task, and perform weighted fusion on the subjective state assessment value input by the user, the physiological emergency index and the features of each dimension in the multi-dimensional voice features to obtain the initial emotion score; The initial emotion score is corrected based on the concealment penalty factor to obtain a corrected emotion score, and the psychological state assessment result of the user in the psychological assessment dimension to which the psychological assessment guidance task belongs is generated based on the corrected emotion score. Different psychological assessment dimensions correspond to different feature weight matrices. The feature weight matrix corresponding to any psychological assessment dimension is obtained by iteratively optimizing the subjective state assessment value input by the sample user during voice interaction during the sample psychological assessment guidance task of that psychological assessment dimension, the multi-dimensional voice features of the corresponding sample voice signal, the physiological emergency index of the sample physiological detection signal corresponding to the sample user, and the initial emotion score of the sample corresponding to the sample user as training samples through machine learning algorithms.

2. The psychological state assessment method according to claim 1, characterized in that, Multi-dimensional speech features are obtained by performing multi-dimensional feature extraction on the speech signal, including: The semantic sentiment valence features of the speech signal are extracted through natural speech processing; High-dimensional acoustic representation vectors of the speech signal are extracted using a pre-trained speech deep neural network model. The speech signal is separated into a foreground human voice segment and a background silent segment; Extract the acoustic dynamics features of the foreground human voice segment; Based on the duration ratio of the speech signal to the background silence segment, the psychomotor retardation characteristics are determined. The high-dimensional acoustic representation vector, the semantic emotional valence feature, the acoustic dynamics feature, and the psychomotor lag feature are used as the multi-dimensional speech features.

3. The psychological state assessment method according to claim 2, characterized in that, The psychomotor retardation is characterized by Rpause=(Ttotal) Tspeech / Ttotal, where Ttotal represents the total duration of the speech signal, Tspeech represents the duration of the foreground human voice segment, and Ttotal... Tspeech represents the duration of the background silence segment.

4. The psychological state assessment method according to claim 1, characterized in that, The consistency verification of the semantic-emotional valence features, the acoustic-dynamic features, and the psychological states indicated by the physiological stress indicators includes: Determine whether the psychological state indicated by the semantic emotional valence features, the acoustic dynamic features, and the physiological stress indicators is consistent; If the psychological states indicated by the semantic emotional valence features, the acoustic dynamic features, and the physiological emergency indicators are all consistent, then the consistency check is deemed to have passed. Otherwise, the consistency check is deemed to have failed; The generation of the masking punishment factor based on the physiological stress indicators includes: Based on the aforementioned physiological emergency indicators, a masking penalty factor is generated through a preset continuous mapping function or neural network model.

5. The psychological state assessment method according to claim 1, characterized in that, The psychological assessment guidance task is presented based on a preset trigger strategy, including: When the set daily task trigger time is reached or a user-initiated task trigger request is received, determine whether a new psychological state assessment cycle has been reached and whether the psychological state assessment results of the previous consecutive days are all abnormal. If a new psychological state assessment cycle has not yet been reached and at least one day in the previous consecutive days of psychological state assessment results shows that the psychological state is normal, then the subjective state assessment value entered by the user will trigger and display a psychological assessment guidance task for one psychological assessment dimension. If a new psychological state assessment cycle is reached or the psychological state assessment results for several consecutive days in the past are all abnormal, the subjective state assessment value input by the user will trigger psychological assessment guidance tasks for multiple psychological assessment dimensions. These tasks will be displayed sequentially to collect the user's voice signals and physiological detection signals when responding to the multiple psychological assessment guidance tasks and engaging in voice interaction, in order to conduct psychological state assessments for multiple psychological assessment dimensions.

6. The psychological state assessment method according to claim 5, characterized in that, After conducting psychological state assessments across multiple dimensions, the method further includes: Based on the dispersion of the corrected emotion scores corresponding to multiple psychological assessment dimensions, the user's emotional stability and psychological energy indicators are determined. A circular closed curve is generated, and the radial perturbation amplitude of the edge of the circular closed curve is modulated using the emotional stability index to obtain the modulated closed curve. The psychological energy index is mapped to warm and cool color temperature values, and the modulated closed curve is color-rendered to obtain a user's psychological state profile.

7. The psychological state assessment method according to claim 1, characterized in that, After generating the user's psychological state assessment result for the psychological assessment dimension to which the psychological assessment guidance task belongs based on the modified emotion score, the method further includes: Based on the modified emotion score, a corresponding delivery message is generated.

8. A psychological state assessment device, characterized in that, include: The triggering unit is used to display the psychological assessment guidance task based on the subjective state evaluation value input by the user and the preset triggering strategy. The acquisition unit is used to acquire the voice signals of the user when responding to the psychological assessment guidance task and the user's physiological detection signals. The first extraction unit is used to extract multi-dimensional features from the speech signal to obtain multi-dimensional speech features, the multi-dimensional speech features including semantic sentiment valence features and acoustic dynamic features; The second extraction unit is used to extract physiological emergency indicators characterizing the state of the nervous system from the physiological detection signals. A consistency verification unit is used to verify the consistency of the semantic emotional valence features, the acoustic dynamic features, and the psychological state indicated by the physiological emergency indicators. The judgment and generation unit is used to determine the presence of emotional concealment behavior when the consistency check fails and to generate a concealment punishment factor based on the physiological emergency indicators. The computation unit is used to retrieve the feature weight matrix corresponding to the psychological assessment dimension of the psychological assessment guidance task, and to perform weighted fusion of the subjective state assessment value input by the user, the physiological emergency index and the features of each dimension in the multi-dimensional voice features to obtain an initial emotion score. The correction and generation unit is used to correct the initial emotion score based on the concealment penalty factor to obtain a corrected emotion score, and generate the psychological state assessment result of the user in the psychological assessment dimension to which the psychological assessment guidance task belongs based on the corrected emotion score. Different psychological assessment dimensions correspond to different feature weight matrices. The feature weight matrix corresponding to any psychological assessment dimension is obtained by iteratively optimizing the subjective state assessment value input by the sample user during voice interaction during the sample psychological assessment guidance task of that psychological assessment dimension, the multi-dimensional voice features of the corresponding sample voice signal, the physiological emergency index of the sample physiological detection signal corresponding to the sample user, and the initial emotion score of the sample corresponding to the sample user as training samples through machine learning algorithms.

9. An electronic device, characterized in that, The device includes a memory, a processor, and a transceiver that are sequentially and communicatively connected. The memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the psychological state assessment method as described in any one of claims 1 to 7.

10. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or the instructions are executed by the computer, they implement the psychological state assessment method as described in any one of claims 1 to 7.