Intelligent question and answer interaction method and device for psychological health education, equipment and medium

By constructing a temporal model of potential emotional states based on dialogue data and dynamically adjusting security strategies, the limitations and privacy issues of emotion analysis in existing mental health systems are resolved, enabling continuous tracking of users' emotional states and personalized security protection.

CN121709154APending Publication Date: 2026-03-20ZHONGNAN PRIMARY SCHOOL HECHUAN DISTRICT CHONGQING
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511873020.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing mental health service systems suffer from several limitations in emotion recognition. Their emotion analysis is limited to the surface level and cannot capture continuous changes in users' emotional states. Their security mechanisms lack personalized adaptability, and the collection of biosignals infringes on user privacy, thus restricting the system's widespread adoption and usability.

Method used

By collecting users' historical dialogue data, a potential emotional state temporal model is constructed using a pre-trained language model, variational autoencoder, and recurrent neural network. This model dynamically tracks users' emotional states and adjusts security policy parameters to generate personalized responses, thereby achieving emotion recognition and security protection.

Benefits of technology

Without relying on biosignals, it achieves continuous tracking of users' long-term emotional states and adaptive personalized security strategies, improving the accuracy of emotion recognition and the security protection capabilities of mental health interaction systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121709154A_ABST
    Figure CN121709154A_ABST
Patent Text Reader

Abstract

The invention relates to an intelligent question and answer interaction method and device for psychological health education, equipment and a medium. According to the method, historical dialogue data of a user is collected and subjected to text preprocessing and serialization, and an instant emotion vector is extracted based on a pre-training language model to capture current emotion expression of the user; a potential emotional state is decoupled from a dialogue history by using a time sequence model constructed by a variational auto-encoder and a recurrent neural network, a cross-dialogue evolution rule of the potential emotional state is tracked to generate a state vector sequence, a historical emotional baseline is calculated on the basis, and a risk detection threshold and a response generation strategy are dynamically adjusted through statistical deviation analysis; finally, a personalized response is generated in combination with real-time user input and the emotional state vector, continuous tracking of the long-term emotional state of the user and personalized security strategy self-adaption in a pure dialogue environment independent of biological signals are achieved, and the accuracy of emotion recognition and the security protection capability of a mental health interaction system are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of artificial intelligence, and particularly relates to an intelligent question and answer interaction method and device for mental health education, equipment and a medium. BACKGROUND

[0002] In the current field of mental health services, intelligent question and answer systems have become an important tool for providing instant psychological support and education. These systems are usually based on natural language processing technology and can analyze and respond to user input text in real time. However, there are several significant technical bottlenecks in the practical application of existing technologies. The emotional recognition capabilities of most systems remain at the surface analysis level, and only classify the emotions of the user's current single round of dialogue. This "instant" analysis is like a snapshot and cannot capture the continuous change process of the user's emotional state. In fact, the user's real psychological state is often latent and continuously evolving, such as the accumulation of depressive emotions or the fluctuation of anxiety levels, which have obvious time sequence characteristics, and the existing single round analysis mode cannot grasp this long-term evolution rule.

[0003] Another outstanding problem is the lack of personalized adaptability of the system's security protection mechanism. Existing risk detection models usually use fixed sensitivity thresholds, regardless of how the user's emotional state changes, the system uses a uniform judgment standard. This rigid processing method can easily produce two extreme cases: when the user's emotional state deteriorates continuously, the system may not be able to identify the risk in time because the threshold is set too high; while when the user is in normal emotional fluctuations, it may produce false positives because the threshold is too sensitive. Both of these situations can seriously affect user experience and system credibility.

[0004] In addition, in order to improve the accuracy of emotion recognition, some systems have begun to try to introduce multiple biological signal data, such as capturing facial expressions through a camera and monitoring heart rate changes through wearable devices. Although this multi-modal fusion method has improved the recognition accuracy to some extent, it has also brought new problems. The collection of biological signals requires special hardware device support, which not only increases the deployment cost of the system, but more importantly involves the sensitive issue of user privacy protection. In many application scenarios, users have concerns about the collection of personal physiological data, and this technical route actually limits the popularity and usability of the system.

[0005] Therefore, the industry urgently needs a technological solution that can achieve deep emotional state tracking and personalized security protection solely through pure dialogue interaction, while protecting user privacy. This solution needs to overcome the limitations of existing technologies, not only accurately understanding the user's current emotional expression but also uncovering potential patterns of emotional evolution from continuous dialogue history, and establishing an adaptive security mechanism based on this. The ideal technological path should achieve a deep understanding and intelligent response to the user's psychological state through more advanced algorithmic models, without increasing additional hardware burden or infringing on user privacy. Summary of the Invention

[0006] Therefore, it is necessary to provide an intelligent question-and-answer interaction method, device, equipment, and medium for mental health education to address the aforementioned technical problems.

[0007] Firstly, this application provides an intelligent question-and-answer interaction method for mental health education, including:

[0008] S1. Collect users' historical dialogue data, perform text cleaning and word segmentation on the historical dialogue data, and generate serialized user input data; based on the serialized user input data, perform emotion classification processing through a pre-trained language model to obtain the instant emotion vector of each user input.

[0009] S2. Based on serialized user input data and real-time emotion vectors, latent emotion state temporal modeling is performed using a pre-trained variational autoencoder and recurrent neural network model to construct a latent emotion state temporal model; the latent emotion state temporal model is used to dynamically track and update the user's latent emotion state across sessions, generating a sequence of latent emotion state vectors.

[0010] S3. Based on the potential emotional state vector sequence, calculate the user's historical emotional baseline, including the historical mean vector and the historical standard deviation vector; dynamically adjust the safety policy parameters according to the statistical deviation between the current potential emotional state vector in the potential emotional state vector sequence and the historical emotional baseline; wherein, the safety policy parameters include risk detection threshold and response generation strategy;

[0011] S4. Based on the current user input, the immediate emotion vector, and the current potential emotion state vector, combined with the adjusted safety policy parameters, a personalized response is generated through a neural generative model or template library.

[0012] Secondly, this application also provides an intelligent question-and-answer interactive device for mental health education, used to implement the method described in the first aspect, the device comprising:

[0013] The user emotion perception preprocessing module is configured to collect historical dialogue data of the user, perform text cleaning and word segmentation processing on the historical dialogue data, and generate serialized user input data; based on the serialized user input data, perform emotion classification processing through a pre-trained language model to obtain an instant emotion vector of each user input;

[0014] The emotion state dynamic tracking module is configured to perform latent emotion state time series modeling on the serialized user input data and the instant emotion vector through a pre-trained variational autoencoder and a recurrent neural network model, and construct a latent emotion state time series model; and perform dynamic tracking and updating on a latent emotion state of the user across sessions by using the latent emotion state time series model, and generate a latent emotion state vector sequence.

[0015] The security policy optimization module is configured to calculate a user historical emotion baseline including a historical mean vector and a historical standard deviation vector based on the latent emotion state vector sequence; and dynamically adjust security policy parameters according to a statistical deviation of a current latent emotion state vector in the latent emotion state vector sequence from the historical emotion baseline, wherein the security policy parameters include a risk detection threshold and a response generation strategy.

[0016] The personalized response generation engine is configured to generate a personalized response through a neural generation model or a template library based on the current user input, the instant emotion vector, and the current latent emotion state vector, and in combination with the adjusted security policy parameters.

[0017] In a third aspect, the present application further provides a computer device, comprising a memory and a processor, the memory stores a computer program, and the processor implements the intelligent question and answer interaction method for mental health education according to the first aspect when executing the computer program.

[0018] In a fourth aspect, the present application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the intelligent question and answer interaction method for mental health education according to the first aspect.

[0019] The intelligent question and answer interaction method, device, equipment and medium for mental health education provided by the application, through collecting historical dialogue data of a user and performing text preprocessing and serialization on the historical dialogue data, extract an instant emotion vector based on a pre-trained language model to capture the current emotional expression of the user, and then use a time sequence model constructed by a variational autoencoder and a recurrent neural network to decouple the potential emotional state from the dialogue history and track the cross-session evolution rule to generate a state vector sequence, on the basis of which, a historical emotion baseline is calculated and a risk detection threshold and a response generation strategy are dynamically adjusted through statistical deviation analysis, and finally, personalized responses are generated in combination with real-time user input and the emotion state vector, realizing continuous tracking of the long-term emotional state of the user and self-adaptation of the personalized safety strategy in a pure dialogue environment without relying on biological signals, and effectively improving the accuracy of emotion recognition and the safety protection capability of the mental health interaction system. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed to be used in the embodiments or the related art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0021] Figure 1 A flowchart of an intelligent question and answer interaction method for mental health education provided by the present application is shown in the figure.

[0022] Figure 2 A flowchart of generating a potential emotional state vector sequence in an alternative embodiment of the present application is shown in the figure.

[0023] Figure 3 A structural diagram of an intelligent question and answer interaction device for mental health education provided by the present application is shown in the figure. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical solutions and advantages of the present application more clear, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not to limit the present application.

[0025] REFERENCE Figure 1 A flowchart of an intelligent question and answer interaction method for mental health education provided by the present application is shown in the figure. The method comprises the following steps:

[0026] S1, collect historical dialogue data of a user, perform text cleaning and word segmentation processing on the historical dialogue data, and generate serialized user input data; based on the serialized user input data, perform emotion classification processing through a pre-trained language model to obtain an instant emotion vector of each user input.

[0027] Specifically, when collecting the historical dialogue data of the user, the dialogue records of the user in different time periods and different interaction scenarios are covered, including text input content, dialogue initiation time, dialogue duration, and other associated information, to ensure the integrity and time sequence continuity of the data. Text cleaning processing is carried out on the interference information that may exist in the historical dialogue data, specifically including removing meaningless special characters (such as garbled symbols, continuous repeated punctuation), filtering redundant spaces and line breaks, correcting misspelled words and syntax errors, and eliminating system prompt information (such as "please enter your question" which is not user-generated content) unrelated to the dialogue content, to ensure that the cleaned data only retains the text content expressed by the user. The word segmentation processing adopts a method combining dictionary matching and statistical learning, loads a general Chinese word segmentation dictionary and a psychology field exclusive word segmentation dictionary (covering psychological terms, emotional description words, etc.), and then splits the cleaned single user dialogue text into discrete word sequences through a word segmentation algorithm, each word corresponding to a unique word identifier; subsequently, according to the time sequence of the dialogue, the word sequences corresponding to each user dialogue are arranged in turn to form time-dimensioned serialized user input data, and the data format needs to meet the input requirements of the subsequent pre-trained language model, that is, each data contains a word identifier sequence and corresponding timestamp information.

[0028] In the emotion classification processing based on the serialized user input data, a pre-trained language model (such as BERT, RoBERTa, etc.) on a large-scale text corpus is selected as a basic model, and the model input format conversion is performed on each user input text in the serialized user input data, the word sequence is mapped to a word embedding vector sequence recognizable by the model, a position embedding vector is introduced to represent the position information of the word in the text, and a segment embedding vector is introduced to distinguish different user input texts. The three embedding vectors are superimposed to form the input feature vector of the model. The core of the emotion classification processing is to fine-tune the output layer of the pre-trained language model, and an emotion classification layer is added after the last fully connected layer of the model. The output dimension of the classification layer is consistent with the number of pre-set emotion categories (such as joy, anxiety, depression, calm, etc.). In the model training process, the mental health dialogue corpus with labeled emotion categories is used as the training set, the difference between the predicted emotion category and the real emotion category is calculated by using the cross-entropy loss function, the gradient descent algorithm is used to iteratively optimize the model parameters, and the emotion classification accuracy of the model on the validation set is stabilized until the threshold value is reached. After training, each user input text in the serialized user input data is input into the fine-tuned pre-trained language model, and the model output layer will generate a probability distribution of the corresponding text in each emotion category. The probability distribution is converted into a fixed-dimensional vector form, which is the immediate emotion vector of each user input. Each dimension of the vector corresponds to a probability value of an emotion category, which can directly reflect the surface emotion state expressed by the user's current input text.

[0029] S2, based on the serialized user input data and the immediate emotion vector, the latent emotion state time series modeling is performed through the pre-trained variational autoencoder and the recurrent neural network model, and the latent emotion state time series model is constructed; the latent emotion state time series model is used to dynamically track and update the user's cross-session latent emotion state, and a latent emotion state vector sequence is generated.

[0030] Specifically, in the construction of the latent emotion state time series model based on the serialized user input data and the immediate emotion vector, first, the two kinds of data are fused, the serialized word embedding vector (extracted from the serialized user input data, obtained by the word embedding layer of the pre-trained language model) corresponding to each user input is spliced with the immediate emotion vector of the user input to form a fusion feature vector with a dimension of (word embedding vector dimension + immediate emotion vector dimension). Each fusion feature vector corresponds to the user input information at a time step, and the fusion feature vector sequence is formed after being arranged in time sequence, which is used as the input data of the subsequent model.

[0031] The pre-trained variational autoencoder (VAE) plays a crucial role in latent space learning during the modeling process. A VAE consists of an encoder and a decoder. The encoder takes each fused feature vector from the fused feature vector sequence as input, maps it to the latent space through a two-layer fully connected neural network, and outputs the mean vector of the latent variables. and standard deviation vector ,in and All dimensions are preset latent space dimensions. , representing the central location and dispersion of the latent variable in each dimension of the space. Based on the principle of variational inference, the mean is obtained by using the reparameterization technique. Standard deviation is The latent vector is obtained by sampling from the normal distribution. ,Right now ,in These are random variables that follow a standard normal distribution (mean 0, variance 1). This process ensures the randomness and differentiability of the latent vector, facilitating backpropagation optimization of the model. The decoder then uses the sampled latent vector... As input, it is mapped back to the dimension of the original fused feature vector through a two-layer fully connected neural network, and the output is the reconstructed fused feature vector. During model training, the reconstruction loss (such as mean squared error loss) is calculated. Combined with the original fused feature vector The differences are investigated, and the distribution of latent variables is constrained to approximate the standard normal distribution through KL divergence loss. The weighted sum of the two losses is used as the total loss function of VAE. The parameters of encoder and decoder are optimized through gradient descent algorithm to complete the pre-training of VAE.

[0032] Recurrent Neural Network (RNN) models employ Long Short-Term Memory (LSTM) networks or Gated Recurrent Units (GRUs) to handle the temporal dependencies of data. Their input is the sequence of latent vectors output by the VAE (in chronological order). The LSTM model controls the transmission and updating of information in the cell state through gating mechanisms: input gate, forget gate, and output gate. The input gate determines which information from the current latent vector needs to be updated to the cell state; the forget gate determines which information from historical cell states needs to be retained; and the output gate determines which information from the cell state needs to be output as the hidden state at the current time step. Specifically, let the... The potential vector at time step is LSTM input gate output Forgot Gate Output Cell state update Updated cell state Output gate output , No. Hidden state of time step ;in, , , These are the weight matrix and bias term of the input gate, respectively. , , These are the weight matrix and bias term of the forget gate, respectively. , , These are the weight matrix and bias term for cell state updates, respectively. , , These are the weight matrix and bias term of the output gate, respectively. is the sigmoid activation function, and tanh is the hyperbolic tangent activation function. For element-wise multiplication, For the first The hidden state of the time step. For the first Cellular state at time step.

[0033] A pre-trained VAE and an LSTM model are jointly trained to form a temporal model of latent emotional states. The VAE maps fused feature vectors to the latent space, capturing latent information in the user input; the LSTM performs temporal modeling on the sequence of latent vectors, learning the dependencies between latent vectors at different time steps. During model training, the total loss of the VAE and the temporal prediction loss of the LSTM (e.g., based on...) are used. Predict the next time step The weighted sum of the mean squared error loss is used as the joint loss function to optimize the overall parameters of the model.

[0034] When using this model to dynamically track and update a user's potential emotional state across sessions, each time a user initiates a new session, the user input in the new session is first processed into serialized user input data and an immediate emotional vector using method S1, and then fused into a new fused feature vector sequence, which is input into the potential emotional state time series model. The model first maps the new fused feature vector into a new potential vector through a VAE, and then inputs this potential vector into the current time step of the LSTM. Combining the LSTM hidden state and cell state of the last time step of the previous session, the hidden state of the current time step is calculated, which is the potential emotional state vector of the current user. As the user session continues, the model repeats the above process for each new time step of user input, updating the potential emotional state vector, and finally forming a sequence of potential emotional state vectors arranged in chronological order. This sequence can completely reflect the evolution trend of the user's potential emotional state across sessions.

[0035] S3. Based on the potential emotional state vector sequence, calculate the user's historical emotional baseline, including the historical mean vector and the historical standard deviation vector; dynamically adjust the safety policy parameters according to the statistical deviation between the current potential emotional state vector in the potential emotional state vector sequence and the historical emotional baseline; wherein, the safety policy parameters include risk detection threshold and response generation strategy.

[0036] Specifically, when constructing a user's historical sentiment baseline, the first step is to determine the time window range for baseline calculation. This range can be selected from the potential sentiment state vectors corresponding to the user's most recent N complete conversations (the value of N can be set according to the conversation frequency and data volume in the actual application scenario; it is generally recommended to have no less than 5 conversations to ensure statistical validity). This vector sequence is denoted as... ,in (d is the dimension of the latent emotional state vector, determined by the dimension of the latent layer of the variational autoencoder in S2). Historical mean vector The calculation logic is to calculate the arithmetic mean of the corresponding dimensions of all vectors in the sequence, and the formula is expressed as:

[0037]

[0038] In the formula, The value of each dimension represents the user's historical average level on that emotional characteristic dimension; This means summing the N potential emotional state vectors element-wise along their dimensions and then dividing by N to obtain the mean.

[0039] Historical standard deviation vector The calculation is used to reflect the degree of fluctuation of a user's historical sentiment across various dimensions, and the formula is:

[0040]

[0041] In the formula, ; The difference in dimensionality between the i-th potential emotional state vector and the historical mean vector is represented by element-wise squaring. Amplify the degree of deviation, then sum the squared differences of all vectors along their dimensions and divide by . (Using unbiased estimation to avoid bias when the sample size is small), and finally restoring to the original data scale through square root operation to obtain the standard deviation of each dimension.

[0042] Statistical bias calculation uses Euclidean distance to measure the current potential emotional state vector. The degree of deviation from the historical sentiment baseline is calculated using the following formula:

[0043]

[0044] In the formula, These are statistical deviation values, without units. This represents the value of the k-th dimension of the current potential emotional state vector; Let k be the value of the k-th dimension of the historical mean vector; Let be the value of the k-th dimension of the historical standard deviation vector; It is a very small positive number (which can be taken as...) ), used to avoid A calculation error occurs when the denominator approaches 0; the denominator part... Standardization eliminates the impact of differences in the dimensions of different emotion dimensions on the deviation calculation. The numerator is the difference between the current dimension value and the historical mean. The final deviation is obtained by squaring, summing, and then taking the square root.

[0045] Security policy parameter adjustment rules are based on statistical deviation values. Deviation threshold from preset , ( The comparison was determined by domain experts in conjunction with the emotional risk level classification standards: when When this occurs, it indicates that the user's current emotional state is within the historical normal fluctuation range, at which point the risk detection threshold is adjusted. Adjust to a higher value (e.g., maintain 1.1 times the initial threshold), and the response generation strategy should focus on mental health education content, emphasizing knowledge dissemination and positive guidance; when When a user's emotions deviate to a moderate degree, the risk detection threshold can be adjusted. The response generation strategy is adjusted to 0.8-1.0 times the initial threshold, and an emotional empathy module is added to express understanding and concern through verbal communication. When user sentiment deviates significantly from historical baselines, there is a potential risk, and a risk detection threshold can be set. Further reduce the threshold to 0.5-0.8 times the initial threshold, and trigger the risk intervention module in the response generation strategy. Include emotional guidance suggestions in the response and prepare an interface to trigger manual intervention (if the risk continues to rise in subsequent conversations, it will be activated).

[0046] S4. Based on the current user input, the immediate emotion vector, and the current potential emotion state vector, combined with the adjusted safety policy parameters, a personalized response is generated through a neural generative model or template library.

[0047] Specifically, when generating personalized responses based on the current user input, the immediate emotion vector, the current potential emotion state vector, and the adjusted safety policy parameters, two implementation paths are adopted: a neural generative model and a template library. The two paths can be flexibly switched or combined according to the response speed requirements and generation quality requirements of the actual application scenario.

[0048] For the neural generative model path, the first step is to construct an input fusion layer to convert the current user input text into text vectors. (Implemented through the text encoding layer of a pre-trained language model, such as BERT's [CLS] vector, with dimensions consistent with the immediate sentiment vector), then... Real-time emotion vector Current potential emotional state vector Perform vector concatenation to obtain the fused input vector. (";" indicates concatenation by column) The dimension is the sum of the three dimensions. The input is fed into a Transformer-based generative model, with adjusted security policy parameters introduced as constraints. At the masking layer of the attention mechanism, the risk detection threshold is applied. Adjust the attention weight for words related to negative emotions (when) When the risk level is low, the attention weight for risky words such as "depressed" and "despair" is increased to ensure that the response specifically covers the risk points. At the generation probability distribution layer, the style bias of the generated text is adjusted according to the response generation strategy (e.g., under the empathy strategy, the generation probability of empathetic phrases such as "I understand your feelings" and "I can feel your distress" is increased; under the education strategy, the generation probability of popular science phrases such as "According to psychological research" and "Common adjustment methods include" is increased). During the generation process, the Beam Search algorithm is used to optimize the coherence and accuracy of the generated text. The perplexity index is used to evaluate the rationality of the generated text in real time. If the perplexity exceeds a preset threshold (which can be 50), a regeneration mechanism is triggered, and the attention weight is adjusted before regeneration until the perplexity requirement is met.

[0049] For the template library path, a multi-dimensional template classification system is first constructed, with templates divided into three dimensions: "emotion type - problem type - strategy type". The emotion type dimension corresponds to the classification result of the immediate emotion vector (e.g., anxiety, depression, calmness, etc.); the problem type dimension corresponds to the current user's input problem category (e.g., emotion regulation counseling, psychological knowledge inquiry, venting about life troubles, etc., classified through a text classification model). (Classified); the strategy type dimension corresponds to the adjusted response generation strategy (e.g., educational, empathic, intervention). Each template in the template library contains a fixed text segment and dynamic parameter placeholders. For example, the empathic-anxiety-emotion regulation counseling template is: "The anxiety you are currently experiencing is very common. Based on your recent emotional changes, we suggest trying {regulation method}. This method can help you alleviate anxiety in {scenario description}. Would you like to talk more about your current specific feelings?", where "{regulation method}" and "{scenario description}" are dynamic parameter placeholders.

[0050] In the template matching and parameter filling stage, the first step is to select the 3-5 candidate templates with the highest matching degree from the template library based on the three-dimensional labels of "emotion type-problem type-strategy type" (matching degree is calculated by label similarity, using the Jaccard coefficient, formula is...). Where A is the set of 3D labels corresponding to the input data, and B is the set of 3D labels for the template (selecting the candidate template with the highest Jaccard coefficient). Then, based on the current potential emotion state vector... Parameters are populated against a historical sentiment baseline: for example, the selection of "{modulation method}", if... If the "physiological arousal" value is higher than the historical average, then the "deep breathing relaxation method" (for high physiological arousal) is filled in; if the "cognitive entanglement" value is higher than the historical average, then the "cognitive reconstruction method" (for cognitive anxiety) is filled in; the "{scenario description}" is filled in based on high-frequency scenarios mentioned in the user's historical dialogue data (such as work scenarios, academic scenarios), ensuring that the parameters match the user's actual situation. Finally, the filled template text undergoes grammar validation and fluency optimization. Grammatical errors are corrected through the text correction module of a pre-trained language model, and sentence vector similarity calculation (comparing with the sentence vector of the user's input text) ensures that the response is relevant to the user's needs, ultimately outputting personalized response text.

[0051] The aforementioned intelligent question-and-answer interaction method for mental health education collects users' historical dialogue data, preprocesses and serializes it, extracts real-time emotion vectors based on a pre-trained language model to capture users' current emotional expression, and then uses a temporal model constructed from variational autoencoders and recurrent neural networks to decouple potential emotional states from the dialogue history and track their cross-conversation evolution to generate a state vector sequence. Based on this, a historical emotion baseline is calculated, and risk detection thresholds and response generation strategies are dynamically adjusted through statistical bias analysis. Finally, personalized responses are generated by combining real-time user input and emotion state vectors. This method achieves continuous tracking of users' long-term emotional states and adaptive personalized security strategies in a pure dialogue environment that does not rely on biosignals, effectively improving the accuracy of emotion recognition and the security protection capabilities of mental health interaction systems.

[0052] refer to Figure 2 In one optional embodiment, based on serialized user input data and real-time emotion vectors, a potential emotion state temporal model is constructed by performing temporal modeling of the potential emotion state through a pre-trained variational autoencoder and recurrent neural network model; the potential emotion state temporal model is then used to dynamically track and update the user's potential emotion state across sessions, generating a sequence of potential emotion state vectors, including the following steps:

[0053] S11. For each user input in the serialized user input data, calculate the average word embedding through the pre-trained word embedding model to generate an input vector; based on the input vector and the latent emotional state vector of the previous time step, perform nonlinear transformation through the VAE encoder network to output the mean vector and variance vector of the latent variables.

[0054] Specifically, for each user input text in the serialized user input data, each word in the text is first processed by a pre-trained word embedding model (such as the word-level embedding layer of Word2Vec, GloVe, or BERT) to obtain the word embedding vector corresponding to each word (the dimension is set according to the pre-trained model). Then, the average word embedding of the user input text is calculated by averaging the word embedding vectors of all words in the text element-wise to generate the input vector corresponding to the user input (the dimension is consistent with the word embedding vector). This process can convert the variable-length text sequence into a fixed-dimensional vector to meet the input requirements of the subsequent model.

[0055] After obtaining the input vector, the input to the VAE encoder is constructed by combining it with the latent emotion state vector from the previous time step. If the current input is the first user input (without a previous latent emotion state vector), the latent emotion state vector from the previous time step is initialized as a zero vector with the same dimension as the input vector; if data from the previous time step exists, the latent emotion state vector generated at that time step is directly used. The input vector is concatenated with the latent emotion state vector from the previous time step to form the fused input vector of the VAE encoder, which is then input into the pre-trained VAE encoder network.

[0056] The VAE encoder network consists of multiple fully connected layers and activation functions (such as ReLU and LeakyReLU). It extracts and compresses features from the fused input vector through nonlinear transformations. The first fully connected layer maps the fused input vector to an intermediate feature space (the dimension of which can be 1.5-2 times the dimension of the input vector). After the activation function introduces nonlinearity, two parallel fully connected layers output the mean vectors of the latent variables. ) and variance vector ( — The mean vector represents the central location of the latent variable distribution, and the variance vector represents the dispersion of the latent variable distribution. Both have the same dimension and are determined by the dimension of the latent layer during VAE model training, ensuring that the probability distribution of the latent variable can be constructed based on these two vectors.

[0057] S12. Based on the mean vector and variance vector, the current latent variables are sampled using the reparameterization technique to obtain the latent variable sequence; the sampling formula for the current latent variables is:

[0058]

[0059] in, Denotes the current latent variable at time t. Let represent the mean vector at time t. Let represent the variance vector at time t. Describes a random vector sampled from a standard normal distribution. This represents element-wise multiplication.

[0060] Specifically, when obtaining the mean vector at time t ( ) and variance vector ( After that, the reparameterization technique is used to modify the current latent variables ( Sampling is performed by transferring the randomness of the sampling process to a random vector. This enables the sampling operation to be differentiable, thereby meeting the gradient backpropagation requirements during VAE model training.

[0061] In the sampling formula, The dimension and mean vector of the current latent variable obtained by sampling at time t. ), variance vector ( The value of the input is consistent with the latent semantic and emotional fusion features corresponding to the current user input; The mean vector of the VAE encoder output at time t determines the central tendency of the latent variable distribution; Let ε be the variance vector of the VAE encoder output at time t, controlling the dispersion of the latent variable distribution; a larger value indicates higher uncertainty of the latent variable. ϵ represents the variance vector from the standard normal distribution (…). The random vector obtained by independent sampling in ) has dimensions and , Consistency, its randomness introduces reasonable uncertainty into the latent variables, avoiding model overfitting; ⊙ represents element-wise multiplication, i.e. , After multiplying by the corresponding dimension elements of ϵ respectively, then... The product of ϵ and Performing element-wise addition, we eventually obtain .

[0062] For each time step in the serialized user input data (from the first user input to the last user input), the above sampling process is repeated: for each time step t, based on that time step... and Sampling yields the corresponding , all time steps Arranging them in chronological order will generate a sequence of latent variables. T represents the total number of time steps for serializing user input data. This sequence contains latent feature information corresponding to each user input, providing continuous input data for subsequent temporal modeling by the LSTM network, ensuring that the temporal model can capture the evolution of the user's emotional state as the dialogue progresses.

[0063] S13. Based on the latent variable sequence, time series modeling is performed using an LSTM network, and cell states and hidden states are updated, with the hidden states used as the latent emotional state vector.

[0064] Specifically, before conducting temporal modeling with the LSTM network, network pre-training and structure configuration are completed. The dimensions of the latent variables are matched, and the hidden layer dimensions are set according to the needs of emotion representation (64 or 128 dimensions). During training, a mental health dialogue dataset with emotional temporal labels is used, aiming to minimize the error in emotion state prediction. During modeling, the sequence of latent variables is first input into the LSTM in chronological order, while an immediate emotion vector is introduced. The actual input to the LSTM is formed by concatenating these vectors. This approach can integrate latent semantics of the text with surface emotional features, improving the temporal modeling effect.

[0065] LSTM achieves selective memorization and updating of temporal information through forget gates, input gates, and output gates. Its core lies in the dynamic adjustment of cell states and hidden states: the cell state serves as the storage carrier of long-term temporal information, and its update requires combining previous cell states with the new information input at present. The formula is as follows:

[0066]

[0067] In the formula, This is the current cell state vector. This is the preceding cell state vector. (Forget gate output) controls the retention ratio of previous information. (Input gate output) controls the proportion of new information that is included. (Candidate cell state) is the new emotional feature generated from the current input. This is element-wise multiplication.

[0068] The hidden state then filters key information from the cell state and outputs it as the potential emotional state vector for the current time step, using the following formula:

[0069]

[0070] In the formula, This is the current hidden state vector (i.e., the potential emotional state vector). (Output gate output) controls the output ratio of cell state information. Cell state values ​​are mapped to the [-1,1] interval to avoid numerical overflow.

[0071] S14. Repeat steps S11 to S13 to generate a sequence of potential emotional state vectors for the entire dialogue history.

[0072] Specifically, this step iteratively executes the aforementioned temporal modeling steps to cover all user interaction data in order to generate a continuous sequence of potential emotional state vectors. The "dialogue history" contains all user input data across sessions. If it is the first interaction, only the current session data is covered, with the time step increasing from 1 to the total number of rounds in the current session. If there are historical interactions, all historical session data are integrated, and a global time step index is established according to the interaction time order to ensure temporal continuity.

[0073] The iterative process follows these fixed rules: During the first interaction (global time step t=1), since there is no prior latent emotion state vector, the initial vector is set to an all-zero vector or the default hidden state vector of the pre-trained LSTM to avoid bias caused by random initial values. In subsequent time steps (t≥2), when performing the modeling step, the latent emotion state vector generated in the previous time step is called, fused with the current input, and used in the calculation to ensure continuous transmission of temporal information. After each time step's calculation is completed, the generated latent emotion state vector is stored in a cache according to a "time step-session ID" key-value pair structure for easy subsequent retrieval.

[0074] Once the iteration covers all dialogue rounds, the computation stops and the validity of the generated vector sequence is verified: the temporal continuity is judged by calculating the mean Euclidean distance between vectors at adjacent time steps (if the mean is too large, the LSTM parameters need to be checked), and the consistency between the sequence and the user's actual emotion annotation (if any) is compared by cosine similarity (e.g., the similarity should not be less than 0.7). If the verification fails, the parameters in the modeling stage are corrected and the iteration is restarted until a vector sequence that meets the requirements is generated.

[0075] In one optional embodiment, time-series modeling is performed using an LSTM network based on the latent variable sequence, and cell states and hidden states are updated, including the following steps:

[0076] S21. Based on the current latent variables and the hidden state of the previous time step, calculate the input gate vector through the LSTM input gate; the formula for calculating the input gate vector is:

[0077]

[0078] in, This represents the input gating vector at time t. Denotes the current latent variable at time t. This represents the hidden state at time t-1. and This represents the input gate weight matrix. This represents the input gate bias vector. This represents the sigmoid function.

[0079] Specifically, this step revolves around the fusion calculation of current latent variables and historical hidden states. The core is to determine the proportion of new information about the cell state to be included through the input gate. First, the current latent variables at time t are obtained. (Generated from previous reparameterized sampling, carrying the latent semantic and emotional features of the current input) and the hidden state at time t-1. (Storing the temporal information of the previous time step, reflecting the user's previous emotional state), after linearly transforming both with their corresponding weight matrices, the bias vector is superimposed and processed by the sigmoid activation function to obtain the input gating vector. The specific formula is as follows: .

[0080] In the formula, For the current latent variables The input gate weight matrix is ​​used to adjust... The degree of influence on the input gating result is determined by the following dimensions: The dimension and the gating vector dimension are jointly determined (if) If the dimension is d and the gate vector is h, then (in h×d dimensions). Preorder hidden state The input gate weight matrix is ​​used to integrate the influence of historical time series information on the current gate control. It has a dimension of h×h (h is the dimension of the hidden state). The input gate bias vector is used to offset the numerical offset after the linear transformation, and its dimension is h. The sigmoid activation function maps the calculation result to the interval [0,1], where The closer the value is to 1, the more new information carried by the current latent variable is allowed to enter the subsequent cell state update; the closer it is to 0, the more it inhibits the inclusion of new information, thus achieving selective screening of new information.

[0081] S22. Based on the current latent variables and the hidden state of the previous time step, calculate the forgetting gate vector using the LSTM forgetting gate; the formula for calculating the forgetting gate vector is:

[0082]

[0083] in, This represents the forgetting gating vector at time t. and This represents the forget gate weight matrix. This represents the forget gate bias vector.

[0084] Specifically, the core function of this step is to filter the historical cell state information that needs to be retained, avoiding invalid historical data from interfering with the current emotional state modeling. The computational logic is similar to that of the input gate, using time t... and at time t-1 Based on this, a linear transformation is performed using a weight matrix and bias vector specific to the forgetting gate, followed by a sigmoid activation function to output the forgetting gate vector. The formula is: .

[0085] In the formula, for The corresponding forget gate weight matrix, with dimensions and Consistency is used to measure the impact of current latent variables on historical information retention decisions (e.g., when...). When reflecting new changes in user sentiment, (The weighting of old emotional information can be adjusted.) for The corresponding forget gate weight matrix has a dimension of h×h and its function is to determine the validity of the historical cell state by combining the historical hidden state. This is the forget gate bias vector, with dimensions h, used to correct the results of the linear transformation; Also using the sigmoid activation function, The value falls within the interval [0,1] A value close to 1 means that more of the cell state at time t-1 is preserved. The cell state stores only valuable temporal information (such as the user's long-term stable emotional characteristics). When the cell state is close to 0, it means that most of the historical information (such as the user's short-term negative emotions that have been alleviated) has been forgotten, ensuring that the cell state only stores valuable temporal information.

[0086] S23. Based on the current latent variables and the hidden state of the previous time step, calculate the output gating vector through the LSTM output gate; the formula for calculating the output gating vector is:

[0087]

[0088] in, This represents the output gating vector at time t. and This represents the output gate weight matrix. This represents the output gate bias vector.

[0089] Specifically, this step aims to determine the proportion of information in the cell state that needs to be transmitted to the hidden state, laying the foundation for the subsequent generation of potential emotional state vectors. The calculation still uses... and As input, the output gate vector is generated through the combined action of the weight matrix, bias vector, and sigmoid function of the output gate. The formula is: .

[0090] In the formula, for The corresponding output gate weight matrix has the same dimensions as... Consistency is used to control the influence of the current latent variable on the output screening results (e.g., when...). When displaying user emotional fluctuations (Adjustable whether to prioritize outputting cell state information related to fluctuations). for The corresponding output gate weight matrix has a dimension of h×h and its function is to determine the output priority of cell state information by combining the historical hidden state. The output gate bias vector has a dimension of h and is used to balance the values ​​after the linear transformation. sigmoid activation makes The value is in the interval [0,1]. The closer to 1, the more cell state information is allowed to be transmitted to the hidden state; the closer to 0, the less is transmitted, thus achieving the filtering of cell state information output.

[0091] S24. Based on the current latent variables and the hidden state of the previous time step, calculate the candidate cell state values ​​using the tanh activation function; the formula for calculating the candidate cell state values ​​is:

[0092]

[0093] in, This represents the candidate values ​​for the cell state at time t. and Represents the candidate state weight matrix. This represents the candidate state bias vector.

[0094] Specifically, the purpose of this step is to generate new candidate information for cell states to be included, providing "raw materials" for cell state updates. The calculation is performed using... and As input, after a linear transformation specific to the candidate state, the result is mapped to the interval [-1, 1] by the tanh activation function to obtain the candidate cell state values. The formula is: .

[0095] In the formula, for The corresponding candidate state weight matrix, with dimensions and Consistency is used to extract sentiment features from the current latent variables as candidates for new information; for The corresponding candidate state weight matrix has a dimension of h×h and its function is to integrate the temporal features in the historical hidden states so that the candidate information is more in line with the user's emotional evolution pattern. The candidate state bias vector has a dimension of h and is used to correct the linear transformation result. The core function of the tanh activation function is to limit the numerical range of the candidate values ​​to avoid gradient explosion caused by excessively large values ​​during subsequent cell state updates. At the same time, positive and negative values ​​are respectively associated with emotional features in different directions (such as positive values ​​corresponding to positive emotional tendencies and negative values ​​corresponding to negative emotional tendencies), which facilitates the quantitative representation of emotional states in the future.

[0096] S25. Based on the input gating vector, forgetting gating vector, candidate cell states, and the cell state at the previous time step, update the current cell state using element-wise multiplication and addition; the cell state update formula is:

[0097]

[0098] in, This represents the cell state at time t. This represents the cell state at time t-1. This represents element-wise multiplication.

[0099] Specifically, this step is the core of LSTM's storage of long-term temporal information. By fusing historical cell states with current candidate information, dynamic updates of sentiment temporal information are achieved. During computation, the forgetting gating vector generated in S22 is first used... Cell state at time t-1 Perform element-wise multiplication, filter and retain. The effective historical information in the data; and then the input gating vector generated by S21. Candidate cell state values ​​generated by S24 Element-wise multiplication is performed to filter and extract new information to be included; finally, the two results are added together to obtain the current cell state at time t. The formula is: .

[0100] In the formula, The cell state at time t-1 stores long-term emotional temporal information from the previous time step; This is element-wise multiplication, where the values ​​of corresponding dimensions are directly multiplied, enabling the gating vector to filter information dimension by dimension. The entire calculation process is essentially a dynamic balance of "retaining useful old information + incorporating new information after filtering," ensuring... It can not only continuously accumulate the long-term evolution characteristics of user emotions, but also update the latest changes in current emotions in a timely manner, providing comprehensive temporal information support for subsequent hidden state calculations.

[0101] S26. Based on the output gating vector and the current cell state, calculate the current hidden state using the tanh function; the update formula for the current hidden state is:

[0102]

[0103] in, This represents the hidden state at time t.

[0104] Specifically, this step filters cell state information through output gating vectors to generate hidden states that can be directly used as potential emotional state vectors. During calculation, the current cell state generated in S25 is first processed... Applying the tanh activation function maps the values ​​to the [-1,1] interval to avoid numerical overflow and standardize the sentiment features; then it is combined with the output gating vector generated by S23. By performing element-wise multiplication, the key information that needs to be passed to the hidden state is selected, and finally the current hidden state at time t is obtained. The formula is: .

[0105] In the formula, Its function is to bring the extreme values ​​that may exist in the cell state back to a reasonable range, while enhancing the distinguishability of emotional characteristics (such as adjusting the cell state value corresponding to strong negative emotions from a large negative value to close to -1, which facilitates subsequent emotional state judgment). The element-wise multiplication rule further filters out the information most relevant to the current emotional state, ensuring... It can accurately characterize the user's potential emotional state at time t. In subsequent steps... It directly serves as a potential emotional state vector in the construction of the emotional baseline and the adjustment of the security strategy, and is a key output connecting LSTM temporal modeling and subsequent emotional analysis.

[0106] In one optional embodiment, a user historical sentiment baseline, including a historical mean vector and a historical standard deviation vector, is calculated based on a sequence of potential emotional state vectors. The security policy parameters are then dynamically adjusted according to the statistical deviation between the current potential emotional state vector in the sequence and the historical sentiment baseline, including the following steps:

[0107] S31. Based on the potential emotional state vector sequence, calculate the historical mean vector and the historical standard deviation vector, where the formula for calculating the historical mean vector is: The formula for calculating the historical standard deviation vector is: ,in This is a vector of historical means. The historical standard deviation vector, This represents the vector of potential emotional states at time t. Indicates the length of the time series.

[0108] Specifically, the core of this step is to construct a historical emotional baseline reflecting the user's typical emotional level based on the generated sequence of potential emotional state vectors. This is achieved by calculating the historical mean vector and the historical standard deviation vector. The sequence of potential emotional state vectors consists of the potential emotional state vectors at each time step. Composition, in which This vector represents the user's potential emotional state at time t (directly derived from the hidden state output by the previous LSTM network, containing multi-dimensional features of the user's emotions at that time, such as anxiety tendency, depression tendency, etc., with dimensions consistent with those of the hidden layer). The length of the time series is the total number of time steps contained in the vector sequence (covering all valid dialogue rounds of the user's historical interactions; invalid data needs to be removed during the initial cleaning stage to ensure...). All of these correspond to valid vectors.

[0109] Historical mean vector The calculation logic involves averaging all potential emotional state vectors in the sequence at different dimensions, as shown in the formula: .

[0110] In the formula, This means representing all sequences from t=1 to t=T in the sequence. Perform element-wise summation along each dimension (i.e., sum the feature values ​​at all time steps for each emotion feature dimension), then divide by the length of the time series. This yields the average value for each dimension, ultimately forming... This vector reflects the user's long-term average level across various emotional characteristic dimensions and is the core benchmark for judging whether the current emotion deviates from the norm.

[0111] Historical standard deviation vector The calculation is used to quantify the range of fluctuations in a user's historical sentiment across various dimensions. The formula is: .

[0112] In the formula, Represents each time step with historical mean vector The difference in the corresponding dimension (reflecting the degree of deviation between the sentiment at a single time step and the mean). The element-wise square of this difference (amplifying the degree of deviation and eliminating the effect of negative values). Sum the squared differences of all time steps along the dimension and then divide by . We obtain the variance of each dimension, and then convert the variance to standard deviation using the square root operation, ultimately forming... This vector is used for subsequent standardized bias calculations to eliminate the influence of differences in the dimensions of different emotion dimensions on bias judgments.

[0113] S32. Based on the current potential emotional state vector, the historical mean vector, and the historical standard deviation vector, calculate the z-score bias vector, where the formula for calculating the z-score bias vector is: ,in This represents the z-score bias vector. This represents the current potential emotional state vector.

[0114] Specifically, this step uses the z-score normalization method to convert the deviation between the current potential emotional state and the historical emotional baseline into a dimensionless deviation vector, which facilitates subsequent judgment of the degree of emotional deviation using a unified threshold. The input data required for calculation includes the current potential emotional state vector. (That is, the potential emotional state vector corresponding to the user's current dialogue turn, generated by the LSTM hidden state at the latest time step, with dimensions equal to...) Consistent), generated by S31 and z-score bias vector The calculation formula is: .

[0115] In the formula, the molecule part The dimensionality difference between the current potential emotional state vector and the historical mean vector directly reflects the original deviation of the current emotion from the long-term average level; the denominator part... This is a historical standard deviation vector used to standardize the original deviations. Through element-wise division (i.e., dividing the original deviation value of each dimension by the standard deviation of the corresponding dimension), the deviation values ​​of different dimensions are converted into z-score values ​​on a uniform scale, allowing for direct horizontal comparison of deviations across dimensions (for example, a z-score of 2 for a certain dimension means that the current sentiment feature of that dimension deviates from the mean by 2 standard deviations, significantly higher than the normal fluctuation range). The final generated... The values ​​of each dimension of the vector correspond to the degree of standardized deviation of the current sentiment relative to the historical baseline in that feature dimension.

[0116] S33. Based on the z-score bias vector, a preset threshold comparison is used to determine whether a high-sensitivity state is triggered. When a high-sensitivity state is triggered, the risk detection threshold is dynamically reduced, and the response generation strategy identifier is set to template library mode. When a high-sensitivity state is not triggered, the default risk detection threshold is maintained, and the response generation strategy identifier is set to neural generation model mode.

[0117] Specifically, this step determines whether the user is in a state of heightened emotional sensitivity by comparing a preset threshold with the z-score deviation vector, and dynamically adjusts the safety strategy parameters (risk detection threshold and response generation strategy) accordingly to ensure the strategy is adapted to the current emotional state. First, a threshold for judging heightened sensitivity is preset; this threshold is in vector form (and...). The vector dimensions are consistent and set by domain experts in conjunction with mental health risk level standards. For example, the thresholds for each dimension are usually set to ±1.5 or ±2, and the specific values ​​are adjusted according to the risk sensitivity requirements in the application scenario. This is denoted as... .

[0118] The specific logic for determining a highly sensitive state is: dimensional comparison. Vector and The absolute value of a vector if it has at least one dimension. ( for The value of the k-th dimension of the vector. for If the value of the k-th dimension of the vector is found, then a high-sensitivity state is triggered (indicating that the user's current emotion deviates from the historical range in at least one key dimension, posing a potential emotional risk, and the sensitivity of risk detection needs to be improved); if all dimensions are... If the high-sensitivity state is not triggered (the user's current emotion is within the normal historical fluctuation range, and there is no need to adjust the detection sensitivity), it is determined that the high-sensitivity state has not been triggered (the user's current emotion is within the normal historical fluctuation range, and there is no need to adjust the detection sensitivity).

[0119] The safety strategy parameter adjustment rules dynamically change based on the high-sensitivity state judgment results: When a high-sensitivity state is triggered, the risk detection threshold needs to be dynamically reduced. Specifically, the reduction can be 50%-70% of the default threshold (e.g., if the default threshold is 0.6, the adjusted threshold is 0.3-0.42) to improve the system's sensitivity to negative emotions and risky expressions, avoiding overlooking potential risks. Simultaneously, the response generation strategy identifier is set to template library mode. Because template library mode generates responses based on preset standardized empathy and intervention scripts, it ensures the safety and accuracy of the response content, avoids potential expression biases in neural generative models, and better adapts to the emotional support needs in high-sensitivity states.

[0120] When a high-sensitivity state is not triggered, the default risk detection threshold is maintained (this threshold is the initially configured standard threshold, determined by validation using prior training data, balancing recognition accuracy and false alarm rate), and no sensitivity adjustment is required. Simultaneously, the response generation strategy identifier is set to neural generative model mode, as this mode can generate more flexible and natural responses based on the user's personalized input, better aligning with the personalized needs of users in a normal emotional state for mental health education content, thus enhancing the interactive experience. After setting the strategy identifier, the corresponding response generation module will be automatically invoked to complete the subsequent personalized response generation.

[0121] In an alternative embodiment, based on the current user input, the instant emotion vector, and the current potential emotion state vector, combined with the adjusted security policy parameters, a personalized response is generated through a neural generation model or a template library, including the following steps:

[0122] S41. Based on the current user input, through text cleaning and tokenization processing, a standardized input text is generated; based on the standardized input text, instant emotion classification processing is performed through a pre-trained language model, and the current user's instant emotion vector is output.

[0123] Specifically, for the original text input by the current user, first perform text cleaning operations: remove irrelevant special characters (such as punctuation marks, emojis, URL links, etc.) in the text, filter out meaningless stop words (such as "de", "le", "a", etc., which can be screened based on a special stop word list in the field of mental health to avoid deleting emotion-related words), and correct spelling mistakes, pinyin abbreviations, etc. in the text to ensure the normality of the text content. Subsequently, use a tokenization tool (such as jieba tokenization, HanLP, etc.) to perform tokenization processing on the cleaned text, splitting the continuous text into discrete lexical sequences, while retaining compound words with obvious emotion features (such as "feeling low", "fidgety", etc., ensuring that such words are not split through a custom dictionary), and finally forming a standardized input text.

[0124] After obtaining the standardized input text, input it into a pre-trained language model (such as a model optimized for Chinese emotion analysis like BERT, RoBERTa, etc.) for instant emotion classification processing. This model encodes the standardized input text to extract the emotion features in the text: the input layer of the model converts the lexical sequence into word embedding vectors, and after feature extraction by multiple layers of Transformer encoders, the softmax classifier in the output layer outputs the probability distribution of the corresponding emotion categories (such as categories like "calm", "anxious", "depressed", "angry", etc.). Select the feature vector corresponding to the emotion category with the highest probability as the current user's instant emotion vector, and the vector dimension is the same as the hidden layer dimension of the model. The value of each dimension corresponds to the intensity of different emotion features under this emotion category (for example, in the "anxious" category vector, the higher the value of a certain dimension, the more obvious the "fidgety" feature), ensuring that the vector can accurately quantify the current user's surface emotion state.

[0125] S42. Based on the standardized input text and the current user's instant emotion vector, by updating the potential emotion state time series model, the current user's potential emotion state vector is output.

[0126] Specifically, the constructed temporal model of latent emotional states is dynamically updated using the standardized input text generated by S41 and the current user's instantaneous emotion vector as input. First, the standardized input text is converted into a vector form that can be input into the model. A pre-trained word embedding model (consistent with the model used in processing historical data to ensure uniform vector dimensions) is used to calculate the word embedding vector for each word in the text. Then, all word embedding vectors are averaged dimension-wise to obtain the text vector corresponding to the standardized input text. This text vector is then concatenated with the current user's instantaneous emotion vector to form the fused input vector required for model updates. The dimension of the concatenated vector is the sum of the dimensions of the text vector and the instantaneous emotion vector.

[0127] The fused input vector is fed into the latent emotional state temporal model (composed of VAE and LSTM). The VAE encoder combines the latent emotional state vector output by the model at the previous time step (i.e., the hidden state generated by the previous round of interaction) and performs a nonlinear transformation on the fused input vector, outputting the mean vector and variance vector of the latent variable at the current time step. Then, the current latent variable is obtained by sampling through the reparameterization technique. The LSTM network takes this current latent variable as input, calls the cell state and hidden state at the previous time step, updates the cell state through the gating mechanism (forget gate, input gate, output gate), and calculates the hidden state at the current time step. This hidden state is the current user's latent emotional state vector, and its dimension is consistent with the dimension of the LSTM hidden layer. The values ​​of each dimension in the vector comprehensively reflect the current user's potential emotional evolution trend, rather than being limited to the surface emotional expression.

[0128] S43. Based on the current user's potential emotional state vector and historical emotional baseline, the current security status is determined by the security policy adaptive module built based on security policy parameters.

[0129] Specifically, the adaptive security policy module uses the current user's potential emotional state vector output by S42 and the previously constructed user historical emotional baseline (including historical mean vector and historical standard deviation vector) as core inputs, and performs security status judgments in conjunction with the adjusted security policy parameters (risk detection thresholds). First, it calculates the statistical deviation between the current user's potential emotional state vector and the historical emotional baseline: using a previously defined deviation calculation method (such as z-score deviation), the difference between the current vector and the historical mean vector is divided by the historical standard deviation vector to obtain the deviation vector. Then, by calculating the Euclidean distance or L2 norm of the deviation vector, a single-valued deviation index is obtained, eliminating the influence of vector dimension on the judgment result.

[0130] The single-valued deviation index is compared with the risk detection threshold in the security policy parameters. If the deviation index is greater than the risk detection threshold, it indicates that the current user's potential emotional state deviates significantly from the historical normal range, and there is a high probability of emotional risk; the current security status is determined to be a high-sensitivity state. If the deviation index is less than or equal to the risk detection threshold, it indicates that the current user's potential emotional state is within the historical normal fluctuation range, with no obvious emotional risk; the current security status is determined to be a normal state. Throughout the judgment process, the security policy adaptive module will call the latest historical emotional baseline in real time (if new interaction data has updated the baseline, the updated data will be used) to ensure the timeliness and accuracy of the judgment results and avoid misjudgments due to outdated baselines.

[0131] S44. If the current security state is a high-sensitivity state, randomly select a response text from the predefined support template library to generate a personalized response; if the current security state is a normal state, input the current user input and the current user's potential emotional state vector into the neural generation model, and generate a personalized response through beam search.

[0132] Specifically, when S43 determines the current safety status to be a high-sensitivity state, it uses a predefined supportive template library to generate a personalized response. This template library stores response texts categorized by emotional risk type (such as anxiety risk, depression risk, emotional instability risk, etc.). Each category's templates include emotional empathy expressions, basic guidance suggestions, and safety tips (such as "I noticed you are currently experiencing significant emotional fluctuations, which is a normal reaction. Try to breathe slowly and deeply, feeling the breath in and out. If you feel overwhelmed, you can also contact someone you trust for support"). During generation, the corresponding emotional risk category is first determined based on the user's current immediate emotional vector. One or two candidate templates are randomly selected from the templates under that category. Then, combined with the emotional evolution trend reflected in the user's potential emotional state vector, the variable information in the templates (such as "recent emotional changes," "common adjustment methods," etc.) is fine-tuned (if the vector shows that the emotion is continuously deteriorating, the prompt "seek professional help in time" is strengthened in the template), ultimately forming a personalized response that fits the user's current state.

[0133] When S43 determines the current safety state to be normal, a neural generative model combined with beam search is used to generate a personalized response. This neural generative model uses a Transformer decoder as its core structure, and its input is the concatenation of the standardized input text vector from S41, the current user's immediate emotion vector from S41, and the current user's potential emotional state vector from S42. The model training phase has been optimized using mental health education corpora (including popular science knowledge and emotional guidance cases) to ensure the professionalism and applicability of the generated content. During the generation process, a beam search algorithm is used, with a beam width set to 3-5 (balancing generation speed and text quality). At each step of generating candidate words, the algorithm retains the 3-5 candidates with the highest probability until a complete sentence is generated; simultaneously, it uses a perplexity metric to filter in real time, eliminating candidate sentences with a perplexity higher than a preset threshold (usually 50) to ensure the coherence and logic of the generated response, ultimately outputting an optimal personalized response (e.g., for the user's input "how to relieve study pressure," combined with their normal emotional state, generating a response containing specific suggestions such as "develop a segmented study plan" and "appropriate exercise for relaxation").

[0134] The aforementioned intelligent question-and-answer interaction method for mental health education collects users' historical dialogue data, preprocesses and serializes it, extracts real-time emotion vectors based on a pre-trained language model to capture users' current emotional expression, and then uses a temporal model constructed from variational autoencoders and recurrent neural networks to decouple potential emotional states from the dialogue history and track their cross-conversation evolution to generate a state vector sequence. Based on this, a historical emotion baseline is calculated, and risk detection thresholds and response generation strategies are dynamically adjusted through statistical bias analysis. Finally, personalized responses are generated by combining real-time user input and emotion state vectors. This method achieves continuous tracking of users' long-term emotional states and adaptive personalized security strategies in a pure dialogue environment that does not rely on biosignals, effectively improving the accuracy of emotion recognition and the security protection capabilities of mental health interaction systems.

[0135] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0136] Based on the same inventive concept, this application also provides an apparatus for implementing the intelligent question-and-answer interaction method for mental health education as described above. The solution provided by this apparatus is similar to the solution described in the above method; therefore, the specific limitations of one or more embodiments of the intelligent question-and-answer interaction apparatus for mental health education provided below can be found in the limitations of the intelligent question-and-answer interaction method for mental health education described above, and will not be repeated here.

[0137] In one exemplary embodiment, such as Figure 3 As shown, a smart question-and-answer interactive device 30 for mental health education is provided to implement the methods in the above-described method embodiments. The device includes:

[0138] The user emotion perception preprocessing module 31 is used to collect users' historical dialogue data, perform text cleaning and word segmentation on the historical dialogue data, and generate serialized user input data; based on the serialized user input data, emotion classification is performed through a pre-trained language model to obtain the instantaneous emotion vector of each user input.

[0139] The emotional state dynamic tracking module 32 is used to perform temporal modeling of potential emotional states based on serialized user input data and real-time emotional vectors, and to construct a potential emotional state temporal model. The potential emotional state temporal model is used to dynamically track and update the user's potential emotional states across sessions, and generate a sequence of potential emotional state vectors.

[0140] The security strategy optimization module 33 is used to calculate the user's historical emotional baseline, including the historical mean vector and the historical standard deviation vector, based on the potential emotional state vector sequence; and to dynamically adjust the security strategy parameters according to the statistical deviation between the current potential emotional state vector in the potential emotional state vector sequence and the historical emotional baseline; wherein, the security strategy parameters include risk detection threshold and response generation strategy.

[0141] The personalized response generation engine 34 is used to generate personalized responses based on the current user input, the immediate emotion vector, and the current potential emotion state vector, combined with adjusted safety policy parameters, through a neural generative model or template library.

[0142] Embodiments of this application also provide a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the aforementioned method embodiments.

[0143] Embodiments of this application also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described method embodiments.

[0144] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0145] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.

Claims

1. An intelligent question-and-answer interactive method for mental health education, characterized in that, The method includes: S1. Collect users' historical dialogue data, perform text cleaning and word segmentation on the historical dialogue data, and generate serialized user input data; based on the serialized user input data, perform emotion classification processing through a pre-trained language model to obtain the instantaneous emotion vector of each user input. S2. Based on the serialized user input data and the instantaneous emotion vector, latent emotion state temporal modeling is performed using a pre-trained variational autoencoder and recurrent neural network model to construct a latent emotion state temporal model; the latent emotion state temporal model is used to dynamically track and update the user's latent emotion state across sessions to generate a sequence of latent emotion state vectors. S3. Based on the potential emotional state vector sequence, calculate the user's historical emotional baseline, including the historical mean vector and the historical standard deviation vector; dynamically adjust the security strategy parameters according to the statistical deviation between the current potential emotional state vector in the potential emotional state vector sequence and the historical emotional baseline; wherein, the security strategy parameters include a risk detection threshold and a response generation strategy; S4. Based on the current user input, the instantaneous emotion vector, and the current potential emotion state vector, combined with the adjusted security strategy parameters, a personalized response is generated through a neural generative model or template library.

2. The method according to claim 1, characterized in that, The process involves using the serialized user input data and the instantaneous emotion vector to perform temporal modeling of the latent emotion state through a pre-trained variational autoencoder and recurrent neural network model, thereby constructing a temporal model of the latent emotion state. The latent emotional state time series model is used to dynamically track and update the user's latent emotional state across sessions, generating a sequence of latent emotional state vectors, including: S11. For each user input in the serialized user input data, calculate the average word embedding through a pre-trained word embedding model to generate an input vector; based on the input vector and the latent emotional state vector of the previous time step, perform a nonlinear transformation through a VAE encoder network to output the mean vector and variance vector of the latent variables. S12. Based on the mean vector and the variance vector, the current latent variables are sampled using a reparameterization technique to obtain a sequence of latent variables; the sampling formula for the current latent variables is: in, This represents the current latent variable at time t. Let represent the mean vector at time t. Let represent the variance vector at time t. Describes a random vector sampled from a standard normal distribution. Represents element-wise multiplication; S13. Based on the latent variable sequence, perform time-series modeling through an LSTM network, and update the cell state and hidden state, using the hidden state as a latent emotional state vector. S14. Repeat steps S11 to S13 to generate the sequence of potential emotional state vectors for the entire dialogue history.

3. The method according to claim 2, characterized in that, The step of performing time-series modeling based on the latent variable sequence using an LSTM network and updating cell states and hidden states includes: S21. Based on the current latent variables and the hidden state at the previous time step, calculate the input gating vector through the LSTM input gate; the formula for calculating the input gating vector is: in, This represents the input gating vector at time t. Denotes the current latent variable at time t. This represents the hidden state at time t-1. and This represents the input gate weight matrix. This represents the input gate bias vector. Represents the sigmoid function; S22. Based on the current latent variables and the hidden state of the previous time step, calculate the forgetting gate vector using an LSTM forgetting gate; the formula for calculating the forgetting gate vector is: in, This represents the forgetting gating vector at time t. and This represents the forget gate weight matrix. Represents the forget gate bias vector; S23. Based on the current latent variables and the hidden state of the previous time step, calculate the output gating vector through the LSTM output gate; the formula for calculating the output gating vector is: in, This represents the output gating vector at time t. and This represents the output gate weight matrix. This represents the output gate bias vector; S24. Based on the current latent variables and the hidden state of the previous time step, calculate the candidate values ​​of the cell state using the tanh activation function; the formula for calculating the candidate values ​​of the cell state is: in, This represents the candidate values ​​for the cell state at time t. and Represents the candidate state weight matrix. Represents the candidate state bias vector; S25. Based on the input gating vector, the forgetting gating vector, the candidate cell state value, and the cell state at the previous time step, update the current cell state using element-wise multiplication and addition; the cell state update formula is: in, This represents the cell state at time t. This represents the cell state at time t-1. Represents element-wise multiplication; S26. Based on the output gating vector and the current cell state, calculate the current hidden state using the tanh function; the update formula for the current hidden state is: in, This represents the hidden state at time t.

4. The method according to claim 1, characterized in that, The process involves calculating a user's historical emotional baseline, including a historical mean vector and a historical standard deviation vector, based on the potential emotional state vector sequence; and dynamically adjusting security policy parameters according to the statistical deviation between the current potential emotional state vector in the potential emotional state vector sequence and the historical emotional baseline, including: S31. Based on the potential emotional state vector sequence, calculate the historical mean vector and the historical standard deviation vector, wherein the historical mean vector is calculated using the following formula: The formula for calculating the historical standard deviation vector is as follows: ,in Let the historical mean vector be... Let be the historical standard deviation vector. This represents the vector of potential emotional states at time t. Indicates the length of the time series; S32. Based on the current potential emotional state vector, the historical mean vector, and the historical standard deviation vector, calculate the z-score bias vector, wherein the formula for calculating the z-score bias vector is: ,in This represents the z-score deviation vector. This represents the current potential emotional state vector; S33. Based on the z-score deviation vector, determine whether a high-sensitivity state is triggered by comparing a preset threshold. When the high-sensitivity state is triggered, dynamically reduce the risk detection threshold and set the response generation strategy identifier to template library mode. When the high-sensitivity state is not triggered, maintain the default risk detection threshold and set the response generation strategy identifier to neural generation model mode.

5. The method according to any one of claims 1 to 4, characterized in that, The process of generating personalized responses based on the current user input, the immediate emotion vector, and the current potential emotion state vector, combined with the adjusted safety policy parameters, through a neural generative model or template library, includes: S41. Based on the current user input, generate standardized input text through text cleaning and word segmentation; based on the standardized input text, perform real-time emotion classification processing through the pre-trained language model, and output the current user's real-time emotion vector. S42. Based on the standardized input text and the current user's real-time emotion vector, the current user's potential emotion state vector is output by updating the potential emotion state time series model. S43. Based on the current user's potential emotional state vector and the historical emotional baseline, the current security status is determined by the security policy adaptive module constructed based on the security policy parameters. S44. If the current security state is a high-sensitivity state, randomly select a response text from a predefined support template library to generate a personalized response; if the current security state is a normal state, input the current user input and the current user's potential emotional state vector into the neural generation model, and generate a personalized response through beam search.

6. An intelligent question-and-answer interactive device for mental health education, used to implement the method according to any one of claims 1 to 5, characterized in that, The device includes: The user emotion perception preprocessing module is used to collect users' historical dialogue data, perform text cleaning and word segmentation on the historical dialogue data, and generate serialized user input data; based on the serialized user input data, emotion classification is performed through a pre-trained language model to obtain the instantaneous emotion vector of each user input. The emotional state dynamic tracking module is used to perform temporal modeling of potential emotional states based on the serialized user input data and the instantaneous emotional vector, and to construct a potential emotional state temporal model through a pre-trained variational autoencoder and recurrent neural network model; and to dynamically track and update the user's potential emotional states across sessions using the potential emotional state temporal model, thereby generating a sequence of potential emotional state vectors. The security strategy optimization module is used to calculate a user's historical emotional baseline, including a historical mean vector and a historical standard deviation vector, based on the potential emotional state vector sequence; and to dynamically adjust security strategy parameters according to the statistical deviation between the current potential emotional state vector in the potential emotional state vector sequence and the historical emotional baseline; wherein, the security strategy parameters include a risk detection threshold and a response generation strategy; A personalized response generation engine is used to generate personalized responses based on the current user input, the instantaneous emotion vector, and the current potential emotion state vector, combined with the adjusted security policy parameters, through a neural generative model or template library.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 5.