Psychological counseling-based ai question and answer expert model construction method and system, and medium

By using a dual-channel pickup strategy to separate and process speech content and emotional acoustic features, the problem of loss of emotional information in traditional speech processing is solved, and more accurate psychological state analysis and crisis state recognition are achieved.

CN121171262BActive Publication Date: 2026-02-13CHENGDU SHANMEI TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511695344.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-19
Publication Date
2026-02-13
Estimated Expiration
2045-11-19

AI Technical Summary

Technical Problem

Traditional voice processing systems inadvertently filter out subtle auditory cues carrying important emotional information in psychological support scenarios due to noise reduction, making it difficult for AI models to accurately judge the user's deep psychological state, resulting in the loss of emotional information and inaccurate psychological analysis.

Method used

A dual-channel pickup strategy is adopted, in which the main pickup channel acquires speech content and processes it in parallel, while the auxiliary emotion pickup channel acquires emotional acoustic features. The speech content and emotional acoustic features are processed separately and then fused together to improve the ability to perceive psychological states.

Benefits of technology

It significantly enhances the AI ​​model's ability to perceive users' deep psychological states, ensures the complete transmission of emotional information, and improves the accuracy of psychological analysis and the effectiveness of crisis identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121171262B_ABST
    Figure CN121171262B_ABST
Patent Text Reader

Abstract

The application discloses an AI question and answer expert model construction method and system based on psychological counseling and a medium, relates to the technical field of artificial intelligence and voice analysis, and comprises the following steps: acquiring double-channel voice signals in a user psychological counseling voice interaction process, including a first-channel voice signal and a second-channel voice signal; processing the first-channel voice signal based on a main voice pickup channel to acquire voice content, and simultaneously processing the second-channel voice signal based on an auxiliary emotion voice pickup channel to acquire emotional acoustic features; converting the voice content into a text stream after voice recognition, and simultaneously forming an emotional feature stream after extracting and quantizing the emotional acoustic features; fusing the text stream and the emotional feature stream to obtain a fused information stream; and analyzing a user psychological state according to the fused information stream. The application solves the problem that emotional information loss leads to inaccurate psychological analysis in the prior art, and lays a foundation for accurately performing subsequent crisis state recognition.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence and speech analysis, in particular to an AI question and answer expert model construction method and system based on psychological counseling and a medium. BACKGROUND

[0002] In the field of artificial intelligence application, especially in the context of psychological support which requires high sensitivity, it is crucial to accurately capture and interpret the subtle changes in human emotions. Traditional speech processing systems, although excellent in converting spoken language into text, often inadvertently filter out weak sound clues that carry important emotional information. This loss of information can severely limit the ability of artificial intelligence to understand the user's deep psychological state, especially when users communicate in various noisy environments through ordinary consumer-grade recording devices.

[0003] When constructing an AI question and answer expert model that provides psychological counseling services, the core goal is to facilitate natural communication with users, so from the beginning of design, the model is planned to receive both text and voice forms of questions. In order to ensure the reliability of voice input in various actual use scenarios, especially in complex acoustic environments that users may be in, the speech-to-text (STT) module of the AI question and answer expert model uses the sound data enhancement method commonly used in the industry during the training phase to improve the recognition accuracy of the user's voice content in the presence of background noise.

[0004] However, although the above robustness training against noise environment has been carried out, when users communicate with AI through ordinary built-in microphones of smartphones in actual noisy environments, the microphones of smartphones are usually integrated with basic noise reduction processing functions, and the STT module also contains more complex noise reduction algorithms. These noise reduction mechanisms, while effectively filtering out background noise and improving voice clarity, often inadvertently weaken or remove some subtle emotional expressions in the user's voice that are crucial for psychological state judgment. For example, when expressing inner struggle, the user's voice may have slight tremors, short hesitation pauses, subtle changes in speech rate, or low sighs, etc. These non-verbal details that reflect the true psychological state, due to their acoustic characteristics may be similar to some components of background noise, or their energy is low, are easily mistaken by the noise reduction algorithm as useless background sound and are processed together.

[0005] While existing speech-to-text modules may accurately identify the user's spoken content at the text level, they completely lose the hesitation, low tone, and even suppressed emotion accompanying the user's speech. Ultimately, what is delivered to the core AI psychological counseling model is a text that appears emotionally neutral. The genuine emotional information originally contained in the user's speech has undergone irreversible loss or distortion during the speech-to-text conversion process. When the core AI psychological counseling model receives this "emotionally filtered" text, its internal psychological analysis mechanism faces a significant challenge, struggling to accurately determine the user's true psychological state and thus ignoring the user's potential pain, struggle, or negative emotions. This failure to consider the user's true emotions further directly impacts the effectiveness of the built-in rules used in the model to identify the user's crisis state, as these rules largely rely on capturing the intensity, duration, and specific acoustic features of the negative emotions expressed in the user's voice.

[0006] In view of the above, this application is hereby submitted. Summary of the Invention

[0007] The technical problem this invention addresses is that traditional speech processing systems in psychological support scenarios inadvertently filter out subtle auditory cues carrying important emotional information during noise reduction, leading to inaccurate psychological analysis and difficulty in accurately judging the user's deep psychological state by the constructed AI model. This results in the loss of emotional information. The purpose of this invention is to provide a method, system, and medium for constructing an AI question-answering expert model based on psychological counseling. This invention innovatively introduces a dual-channel sound pickup strategy, separating the processing of speech content and emotional acoustic features. This effectively avoids the accidental damage to emotional information caused by noise reduction in traditional speech processing, thereby significantly improving the AI ​​model's ability to perceive the user's deep psychological state. It solves the problem of inaccurate psychological analysis due to the loss of emotional information in existing technologies, laying the foundation for accurate subsequent crisis state identification.

[0008] This invention is achieved through the following technical solution:

[0009] In a first aspect, the present invention provides a method for constructing an AI question-answering expert model based on psychological counseling, the method comprising:

[0010] Acquire dual-channel voice signals during the user's psychological counseling voice interaction process. The dual-channel voice signals include a first-channel voice signal and a second-channel voice signal.

[0011] The first channel speech signal is processed based on the main pickup channel to obtain speech content, while the second channel speech signal is processed based on the auxiliary emotion pickup channel to obtain emotional acoustic features.

[0012] The voice content is converted into a text stream after voice recognition, and the emotional acoustic features are extracted and quantified to form an emotional feature stream;

[0013] The text stream and the emotional feature stream are fused to obtain a fused information stream;

[0014] According to the fused information stream, the user's psychological state is analyzed.

[0015] Further, the dual-channel voice signal is obtained, including:

[0016] According to the environmental acoustic scene, the microphone pickup and preprocessing strategy is selected; the environmental acoustic scene refers to the acoustic environment in which the user performs the voice interaction during the psychological counseling;

[0017] According to the microphone pickup and preprocessing strategy, the dual-channel voice signal is obtained through the microphone array; the microphone pickup and preprocessing strategy refers to dynamically adjusting the pickup mode (such as omnidirectional, cardioid, and super cardioid), gain, noise reduction algorithm, and echo cancellation of the microphone according to different environmental acoustic scenes, so as to optimize the capture quality of the voice signal.

[0018] Further, the second channel voice signal is processed based on the auxiliary emotional pickup channel to obtain emotional acoustic features, including:

[0019] The sound macroscopic parameters of the second channel voice signal are obtained, including short-term average energy and instantaneous speech rate;

[0020] According to the change of the sound macroscopic parameters, the adjustment of the user's vocalization manner is judged to obtain an adjustment result;

[0021] According to the adjustment result, the sensitivity factor of the auxiliary emotional pickup channel in different types of frequency bands is dynamically adjusted to optimize the sound macroscopic parameters;

[0022] The sound microcosmic parameters of the second channel voice signal are obtained and tracked, including the resonance peak change, the jitter and tremor parameters of the fundamental frequency;

[0023] According to the sound macroscopic parameters and the sound microcosmic parameters, multi-dimensional emotional acoustic features are generated.

[0024] The above technical solutions can more finely capture the subtle changes in the user's vocalization manner, and dynamically adjust the sensitivity factor of the pickup channel, so as to more accurately extract emotional acoustic features, and further improve the accuracy and robustness of emotion recognition.

[0025] Further, the sound microcosmic parameters of the second channel voice signal are obtained and tracked, including:

[0026] establishing a user pronunciation feature benchmark, the user pronunciation feature benchmark comprising an average fundamental frequency, a fundamental frequency standard deviation, a jitter rate, a shimmer rate, and a typical distribution range of formants of the user in a neutral emotional state;

[0027] obtaining sound micro-parameters of the second-channel voice signal, and comparing the sound micro-parameters with the user pronunciation feature benchmark to obtain a comparison result;

[0028] According to the comparison result, the micro-parameters of the current second-channel voice signal are attributed to a corresponding feature type; the feature type includes physiological or habitual vocalization features, emotional acoustic features.

[0029] Further, obtaining and tracking the sound micro-parameters of the second-channel voice signal further comprises:

[0030] The pronunciation intelligibility of the user is evaluated by analyzing the spectral tilt and the vowel space size of the second-channel voice signal of the user to obtain a pronunciation intelligibility evaluation result;

[0031] According to the pronunciation intelligibility evaluation result, the recognition threshold of the emotional acoustic feature is adjusted; including:

[0032] By analyzing whether the current voice text content contains negative emotional keywords, and whether there are physiological acoustic signals related to emotional distress in the emotional feature stream obtained by the auxiliary emotion pickup channel, the association between the pronunciation intelligibility evaluation result and the user emotional distress is determined.

[0033] According to the association, the recognition threshold of the emotional acoustic feature is adjusted; when the association indicates that low intelligibility is an expression of emotional distress, the recognition threshold of the emotional acoustic feature is maintained at a sensitive level, or the recognition threshold of the emotional acoustic feature is reduced.

[0034] The above technical solutions, by establishing a personalized pronunciation feature benchmark and evaluating pronunciation intelligibility, can more accurately distinguish physiological vocalization features and emotional acoustic features, and dynamically adjust the recognition threshold according to the pronunciation intelligibility, effectively avoiding misjudgment, making the extraction of emotional features more accurate. In addition, the pronunciation intelligibility is associated with the user emotional distress, and the recognition threshold of the emotional acoustic feature is intelligently adjusted accordingly, so that the system can still maintain sensitivity to emotional signals when the user emotional distress causes unclear pronunciation, avoiding emotional information missing caused by decreased pronunciation intelligibility.

[0035] Further, determining the association between the pronunciation intelligibility evaluation result and the user emotional distress comprises:

[0036] Analyzing the emotional tendency of the current voice text content, identifying ambiguous or indirect negative emotional expressions in the current voice text content;

[0037] Perform weak signal enhancement processing on the emotional feature flow, and locally amplify the physiological acoustic parameters in the emotional feature flow;

[0038] Analyze user historical dialogue data to establish a correlation pattern between low clarity, ambiguous or indirect negative emotional expression and emotional distress and physiological acoustic signal suppression;

[0039] According to the correlation pattern, calculate the potential correlation strength between low clarity and deep emotional distress;

[0040] According to the potential correlation strength, judge the correlation between low clarity and user deep emotional distress;

[0041] Among them, analyzing user historical dialogue data to establish a correlation pattern between low clarity, ambiguous or indirect negative emotional expression and emotional distress and physiological acoustic signal suppression comprises:

[0042] After each psychological counseling voice interaction of the user ends, extract the features of low clarity, ambiguous or indirect negative emotional expression and physiological acoustic signal suppression in this voice interaction;

[0043] Combine the user psychological state analysis result of this voice interaction to update the correlation pattern established in the user historical psychological counseling voice interaction data;

[0044] According to the context information of this voice interaction, adjust the weight of this voice interaction data in the correlation pattern update, and establish and maintain multiple context-specific correlation patterns for the user;

[0045] When calculating the potential correlation strength between low clarity and deep emotional distress, according to the current interaction context, select the corresponding context-specific correlation pattern for matching.

[0046] The above technical solutions, the present application establishes the correlation pattern between low clarity and deep emotional distress through multi-dimensional analysis (text emotion, weak physiological acoustic signal enhancement, historical dialogue data), and calculates the potential correlation strength, so as to more deeply excavate the potential emotional distress of the user, and improves the recognition ability of the implicit emotional problem. In addition, by continuously updating and maintaining the context-specific correlation pattern, and matching according to the current interaction context, the system can more accurately evaluate the correlation between low clarity and deep emotional distress according to the emotional expression characteristics of the user in different contexts, and improve the adaptability and accuracy of the model.

[0047] Further, multiple context-specific correlation patterns are established and maintained for the user, comprising:

[0048] The context information of the current voice interaction is analyzed to identify context keywords and context acoustic clues; the analysis includes: analyzing the text content of the current voice interaction through natural language processing technology to identify context keywords; analyzing the voice intonation characteristics of the current voice interaction to identify context acoustic clues;

[0049] According to the context keywords and context acoustic clues, the matching degrees of the current voice interaction and each context-specific association mode are calculated;

[0050] The context-specific association mode with the highest matching degree is selected from the calculated matching degrees, or when multiple matching degrees are close, multiple context-specific association modes are fused for judgment; wherein, fusing multiple context-specific association modes for judgment includes:

[0051] The emotional feature parameters inside each context-specific association mode are cross-compared to identify conflict parameters and redundant parameters;

[0052] The conflict parameters are prioritized or weighted average processed, and the redundant parameters are de-duplicated or feature selected;

[0053] According to the complexity and contradiction degree of the current emotional expression of the user, the parameter processing strategy of the priority sorting, weighted average processing, de-duplication or feature selection is adjusted to obtain the processed emotional feature parameters;

[0054] And the processed emotional feature parameters are fused to obtain a fusion judgment result.

[0055] The above technical solutions, the present application comprehensively analyzes the text content and voice intonation characteristics to identify the context, and calculates the matching degrees with each context-specific association mode, so as to more accurately select or fuse the context-specific association mode, and further improves the accuracy and context adaptability of the judgment of user emotional distress. In addition, when the matching degrees of multiple context-specific association modes are close, these modes can be intelligently fused, the emotional feature parameters inside each context-specific association mode are cross-compared, and the parameter processing strategy is dynamically adjusted, so as to effectively solve the conflict and redundancy problems between modes, and thus obtain a more comprehensive and accurate fusion judgment result.

[0056] Further, according to the complexity and contradiction degree of the current emotional expression of the user, the parameter processing strategy of the priority sorting, weighted average processing, de-duplication or feature selection is adjusted to obtain the processed emotional feature parameters, including:

[0057] By analyzing the deviation degree and duration of the physiological acoustic parameters in the emotional feature flow, and the density and emotional intensity of the negative emotional keywords in the text, the intensity of the user's emotional expression is evaluated to obtain the emotional intensity;

[0058] The consistency index is obtained by calculating the difference between the text sentiment polarity and the emotion polarity reflected by the physiological acoustic signal.

[0059] According to the emotion intensity and the consistency index, the parameter processing strategy of the priority sorting, the weighted average processing, the deduplication or the feature selection is dynamically adjusted, and when the emotion intensity is high and the consistency is low, the weight of the emotion feature stream is increased.

[0060] The above technical scheme, the application can dynamically adjust the parameter processing strategy according to the complexity and contradiction degree of the user emotion expression, especially when the emotion intensity is high but the text and the voice emotion are inconsistent, the weight of the emotion feature stream is increased, so that the real emotion of the user can be more accurately captured, and the judgment deviation caused by the surface information misleading is avoided.

[0061] In the second aspect, the application further provides an AI question and answer expert model construction system based on psychological counseling, which comprises:

[0062] The double-channel voice acquisition unit is used for acquiring double-channel voice signals in the user psychological counseling voice interaction process, and the double-channel voice signals comprise first-channel voice signals and second-channel voice signals.

[0063] The preprocessing unit is used for processing the first-channel voice signals based on the main sound pickup channel to obtain voice content, and processing the second-channel voice signals based on the auxiliary emotion sound pickup channel to obtain emotion acoustic features.

[0064] The conversion unit is used for converting the voice content into a text stream after voice recognition, and converting the emotion acoustic features into an emotion feature stream after extraction and quantization.

[0065] The fusion unit is used for fusing the text stream and the emotion feature stream to obtain a fused information stream.

[0066] The psychological state analysis unit is used for analyzing the user psychological state according to the fused information stream.

[0067] In the third aspect, the application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the AI question and answer expert model construction method based on psychological counseling.

[0068] Compared with the prior art, the application has the following advantages and beneficial effects:

[0069] 1. The AI question and answer expert model construction method, system and medium based on psychological counseling, by perceiving the environmental acoustic context and selecting appropriate microphone pickup and preprocessing strategies, innovatively adopts a dual-channel speech signal capture mechanism. Among them, the main pickup channel focuses on the acquisition of speech content, while the auxiliary emotional pickup channel is specially designed to process speech signals to obtain emotional acoustic features. This parallel processing method ensures that the text stream corresponding to the speech content and the emotional feature stream corresponding to the emotional acoustic features can be output simultaneously. Subsequently, the two information streams are received and fused, and the user psychological state analysis is performed, so that the user psychological state analysis can consider both the language content and the non-verbal emotional expression, thereby obtaining a more comprehensive and accurate user psychological portrait. This multi-modal fusion method significantly improves the depth and accuracy of user psychological state analysis, especially in crisis state recognition, it can detect potential risks earlier and more accurately, providing more reliable technical support for psychological counseling services.

[0070] 2. The AI question and answer expert model construction method, system and medium based on psychological counseling effectively solves the problem that the traditional noise reduction processing in the prior art weakens or removes the weak but crucial emotional performance in the user's voice while improving the clarity of the speech. By separating and capturing and processing the speech content and emotional acoustic features, especially through the auxiliary emotional pickup channel, the emotional acoustic features are specially optimized, thereby avoiding the misjudgment and filtering of emotional information by the noise reduction algorithm, so that the real emotional information contained in the user's speech can be completely preserved and transmitted. This enables the core AI psychological counseling model to receive more comprehensive and real user information, significantly improving its ability to accurately judge the user's real psychological condition, overcoming the shortcomings of the prior art AI model that does not consider the user's real emotions. Ultimately, the present application can more effectively identify the user's potential pain, struggle or negative emotions, thereby enhancing the effectiveness of the built-in rules in the model for identifying the user's crisis state, providing more accurate and reliable technical support for the AI question and answer expert model for psychological counseling. BRIEF DESCRIPTION OF DRAWINGS

[0071] The drawings described herein are used to provide further understanding of the embodiments of the present application, form a part of the present application, and do not constitute a limitation of the embodiments of the present application. In the drawings:

[0072] Figure 1 The flowchart of the AI question and answer expert model construction method based on psychological counseling of the present application;

[0073] Figure 2 The block diagram of the AI question and answer expert model construction system based on psychological counseling of the present application. DETAILED DESCRIPTION

[0074] In order to make the objects, technical solutions and advantages of the present application clearer, further detailed description will be given below in combination with embodiments and drawings, the illustrative embodiments and their description of the present application are only used to explain the present application, and do not limit the present application.

[0075] The existing voice processing system often filters out the weak sound clues carrying important emotional information when converting spoken language into text, making it difficult for artificial intelligence to understand the user's deep psychological state. Especially when the user communicates in various noisy environments through ordinary consumer-level recording equipment, the noise reduction processing function of the built-in microphone and voice-to-text module of the smartphone, while improving the clarity of the voice, inadvertently weakens or removes the emotional performance in the user's voice that is weak but crucial to the judgment of the psychological state, such as slight tremors, brief hesitation pauses, subtle changes in speech speed, or low sighs. These non-verbal details that reflect the true psychological state are easily misjudged as useless background noise and are processed together, as their acoustic characteristics may be similar to background noise or have low energy, resulting in the core AI psychological counseling model receiving emotion-neutral text, making it difficult to accurately judge the user's true psychological state, and thus affecting the effectiveness of subsequent crisis state recognition.

[0076] Therefore, the present application proposes an AI question and answer expert model construction method based on psychological counseling, which innovatively introduces a dual-channel sound pickup strategy to capture and retain emotional acoustic features while obtaining voice content, and the voice content and emotional acoustic features are separated in different channels, effectively avoiding the damage of noise reduction processing to emotional information in traditional voice processing; and the emotional acoustic features and text content are fused for psychological state analysis, thereby significantly improving the perception ability of the AI model to the user's deep psychological state, solving the problem of inaccurate psychological analysis caused by loss of emotional information in the prior art, and laying a foundation for accurate subsequent crisis state recognition.

[0077] Embodiment 1

[0078] As shown in Figure 1 , the AI question and answer expert model construction method based on psychological counseling of the present application comprises:

[0079] Step 1, obtaining dual-channel voice signals in the user's psychological counseling voice interaction process, the dual-channel voice signals comprising a first-channel voice signal and a second-channel voice signal;

[0080] In this embodiment, step 1 specifically comprises:

[0081] perceiving the environmental acoustic scene, and selecting a microphone sound pickup and preprocessing strategy according to the environmental acoustic scene;

[0082] obtaining the dual-channel voice signals through a microphone array according to the microphone sound pickup and preprocessing strategy;

[0083] Note: The environmental acoustic scene refers to the acoustic environment in which the user is located during the process of voice interaction for psychological counseling. For example, a quiet indoor environment, a noisy street, an office, etc. Its characteristics include background noise type, noise intensity, reverberation, etc.

[0084] The microphone pickup and preprocessing strategy refers to dynamically adjusting the pickup mode (such as omnidirectional, cardioid, and super cardioid), gain, noise reduction algorithm, and echo cancellation of the microphone according to different environmental acoustic scenes to optimize the capture quality of the voice signal.

[0085] The dual-channel voice signal refers to simultaneously capturing two or more independent voice signal streams through a microphone array, where the first channel serves as the main pickup channel, focusing on the clarity of voice content, and the second channel serves as the auxiliary emotional pickup channel, focusing on the sensitivity of emotional acoustic features.

[0086] Step 2: Process the first channel voice signal based on the main pickup channel to obtain the voice content, and simultaneously process the second channel voice signal based on the auxiliary emotional pickup channel to obtain the emotional acoustic features.

[0087] Note: Voice content refers to the textual information expressed in the user's spoken language.

[0088] Emotional acoustic features refer to non-verbal acoustic parameters in the voice signal that reflect emotional states, such as speech rate, pitch, volume, timbre, formant, pitch jitter, and tremor, etc.

[0089] Step 3: Convert the voice content into a text stream through speech recognition, and simultaneously form an emotional feature stream by extracting and quantizing the emotional acoustic features.

[0090] Note: The text stream refers to the sequence of text converted from the voice content through the speech recognition (ASR) system.

[0091] The emotional feature stream refers to the data sequence formed by extracting and quantizing the emotional acoustic features.

[0092] Step 4: Fuse the text stream and the emotional feature stream to obtain the fused information stream.

[0093] Step 5: Based on the fused information stream, perform user psychological state analysis to obtain the user psychological state analysis result; use the user psychological state analysis result to assist in subsequent crisis state recognition.

[0094] Note: User psychological state analysis refers to evaluating the user's mental health status, emotional tendency, stress level, etc. by comprehensively analyzing the text stream and the emotional feature stream.

[0095] The crisis state recognition refers to judging whether the user has a psychological crisis situation such as a suicide risk, a serious depression, an anxiety attack and the like needing an emergency intervention according to the user psychological state analysis result.

[0096] In the implementation of the present application, the processes of steps 1 to 5 are taken as the AI question and answer expert model construction process to analyze the user psychological state. Specifically:

[0097] First, the environmental acoustic scene is perceived, and a microphone pickup and preprocessing strategy is selected according to the environmental acoustic scene. For example, it can be identified through the built-in environmental sound sensor or by analyzing the initial captured background noise spectrum that the current environment is a quiet indoor or a noisy outdoor. If the environment is noisy, a more directional pickup mode can be selected, and a more aggressive noise reduction algorithm can be enabled; if the environment is quiet, an omnidirectional pickup mode can be selected, and a more moderate noise reduction strategy can be adopted to maximize the preservation of sound details.

[0098] Secondly, according to the selected microphone pickup and preprocessing strategy, a dual-channel voice signal is captured through a microphone array. The microphone array can be an array with two independent microphones in a physical sense, or a single microphone array that simulates a dual-channel effect through digital signal processing. The main pickup channel mainly focuses on the clarity of the voice content, and its processing flow may include traditional noise reduction, echo cancellation, voice enhancement, etc. to ensure the accuracy of voice recognition. The auxiliary emotion pickup channel focuses on capturing emotional acoustic features, and its processing strategy will pay more attention to preserving the weak details of the sound, for example, it may use a lighter noise reduction or specifically enhance the emotion-related frequency band.

[0099] Further, the text stream corresponding to the voice content and the emotion feature stream corresponding to the emotional acoustic features are output in parallel. The voice signal processed by the main pickup channel is sent to the speech recognition (ASR) module to generate the text stream; the voice signal processed by the auxiliary emotion pickup channel is sent to the emotional acoustic feature extraction module, which analyzes the pitch, volume, speech rate, timbre, formant, fundamental frequency jitter and tremor of the voice, and generates the emotion feature stream. The two data streams are generated in parallel to ensure the time synchronization of the information.

[0100] Then, the text stream and the emotion feature stream are fused for user psychological state analysis. The fusion process can use various technologies, for example, the emotional words and syntax structure in the text stream and the acoustic parameters in the emotion feature stream can be weighted and combined, or a deep learning model (such as a multi-modal Transformer) can be used to learn the complex correlation between text and acoustic features. Through this fusion, the user's psychological state can be evaluated more comprehensively and accurately, for example, to identify a situation where the text content expresses calm but the voice reveals anxiety.

[0101] The user psychological state analysis result obtained through the above process assists in subsequent crisis state recognition. For example, if the psychological state analysis result shows that the user has severe depression, suicidal tendencies, or extreme anxiety, the present application will immediately trigger a crisis warning mechanism, such as sending an alarm to a psychological counselor or providing emergency psychological support resources.

[0102] In some embodiments, the second channel speech signal is processed based on the auxiliary emotion pickup channel to obtain emotional acoustic features. However, in the implementation process, if the dynamic changes in the user's vocalization manner are not effectively perceived and adaptively adjusted, the extraction accuracy of the emotional acoustic features may be insufficient, thereby affecting the accuracy of subsequent psychological state analysis. In this regard, the present application further proposes specific steps for processing the second channel speech signal based on the auxiliary emotion pickup channel to obtain emotional acoustic features, to more accurately capture the user's emotions. Specifically, the specific steps for processing the second channel speech signal based on the auxiliary emotion pickup channel to obtain emotional acoustic features are:

[0103] Obtaining sound macro parameters of the second channel speech signal, the sound macro parameters including short-term average energy and instantaneous speech rate;

[0104] According to the changes in the sound macro parameters, judging the adjustment of the user's vocalization manner to obtain an adjustment result;

[0105] According to the adjustment result, dynamically adjusting the sensitivity factors of different types of frequency bands of the auxiliary emotion pickup channel to optimize the sound macro parameters;

[0106] Obtaining and tracking sound micro parameters of the second channel speech signal, the micro parameters including formant variation, pitch jitter, and tremor parameters;

[0107] According to the sound macro parameters and the sound micro parameters, generating multi-dimensional emotional acoustic features.

[0108] Specifically, by monitoring the short-term average energy and the instantaneous speech rate of the second-channel speech signal, the macroscopic features of the user's speech are obtained in real time. The short-term average energy reflects the loudness or intensity of the speech, and the instantaneous speech rate indicates the speed of the user's speech. Changes in these sound macroscopic parameters are often closely related to the user's emotional state or adjustment of the vocal habit. For example, when the user is emotionally excited, the speech energy may increase and the speech rate may increase; when the user is emotionally depressed or hesitant, the speech energy may decrease and the speech rate may decrease. Among them, according to the changes of the sound macroscopic parameters, the adjustment of the user's vocal manner is to analyze the trend and amplitude of these sound macroscopic parameters to identify whether the user is changing the normal vocal mode. For example, sustained low energy and slow speech rate may indicate that the user is whispering or hesitating, and sudden high energy and fast speech rate may indicate that the user is emotionally excited or the speech rate is accelerated. In practical applications, according to the adjustment result, the sensitivity factor of different types of frequency bands of the auxiliary emotional pickup channel is dynamically adjusted to optimize the sound macroscopic parameters, the purpose of which is to optimize the capture effect of the emotional acoustic features. For example, when it is judged that the user is whispering, the sensitivity factor to the low frequency band can be increased to better capture the weak physiological acoustic signals; when the user's speech rate is too fast, the sampling rate or filtering parameters can be adjusted to avoid information loss and focus on the key emotional frequency band.

[0109] In addition, obtaining and tracking the sound microscopic parameters of the second-channel speech signal is the key to in-depth analysis of the microscopic features of the speech. Formants reflect the shape and resonance characteristics of the vocal tract, and their changes are related to vowel pronunciation and emotional state (such as tension, relaxation). The jitter and shimmer parameters of the fundamental frequency respectively quantify the slight instability of the fundamental frequency period and amplitude, which are often considered as physiological indicators of emotional stress, tension or fatigue. By accurately tracking these sound microscopic parameters, the user's emotional state can be more detailed.

[0110] Finally, according to the sound macroscopic parameters and the sound microscopic parameters, the emotional acoustic features and their corresponding emotional feature streams are generated, which integrates and outputs all the emotional-related acoustic information obtained by the above monitoring, judgment, adjustment and tracking, forms a continuous, multi-dimensional emotional feature representation, and is used for subsequent psychological state analysis.

[0111] The above technical scheme can timely perceive macro changes in the user's vocalization mode by monitoring the short-term average energy and instantaneous speech rate of the second channel voice signal in real time. Based on these changes, the application can determine whether the user has adjusted his vocalization mode, such as changing from normal speech rate to whisper or rapid speech rate. Thus, the specific frequency band sensitivity factor of the auxiliary emotion pickup channel can be dynamically adjusted to ensure that key emotional acoustic information can be captured in the best state under different vocalization situations, avoiding missing or misjudging emotional signals due to improper pickup strategies. At the same time, by tracking the formant changes in the voice signal and the jitter and shimmer parameters of the fundamental frequency, the application can reveal the physiological manifestations of the user's emotions at the microscopic level, such as the tension of the vocal cords and the stability of the pronunciation. This macro and micro combined analysis method makes the extraction of emotional acoustic features more comprehensive and accurate, effectively solving the problem of insufficient emotional feature capture accuracy of traditional methods when facing dynamic changes in user vocalization mode. The emotional feature stream generated by the application can more truly and comprehensively reflect the user's current psychological state, providing a more reliable data basis for subsequent user psychological state analysis and crisis state recognition, thereby improving the professionalism and effectiveness of the AI question and answer expert model in the field of psychological counseling.

[0112] For example: Suppose a user in the process of psychological counseling initially communicates at a normal speech rate and volume. At this time, the auxiliary emotion pickup channel is set to the regular sensitivity factor. When the user starts talking about sensitive topics, his speech rate gradually slows down, his voice becomes low, and even a slight tremolo appears. The application judges that the user's vocalization mode has changed, and he may be in a state of hesitation or depression by monitoring the decrease in short-term average energy and the slowing down of instantaneous speech rate of the voice signal. Then, the low-frequency band sensitivity factor of the auxiliary emotion pickup channel is dynamically increased to better capture the weak physiological acoustic signals in the user's whisper. At the same time, the application starts to more intensively track the formant changes in the voice and the jitter and shimmer parameters of the fundamental frequency. For example, if the fundamental frequency jitter and shimmer parameters are significantly increased, it may indicate that the user's vocal cords are tense, further confirming his emotional distress. Through this dynamic adaptation and fine analysis, the application can more accurately extract emotional acoustic features such as the user's low mood and tension, and integrate them into the emotional feature stream for subsequent psychological state analysis, avoiding missing emotional information due to changes in the user's vocalization mode.

[0113] However, in practical applications, simply tracking these acoustic parameters may not be sufficient to accurately capture the true emotional state of the user. This is because the user's pronunciation characteristics have a high degree of individual variability, and changes in acoustic parameters can be due to non-emotional factors such as physiology, habitual vocalization, or pronunciation clarity. If not distinguished, it may lead to misjudgment or insufficient sensitivity of emotion recognition. If the above problems are not solved, the AI question and answer expert model may affect the reliability of the overall analysis when analyzing the user's psychological state due to inaccurate extraction of emotional acoustic features.

[0114] To this end, the present application further provides specific steps for obtaining and tracking the sound micro-parameters of the second channel voice signal, comprising:

[0115] Establishing a user pronunciation characteristic reference, which includes the user's average fundamental frequency, fundamental frequency standard deviation, jitter rate, tremor rate, and typical distribution range of formant in a neutral emotional state;

[0116] Obtaining the sound micro-parameters of the second channel voice signal and comparing the sound micro-parameters with the user pronunciation characteristic reference to obtain a comparison result;

[0117] According to the comparison result, attributing the micro-parameters of the current second channel voice signal to corresponding characteristic types; the characteristic types include physiological or habitual vocalization characteristics, emotional acoustic characteristics;

[0118] Evaluating the user's pronunciation clarity by analyzing the spectral tilt and vowel space size of the user's second channel voice signal to obtain a pronunciation clarity evaluation result;

[0119] According to the pronunciation clarity evaluation result, adjusting the recognition threshold of the emotional acoustic characteristics; including:

[0120] By analyzing whether the current voice text content contains negative emotional keywords, and whether there are physiological acoustic signals related to emotional distress in the emotional feature stream obtained by the auxiliary emotion pickup channel, the association between the pronunciation clarity evaluation result and the user's emotional distress is determined;

[0121] According to the association, adjusting the recognition threshold of the emotional acoustic characteristics, when the association indicates that low clarity is an expression of emotional distress, maintaining the recognition threshold of the emotional acoustic characteristics at a sensitive level, or reducing the recognition threshold of the emotional acoustic characteristics.

[0122] Specifically, establishing a user personalized pronunciation characteristic reference means that under the user's neutral emotional state, by collecting and analyzing his / her voice data multiple times, a set of acoustic parameters reflecting his / her unique pronunciation habits is constructed. The purpose of establishing a user personalized pronunciation characteristic reference is to provide a reliable individual reference point for subsequent emotional acoustic feature analysis, thereby effectively distinguishing individual differences from emotional changes.

[0123] wherein extracting the formant, jitter, and shimmer parameters of the current speech signal can be understood as using digital signal processing techniques to accurately calculate these key acoustic indicators from the real-time captured speech signal. For example, the formant can be extracted by linear predictive coding (LPC) or the like, while the jitter and shimmer parameters of the fundamental frequency can be obtained by a fundamental frequency detection algorithm (such as autocorrelation method, cepstrum method) combined with perturbation analysis.

[0124] In practical applications, the formant, jitter, and shimmer parameters of the current speech signal are compared with the personalized pronunciation feature benchmark. The specific method can use statistical methods (such as Z-score, T-test) or machine learning models for difference analysis. The purpose is to quantify the degree of deviation between the current pronunciation and the user's normal pronunciation.

[0125] Further, according to the comparison result, the formant, jitter, and shimmer parameters of the current speech signal are attributed to physiological or habitual vocal characteristics, or to emotional acoustic characteristics. Specifically, if the degree of deviation is within the normal fluctuation range of the benchmark, or consistent with known physiological / habitual patterns (such as nasalization when having a cold, hoarseness caused by long-term smoking), it is attributed to physiological or habitual characteristics; if the degree of deviation is significant and consistent with known emotional acoustic patterns (such as increased fundamental frequency when angry, slower speech rate when sad), it is attributed to emotional acoustic characteristics. This can be achieved by a pre-trained classifier or a rule engine.

[0126] In addition, the user's pronunciation clarity is evaluated by analyzing the spectral tilt and vowel space size of the user's second channel speech signal to determine whether the user has pronunciation disorders or uses whispering, mumbling, or other vocalization methods. The spectral tilt can reflect the distribution of speech energy in different frequency bands, and the vowel space size reflects the range of activity of the vocal organs. For example, whispering or mumbling usually results in a larger spectral tilt and a smaller vowel space. The purpose is to identify pronunciation problems that may affect the accuracy of emotional acoustic characteristics.

[0127] Thus, according to the pronunciation clarity evaluation result, the recognition threshold of the emotional acoustic characteristics is adjusted. Specifically, if the pronunciation clarity evaluation result shows that the user's pronunciation clarity is low, for example, there is whispering or mumbling, the recognition threshold of the emotional acoustic characteristics can be adjusted appropriately to make it more cautious or more inclusive when recognizing emotions, in order to avoid misjudgment of emotions due to unclear pronunciation.

[0128] The scheme of the present application introduces user personalized pronunciation feature benchmarks, so that the extraction of emotional acoustic features no longer relies on universal standards, but is based on the comparison of the user's own normal pronunciation patterns, thereby effectively eliminating the interference of individual physiological differences and habitual pronunciation methods on emotion recognition. At the same time, through the evaluation of pronunciation intelligibility, the distortion of emotional acoustic features caused by pronunciation disorders or specific pronunciation methods (such as whispering, mumbling) can be identified and compensated, and the emotion recognition threshold is dynamically adjusted accordingly, ensuring that in various complex pronunciation situations, the recognition of emotional acoustic features can still maintain high accuracy and robustness.

[0129] Through the above technical scheme, the present application can significantly improve the accuracy and personalization level of emotional acoustic feature extraction, effectively distinguish physiological, habitual pronunciation and real emotional expression, and cope with the challenges brought by pronunciation intelligibility changes. This enables the AI question and answer expert model to obtain more detailed and reliable emotional information when analyzing the user's psychological state, thereby improving the quality and effectiveness of psychological counseling.

[0130] For example: Suppose a user's fundamental frequency and jitter parameters are slightly higher than the average level during psychological counseling. If only a universal threshold is used, the present application may misjudge that the user is in a certain emotional agitation state. However, through the scheme of the present application, first, the user's personalized pronunciation feature benchmarks in a neutral emotional state are established, and it is found that the normal fundamental frequency and jitter parameters are slightly high. When the user appears similar parameters again during counseling, the present application will compare them with the personalized pronunciation feature benchmarks. If the deviation is within the normal fluctuation range of the benchmarks, it is attributed to physiological or habitual pronunciation characteristics, rather than emotional acoustic characteristics. In addition, if the user uses a whispering method due to emotional depression, resulting in an increase in speech spectrum tilt and a decrease in vowel space, the present application will evaluate that the pronunciation intelligibility is low. At this time, the recognition threshold of emotional acoustic features will be dynamically adjusted, such as appropriately relaxing the recognition requirements for certain emotional features in a whispering state, or increasing the weight of other non-acoustic cues, to avoid missing the user's deep emotional distress due to unclear pronunciation. In this way, the scheme of the present application can more accurately capture the user's real emotions and avoid misjudgment, improving the accuracy of psychological state analysis.

[0131] Specifically, the association between the current pronunciation intelligibility evaluation result and the user emotional distress is that the present application further explores the deep reasons for low intelligibility on the basis of evaluating the user's pronunciation intelligibility. The judgment is made by comprehensively analyzing data in multiple dimensions. Among them, analyzing whether the current text content contains negative emotion keywords aims to capture the negative emotions or signs of distress that the user may express from the semantic level. For example, the appearance of words such as "anxiety", "depression", "helplessness", "stress" in the text can be used as an indication of negative emotions. At the same time, whether there are physiological acoustic signals related to emotional distress in the emotional feature stream captured by the auxiliary emotion pickup channel means that through the analysis of non-language components in the speech signal, physiological reactions closely related to the emotional state are identified. These physiological acoustic signals can include heart rate variability, respiratory pattern changes, vocal cord tension, laryngeal muscle activity, etc., which often show specific patterns when the user is emotionally distressed.

[0132] According to the results of the above-mentioned association judgment, the recognition threshold of emotional acoustic features is dynamically adjusted. Among them, when the association clearly indicates that low intelligibility is a manifestation of emotional distress, for example, a large number of negative emotion keywords appear in the text, and significant physiological acoustic stress reactions are detected in the emotional feature stream, at this time, in order to ensure that the present application can more sensitively capture the emotional changes of the user, the recognition threshold of emotional acoustic features will be maintained at a sensitive level, or further reduced. Maintaining a sensitive level means that the present application will not easily ignore potential emotional signals due to low intelligibility; lowering the recognition threshold allows even weak emotional acoustic features to be identified, thereby improving the ability to identify the user's deep emotional distress.

[0133] The above technical solutions, by associating the pronunciation intelligibility evaluation result with the user's emotional distress, the present application can more accurately understand the reasons behind the user's low intelligibility. By analyzing negative emotion keywords in the text content and physiological acoustic signals in the emotional feature stream, the present application can distinguish between low intelligibility caused by physiological factors (such as cold, unclear pronunciation) and low intelligibility caused by emotional distress (such as anxiety, depression leading to slow speech and low voice). It is precisely because of this multi-dimensional and situational judgment that the recognition threshold of emotional acoustic features can be more intelligently adjusted, thereby avoiding the misjudgment that may be caused by single-dimensional judgment, ensuring that the present application can maintain high sensitivity when the user is emotionally distressed, and timely capture key emotional clues. The present application can more effectively provide personalized and accurate psychological counseling services for users, improving the application value of AI question and answer expert models in the field of mental health.

[0134] For example, assume a user is in a psychological counseling process, and the articulation evaluation result shows a low articulation. If only the articulation is considered for threshold adjustment, the application may simply increase the recognition threshold, considering the speech quality as poor, and thus may miss some subtle emotional signals. However, according to the application scheme, the application will further determine whether this low articulation is associated with emotional distress. Specifically, the application will analyze the current text content of the user, for example, the user may say "I have been feeling very depressed, and I can't even speak with enthusiasm" and other keywords containing negative emotions. At the same time, the auxiliary emotional pickup channel captures physiological acoustic signals such as increased fundamental frequency jitter and significantly slowed speech speed that are related to emotional distress. When these information are combined, indicating that the low articulation is an expression of the user's emotional distress, the application will not increase the recognition threshold, but will maintain its sensitivity level, or even lower the recognition threshold. For example, the threshold for recognizing slight emotional fluctuations is lowered from 0.5 to 0.3, so that even the user's weak sigh or slight change in tone can be recognized as potential emotional signals by the application. In this way, the application can more timely and accurately capture the user's deep psychological state, thereby providing more targeted psychological support.

[0135] However, in actual application, the user's emotional expression may not always be direct and obvious, and sometimes the deep emotional distress is expressed through ambiguous, indirect language or suppression of physiological acoustic signals, which makes the judgment relying only on keywords and direct physiological signals may have limitations, and it is difficult to accurately capture the user's potential and deeper emotional state.

[0136] To this end, the application further provides specific steps for determining the association between the articulation evaluation result and the user's emotional distress, including:

[0137] analyzing the sentiment tendency of the current speech text content to identify ambiguous or indirect negative emotional expressions in the current speech text content;

[0138] performing weak signal enhancement processing on the emotional feature stream to locally amplify physiological acoustic parameters in the emotional feature stream;

[0139] analyzing the user's historical dialogue data to establish an association pattern between low articulation, ambiguous or indirect negative emotional expression, and physiological acoustic signal suppression and emotional distress;

[0140] According to the association pattern, calculate the potential association strength between low articulation and deep emotional distress;

[0141] According to the potential association strength, determine the association between low articulation and the user's deep emotional distress.

[0142] Specifically, analyzing the sentiment orientation of the text content aims to identify potential ambiguous or indirect negative emotional expressions within the text content. This can be understood as a deep semantic analysis of non-directly expressed emotions such as sadness, anxiety, or helplessness in the user's speech, for example, by identifying metaphors, irony, euphemisms, or contextual contexts to infer their true emotions. The purpose is to make up for the shortcomings of relying solely on negative keyword recognition and more comprehensively capture the user's emotional state.

[0143] Among them, the weak signal enhancement processing of the emotional feature flow locally amplifies the physiological acoustic parameters in the emotional feature flow, which means that through a specific signal processing algorithm, the strength of physiological acoustic signals that are difficult to directly detect due to the user's deliberate suppression or weak physiological response is improved. These physiological acoustic parameters may include heart rate variability, breathing pattern, skin electrical response, and other subtle changes reflected in speech. The purpose is to reveal the physiological stress response that users may hide under the surface of calmness, providing more sensitive biological evidence for the identification of emotional distress.

[0144] In practical applications, analyzing user historical conversation data to establish the correlation pattern between low clarity, ambiguous or indirect negative emotional expression and physiological acoustic signal suppression and emotional distress means that through long-term accumulation of user interaction data, a multi-dimensional correlation model is constructed. This model maps the user's pronunciation clarity, text expression method (ambiguous or indirect negative emotion) and physiological acoustic signal suppression degree in different situations to the final diagnosed or identified emotional distress state. For example, machine learning algorithms such as support vector machines (SVM), neural networks or decision trees can be used to train these features to learn their complex nonlinear relationship with emotional distress. The purpose is to use big data and artificial intelligence technology to learn from individual historical behavior, improving the prediction accuracy of deep emotional distress.

[0145] Further, according to the correlation pattern, the potential correlation strength between low clarity and deep emotional distress is calculated. This can be understood as using the established correlation pattern to quantitatively analyze the observed low clarity, ambiguous expression and physiological signal suppression in the current voice interaction, thereby obtaining a numerical value representing the possibility or strength of the correlation between these features and the user's deep emotional distress. For example, a probability value between 0 and 1 can be output, where a higher value indicates a stronger correlation. The purpose is to provide a quantitative indicator to assist in subsequent judgment.

[0146] Thus, according to the potential correlation strength, the correlation between low intelligibility and user deep emotional distress is judged. This means that when the calculated potential correlation strength reaches the preset threshold, the present application can consider that the currently observed low intelligibility and the like phenomenon is not only a simple pronunciation problem or a surface emotion, but also has a significant correlation with the user's deeper emotional distress. The purpose is to provide a deeper and more accurate basis for crisis state recognition.

[0147] The technical scheme of the present application effectively solves the limitations of traditional methods in identifying user deep emotional distress through multi-dimensional and deep analysis. First, by analyzing the sentiment tendency of the text content, the user's speech can capture those indirect and implicit negative emotional expressions, making up for the shortcomings of relying only on keywords. Second, the weak signal enhancement processing of the emotional feature flow makes the physiological and acoustic parameters that are difficult to detect due to user suppression or weak physiological response visible, thereby revealing the user's potential physiological stress state. Further, by analyzing the user's historical dialogue data, the correlation pattern between low intelligibility, ambiguous expression and physiological acoustic signal suppression and emotional distress is established, so that the present application can learn and identify unique correlation rules from individual long-term behavior, thereby avoiding one-sided judgment of a single event. Finally, based on the established correlation patterns, the potential correlation strength between low intelligibility and deep emotional distress is calculated, and the judgment is made accordingly, so that the present application can provide a quantitative and personalized evaluation, thereby more accurately identifying the user's deep emotional distress and providing a more reliable basis for subsequent crisis state recognition.

[0148] The above technical scheme, the present application can significantly improve the recognition ability of user deep emotional distress. Specifically, by identifying ambiguous or indirect negative emotions in the text and enhancing the weak physiological and acoustic signals in the emotional feature flow, the present application can penetrate the user's surface disguise or unconscious cover-up and capture a more real emotional state. In addition, personalized correlation patterns are established using user historical dialogue data, and the potential correlation strength is calculated accordingly, so that the judgment result is more individual-specific and accurate, avoiding a one-size-fits-all judgment method. Thus, the present application can more accurately identify the user's potential psychological crisis, lay a solid foundation for timely intervention and effective psychological support, and effectively improve the professionalism and effectiveness of AI psychological counseling.

[0149] For example, assume a user is in a counseling session, and their text content does not directly use negative keywords such as "pain", "despair", etc., but repeatedly mentions "feeling a bit empty", "not feeling motivated", "it's like nothing matters" and other ambiguous expressions. At the same time, their speech clarity has slightly decreased, the speech rate has slowed down, and the physiological acoustic parameters of weak heart rate variability increase and irregular breathing patterns are detected in the emotional feature stream, but these signal strengths are not strong enough to be directly recognized as strong emotions by conventional methods.

[0150] At this time, the present solution will first analyze the sentiment orientation of these text contents, identifying the indirect negative emotions hidden behind expressions such as "empty", "not feeling motivated", "doesn't matter". Then, the weak signal enhancement processing is performed on the emotional feature stream, so that the originally weak heart rate variability increase and irregular breathing pattern physiological acoustic parameters are locally amplified, making them identifiable. Subsequently, the present invention will retrieve the user's historical dialogue data and find that in the past few counseling sessions, when similar low clarity, ambiguous expressions and physiological acoustic signal suppression occur, they are often accompanied by subsequent diagnoses of mild depression or anxiety, thereby establishing an association pattern between these features and deep emotional distress.

[0151] Based on this association pattern, the present invention will calculate the potential association strength between the current features and the user's deep emotional distress. For example, the calculation result shows that the association strength is 0.85 (higher than the preset threshold 0.7). According to this high potential association strength, the present invention can judge that the low clarity, ambiguous expression and weak physiological signal suppression exhibited by the current user are significantly associated with deep emotional distress, rather than simply fatigue or pronunciation habits. Thus, the present invention can provide the AI Q&A expert model with a deeper analysis of the user's psychological state, prompting it to adopt more targeted intervention strategies in subsequent voice interactions, such as guiding the user to explore the source of their "emptiness" or suggesting more professional psychological assessment.

[0152] However, in actual application, the user's psychological state and emotional expression method are not fixed, and the meaning of their low clarity or ambiguous expression may vary depending on the situation. Moreover, a single, static association pattern may not accurately reflect the dynamic changes in the user's deep emotional distress over time. If these problems are not addressed, it may lead to misjudgment of the user's psychological state, thereby affecting the counseling effect of the AI Q&A expert model. To address this, the present invention further proposes a more dynamic and contextual association pattern establishment and updating mechanism to improve the accuracy and adaptability of identifying the user's deep emotional distress.

[0153] Specifically, the specific steps for analyzing the user's historical dialogue data to establish an association pattern between low clarity, ambiguous or indirect negative emotional expression and physiological acoustic signal suppression and emotional distress include:

[0154] extracting the features of low intelligibility, ambiguous or indirect negative emotion expression and physiological acoustic signal suppression in this voice interaction after each psychological counseling voice interaction of the user ends;

[0155] updating the established association patterns in the historical psychological counseling voice interaction data of the user in combination with the analysis result of the psychological state of the user in this voice interaction;

[0156] adjusting the weight of the voice interaction data in the association pattern update according to the context information of this voice interaction, and establishing and maintaining multiple context-specific association patterns for the user;

[0157] When calculating the potential association strength between low intelligibility and deep emotional distress, the corresponding context-specific association pattern is selected for matching according to the current interaction context.

[0158] Among them, after each psychological counseling voice interaction of the user ends, the present invention will identify and quantify specific indicators related to emotional distress from the voice content and emotional acoustic features. Specifically, the low intelligibility features can include but are not limited to slow speech, unclear pronunciation, reduced volume, etc.; the ambiguous or indirect negative emotion expression features can refer to the ambiguous expressions, metaphors, irony, etc. in the text; the physiological acoustic signal suppression features may include reduced heart rate variability, abnormal breathing pattern, etc. The extraction of these features aims to provide fine-grained data input for subsequent association pattern update.

[0159] Further, the psychological state analysis result of each interaction, such as the type and severity of emotional distress diagnosed for the user, is used as a supervisory signal to guide the adjustment of the existing association pattern. This ensures that the association pattern can evolve dynamically with the changes in the actual psychological state of the user, improving its timeliness and accuracy. In practical applications, the update of the association pattern can use machine learning algorithms such as incremental learning or online learning, enabling the model to continuously learn from new data.

[0160] In addition, the context information can include the topic of the conversation, the time, the place, the participants, the user's emotional baseline, etc. For example, in a context where the user's emotions fluctuate greatly, the data of this interaction may be given a higher weight to more quickly reflect the current psychological state of the user; while in a context where the emotions are stable, the weight may be lower to avoid over-sensitivity. This weight adjustment mechanism makes the update of the association pattern more flexible and context-adaptive.

[0161] Considering that users may exhibit different emotional expression patterns under different situations (e.g., work stress, family conflicts, social anxiety, etc.), the present application is not limited to a single association pattern. Instead, it establishes and maintains multiple independent association patterns optimized for specific situations based on the identified situation information. For example, there can be an "emotional association pattern under work stress" and an "emotional association pattern under family relationship".

[0162] Thus, when evaluating the potential association strength between low intelligibility and deep emotional distress in the current voice interaction, the present application first identifies the situation in which the current interaction is located. Then, from the multiple situation-specific association patterns established, the most matching pattern is selected for analysis. This matching mechanism ensures that in different situations, the most relevant knowledge and experience can be used to interpret the user's acoustic and textual features, thereby improving the accuracy of association strength calculation.

[0163] The present application scheme effectively solves the limitations of traditional single association pattern in the face of complex and variable psychological state of users by introducing dynamic updating, situation weight adjustment and multiple situation-specific association patterns. Specifically, after each psychological counseling voice interaction, the present application extracts the key features of this interaction and combines the psychological state analysis results of this interaction to incrementally update the established association pattern. This continuous learning mechanism enables the association pattern to timely reflect the latest changes in the user's psychological state, avoiding misjudgment due to outdated patterns. At the same time, the data weight is adjusted according to the situation information of this interaction, ensuring that in critical situations (e.g., when the user's emotions fluctuate sharply), the data can have a greater impact on pattern updating, thereby improving the sensitivity and adaptability of the pattern to the current situation. Further, by establishing and maintaining multiple situation-specific association patterns for users, the present application can fine-tune the modeling of unique emotional expression patterns users may exhibit under different life situations (e.g., work, family, social interaction, etc.). When calculating the potential association strength, the present application intelligently selects the most matching pattern for analysis based on the current interaction situation, thereby avoiding the forced application of a general pattern that is not suitable for the current situation to a specific situation, significantly improving the recognition accuracy and situational adaptability of the association between low intelligibility, ambiguous or indirect negative emotional expression and deep emotional distress.

[0164] The above technical scheme can realize more refined, dynamic and situational analysis of the user's psychological state. The mode update after each interaction enables the AI question and answer expert model to continuously learn and adapt to the evolution of the user's psychological state, improving the accuracy of long-term tracking and evaluation. The context weight adjustment mechanism ensures that in critical or sensitive situations, the invention can more quickly and accurately capture subtle clues of user emotional changes. Most importantly, by establishing and maintaining multiple context-specific association patterns, the invention overcomes the limitations of a single pattern, enabling it to select the most appropriate analysis perspective when faced with diverse emotional expressions by users in different contexts, thereby significantly improving the accuracy and robustness of the association patterns between low clarity, ambiguous or indirect negative emotional expressions and physiological acoustic signal suppression and deep emotional distress, providing a more reliable basis for user psychological state evaluation for the AI question and answer expert model, and further improving the effectiveness of crisis identification.

[0165] For example, assume that a user, Li, is receiving psychological counseling. In the first counseling session, Li slows down his speech significantly when talking about work pressure and uses ambiguous expressions such as "a bit tired" and "not too good" multiple times. After this interaction, the invention extracts these low clarity and ambiguous expression features and, combined with the psychological counselor's preliminary assessment of Li's work pressure, updates Li's association pattern for the first time. In subsequent counseling sessions, when Li talks about family relationships, although there is a slowdown in speech, his text expression is more direct, and the physiological acoustic signal shows slight tension rather than depression. At this time, the invention will identify the current context as "family relationship" and adjust the weight of the interaction data in the association pattern update according to the context information. At the same time, the invention will establish a "family relationship context-specific association pattern" for Li. When Li again appears low clarity expression in a counseling session and the text content involves work, the invention will select the "work pressure context-specific association pattern" for matching according to the current interaction context, and calculate the potential association strength between low clarity and deep emotional distress. If the association strength between low clarity and physiological acoustic signal suppression in the work context is significantly higher than that in the family context, the invention will be more inclined to judge that Li's low clarity expression is related to deep emotional distress related to work. This dynamic updating and context matching mechanism enables the invention to more accurately understand Li's emotional expression in different contexts, avoiding misjudgment of emotional expression under work pressure as a family problem, thereby improving the accuracy of psychological state analysis.

[0166] However, in its implementation process, how to accurately identify the current interaction context and effectively select or fuse the context-specific association patterns to ensure the accuracy and robustness of the psychological state analysis is still a problem that needs further refinement. If the context recognition is inaccurate or the pattern selection is improper, it may lead to misjudgment of the user's psychological state, especially in complex or ambiguous emotional expression scenarios.

[0167] To this end, the present application further proposes the specific steps of establishing and maintaining multiple context-specific association patterns for the user according to the context information, which analyzes the text content and speech intonation features of the current voice interaction to identify context keywords and context acoustic cues, and calculates the matching degree with each context-specific association pattern based on these cues, thereby realizing more accurate pattern selection or fusion judgment.

[0168] The specific steps of establishing and maintaining multiple context-specific association patterns for the user include:

[0169] Analyzing the context information of the current voice interaction to identify context keywords and context acoustic cues; the analysis includes analyzing the text content of the current voice interaction to identify context keywords through natural language processing technology, and analyzing the speech intonation features of the current voice interaction to identify context acoustic cues;

[0170] According to the context keywords and context acoustic cues, calculate the matching degree of the current voice interaction with each context-specific association pattern;

[0171] Select the context-specific association pattern with the highest matching degree from the calculated matching degrees, or when multiple matching degrees are close, then fuse multiple context-specific association patterns for judgment.

[0172] Specifically, analyzing the text content of the current voice interaction to identify context keywords means that through natural language processing (NLP) technology, the semantic analysis of the textual information expressed by the user in the current voice dialogue is performed, and the words or phrases that can indicate specific contexts are extracted. For example, when the words "work pressure", "family conflicts", "academic anxiety" appear in the text, they can be identified as keywords of "workplace context", "family context", and "learning context", respectively. The purpose is to understand the specific background or focus of the user from the semantic level.

[0173] Among them, analyzing the speech intonation features of the current voice interaction to identify context acoustic cues can be understood as analyzing the prosody, pitch, speech rate, volume, and other acoustic parameters of the voice signal to identify non-verbal information related to specific contexts. For example, the acceleration of speech rate may be related to tense or excited contexts, and the low tone may be related to depressed or tired contexts. The purpose is to supplement and verify the context information from the non-verbal level to improve the accuracy of context recognition.

[0174] In actual applications, the matching degrees of the current voice interaction and the context-specific association modes are calculated according to the context keywords and the context acoustic cues, specifically, the recognized context keywords and the context acoustic cues are taken as input features, and similarity calculation is performed on the features vectors representing different context-specific association modes which are pre-trained. For example, cosine similarity, Euclidean distance and the like can be used to quantify the fitting degree of the current voice interaction and each context mode. The purpose is to provide a quantitative basis for subsequent mode selection.

[0175] Further, the context-specific association mode with the highest matching degree is selected from the calculated matching degrees, that is, after the matching degrees of the current voice interaction and all context-specific association modes are calculated, the mode with the largest matching degree value is directly selected as the applicable mode of the current voice interaction. Or, when multiple matching degrees are close, multiple context-specific association modes are fused for judgment, that is, when the matching degrees of two or more context modes are not significantly different and it is difficult to clearly select a single mode, the present application will no longer be limited to a single mode, but will comprehensively consider and fuse these modes with close matching degrees to form a more comprehensive and robust psychological state judgment. This fusion judgment can effectively deal with the complexity of user emotional expression and the ambiguity of context.

[0176] The present application can more finely recognize the context information of the current voice interaction by introducing the analysis of the text content and the voice tone features of the current voice interaction. Specifically, the recognition of the context keywords captures the core problem concerned by the user from the semantic level, and the recognition of the context acoustic cues supplements the expression of emotion and context from the non-verbal level, and the combination of the two makes the understanding of the context of the current voice interaction more comprehensive and accurate. It is precisely due to the accurate capture of the context information that it is possible to subsequently calculate the matching degrees of the current voice interaction and the context-specific association modes. By quantifying the matching degrees, the present application can objectively evaluate the association degree of the current voice interaction and different historical context modes, thereby avoiding the deviation that may be caused by subjective judgment or single mode selection. When the context-specific association mode with the highest matching degree is selected, it can ensure that the most relevant historical experience is used for analysis in a clear context. When multiple modes have close matching degrees, the judgment is made by fusing multiple modes, which can effectively handle the situation of ambiguous context or complex emotion, avoid misjudgment due to the limitations of a single mode, and thus improve the accuracy and adaptability of psychological state analysis.

[0177] The above technical scheme can significantly improve the accuracy and robustness of identifying deep emotional distress of a user in a complex psychological counseling scene. Compared with a basic scheme that only relies on historical data to update the association mode, the present application can more accurately identify the current situation by analyzing the text content and voice tone features of the current voice interaction in real time, thereby ensuring that the selected situation-specific association mode is highly consistent with the actual situation of the user. This fine-grained situation matching mechanism effectively avoids misjudgment of the psychological state caused by inaccurate situation recognition, especially in the case of ambiguous emotional expression or changing situation of the user, by calculating the matching degree and selecting or fusing the mode, the potential psychological distress of the user can be more comprehensively and deeply understood. Therefore, the present application has made significant progress in the accuracy, adaptability and robustness of user psychological state analysis.

[0178] For example, assume that a user is having a psychological counseling voice interaction, and the text content frequently appears the keywords of "insomnia", "anxiety", "low work efficiency", etc., and the voice tone shows the emotional acoustic cues of fast speech rate, high pitch and slight tremor.

[0179] Firstly, the present application analyzes the text content of the current voice interaction and identifies the situation keywords of "insomnia", "anxiety", "low work efficiency", etc., which may point to "workplace stress" or "life distress" situations. At the same time, the present application analyzes the voice tone features of the current voice interaction and identifies the situation acoustic cues of fast speech rate, high pitch and tremor, which may further strengthen the emotional state of "anxiety" or "nervousness" and be highly related to "stress situation".

[0180] Then, the present application calculates the matching degree of the current voice interaction with each situation-specific association mode established in the user's historical data according to the identified situation keywords and situation acoustic cues. For example, the present application may find that the matching degree of the current voice interaction with the "workplace stress situation mode" is 0.85, the matching degree with the "family conflict situation mode" is 0.30, and the matching degree with the "academic anxiety situation mode" is 0.25.

[0181] In this case, since the matching degree of the "workplace stress situation mode" is the highest, this mode will be selected as the applicable mode for the current voice interaction for subsequent calculation of the potential association strength of low intelligibility and deep emotional distress.

[0182] However, in another case, if it is found that the current voice interaction matches the "workplace stress situation pattern" with a degree of 0.70 and the "interpersonal relationship distress situation pattern" with a degree of 0.68, the degrees of matching are close. At this time, the present application will no longer simply select a single pattern, but will fuse the "workplace stress situation pattern" and the "interpersonal relationship distress situation pattern" for judgment. This fusion judgment can more comprehensively consider the complex psychological distress that the user may face, such as interpersonal tension caused by workplace stress, thereby providing more accurate psychological state analysis.

[0183] However, in actual operation, simply fusing these patterns can cause conflicts or redundancies in the emotional feature parameters between different patterns, thereby affecting the accuracy and reliability of the final psychological state analysis. If the above problems are not solved, the fusion result can not truly reflect the deep emotional state of the user, and can even cause misjudgment, thereby affecting the effectiveness of crisis state recognition.

[0184] To this end, the present application further proposes specific steps for fusing multiple situation-specific association patterns for judgment, including:

[0185] Cross-comparing the emotional feature parameters within each situation-specific association pattern to identify conflict parameters and redundant parameters;

[0186] Prioritizing or weightedly averaging the conflict parameters, and de-duplicating or feature-selecting the redundant parameters;

[0187] According to the complexity and contradiction degree of the current emotional expression of the user, adjusting the parameter processing strategies of prioritization, weighted averaging, de-duplication or feature selection to obtain processed emotional feature parameters;

[0188] And fusing the processed emotional feature parameters to obtain a fusion judgment result.

[0189] Specifically, the emotional feature parameters within each situation-specific association pattern are cross-compared to find conflict parameters with significantly different parameter values or their indicated emotional tendencies, and redundant parameters that repeatedly appear in different patterns and have the same meaning. For example, one pattern can indicate that the user is in mild anxiety, and another pattern can indicate that the user is in moderate depression, at which time the physiological acoustic parameters related to anxiety and depression (such as speech rate, pitch) can conflict. The conflict parameter refers to a parameter that gives mutually contradictory or significantly deviating evaluation results for the same emotional dimension or physiological indicator in different situation-specific association patterns. The redundant parameter refers to a parameter that gives the same or highly similar evaluation results for the same emotional dimension or physiological indicator in different situation-specific association patterns, and has a low information gain.

[0190] The priority ranking or weighted average processing of the conflict parameters aims to effectively solve the emotional information conflicts between different modes. The priority ranking can give higher weights to specific situational modes or specific emotional features according to preset rules or expert knowledge, for example, when identifying crisis states, parameters related to high-risk emotions such as despair and suicidal ideation can be given higher priority. The weighted average processing is a numerical compromise of conflict parameters according to the matching degree, reliability or preset weight of each mode to obtain a comprehensive evaluation value.

[0191] In practical applications, the purpose of de-duplication or feature selection of redundant parameters is to simplify the emotional feature set and avoid information overload and waste of computing resources. De-duplication directly removes duplicate parameters, while feature selection selects the most representative parameters from multiple similar parameters according to their information amount, discriminability or relevance to the target emotional state.

[0192] Further, according to the complexity and contradiction of the user's current emotional expression, the parameter processing strategy of priority ranking, weighted average processing, de-duplication or feature selection is adjusted, which aims to make the fusion process more adaptive and intelligent. For example, when the user's emotional expression is complex and contradictory, a more refined weighted average strategy or more stringent feature selection may be needed to avoid misjudgment.

[0193] The present solution cross-compares emotional feature parameters within each situational-specific correlation mode, which can effectively identify potential conflict parameters and redundant parameters. Therefore, the priority ranking or weighted average processing of conflict parameters can effectively solve the contradictory information between different modes and ensure the logical consistency of the fusion result. At the same time, de-duplication or feature selection of redundant parameters can optimize the feature set, improve processing efficiency and reduce unnecessary interference. Further, the parameter processing strategy is dynamically adjusted according to the complexity and contradiction of the user's current emotional expression, so that the fusion process can flexibly adapt to the subtle changes of the user's emotions, thereby obtaining accurate and reliable psychological state analysis results even in the case of multiple situational information interweaving.

[0194] The above technical solutions can effectively solve the conflict and redundancy problems that may occur during the fusion of multiple situational-specific correlation modes, significantly improving the accuracy and robustness of user psychological state analysis. Especially when the user's emotional expression is complex or ambiguous, this solution can accurately capture and understand the user's deep emotional distress through refined parameter processing strategies, thereby providing a more solid data foundation for subsequent crisis state identification, avoiding misjudgment due to information conflicts or redundancies, and improving the practicality and reliability of AI question and answer expert models in the field of psychological counseling.

[0195] For example, assume that a user's voice signal and text content simultaneously activate two context-specific association patterns, "work stress context" and "family conflict context", in a voice interaction for psychological counseling. In the "work stress context" pattern, the user exhibits faster speech rate and higher pitch (indicating anxiety), while in the "family conflict context" pattern, the user exhibits lower tone and increased pauses (indicating depression). At this point, a consistency check on the emotional feature parameters within these two patterns identifies a conflict between the speech rate and pitch parameters. For example, the "work stress context" pattern might attribute the high speech rate to anxiety, while the "family conflict context" pattern might attribute the low speech rate to depression. Meanwhile, both patterns might include the "voice energy" parameter, which could be a redundant parameter.

[0196] According to the cross-comparison of emotional feature parameters within each context-specific association pattern, conflicting parameters can be processed. For example, if the current text content contains more negative keywords related to "family conflict" and the auxiliary emotional pickup channel captures more intense physiological acoustic signals pointing to depression, the parameters related to depression in the "family conflict context" pattern can be given higher priority. Alternatively, the speech rate and pitch parameters can be weighted and averaged to obtain a comprehensive emotional indicator based on the matching degree of the two patterns. For redundant parameters, such as "voice energy", de-duplication processing can be performed to retain only the most representative parameter.

[0197] Further, if the evaluation finds that the user's current emotional expression has high complexity (e.g., there is a certain degree of inconsistency between the text content and the emotional feature stream), the parameter processing strategy can be adjusted. For example, increase the weight of the emotional feature stream to focus more on the deep emotional state reflected by the physiological acoustic signals, and interpret the ambiguous expression in the text content more cautiously. In this way, the processed emotional feature parameters can more accurately and comprehensively reflect the user's true psychological state under multiple contexts, ultimately obtaining a fusion judgment result, such as the user being in a "mild depression with anxiety" state, thereby providing a more detailed basis for crisis state recognition.

[0198] When fusing multiple context-specific association patterns for judgment, although the parameter processing strategy is adjusted according to the complexity and contradiction of the user's current emotional expression, there is a lack of specific guidance and quantitative standards for how to accurately assess this complexity and contradiction, and how to dynamically adjust based on it. This may result in inaccurate parameter adjustment, which cannot fully respond to subtle changes in user emotional expression and potential contradictory information, thereby affecting the accuracy of the final psychological state analysis.

[0199] In response, this invention further proposes a method for adjusting parameter processing strategies such as priority ranking, weighted averaging, deduplication, or feature selection based on the complexity and contradiction of the user's current emotional expression. The aim is to achieve more refined parameter adjustment by quantitatively evaluating the intensity and consistency of emotions, thereby improving the accuracy and robustness of emotional feature fusion.

[0200] Based on the complexity and contradiction of the user's current emotional expression, the parameter processing strategies, such as priority ranking, weighted averaging, deduplication, or feature selection, are adjusted to obtain the processed emotional feature parameters. Specific steps include:

[0201] By analyzing the deviation and duration of physiological acoustic parameters in the emotional feature stream, as well as the density and emotional intensity of negative emotional keywords in the text, the intensity of user emotional expression is assessed, and emotional intensity is obtained.

[0202] By calculating the difference between the emotional polarity of the text and the emotional polarity reflected by the physiological acoustic signal, the consistency between the text content and the emotional feature flow is evaluated, and a consistency index is obtained.

[0203] Based on the emotion intensity and consistency index, the parameter processing strategies of priority ranking, weighted average processing, deduplication, or feature selection are dynamically adjusted. When the emotion intensity is high and the consistency is low, the weight of the emotion feature stream is increased.

[0204] Specifically, assessing the intensity of a user's emotional expression involves analyzing physiological acoustic parameters in the emotional feature stream, such as tone, speech rate, volume, fundamental frequency jitter, and tremor, to determine the degree and duration of deviation from a normal baseline. Simultaneously, it combines the frequency, density, and emotional polarity of negative emotional keywords in the text content to comprehensively quantify the intensity of the user's emotion. For example, when physiological acoustic parameters deviate significantly from the baseline for an extended period, and the text contains a high density of strongly negative words, the emotional intensity can be determined to be high.

[0205] Assessing the consistency between text content and emotional feature stream can be understood as determining the internal and external consistency of a user's expression by calculating the difference between the text's emotional polarity and the emotional polarity reflected in the physiological acoustic signal. Textual emotional polarity can be obtained through sentiment analysis of the text content using natural language processing techniques, while the emotional polarity reflected in the physiological acoustic signal is obtained through pattern recognition and classification of physiological acoustic parameters in the emotional feature stream. When the two polarities are opposite or significantly different, it indicates that the user's emotional expression may be contradictory or repressed.

[0206] In practical applications, according to the emotional intensity and consistency index, the parameter processing strategy of priority sorting, weighted average processing, deduplication or feature selection is dynamically adjusted. Specifically, when the evaluation result shows that the user's emotional intensity is high and the consistency between the text content and the emotional feature flow is low, it usually indicates that the user may have deep emotional distress or depression, and his text expression may not fully reflect his true psychological state. In this case, in order to more accurately capture the user's potential emotional information, the weight of the emotional feature flow in the subsequent fusion judgment is increased, so that it occupies a more important position in the final psychological state analysis. On the contrary, if the emotional intensity is low and the consistency is high, the weight of the emotional feature flow can be appropriately reduced or the default weight can be maintained.

[0207] The present application solves the problem of lack of fine basis for parameter adjustment strategy in the prior art by introducing quantitative evaluation of user emotional expression intensity and consistency between text and emotional feature flow. Specifically, by analyzing the deviation degree and duration of physiological acoustic parameters in the emotional feature flow, as well as the density and emotional intensity of negative emotional keywords in the text, the intensity of user emotion can be evaluated comprehensively and objectively. At the same time, by calculating the difference between the text emotional polarity and the emotional polarity reflected by the physiological acoustic signal, the internal and external contradictions of user emotional expression can be effectively identified, such as insincerity or emotional depression. It is precisely because of the accurate quantification of emotional intensity and consistency that the present application can dynamically and intelligently adjust the parameter processing strategy of priority sorting, weighted average processing, deduplication or feature selection according to the true complexity and contradiction degree of the user's current emotion. For example, when high emotional intensity and low consistency are identified, the present application can actively increase the weight of the emotional feature flow, so that in the case of possible verbal disguise or emotional depression of the user, its true psychological state can still be effectively captured and analyzed by physiological acoustic signals, avoiding false judgments caused by the misleading of a single information source.

[0208] The above technical solutions, the present application can realize a deeper understanding of user emotional expression and more accurate psychological state analysis. By quantitatively evaluating the emotional intensity and consistency between the text and the emotional feature flow, the present application can more accurately identify the true state of the user's emotion, especially in the case of complex, contradictory or suppressed expression of the user's emotion. This ability to dynamically adjust the parameter processing strategy makes the fusion process of emotional features more adaptive and robust, significantly improving the accuracy and reliability of the AI question and answer expert model in analyzing the user's psychological state in the psychological counseling scenario, so as to more timely and effectively identify the potential crisis state and provide more targeted and personalized psychological support for the user.

[0209] For example, assume that a user is in a psychological counseling voice interaction, and the text content expresses relatively calm, such as "I feel a little tired recently, but there is no big problem." However, through the analysis of the emotion feature stream captured by the auxiliary emotion pickup channel, it is found that there is obvious pitch jitter and tremor in the user's voice tone, the speech speed is slow and the volume is low, and the physiological acoustic parameters deviate from the normal benchmark for a long time, indicating that the emotional intensity is high. At the same time, the text sentiment polarity analysis result is neutral to negative, while the emotion polarity reflected by the physiological acoustic signal is obviously negative emotion, and there is a significant difference between the two, that is, the consistency is low.

[0210] The present application will identify that the current emotional intensity of the user is high and the consistency between the text and the emotion feature stream is low. In this case, the present application will dynamically adjust the processing strategy of the emotion feature parameter, which is specifically manifested as increasing the weight of the emotion feature stream in the subsequent fusion judgment. For example, when the emotion feature parameters from different context-specific association modes are fused, the physiological acoustic parameters (such as pitch jitter, tremor, speech speed, etc.) in the emotion feature stream will be given higher priority or greater weighting coefficients. Thus, even if the text content fails to fully reveal the user's deep emotional distress, the present application can more accurately judge the user's possible anxiety or depression emotion through the key analysis of the emotion feature stream, thereby providing more reliable input for the psychological state analysis module and helping the crisis recognition module to more timely identify potential psychological crisis.

[0211] Embodiment 2

[0212] As shown in Figure 2 The difference between the present embodiment and embodiment 1 is that the present embodiment provides an AI question and answer expert model construction system based on psychological counseling, which corresponds one-to-one with the AI question and answer expert model construction method based on psychological counseling of embodiment 1; the system comprises:

[0213] A dual-channel voice acquisition unit is configured to acquire dual-channel voice signals in a user's psychological counseling voice interaction process, and the dual-channel voice signals include first-channel voice signals and second-channel voice signals.

[0214] A preprocessing unit is configured to process the first-channel voice signals based on a main pickup channel to obtain voice content, and simultaneously process the second-channel voice signals based on an auxiliary emotion pickup channel to obtain emotion acoustic features.

[0215] A conversion unit is configured to convert the voice content into a text stream after voice recognition, and simultaneously form an emotion feature stream after extracting and quantizing the emotion acoustic features.

[0216] A fusion unit is configured to fuse the text stream and the emotion feature stream to obtain a fused information stream.

[0217] The mental state analysis unit is configured to analyze the user's mental state based on the fused information stream to obtain a user mental state analysis result.

[0218] As a further implementation, the system further comprises: identifying a crisis state based on the user mental state analysis result.

[0219] The execution process of each unit can be performed according to the AI question and answer expert model construction method based on psychological counseling in Embodiment 1, and will not be described again in this embodiment.

[0220] Meanwhile, the application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the AI question and answer expert model construction method based on psychological counseling.

[0221] The above specific embodiments further illustrate the purpose, technical solutions and beneficial effects of the application. It should be understood that the above description is only a specific embodiment of the application and is not intended to limit the protection scope of the application. Any modification, equivalent replacement, improvement, etc. within the spirit and principles of the application should be included in the protection scope of the application.

Claims

1. A method for constructing an AI question-answering expert model based on psychological counseling, characterized in that, The method includes: Acquire dual-channel voice signals during the user's psychological counseling voice interaction process, wherein the dual-channel voice signals include a first-channel voice signal and a second-channel voice signal; The first channel speech signal is processed based on the main pickup channel to obtain speech content, and the second channel speech signal is processed based on the auxiliary emotion pickup channel to obtain emotional acoustic features. The speech content is converted into a text stream after speech recognition, and the emotional acoustic features are extracted and quantized to form an emotional feature stream. The text stream and the emotion feature stream are fused to obtain a fused information stream; Based on the fused information flow, user psychological state analysis is performed; The second channel speech signal is processed based on the auxiliary emotion pickup channel to obtain emotional acoustic features, including: Acquire the macroscopic sound parameters of the second channel speech signal, the macroscopic sound parameters including short-term average energy and instantaneous speech rate; Based on the changes in the macroscopic sound parameters, the adjustment of the user's vocalization method is determined, and the adjustment result is obtained; Based on the adjustment results, the sensitivity factors of different frequency bands in the auxiliary emotion pickup channel are dynamically adjusted to optimize the macroscopic sound parameters; Acquire and track the sound micro-parameters of the second channel speech signal, including formant variation, fundamental frequency jitter and tremor parameters; Based on the macroscopic and microscopic sound parameters, multidimensional emotional acoustic features are generated. Acquire and track the microscopic acoustic parameters of the second channel speech signal, including: Establish a user pronunciation feature benchmark, which includes the user's average fundamental frequency, fundamental frequency standard deviation, jitter rate, tremor rate, and typical distribution range of formants in a neutral emotional state; The sound micro-parameters of the second channel speech signal are obtained, and the sound micro-parameters are compared with the user pronunciation feature benchmark to obtain the comparison result; Based on the comparison results, the microscopic parameters of the current second channel speech signal are attributed to corresponding feature types; the feature types include physiological or habitual vocalization features and emotional acoustic features. Acquiring and tracking the sound micro-parameters of the second channel speech signal also includes: The clarity of a user's speech is assessed by analyzing the spectral tilt and vowel space size of the user's second-channel speech signal, and the speech clarity assessment results are obtained. Based on the pronunciation clarity assessment results, adjust the recognition threshold of emotional acoustic features; including: By analyzing whether the current speech text content contains negative emotion keywords, and whether there are physiological acoustic signals related to emotional distress in the emotional feature stream obtained by the auxiliary emotional pickup channel, the correlation between the speech clarity assessment result and the user's emotional distress is determined. Based on the association, the recognition threshold of the emotional acoustic features is adjusted. When the association indicates that low resolution is an expression of emotional distress, the recognition threshold of the emotional acoustic features is maintained at a sensitive level, or the recognition threshold of the emotional acoustic features is reduced.

2. The method for constructing an AI question-answering expert model based on psychological counseling according to claim 1, characterized in that, Acquire dual-channel audio signals, including: Microphone pickup and preprocessing strategies are selected based on the ambient acoustic scenario; the ambient acoustic scenario refers to the acoustic environment in which the user is located during the voice interaction process of psychological counseling. According to the microphone pickup and preprocessing strategy, dual-channel speech signals are acquired through a microphone array; the microphone pickup and preprocessing strategy refers to dynamically adjusting the microphone pickup mode, gain, noise reduction algorithm and echo cancellation parameters according to different environmental acoustic scenarios to optimize the capture quality of the speech signal.

3. The method for constructing an AI question-answering expert model based on psychological counseling according to claim 1, characterized in that, Determining the correlation between the pronunciation clarity assessment results and the user's emotional distress includes: Analyze the sentiment tendency of the current audio-text content and identify vague or indirect negative emotional expressions in the current audio-text content; Weak signal enhancement processing is applied to the emotional feature stream to locally amplify the physiological acoustic parameters in the emotional feature stream; Analyze users' historical dialogue data to establish a correlation pattern between the low resolution, the vague or indirect expression of negative emotions, the inhibition of physiological acoustic signals, and emotional distress. Based on the association pattern, calculate the potential association strength between low clarity and deep emotional distress; Based on the strength of the potential association, determine the association between low resolution and the user's deep emotional distress; This includes analyzing user history dialogue data to establish a correlation pattern between the low resolution, the ambiguous or indirect negative emotional expression, and the inhibition of physiological acoustic signals and emotional distress, including: After each psychological counseling voice interaction with the user, the features of low-resolution, vague or indirect negative emotional expressions and physiological acoustic signal suppression in this voice interaction are extracted. Based on the analysis results of the user's psychological state in this voice interaction, the established association patterns in the user's historical psychological counseling voice interaction data are updated; Based on the contextual information of this voice interaction, the weight of this voice interaction data in the association pattern update is adjusted, and multiple context-specific association patterns are established and maintained for the user. When calculating the potential correlation strength between low resolution and deep emotional distress, a corresponding context-specific correlation pattern is selected for matching based on the current interaction context.

4. The method for constructing an AI question-answering expert model based on psychological counseling according to claim 3, characterized in that, To establish and maintain multiple context-specific association patterns for users, including: The contextual information of this voice interaction is analyzed to identify contextual keywords and contextual acoustic cues; the analysis includes: analyzing the text content of this voice interaction to identify contextual keywords; and analyzing the intonation features of this voice interaction to identify contextual acoustic cues. Based on the contextual keywords and contextual acoustic cues, calculate the matching degree between this voice interaction and the context-specific association patterns; The context-specific association pattern with the highest matching degree is selected from the calculated matching degree, or when multiple matching degrees are close, multiple context-specific association patterns are fused for judgment; wherein, fusing multiple context-specific association patterns for judgment includes: Cross-comparison of emotional feature parameters within context-specific association patterns is performed to identify conflicting and redundant parameters. The conflict parameters are prioritized or weighted averaged, and the redundant parameters are deduplicated or feature-selected. Based on the complexity and contradiction of the user's current emotional expression, the parameter processing strategies of priority sorting, weighted average processing, deduplication, or feature selection are adjusted to obtain the processed emotional feature parameters. The processed emotional feature parameters are then fused to obtain a fusion judgment result.

5. The method for constructing an AI question-answering expert model based on psychological counseling according to claim 4, characterized in that, Based on the complexity and contradiction of the user's current emotional expression, the parameter processing strategies of priority sorting, weighted averaging, deduplication, or feature selection are adjusted to obtain processed emotional feature parameters, including: By analyzing the deviation and duration of physiological acoustic parameters in the emotional feature stream, as well as the density and emotional intensity of negative emotional keywords in the text, the intensity of user emotional expression is assessed, and emotional intensity is obtained. By calculating the difference between the emotional polarity of the text and the emotional polarity reflected by the physiological acoustic signal, the consistency between the text content and the emotional feature flow is evaluated, and a consistency index is obtained. Based on the emotion intensity and the consistency index, the parameter processing strategies for priority sorting, weighted average processing, and deduplication or feature selection are dynamically adjusted, wherein when the emotion intensity is high and the consistency is low, the weight of the emotion feature stream is increased.

6. A system for constructing an AI question-answering expert model based on psychological counseling, characterized in that, The system includes: A dual-channel voice acquisition unit is used to acquire dual-channel voice signals during the user's psychological counseling voice interaction process, wherein the dual-channel voice signals include a first-channel voice signal and a second-channel voice signal. The preprocessing unit is used to process the first channel speech signal based on the main pickup channel to obtain speech content, and at the same time, to process the second channel speech signal based on the auxiliary emotion pickup channel to obtain emotional acoustic features. The conversion unit is used to convert the speech content into a text stream after speech recognition, and at the same time, to extract and quantize the emotional acoustic features to form an emotional feature stream. The fusion unit is used to fuse the text stream and the emotion feature stream to obtain a fused information stream; The psychological state analysis unit is used to perform user psychological state analysis based on the fused information flow. The second channel speech signal is processed based on the auxiliary emotion pickup channel to obtain emotional acoustic features, including: Acquire the macroscopic sound parameters of the second channel speech signal, the macroscopic sound parameters including short-term average energy and instantaneous speech rate; Based on the changes in the macroscopic sound parameters, the adjustment of the user's vocalization method is determined, and the adjustment result is obtained; Based on the adjustment results, the sensitivity factors of different frequency bands in the auxiliary emotion pickup channel are dynamically adjusted to optimize the macroscopic sound parameters; Acquire and track the sound micro-parameters of the second channel speech signal, including formant variation, fundamental frequency jitter and tremor parameters; Based on the macroscopic and microscopic sound parameters, multidimensional emotional acoustic features are generated. Acquire and track the microscopic acoustic parameters of the second channel speech signal, including: Establish a user pronunciation feature benchmark, which includes the user's average fundamental frequency, fundamental frequency standard deviation, jitter rate, tremor rate, and typical distribution range of formants in a neutral emotional state; The sound micro-parameters of the second channel speech signal are obtained, and the sound micro-parameters are compared with the user pronunciation feature benchmark to obtain the comparison result; Based on the comparison results, the microscopic parameters of the current second channel speech signal are attributed to corresponding feature types; the feature types include physiological or habitual vocalization features and emotional acoustic features. Acquiring and tracking the sound micro-parameters of the second channel speech signal also includes: The clarity of a user's speech is assessed by analyzing the spectral tilt and vowel space size of the user's second-channel speech signal, and the speech clarity assessment results are obtained. Based on the pronunciation clarity assessment results, adjust the recognition threshold of emotional acoustic features; including: By analyzing whether the current speech text content contains negative emotion keywords, and whether there are physiological acoustic signals related to emotional distress in the emotional feature stream obtained by the auxiliary emotional pickup channel, the correlation between the speech clarity assessment result and the user's emotional distress is determined. Based on the association, the recognition threshold of the emotional acoustic features is adjusted. When the association indicates that low resolution is an expression of emotional distress, the recognition threshold of the emotional acoustic features is maintained at a sensitive level, or the recognition threshold of the emotional acoustic features is reduced.

7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for constructing an AI question-answering expert model based on psychological counseling as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Good friend recommendation method and apparatus

    CN105159979A

  • Noise reduction method and device, terminal and storage medium

    CN116312586A

  • Speaker role recognition method and device based on double-recording system

    CN118447855A

  • Voiceprint recognition method based on smart home scene and related equipment

    CN120431935A

  • Multi-mode psychological counseling system based on artificial intelligence

    CN120690385A