Psychological counseling system and method based on AI model and voice communication

The psychological counseling system, which uses multiple modules working together, solves the problem of insufficient authenticity and immediacy of interaction in existing psychological counseling systems, realizes real-time assessment of emotional states, personalized mental health services, and data closed-loop optimization, and improves the system's automation level and user satisfaction.

CN120748723APending Publication Date: 2025-10-03SHANGHAI ADI FEIXIU INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510858463.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing psychological counseling systems do not make it easy for users to obtain mental health support when using voice and AI models for conversations, and the authenticity and immediacy of the interactions are low.

Method used

It uses interactive voice response module, emotion classification algorithm, hot word recognition algorithm, intelligent interruption algorithm, manual consultation module, information acquisition module, historical conversation data acquisition module, intelligent psychological assessment module, AI big model of experts in the psychological field, consultation data storage module, emotional health analysis report generation module and other modules to work together to achieve real-time assessment of emotional status, risk grading, personalized mental health services and data closed-loop optimization.

Benefits of technology

It improves the automation level and security of the psychological counseling system, ensures the timeliness of crisis intervention, enhances the naturalness of conversations and user participation, provides highly personalized mental health services, and improves service coverage and professionalism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748723A_ABST
    Figure CN120748723A_ABST
Patent Text Reader

Abstract

The invention discloses a psychological counseling system and method based on an AI model and voice communication. The psychological counseling system comprises an interactive voice response module; a high risk monitoring module; a psychological counseling AI component; a manual consultation module; an information acquisition module; and a historical dialogue data acquisition module. The psychological counseling method and the psychological counseling device solve the problems that in an existing psychological counseling process, a user cannot conveniently use voice and an AI model to carry out dialogue, a channel for people to obtain psychological health support is limited, and the authenticity and instantaneity of interaction are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and specifically to a psychological consultation system and consultation method based on AI models and voice calls. Background Art

[0002] Psychological counseling based on AI models and voice calls is a form of service that combines advanced artificial intelligence technology with traditional voice communication methods to provide users with mental health support. This service is usually carried out through telephone, Internet voice calls, etc., and users can talk to an AI-driven psychological counseling assistant to get immediate mental health support. Because they are AI-based systems, they are able to provide support without time restrictions, which is very useful for people who need help at any time. Using AI for counseling can help users discuss sensitive topics without revealing their identities, thereby adding a layer of privacy protection. Modern AI psychological counseling tools use NLP technology to understand and respond to users' words, which makes the interaction more natural and smooth, and can better understand the user's emotional state. By analyzing the content of the conversation, the AI ​​system is able to provide personalized mental health advice and support, adjusting its response according to each user's unique situation.

[0003] During the existing psychological counseling process, it is not convenient for users to use voice and AI models to communicate, which limits people's channels for obtaining mental health support and reduces the authenticity and immediacy of the interaction. Summary of the Invention

[0004] The purpose of the present invention is to provide a psychological counseling system and counseling method based on AI models and voice calls, which solves the problem that during the existing psychological counseling process, it is inconvenient for users to use voice and AI models to communicate, which limits people's channels for obtaining mental health support and reduces the authenticity and immediacy of the interaction.

[0005] To achieve the above objectives, the present invention provides the following technical solutions: a psychological consultation system based on AI models and voice calls, comprising:

[0006] Interactive Voice Response Module;

[0007] High-risk monitoring module;

[0008] Psychological counseling AI component;

[0009] Manual consultation module;

[0010] Information acquisition module;

[0011] Historical conversation data acquisition module;

[0012] Intelligent psychological assessment module;

[0013] AI big model of experts in the field of psychology;

[0014] Consulting data storage module;

[0015] Data processing and recycling module;

[0016] Emotional health analysis report generation module.

[0017] As a preferred psychological consultation system based on AI model and voice call of the present invention, the interactive voice response module receives the user's voice call and transfers it to manual consultation or AI model consultation according to the user's needs. When identifying the user's needs, the interactive voice response module converts the user's voice input into text through voice recognition, analyzes the user's intention in combination with natural language processing technology, and transfers it to manual consultation or AI model consultation according to preset rules. When transferring to manual consultation, the call is transferred to an idle customer service in real time through the API. The manual consultation transfer needs to be triggered when the user explicitly requests it. After the triggering, customer service resources are allocated through queue management.

[0018] As a preferred psychological counseling system based on AI model and voice call of the present invention, the high-risk monitoring module is turned on when the user's phone is connected, and determines whether the call should be transferred to manual according to the user's voice emotion. When it is determined that the user is a high-risk user, it is automatically transferred to manual. When it is determined that the user's emotion is normal, it is transferred to AI model consultation. The high-risk monitoring module includes a voice emotion detection module and a hot word recognition system. The voice emotion detection module detects the user's emotion, and the hot word recognition system detects high-risk words in the user's voice;

[0019] When performing emotion recognition, an emotion classification algorithm is used. The emotion classification algorithm is as follows:

[0020]

[0021] Emotion is the user's emotion category;

[0022] Features are features extracted from speech;

[0023] P(Features|Emotion) is the probability of these features appearing under a certain emotion;

[0024] P(Emotion) is the prior probability of a certain emotion;

[0025] P(Features) is the overall probability of the feature;

[0026] When performing hot word recognition, a hot word recognition algorithm is used. The hot word recognition algorithm is as follows:

[0027] R=wi∈T∑TF-IDE(wi)·Severity(wi)

[0028] R is the risk score of the hot words in the text;

[0029] T is the word set in the current speech transcription text;

[0030] wi is a word in the hot word dictionary that appears in the text;

[0031] TF-IDF(wi) is the TF-IDF value of word wi, which measures its importance in the text;

[0032] Severity(wi) is the risk severity of word wi, and the manually defined value range is generally between 0 and 1.

[0033] 5. As a preferred embodiment of the psychological counseling system based on an AI model and voice communication of the present invention, when the user's emotions are judged to be normal, the psychological counseling AI component communicates with the user. During the communication process, if a high-risk judgment is triggered, manual consultation is transferred. The psychological counseling AI component includes an intelligent segmentation module and an interruption module. The intelligent segmentation module segments the output voice, and the interruption module identifies whether the user needs to interrupt the AI ​​output voice.

[0034] When interrupting, the intelligent interruption algorithm is used. The intelligent interruption algorithm is as follows:

[0035] P(Interrupt|Input)=σ(α·Energy+β·Overlap+γ·WakeWord+δ·IntentChange)

[0036] Interrupt indicates whether the user interrupted the AI's current speech;

[0037] Input is the current audio input stream;

[0038] Energy is the energy intensity of the current audio input, indicating the intensity of the user's speech;

[0039] Overlap indicates whether the user's voice overlaps with the AI's output voice.

[0040] WakeWord indicates whether the user says the interruption word;

[0041] α, β, γ, δ: weight parameters corresponding to the features Energy, Overlap, WakeWord, and IntentChange, respectively. These parameters can be set manually or obtained through training;

[0042] The Sigmoid function maps the score of the linear combination to a probability between 0 and 10\sim 10 and 1.

[0043] As a preferred psychological counseling system based on AI model and voice call of the present invention, the manual counseling module conducts manual psychological counseling communication for transferred users. The manual counseling module matches different psychological counselors according to user needs and priorities, and monitors the current number of tasks and idle time of psychological counselors in real time, giving priority to idle or low-load psychological counselors. When the user is a high-risk user, he / she enters the advanced priority consultation queue and is given priority to idle psychological counselors. When the user is an ordinary consultation user, he / she enters the queuing consultation queue and is assigned to psychological counselors in the order of the queue.

[0044] As a preferred embodiment of the psychological consultation system based on the AI ​​model and voice call of the present invention, the information acquisition module obtains the user's personal information based on the conversation content and constructs a user portrait based on the personal information;

[0045] When acquiring user information, a user information identification algorithm is used. User information identification is as follows:

[0046] P(Transcript|Speech)=∑ProfileP(Transcript|Speech,Profile)·∑ProfileP(Profile|Speech)

[0047] Speech is the voice signal input, representing the original voice data spoken by the user;

[0048] Transcript is the final output of the speech recognition system, which is the text converted from speech.

[0049] Profile is the basic information of the user;

[0050] P(Transcript|Speech) is the probability of outputting a transcript given a speech input, i.e. the standard speech recognition probability;

[0051] P(Transcript|Speech, Profile) is the probability of speech-to-text conversion, taking into account basic information. By using user profiles, we can more precisely adjust the recognition model and improve the recognition accuracy of specific users.

[0052] P(Profile|Speech) is the probability of inferring that the speech belongs to a certain user profile after a given speech input.

[0053] As a preferred psychological counseling system based on AI model and voice calls of the present invention, the historical conversation data acquisition module obtains whether the communicating users have historical consultation data. When there is historical conversation data, the historical data and current conversation data are combined and sent to the AI ​​large model of experts in the psychological field for analysis. When there is no historical conversation data, the user is guided to perform psychological evaluation. When combining historical data, the historical data is hierarchically indexed. When indexing short-term historical data, the complete conversation semantics are retained. When indexing long-term historical data, the keywords of the historical conversation and the analysis results of the end of the conversation are extracted.

[0054] As a preferred embodiment of the psychological consultation system based on AI model and voice call of the present invention, the intelligent psychological assessment module adjusts the questionnaire of the user to obtain the prediction results of the user's mental health, and sends the results to the AI ​​big model of psychology experts for analysis. The AI ​​big model of psychology experts provides corresponding consulting services based on the user's problems, including general psychological knowledge, depression and adolescent psychological problems;

[0055] When performing questionnaire adjustment, a questionnaire adjustment algorithm is used. The questionnaire adjustment algorithm is as follows:

[0056]

[0057] The most likely mental health status prediction result is obtained by maximizing the posterior probability;

[0058] S is the mental health status variable, which represents the user's mental health assessment result;

[0059] P(S) is the prior probability of mental health status, which reflects the estimated probability distribution of different mental states when there is no response;

[0060] P(d i |d <i ,q i ,S) is the probability of generating the th answer given the previous answer history, current question and mental health status.

[0061] As a preferred embodiment of the psychological consultation system based on AI model and voice call of the present invention, the consultation data storage module desensitizes and stores the consultation data of the manual consultation module and the psychology field expert AI big model, and marks important data. The data processing and recycling module processes the data inside the consultation data storage module, filters the available data, and transmits the available data to the psychology field expert AI big model, performs model iterative training on the psychology field expert AI big model, and realizes a closed loop of data utilization.

[0062] The emotional health analysis report generation module receives consultation data from the AI ​​model of experts in the field of psychology, and generates a corresponding analysis report based on the user's conversation consultation data. The analysis report summarizes the key points of the conversation and provides users with mental health guidance. Mental health guidance includes voice prompts for meditation, emotional counseling and psychological treatment plans. The emotional health analysis report generation module includes a user psychological positive reinforcement system. When the user completes the psychological health guidance plan, the user psychological positive reinforcement system provides instant feedback to the user. The instant feedback includes point rewards and achievement badges to encourage users to form mental health treatment habits.

[0063] A psychological consultation method based on an AI model and voice calls, comprising the following steps:

[0064] Step 1: User voice input and semantic analysis processing;

[0065] S1. Voice input processing: When a user calls, the voice recognition module first identifies the language of the voice and uses a language recognition model to recognize the voice. During the recognition process, a deep neural network extracts fixed-length feature vectors from the input voice to capture language-related acoustic features. Subsequently, the PLDA posterior probability model classifies the X-vector embeddings and calculates the probability that the voice segment belongs to each language, thereby identifying the specific language.

[0066] S2. Emotion detection and recognition. After processing the voice input, the system enters the emotion recognition module to identify the user's emotional state. This module enhances the accuracy and robustness of emotion recognition by integrating voice emotion recognition with text emotion recognition. During the recognition process, low-level acoustic features related to emotions are extracted from the voice signal. Acoustic features include Mel-frequency cepstral coefficients, fundamental frequency, energy variation, and intonation. These features are then modeled using TDNN to obtain an acoustic emotion vector.

[0067] The calculation formula of acoustic emotion vector is: s =f audio (s) = TDNN(MFCC(s))

[0068] At the same time, using the text input transcribed by the previous module, the text sentiment vector is extracted with the help of the pre-trained sentiment recognition model;

[0069] The calculation formula of text sentiment vector is: x =f text (x)=RoBERTa emotion (x)

[0070] The system introduces the attention mechanism to perform feature fusion and calculate the fusion vector. The calculation formula is: u f =a·u8(1-a)·u x ,uf is the fusion vector, where the fusion weight a is dynamically adjusted and is calculated using the sigmoid function based on the speech signal-to-noise ratio and text recognizability. The fused feature is fed into a Bi-LSTM classification network with an attention mechanism to extract the global semantic context vector, and the softmax classifier outputs the probability distribution of the emotion category.

[0071] S3, speech-to-text processing: After identifying the language, the system inputs the speech stream into the automatic speech recognition system adapted for that language. Built using the Whisper model, the input speech is preprocessed to extract acoustic features, providing high-quality input for subsequent modeling. Based on the LID recognition results, the system dynamically loads the ASR model for the corresponding language. The ASR model converts the speech feature sequence into a text sequence.

[0072] S4. Semantic Parsing and Demand Understanding. The semantic parsing module uses the BEAT model to identify user intent and extract core psychological demands. If there is semantic ambiguity, the module uses conversational AI to ask supplementary questions and interact with the user to clarify intent. The module converts the user's input text into a vector representation using the semantic model BEAT. Each intent category has a representation vector, which can be obtained through sample aggregation. For each intent category, the cosine similarity with the input vector is calculated;

[0073] S5. Keyword extraction and sentiment analysis: The semantic parsing module extracts keywords from the text and generates psychological appeal labels. The keyword extraction algorithm uses the TF-IDF extraction model and introduces the sentiment recognition module to analyze the emotional tendency of the user input and determine the user's psychological state. After obtaining the state label, the state label is mapped to predefined categories in the knowledge base, including anxiety, depression, autism, self-cognition disorder, intimacy problems, and adolescent adjustment disorder.

[0074] Step 2: User portrait construction and dynamic updating: A collaborative filtering algorithm is used to analyze user historical behavior and construct a preliminary user portrait. This is combined with conversation content and feedback data, and neural collaborative filtering is used to update the portrait label weights in real time. Furthermore, a recurrent neural network is introduced to model user interest trends, dynamically adjusting recommendation strategies and response styles to achieve precise and personalized services.

[0075] Step 3: Psychological model matching and expert model scheduling. A label similarity weighted algorithm is used to calculate the matching degree between user portrait labels and model labels, and the psychological consultation model is selected. The system then loads a large language model and expert rule system fine-tuned for the corresponding psychological classification task, performs model fusion, and comprehensively determines the recommendation priority.

[0076] Step 4: Optimize the feedback mechanism and recommendation strategy. Collect user voice feedback during calls, including emotional tone analysis results, behavioral signals indicating whether to continue the call, and whether to request a transfer to a human customer service representative. Generate structured feedback records for subsequent recommendation strategy adjustments and model training. Use the Q-learning reinforcement learning algorithm to construct a reward function based on user satisfaction and recommendation effectiveness. Based on the state transition of voice feedback, dynamically optimize the recommendation path to improve the match between the consultation content and the user's psychological needs. For first-time callers, the system launches a guided voice Q&A module to obtain their basic psychological preferences, topics of concern, and tendencies. Build an initial user profile and load the psychological AI basic model to provide targeted suggestions.

[0077] Step 5: Data storage and model training data reflow. Use regular rules and named entity recognition models to detect and mask privacy fields in user call content. Convert desensitized speech recognition text and behavior records into structured data and store them in a database. Recognize and simultaneously desensitize user speech data to form one-to-one corresponding speech and text training sample pairs. Ensure training data quality and privacy compliance. Regularly use accumulated high-quality conversation data for multi-model joint training, including intent recognition models, sentiment analysis models, and recommendation generation models. Use incremental training and model distillation methods to optimize performance. After verification, the trained model is deployed in a production environment for intelligent psychological response services.

[0078] Step 6. After the mental health consultation is completed, the system generates a corresponding report based on the user's conversation content to summarize the key points of the conversation and the user's emotional patterns, and provides mental health guidance to the user based on the key points of the conversation and emotional patterns. Mental health guidance includes voice prompts for meditation, emotional counseling, and psychotherapy plans.

[0079] Compared with the prior art, the present invention has the following beneficial effects:

[0080] 1. The present invention significantly improves the automation level and security of the psychological counseling system through the collaborative work of multiple modules. The interactive voice response module combines the emotion classification algorithm and the hot word recognition algorithm to realize real-time assessment and risk grading of the user's emotional state. When a high-risk user is detected, the system can automatically transfer to manual consultation to ensure the timeliness of crisis intervention; for users with normal emotions, AI provides efficient support, reducing labor costs and improving service coverage. In addition, the intelligent interruption algorithm dynamically determines whether the user needs to interrupt the AI ​​output through weighted analysis of multi-dimensional features such as audio energy, voice overlap, and wake-up words, thereby enhancing the naturalness of the conversation and user participation. The closed-loop design of data storage and model training further optimizes system performance, enabling the AI ​​model to continuously iterate and adapt to the needs of complex scenarios.

[0081] 2. The present invention realizes highly personalized mental health services through multimodal data fusion and dynamic user portrait construction. The semantic parsing module combines the BEAT model and sentiment analysis technology to accurately extract user intentions and psychological demands, generate labeled psychological states, and map them to a predefined knowledge base to provide direction for subsequent consultations. The user portrait system uses collaborative filtering and recurrent neural networks to update user interest changes and psychological needs in real time, dynamically adjust recommendation strategies and response styles, and ensure that services fit individual user differences. The psychological model matching algorithm uses label similarity weighting to screen the most suitable AI consulting model or expert rule system to improve the professionalism and pertinence of intervention. At the same time, the intelligent psychological assessment module generates prediction results of mental health status based on questionnaire adjustment algorithms and historical conversation data, and provides scientific advice in combination with expert AI models. The reinforcement learning-driven feedback mechanism dynamically optimizes the recommendation path through user behavior data, continuously improving user satisfaction and consulting effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] Figure 1 This is a system diagram of the present invention. DETAILED DESCRIPTION

[0083] Example 1

[0084] See also Figure 1 , a psychological consultation system based on AI models and voice calls, including:

[0085] The interactive voice response module receives incoming user voice calls and, based on user needs, directs them to manual or AI model consultation. Upon identifying user needs, the module converts user voice input into text through voice recognition, analyzes user intent using natural language processing technology, and directs them to manual or AI model consultation based on preset rules. When transferring to manual consultation, the call is transferred to an available customer service representative in real time via the API. Manual consultation transfers must be triggered by explicit user requests, and after triggering, customer service resources are allocated through queue management.

[0086] The high-risk monitoring module is activated when a user's call is connected, and determines whether to transfer the call to a human operator based on the user's voice emotion. If the user is determined to be high-risk, the call is automatically transferred to a human operator. If the user is determined to be in a normal mood, the call is transferred to an AI model consultation. The high-risk monitoring module includes a voice emotion detection module and a hot word recognition system. The voice emotion detection module detects user emotions, and the hot word recognition system detects high-risk words in the user's voice.

[0087] When performing emotion recognition, an emotion classification algorithm is used. The emotion classification algorithm is as follows:

[0088]

[0089] Emotion is the user's emotion category;

[0090] Features are features extracted from the language;

[0091] P(Features|Emotion) is the probability of these features appearing under a certain emotion;

[0092] P(Emotion) is the prior probability of a certain emotion;

[0093] P(Features) is the overall probability of the feature;

[0094] When performing hot word recognition, a hot word recognition algorithm is used. The hot word recognition algorithm is as follows:

[0095] R=wi∈T∑TF-IDE(wi)·Severity(wi)

[0096] R is the risk score of the hot words in the text;

[0097] T is the word set in the current speech transcription text;

[0098] wi is a word in the hot word dictionary that appears in the text;

[0099] TF-IDF(wi) is the TF-IDF value of word wi, which measures its importance in the text;

[0100] Severity (wi) is the risk severity of word wi, and the manually defined value range is generally between 0 and 1;

[0101] The psychological counseling AI component communicates with the user when the user's emotions are judged to be normal. During the communication process, if a high-risk judgment is triggered, manual consultation is transferred. The psychological counseling AI component includes an intelligent segmentation module and an interruption module. The intelligent segmentation module segments the output speech, and the interruption module identifies whether the user needs to interrupt the AI ​​output speech.

[0102] When interrupting, the intelligent interruption algorithm is used. The intelligent interruption algorithm is as follows:

[0103] P(Interrupt|Input)=σ(α·Energy+β·Overlap+γ·WakeWord+δ·IntentChange)

[0104] Interrupt indicates whether the user interrupted the AI's current speech;

[0105] Input is the current audio input stream;

[0106] Energy is the energy intensity of the current audio input, indicating the intensity of the user's speech;

[0107] Overlap is whether the user's voice overlaps with the AI's output voice;

[0108] WakeWord indicates whether the user says the interruption word;

[0109] α, β, γ, δ: weight parameters corresponding to the features Energy, Overlap, WakeWord, and IntentChange, respectively. These parameters can be set manually or obtained through training;

[0110] The Sigmoid function maps the score of the linear combination to a probability between 0 and 10\sim 10 and 1;

[0111] Manual consultation module: This module provides manual psychological consultation and communication for transferred users. It matches different psychological counselors according to user needs and priorities, and monitors the current number of tasks and idle time of psychological counselors in real time, giving priority to idle or low-load psychological counselors. When the user is a high-risk user, he / she enters the advanced priority consultation queue and is given priority to idle psychological counselors. When the user is an ordinary consultation user, he / she enters the queuing consultation queue and is assigned to psychological counselors in the order of the queue.

[0112] Information acquisition module, which obtains the user's personal information based on the conversation content and constructs a user profile based on the personal information;

[0113] When acquiring user information, a user information identification algorithm is used. User information identification is as follows:

[0114] P(Transcript|Speech)=ΣProfileP(Transcript|Speech,Profile)·ΣProfileP(Profile|Speech)

[0115] Speech is the voice signal input, representing the original voice data spoken by the user;

[0116] Transcript is the final output of the speech recognition system, which is the text converted from speech.

[0117] Profile is the basic information of the user;

[0118] P(Transcript|Speech) is the probability of outputting a transcript given a speech input, i.e. the standard speech recognition probability;

[0119] P(Transcript|Speech, Profile) is the probability of speech-to-text conversion, taking into account the user's basic information. By using the user profile, the recognition model can be adjusted more precisely to improve the recognition accuracy of specific users.

[0120] P(Profile|Speech) is the probability of inferring that a speech belongs to a certain user profile after a given speech input;

[0121] The historical conversation data acquisition module obtains historical consultation data from the communicating user. If there is historical conversation data, it combines the historical data with the current conversation data and sends it to the AI ​​model of psychology experts for analysis. If there is no historical conversation data, it guides the user to conduct a psychological assessment. When combining historical data, the historical data is hierarchically indexed. When indexing short-term historical data, the complete conversation semantics are retained. When indexing long-term historical data, the keywords of the historical conversation and the analysis results of the end of the conversation are extracted.

[0122] Intelligent psychological assessment module: The intelligent psychological assessment module adjusts the user's questionnaire to obtain the predicted results of the user's mental health, and sends the results to the AI ​​big model of psychology experts for analysis. The AI ​​big model of psychology experts provides corresponding consulting services based on the user's problems, including general psychological knowledge, depression and adolescent psychological problems;

[0123] When performing questionnaire adjustment, a questionnaire adjustment algorithm is used. The questionnaire adjustment algorithm is as follows:

[0124]

[0125] The most likely mental health status prediction result is obtained by maximizing the posterior probability;

[0126] S is the mental health status variable, which represents the user's mental health assessment result;

[0127] P(S) is the prior probability of mental health status, which reflects the estimated probability distribution of different mental states when there is no response;

[0128] P(d i |d <i ,q i ,S) is the probability of generating the th answer given the previous answer history, current question and mental health status;

[0129] AI big model of experts in the field of psychology;

[0130] Consulting data storage module;

[0131] Data processing and recycling module;

[0132] The consulting data storage module processes and stores the consulting data of the manual consulting module and the AI ​​big model of experts in the psychological field, and labels important data. The data processing and recycling module re-processes the data within the consulting data storage module, filters the available data, and transmits the available data to the AI ​​big model of experts in the psychological field. It performs iterative training on the AI ​​big model of experts in the psychological field to achieve a closed loop of data utilization.

[0133] Emotional health analysis report generation module, the emotional health analysis report generation module receives the consultation data within the AI ​​large model of experts in the psychological field, and generates a corresponding analysis report based on the user's conversation consultation data. The analysis report summarizes the key points of the conversation and provides users with mental health guidance. Mental health guidance includes voice prompts for meditation, emotional counseling and psychological treatment plans. The emotional health analysis report generation module includes a user psychological positive reinforcement system. When the user completes the psychological health guidance plan, the user psychological positive reinforcement system provides instant feedback to the user. The instant feedback includes point rewards and achievement badges to encourage users to form mental health treatment habits.

[0134] Example 2

[0135] A psychological consultation method based on an AI model and voice calls, comprising the following steps:

[0136] Step 1: User voice input and semantic analysis processing;

[0137] S1. Voice input processing: When a user calls, the voice recognition module first identifies the language of the voice and uses a language recognition model to recognize the voice. During the recognition process, a deep neural network extracts fixed-length feature vectors from the input voice to capture language-related acoustic features. Subsequently, the PLDA posterior probability model classifies the X-vector embeddings and calculates the probability that the voice segment belongs to each language, thereby identifying the specific language.

[0138] S2. Emotion detection and recognition. After processing the voice input, the system enters the emotion recognition module to identify the user's emotional state. This module enhances the accuracy and robustness of emotion recognition by integrating voice emotion recognition with text emotion recognition. During the recognition process, low-level acoustic features related to emotions are extracted from the voice signal. Acoustic features include Mel-frequency cepstral coefficients, fundamental frequency, energy variation, and intonation. These features are then modeled using TDNN to obtain an acoustic emotion vector.

[0139] The calculation formula of acoustic emotion vector is: s =f audio (s) = TDNN(MFCC(s))

[0140] At the same time, using the text input transcribed by the previous module, the text sentiment vector is extracted with the help of the pre-trained sentiment recognition model;

[0141] The calculation formula of text sentiment vector is: x =f text (x)=RoBERTa emotion (x)

[0142] The system introduces the attention mechanism to perform feature fusion and calculate the fusion vector. The calculation formula is: u f =a·u8(1-a)·u x ,u f is the fusion vector, where the fusion weight a is dynamically adjusted and is calculated using the sigmoid function based on the speech signal-to-noise ratio and text recognizability. The fused feature is fed into a Bi-LSTM classification network with an attention mechanism to extract the global semantic context vector, and the softmax classifier outputs the probability distribution of the emotion category.

[0143] S3, speech-to-text processing: After identifying the language, the system inputs the speech stream into the automatic speech recognition system adapted for that language. Built using the Whisper model, the input speech is preprocessed to extract acoustic features, providing high-quality input for subsequent modeling. Based on the LID recognition results, the system dynamically loads the ASR model for the corresponding language. The ASR model converts the speech feature sequence into a text sequence.

[0144] S4. Semantic Parsing and Demand Understanding. The semantic parsing module uses the BEAT model to identify user intent and extract core psychological demands. If there is semantic ambiguity, the module uses conversational AI to ask supplementary questions and interact with the user to clarify intent. The module converts the user's input text into a vector representation using the semantic model BEAT. Each intent category has a representation vector, which can be obtained through sample aggregation. For each intent category, the cosine similarity with the input vector is calculated;

[0145] S5. Keyword extraction and sentiment analysis: The semantic parsing module extracts keywords from the text and generates psychological appeal labels. The keyword extraction algorithm uses the TF-IDF extraction model and introduces the sentiment recognition module to analyze the emotional tendency of the user input and determine the user's psychological state. After obtaining the state label, the state label is mapped to predefined categories in the knowledge base, including anxiety, depression, autism, self-cognition disorder, intimacy problems, and adolescent adjustment disorder.

[0146] Step 2: User portrait construction and dynamic updating: A collaborative filtering algorithm is used to analyze user historical behavior and construct a preliminary user portrait. This is combined with conversation content and feedback data, and neural collaborative filtering is used to update the portrait label weights in real time. Furthermore, a recurrent neural network is introduced to model user interest trends, dynamically adjusting recommendation strategies and response styles to achieve precise and personalized services.

[0147] Step 3: Psychological model matching and expert model scheduling. A label similarity weighted algorithm is used to calculate the matching degree between user portrait labels and model labels, and the psychological consultation model is selected. The system then loads a large language model and expert rule system fine-tuned for the corresponding psychological classification task, performs model fusion, and comprehensively determines the recommendation priority.

[0148] Step 4: Optimize the feedback mechanism and recommendation strategy. Collect user voice feedback during calls, including emotional tone analysis results, behavioral signals indicating whether to continue the call, and whether to request a transfer to a human customer service representative. Generate structured feedback records for subsequent recommendation strategy adjustments and model training. Use the Q-learning reinforcement learning algorithm to construct a reward function based on user satisfaction and recommendation effectiveness. Based on the state transition of voice feedback, dynamically optimize the recommendation path to improve the match between the consultation content and the user's psychological needs. For first-time callers, the system launches a guided voice Q&A module to obtain their basic psychological preferences, topics of concern, and tendencies. Build an initial user profile and load the psychological AI basic model to provide targeted suggestions.

[0149] Step 5: Data storage and model training data reflow. Use regular rules and named entity recognition models to detect and mask privacy fields in user call content. Convert desensitized speech recognition text and behavior records into structured data and store them in a database. Recognize and simultaneously desensitize user speech data to form one-to-one corresponding speech and text training sample pairs. Ensure training data quality and privacy compliance. Regularly use accumulated high-quality conversation data for multi-model joint training, including intent recognition models, sentiment analysis models, and recommendation generation models. Use incremental training and model distillation methods to optimize performance. After verification, the trained model is deployed in a production environment for intelligent psychological response services.

[0150] Step 6. After the mental health consultation is completed, the system generates a corresponding report based on the user's conversation content to summarize the key points of the conversation and the user's emotional patterns, and provides mental health guidance to the user based on the key points of the conversation and emotional patterns. Mental health guidance includes voice prompts for meditation, emotional counseling, and psychotherapy plans.

[0151] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A psychological consultation system based on AI model and voice call, characterized by: include: Interactive Voice Response Module; High-risk monitoring module; Psychological counseling AI component; Manual consultation module; Information acquisition module; Historical conversation data acquisition module; Intelligent psychological assessment module; AI big model of experts in the field of psychology; Consulting data storage module; Data processing and recycling module; Emotional health analysis report generation module.

2. The psychological consultation system based on AI model and voice communication according to claim 1, characterized in that: The interactive voice response module receives the user's incoming voice call and transfers it to manual consultation or AI model consultation based on the user's needs. When identifying the user's needs, the interactive voice response module converts the user's voice input into text through voice recognition, combines natural language processing technology to analyze the user's intention, and transfers it to manual consultation or AI model consultation according to preset rules. When transferring to manual consultation, the call will be transferred to an idle customer service representative in real time through the API. The manual consultation transfer must be triggered when the user explicitly requests it. After the trigger, customer service resources will be allocated through queue management.

3. The psychological consultation system based on AI model and voice communication according to claim 2, characterized in that: The high-risk monitoring module is activated when a user's call is connected, and determines whether to transfer the call to a human operator based on the user's voice emotion. If the user is determined to be high-risk, the call is automatically transferred to a human operator. If the user is determined to be in a normal mood, the call is transferred to an AI model consultation. The high-risk monitoring module includes a voice emotion detection module and a hot word recognition system. The voice emotion detection module detects user emotions, and the hot word recognition system detects high-risk words in the user's voice. When performing emotion recognition, an emotion classification algorithm is used. The emotion classification algorithm is as follows: Emotion is the user's emotion category; Features are features extracted from speech; P(Features|Emotion) is the probability of these features appearing under a certain emotion; P(Emotion) is the prior probability of a certain emotion; P(Features) is the overall probability of the feature; When performing hot word recognition, a hot word recognition algorithm is used. The hot word recognition algorithm is as follows: R=wi∈T∑TF-IDE(wi)·Severity(wi) R is the risk score of the hot words in the text; T is the word set in the current speech transcription text; wi is a word in the hot word dictionary that appears in the text; TF-IDF(wi) is the TF-IDF value of word wi, which measures its importance in the text; Severity(wi) is the risk severity of word wi, and the manually defined value range is generally between 0 and 1.

4. The psychological consultation system based on AI model and voice communication according to claim 3, characterized in that: When the user's emotions are judged to be normal, the psychological counseling AI component communicates with the user. During the communication process, if a high-risk judgment is triggered, manual consultation is transferred. The psychological counseling AI component includes an intelligent segmentation module and an interruption module. The intelligent segmentation module segments the output voice, and the interruption module identifies whether the user needs to interrupt the AI ​​output voice. When interrupting, the intelligent interruption algorithm is used. The intelligent interruption algorithm is as follows: P(Interrupt|Input)=σ(α·Energy+β·Overlap+γ·WakeWord+δ·IntentChange) Interrupt indicates whether the user interrupted the AI's current speech; Input is the current audio input stream; Energy is the energy intensity of the current audio input, indicating the intensity of the user's speech; Overlap indicates whether the user's voice overlaps with the AI's output voice. WakeWord indicates whether the user says the interruption word; α, β, γ, δ: weight parameters corresponding to the features Energy, Overlap, WakeWord, and IntentChange, respectively. These parameters can be set manually or obtained through training; The Sigmoid function maps the score of the linear combination to a probability between 0 and 10\sim 10 and 1.

5. The psychological consultation system based on AI model and voice communication according to claim 4, characterized in that: The manual consultation module provides manual psychological consultation and communication for transferred users. The manual consultation module matches different psychological counselors according to user needs and priorities, and monitors the current number of tasks and idle time of psychological counselors in real time, giving priority to idle or low-load psychological counselors. When the user is a high-risk user, he / she enters the advanced priority consultation queue and is given priority to idle psychological counselors. When the user is an ordinary consultation user, he / she enters the queuing consultation queue and is assigned to psychological counselors according to the queue order.

6. The psychological consultation system based on AI model and voice communication according to claim 5, characterized in that: The information acquisition module obtains the user's personal information based on the conversation content and constructs a user profile based on the personal information; When acquiring user information, a user information identification algorithm is used. User information identification is as follows: P(Transcript|Speech)=∑ProfileP(Transcript|Speech,Profile)·∑ProfileP(Profile|Speech) Speech is the voice signal input, representing the original voice data spoken by the user; Transcript is the final output of the speech recognition system, which is the text converted from speech. Profile is the basic information of the user; P(Transcript|Speech) is the probability of outputting a transcript given a speech input, i.e. the standard speech recognition probability; P(Transcript|Speech, Profile) is the probability of speech-to-text conversion, taking into account basic user information. By using user profiles, we can more precisely adjust the recognition model and improve the recognition accuracy of specific users. P(Profile|Speech) is the probability of inferring that a speech belongs to a certain user profile after a given speech input.

7. The psychological consultation system based on AI model and voice communication according to claim 6, characterized in that: The historical conversation data acquisition module obtains whether the communicating user has historical consultation data. When there is historical conversation data, it combines the historical data and current conversation data and sends it to the AI ​​model of experts in the psychological field for analysis. When there is no historical conversation data, it guides the user to conduct a psychological assessment. When combining historical data, the historical data is hierarchically indexed. When indexing short-term historical data, the complete conversation semantics are retained. When indexing long-term historical data, the keywords of the historical conversation and the analysis results of the end of the conversation are extracted.

8. The psychological consultation system based on AI model and voice communication according to claim 7, characterized in that: The intelligent psychological assessment module conducts questionnaire adjustments on users to obtain predicted results of their mental health. The results are then sent to the AI ​​big model of psychology experts for analysis. The AI ​​big model of psychology experts provides corresponding consulting services based on the user's problems, including general psychological knowledge, depression, and adolescent psychological problems. When performing questionnaire adjustment, a questionnaire adjustment algorithm is used. The questionnaire adjustment algorithm is as follows: The most likely mental health status prediction result is obtained by maximizing the posterior probability; S is the mental health status variable, which represents the user's mental health assessment result; P(S) is the prior probability of mental health status, which reflects the estimated probability distribution of different mental states when there is no response; P(d i |d <i ,q i ,S) is the probability of generating the th answer given the previous answer history, current question and mental health status.

9. The psychological consultation system based on AI model and voice communication according to claim 8, characterized in that: The consultation data storage module desensitizes and stores the consultation data from the manual consultation module and the psychology expert AI model, and labels important data. The data processing and recycling module processes the data within the consultation data storage module, filters available data, and transmits the available data to the psychology expert AI model. It then iterates and trains the psychology expert AI model to achieve a closed loop of data utilization. The emotional health analysis report generation module receives consultation data from the AI ​​model of experts in the field of psychology, and generates a corresponding analysis report based on the user's conversation consultation data. The analysis report summarizes the key points of the conversation and provides users with mental health guidance. Mental health guidance includes voice prompts for meditation, emotional counseling and psychological treatment plans. The emotional health analysis report generation module includes a user psychological positive reinforcement system. When the user completes the psychological health guidance plan, the user psychological positive reinforcement system provides instant feedback to the user. The instant feedback includes point rewards and achievement badges to encourage users to form mental health treatment habits.

10. A psychological consultation method based on an AI model and voice calls, applicable to the psychological consultation system based on an AI model and voice calls according to claim 9, characterized in that: The following steps are involved: Step 1: User voice input and semantic analysis processing; S1. Voice input processing: When a user calls, the voice recognition module first identifies the language of the voice and uses a language recognition model to recognize the voice. During the recognition process, a deep neural network extracts fixed-length feature vectors from the input voice to capture language-related acoustic features. Subsequently, the PLDA posterior probability model classifies the X-vector embeddings and calculates the probability that the voice segment belongs to each language, thereby identifying the specific language. S2. Emotion detection and recognition. After processing the voice input, the system enters the emotion recognition module to identify the user's emotional state. This module enhances the accuracy and robustness of emotion recognition by integrating voice emotion recognition with text emotion recognition. During the recognition process, low-level acoustic features related to emotions are extracted from the voice signal. Acoustic features include Mel-frequency cepstral coefficients, fundamental frequency, energy variation, and intonation. These features are then modeled using TDNN to obtain an acoustic emotion vector. The calculation formula of acoustic emotion vector is: s =f audio (s) = TDNN(MFCC(s)) At the same time, using the text input transcribed by the previous module, the text sentiment vector is extracted with the help of the pre-trained sentiment recognition model; The calculation formula of text sentiment vector is: x =f text (x)=RoBERTa emotion (x) The system introduces the attention mechanism to perform feature fusion and calculate the fusion vector. The calculation formula is: u f =a·u8(1-a)·u x ,u f is the fusion vector, where the fusion weight a is dynamically adjusted and is calculated using the sigmoid function based on the speech signal-to-noise ratio and text recognizability. The fused feature is fed into a Bi-LSTM classification network with an attention mechanism to extract the global semantic context vector, and the softmax classifier outputs the probability distribution of the emotion category. S3, speech-to-text processing: After identifying the language, the system inputs the speech stream into the automatic speech recognition system adapted for that language. Built using the Whisper model, the input speech is preprocessed to extract acoustic features, providing high-quality input for subsequent modeling. Based on the LID recognition results, the system dynamically loads the ASR model for the corresponding language. The ASR model converts the speech feature sequence into a text sequence. S4. Semantic Parsing and Demand Understanding. The semantic parsing module uses the BEAT model to identify user intent and extract core psychological demands. If there is semantic ambiguity, the module uses conversational AI to ask supplementary questions and interact with the user to clarify intent. The module converts the user's input text into a vector representation using the semantic model BEAT. Each intent category has a representation vector, which can be obtained through sample aggregation. For each intent category, the cosine similarity with the input vector is calculated; S5. Keyword extraction and sentiment analysis: The semantic parsing module extracts keywords from the text and generates psychological appeal labels. The keyword extraction algorithm uses the TF-IDF extraction model and introduces the sentiment recognition module to analyze the emotional tendency of the user input and determine the user's psychological state. After obtaining the state label, the state label is mapped to predefined categories in the knowledge base, including anxiety, depression, autism, self-cognition disorder, intimacy problems, and adolescent adjustment disorder. Step 2: User portrait construction and dynamic updating: A collaborative filtering algorithm is used to analyze user historical behavior and construct a preliminary user portrait. This is combined with conversation content and feedback data, and neural collaborative filtering is used to update the portrait label weights in real time. Furthermore, a recurrent neural network is introduced to model user interest trends, dynamically adjusting recommendation strategies and response styles to achieve precise and personalized services. Step 3: Psychological model matching and expert model scheduling. A label similarity weighted algorithm is used to calculate the matching degree between user portrait labels and model labels, and the psychological consultation model is selected. The system then loads a large language model and expert rule system fine-tuned for the corresponding psychological classification task, performs model fusion, and comprehensively determines the recommendation priority. Step 4: Optimize the feedback mechanism and recommendation strategy. Collect user voice feedback during calls, including emotional tone analysis results, behavioral signals indicating whether to continue the call, and whether to request a transfer to a human customer service representative. Generate structured feedback records for subsequent recommendation strategy adjustments and model training. Use the Q-learning reinforcement learning algorithm to construct a reward function based on user satisfaction and recommendation effectiveness. Based on the state transition of voice feedback, dynamically optimize the recommendation path to improve the match between the consultation content and the user's psychological needs. For first-time callers, the system launches a guided voice Q&A module to obtain their basic psychological preferences, topics of concern, and tendencies. Build an initial user profile and load the psychological AI basic model to provide targeted suggestions. Step 5: Data storage and model training data reflow. Use regular rules and named entity recognition models to detect and mask privacy fields in user call content. Convert desensitized speech recognition text and behavior records into structured data and store them in a database. Recognize and simultaneously desensitize user speech data to form one-to-one corresponding speech and text training sample pairs. Ensure training data quality and privacy compliance. Regularly use accumulated high-quality conversation data for multi-model joint training, including intent recognition models, sentiment analysis models, and recommendation generation models. Use incremental training and model distillation methods to optimize performance. After verification, the trained model is deployed in a production environment for intelligent psychological response services. Step 6. After the mental health consultation is completed, the system generates a corresponding report based on the user's conversation content to summarize the key points of the conversation and the user's emotional patterns, and provides mental health guidance to the user based on the key points of the conversation and emotional patterns. Mental health guidance includes voice prompts for meditation, emotional counseling, and psychotherapy plans.

Citation Information

Cited By

  • Emotion recognition method and device based on audio data and electronic equipment

    CN122024773A