A psychological counseling real-time speech recognition method based on multi-modal data

By constructing a gender-specific speech psychological feature mapping baseline and a multi-dimensional emotion analysis model, and combining pronunciation habits and real data, the fundamental frequency mean amplitude is dynamically corrected, and the location of emotion focus and the tracking of change trends are optimized. This solves the problems of insufficient accuracy in individual differences and real-time performance in existing technologies, and realizes the quantitative assessment of mental health and the identification of language disorders.

CN121337358BActive Publication Date: 2026-02-13MEDICAL HEALTHCARE DIGITAL TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511894311.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-02-13
Estimated Expiration
2045-12-16

AI Technical Summary

Technical Problem

Existing voice recognition technologies for psychological counseling have significant limitations in terms of individual differences, quantitative assessment, real-time performance, and accuracy. They fail to effectively integrate multimodal data, leading to misjudgments and inaccurate assessments of emotions.

Method used

By constructing a gender-specific speech psychological feature mapping baseline, combining pronunciation habit details and real data, dynamically correcting the fundamental frequency mean amplitude, constructing a multi-dimensional emotion analysis model, optimizing emotion focus localization and change trend tracking, designing a comprehensive stability value calculation method, constructing a three-level processing mechanism for language barriers, and conducting feature adaptation and cross-validation.

Benefits of technology

It achieves accurate emotion classification and trend tracking, provides quantitative assessment of mental health, solves misjudgments caused by individual pronunciation differences, and improves the real-time and accuracy of psychological counseling, especially the ability to identify language disorders and severe emotional disturbances.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121337358B_ABST
    Figure CN121337358B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of speech recognition, and discloses a psychological counseling real-time speech recognition method based on multi-modal data, comprising: constructing a gender-specific speech psychological feature mapping baseline, analyzing and judging the emotional stability tendency degree, and combining the amplitude peak value proportion to judge the expression tendency degree; dynamically correcting the fundamental frequency mean amplitude to solve the psychological feature misjudgment caused by individual pronunciation difference; detecting the stress position, key words, speech speed and pause features through the fundamental frequency mutation and amplitude mutation, constructing a multi-dimensional emotion analysis model, and optimizing the emotion focus positioning, emotion type judgment and change trend tracking problems; designing a comprehensive stability value calculation method, simultaneously reflecting the emotional stability and the influence degree of key words, and providing a psychological health evaluation quantitative index; constructing a three-level processing mechanism, and through the fusion of the exclusive baseline and multi-modal evidence, performing feature adaptation, cross-validation and decision correction.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech recognition, in particular to a psychological counseling real-time speech recognition method based on multi-modal data. BACKGROUND

[0002] With the penetration of artificial intelligence technology in the field of mental health, speech recognition and emotion analysis technology has gradually become an important tool for psychological counseling decision-making. The existing technology extracts the acoustic features of the voice, combines with the machine learning model to realize emotion recognition and preliminary judgment of psychological state, but the existing machine learning model still has significant limitations in the professional adaptation of the psychological counseling scene, and it is difficult to meet the needs of clinical counseling for accuracy, individual differences and quantitative evaluation,

[0003] In the prior art, the existing speech psychological feature analysis technology generally adopts a general model approach and does not design for the physiological voice differences between men and women. The existing technology uses a unified feature threshold and does not consider the influence of gender differences on feature distribution. The pronunciation pattern of the psychological counseling object is strongly related to the psychological state, but the existing technology does not quantitatively analyze the phoneme-level pronunciation features, and only processes pronunciation errors through a speech recognition error correction mechanism, resulting in feature distortion. The existing technology does not fuse the real background information of the counseling object. Age growth will naturally cause the vibration frequency of the vocal cords to decrease, and there is a lack of dynamic correction mechanism. It relies on a single feature dimension and identifies emotion types through keyword matching, ignoring the indication of stress position on emotional focus. It quantifies the timing features such as speech rate and pause, but cannot associate the semantics of keywords, resulting in keywords spoken by different emotions being classified into the same emotion level. In the balance between real-time and accuracy, enlarging the sliding window to ensure feature integrity increases the real-time response delay. Reducing the window size to improve real-time performance increases the emotion misjudgment rate due to incomplete semantic units, lacks coherent tracking of emotion change trends, and only classifies emotions for a single window, making it difficult to assist consultants in grasping the counseling pace. The existing technology for evaluating mental health status is mostly limited to qualitative description, lacking a quantitative index system. Consultants develop intervention frequency and programs based on evaluation results, but qualitative descriptions cannot provide specific references. Emotion intensity scores are introduced, but they do not combine keywords with long-term emotional impact and overall stability of emotional fluctuations, resulting in scores only reflecting instantaneous state and lacking comprehensive evaluation value.

[0004] Therefore, it is necessary to provide a psychological counseling real-time speech recognition method based on multi-modal data. SUMMARY

[0005] The purpose of the present application is to provide a psychological counseling real-time speech recognition method based on multi-modal data. To solve the above-mentioned problems of the prior art, the present application realizes the following technical solutions:

[0006] The embodiment of the application provides a psychological counseling real-time voice recognition method based on multi-modal data, and specifically comprises the following steps:

[0007] Step one: by constructing a gender-specific voice psychological feature mapping baseline, the emotional stability tendency degree is initially analyzed and judged through the mean fundamental frequency and the coefficient of variation of fundamental frequency, and the expression tendency degree is judged by combining the amplitude peak value ratio;

[0008] Step two: combining pronunciation habit details and real data, dynamically correcting the mean fundamental frequency amplitude, solving the psychological feature misjudgment caused by individual pronunciation difference;

[0009] Step three: by detecting the accent position, key word, speech speed and pause feature through the fundamental frequency mutation and amplitude mutation, a multi-dimensional emotion analysis model is constructed, and the problems of emotion focus positioning, emotion type judgment and change trend tracking are optimized;

[0010] Step four: combining the key word and the emotion change correlation law, designing a comprehensive stability value calculation method, reflecting the emotional stability and the influence degree of the key word, and providing a psychological health evaluation quantitative index;

[0011] Step five: for the problems of language barriers and serious emotional disorders, a three-level processing mechanism is constructed, and through the fusion of exclusive baseline and multi-modal evidence, feature adaptation, cross-validation and decision correction are carried out.

[0012] Further, the method for constructing a gender-specific voice psychological feature mapping baseline is:

[0013] The initial dialogue voice of the consultation object is collected by using a double microphone array, the main microphone is placed in front of the visitor to focus on the target voice, the reference microphone is placed in the side rear to collect environmental noise, the sampling rate is uniformly set and the number of quantization is quantified;

[0014] The double microphone signal is denoised by the LMS adaptive filtering algorithm, the non-steady state noise in the consultation room is suppressed, and the pure voice signal is reserved; the mean fundamental frequency and the coefficient of variation of fundamental frequency are calculated respectively according to the physiological range of the fundamental frequency of men and women, and the amplitude peak value ratio is calculated in combination with the amplitude dynamic range of men and women;

[0015] Based on the pure voice signals of male and female visitors of different ages and occupations, the gender-specific mapping rules are trained and optimized by the random forest algorithm, and the initial judgment of emotional stability tendency, emotional sensitivity tendency and expression tendency is realized;

[0016] Further, the method for dynamically correcting the mean fundamental frequency amplitude is:

[0017] The preprocessed voice is phoneme-level disassembled by using a Kaldi voice recognition toolkit, three types of pronunciation habit features are extracted, and the deviation rate of average duration of vowel pronunciation from the standard duration in Mandarin, the standard deviation of consonant initial pronunciation intensity, and the deviation rate of four-tone frequency from the standard tone in Mandarin are calculated.

[0018] The structured information of the consultation object's age, occupation, dialect background, and previous psychological state is acquired, a correction factor is constructed, and the average fundamental frequency is corrected by combining the original average fundamental frequency with the age factor and the occupation factor.

[0019] The corrected average fundamental frequency and the amplitude peak ratio are used to recalculate the coefficient of variation of the fundamental frequency and the amplitude, and the psychological characteristics are updated again in combination with the pronunciation habit markers.

[0020] Further, the method for constructing a multi-dimensional emotion analysis model is:

[0021] The maximum value of the fundamental frequency, the average value of the fundamental frequency, the maximum value of the amplitude, and the average value of the amplitude are extracted by setting a 3-second sliding window.

[0022] The position of the stress is determined by the joint distribution entropy of the fundamental frequency mutation value and the amplitude mutation value and the maximum values of the two; the stress position vocabulary is matched with the keyword library screened by the TF-IDF algorithm for psychological counseling; the real-time speech speed, the speech speed change trend, the single pause duration, and the pause frequency are calculated by a 3-second window and a 1-second step; the emotion determination rules are trained and optimized by a decision tree algorithm, and the model is constructed in combination with the keywords and multi-dimensional features.

[0023] Based on the multi-dimensional feature construction model, the emotion classification of the consultation object is analyzed and judged, and the emotion change trend is calculated according to the emotion classification to optimize the emotion focus positioning, emotion type judgment, and change trend tracking problems.

[0024] Further, the method for judging the emotion classification of the consultation object is:

[0025] The emotion determination rules are constructed in combination with multi-dimensional features, the rules are trained and optimized by a decision tree algorithm, the keywords, real-time speech speed, single pause duration, and speech speed change trend are combined, the emotion intensity value is analyzed, and the emotion classification of the consultation object is judged according to the emotion intensity value.

[0026] Further, the method for obtaining the emotion change trend is:

[0027] The sum of the difference values of the emotion intensity values of the continuous three windows is obtained, and the emotion change trend is analyzed according to the sum of the difference values of the emotion intensity values.

[0028] If the emotion change trend is greater than 5, the emotion is rapidly warming up positively or the negative emotion is intensified; if the sum of the difference values of the emotion change trend is less than -5, the emotion is rapidly cooling down and the emotion is relieved or turned to neutral; otherwise, the emotion change is stable.

[0029] Further, the method for obtaining the comprehensive stability value is:

[0030] The difference between the maximum value of emotional intensity within 3 seconds after the appearance of the keyword and the average value within 3 seconds before the appearance of the keyword is calculated to define an influence coefficient; the standard deviation of emotional intensity and the number of emotional type jumps are calculated, and the emotional fluctuation degree is obtained by correcting the standard deviation of emotional intensity combined with the number of jumps;

[0031] A two-dimensional interval deduction item table is established combined with mapping, the sum of the influence coefficients of the keywords and the deduction value of the emotional fluctuation degree are accumulated, and the psychological health stability value is obtained by subtracting the total deduction value from the preset full score;

[0032] Further, the method for evaluating and quantifying psychological health is:

[0033] The emotional stability value is standardized and divided into stable, relatively stable, unstable, and high-risk types;

[0034] If the number of emotional type jumps is greater than the frequent threshold or the absolute value of the keyword influence coefficient is greater than the conflict threshold, suspicious marking of the text containing the keyword, the appearance time, and the emotional change curve is performed; if the emotional stability value is less than 40 points, an early warning prompt is automatically triggered and synchronized to the counselor terminal; a correlation degree model is constructed by fusing the emotional stability value, the pitch position offset, the pause interval change, and the speech rate change rate using the D-S evidence theory, the danger correlation degree is obtained through the BPA function and the synthesis rule, and the danger level is divided combined with the emotional stability value change value;

[0035] Further, the method for adapting features is:

[0036] Voice samples of stuttering counseling objects are collected, abnormal voice sequences are aligned through dynamic time warping algorithm, pathological features are extracted to establish a stuttering type exclusive baseline, and developmental and acquired stuttering are divided; emotional correlation features are analyzed;

[0037] The stuttering segments in the voice are processed using pathological feature masking, the emotional correlation features of non-stuttering segments are retained, and fuzzy logic reasoning is used to correct misjudgments;

[0038] Further, the method for constructing a three-level processing mechanism is:

[0039] For explosive sound, tremolo, and air sound distortion during emotional outburst, a hierarchical enhancement algorithm is used to decompose the voice signal using wavelet threshold denoising and process the high-frequency explosive components using soft threshold processing, adaptively smooth the high-frequency fluctuations of tremolo using Kalman filter, and improve the harmonic-to-noise ratio using spectral subtraction and extract the spectral centroid and mel cepstral coefficients;

[0040] If the voice distortion degree is greater than the preset limit value, a multi-modal cross-validation is started, and the dynamic change rate of 68 facial key points and the distortion voice emotion features are extracted for consistency determination.

[0041] The heart rate variability and the skin conductivity abnormality degree are calculated, and when both reach a high wake-up threshold, the negative emotion determination is strengthened.

[0042] The embodiment of the present application provides a real-time voice recognition system for psychological consultation based on multi-modal data, which specifically comprises the following modules:

[0043] The analysis and judgment module: by constructing a gender-specific voice-psychological feature mapping baseline, the initial analysis and judgment of the emotional stability tendency degree are carried out through the mean fundamental frequency and the fundamental frequency variation coefficient, and the expression tendency degree is judged by combining the amplitude peak value proportion;

[0044] The dynamic correction module: combined with the pronunciation habit details and the real data, the mean fundamental frequency amplitude is dynamically corrected to solve the psychological feature misjudgment caused by individual pronunciation difference;

[0045] The optimization analysis module: by detecting the accent position, the key word, the speech speed and the pause feature through the fundamental frequency mutation and the amplitude mutation, a multi-dimensional emotion analysis model is constructed, and the problems of emotion focus positioning, emotion type judgment and change trend tracking are optimized;

[0046] The quantitative evaluation module: combined with the key word and the emotion change correlation law, a comprehensive stability value calculation method is designed, which can reflect the emotional stability and the influence degree of the key word, and provide a quantitative index for psychological health evaluation;

[0047] The adaptive correction module: for the problems of language barriers and serious emotional disorders, a three-level processing mechanism is constructed, and through the fusion of exclusive baseline and multi-modal evidence, feature adaptation, cross-validation and decision correction are carried out.

[0048] The present application has the following advantages:

[0049] 1. Breakthrough the limitations of traditional general voice feature analysis, set the physiological range difference of male and female fundamental frequency and amplitude, train exclusive mapping rules through random forest, construct correction model combined with pronunciation habits and multi-dimensional real data, solve the misjudgment caused by individual pronunciation difference; adopt the joint distribution entropy judgment of fundamental frequency and amplitude mutation to determine the accent position, fuse the key word library, speech speed and pause feature to construct the decision tree model, design the key word influence coefficient and the emotion fluctuation degree double-dimensional comprehensive stability value calculation method, realize the emotion classification, trend tracking and psychological health quantitative grading;

[0050] 2. For language barriers such as stuttering, through pathological feature masking and emotion feature reconstruction adaptation, combined with multi-modal cross-validation of facial key points and physiological indicators, the problems of distorted voice emotion extraction and crisis misjudgment are solved, and a complete technical system from basic recognition to special scene coverage is formed. Attached Figure Description

[0051] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0052] Figure 1 This is a flowchart of the steps of a real-time speech recognition method for psychological counseling based on multimodal data provided in Embodiment 1 of the present invention;

[0053] Figure 2 This is a schematic diagram of the structure of a real-time speech recognition system for psychological counseling based on multimodal data, provided in Embodiment 2 of the present invention. Detailed Implementation

[0054] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0055] Example 1: As Figure 1 As shown in the figure, the real-time speech recognition method for psychological counseling based on multimodal data provided by this invention specifically includes the following steps:

[0056] Step 1: By constructing a gender-specific speech psychological feature mapping baseline, the degree of emotional stability tendency is judged by the initial analysis of the fundamental frequency mean and fundamental frequency variation coefficient, and the degree of expression tendency is judged by the proportion of amplitude peaks.

[0057] In a specific embodiment, a dual-microphone array is used to collect the initial dialogue duration of the client. The main microphone is placed directly in front of the client to focus on collecting the target speech; the reference microphone is placed to the side and rear to collect ambient noise. The sampling rate is uniformly set, and the bit depth is quantized to ensure the integrity of speech details.

[0058] The dual-microphone signal is denoised using the LMS adaptive filtering algorithm, which tracks changes in noise characteristics in real time, suppresses non-steady-state noise in the consultation room, preserves a clean speech signal with a signal-to-noise ratio of ≥25dB, and provides a basis for feature extraction.

[0059] Core voice feature extraction is performed for males and females respectively, and feature differences caused by physiological voice production mechanism differences of males and females are combined to distinguish;

[0060] Specifically, the male fundamental frequency physiological range is preset as [85Hz, 180Hz], and the fundamental frequency mean F0 is calculated based on the pure speech signal m Reflecting the basic voice production state, the fundamental frequency coefficient of variation CV is calculated F0 Reflecting the degree of emotional fluctuation; the amplitude volume dynamic range is [40dB, 80dB], and the amplitude peak ratio P is calculated by dividing the peak time by the total time amp , wherein the peak time is defined as the time period when the amplitude exceeds 75% of the preset maximum peak, and the peak time reflects the initiative and emotional intensity of expression;

[0061] The female fundamental frequency physiological range is [165Hz, 255Hz], and the fundamental frequency mean F0 is calculated m The fundamental frequency coefficient of variation CV F0 ;

[0062] It should be noted that the female fundamental frequency is higher than the male fundamental frequency, and the female variation amplitude is greater than the male variation amplitude;

[0063] The amplitude dynamic range is [35dB, 75dB], and the amplitude peak threshold is adjusted to 70% of the preset maximum peak, and the amplitude peak ratio P is calculated simultaneously amp Index;

[0064] Based on the pure speech signal of psychological counseling, male and female visitors of different ages and occupations are covered, gender-specific mapping rules are constructed, the rules are optimized through random forest algorithm training, and the initial analysis of the fundamental frequency mean and the fundamental frequency coefficient of variation is used to judge the emotional stability tendency degree, and the amplitude peak ratio is used to judge the expression tendency degree;

[0065] For example, for males, if the fundamental frequency mean is less than 110Hz and the fundamental frequency coefficient of variation is less than 0.15, the initial label is emotion stable tendency; if the fundamental frequency mean is greater than 150Hz or the fundamental frequency coefficient of variation is greater than 0.25, the initial label is emotion sensitive tendency, otherwise, the initial label is emotion unstable tendency; for females, if the fundamental frequency mean is less than 190Hz and the fundamental frequency coefficient of variation is less than 0.2, the initial label is emotion stable tendency; if the fundamental frequency mean is greater than 220Hz or the fundamental frequency coefficient of variation is greater than 0.3, the initial label is emotion sensitive tendency; otherwise, the initial label is emotion unstable tendency; for any counseling object, if the amplitude peak ratio is greater than 0.3, the expression initiative tendency is marked; if the amplitude peak ratio is less than 0.15, the expression passive tendency is marked; otherwise, the expression normal tendency is marked;

[0066] Step two: combine the details of pronunciation habits with real data, dynamically correct the mean amplitude of the fundamental frequency, solve the psychological characteristics misjudgment caused by individual pronunciation differences, and improve the analysis accuracy;

[0067] In specific embodiments, the pre-processed speech is phoneme-level disassembled to implement phoneme segmentation using Kaldi speech recognition toolkit, and three types of pronunciation habit features are extracted to comprehensively depict individual pronunciation patterns;

[0068] Obtain the deviation rate of the average duration of each vowel pronunciation of the consultation object and the standard vowel pronunciation duration of Mandarin, and calculate the vowel duration deviation D by subtracting the standard vowel pronunciation duration from the average vowel pronunciation duration and dividing by the standard vowel pronunciation duration. vow If D vow If the vowel duration deviation is greater than the preset deviation ratio, it is marked that the consultation object has a vowel lengthening habit, and the vowel lengthening habit is present in visitors with emotional depression or hesitation;

[0069] Analyze the consonant initial intensity variation of the consultation object, and obtain the intensity standard deviation SD con of the consonant initial pronunciation intensity of the consultation object. con If the intensity standard deviation SD is greater than 5 dB, it is marked as a consonant initial intensity fluctuation habit, reflecting the stability of airflow control during pronunciation, and a large fluctuation may be related to anxiety, otherwise, it is marked as normal consonant initial pronunciation;

[0070] Analyze the tone accuracy of the consultation object, and obtain the tone deviation rate D tone of the frequency of the four tones of the consultation object, namely, yin ping, yang ping, shang sheng and qu sheng, and the standard tone frequency of Mandarin. tone If the tone deviation rate D is greater than 0.25, it is marked as an unstable tone habit, otherwise, it is marked as a stable tone habit.

[0071] Obtain the structured information of the real data uploaded by the consultation object, which includes but is not limited to age, occupation, dialect background, and previous psychological state, combine with clinical statistical data to build a correction factor to adjust the mean and amplitude of the fundamental frequency, and realize accurate compensation of individual differences.

[0072] For example, the corrected mean fundamental frequency after compensation is calculated as The corrected amplitude after compensation is calculated as ; wherein, is the original mean fundamental frequency, is the original amplitude, is the age factor, for example, 0.05 for [20, 30] years old, 0.08 for [31, 40] years old, 0.12 for [41, 50] years old, and 0.15 for 51 years old and above, combined with tracking analysis of the 5-year speech fundamental frequency change trend of different age groups to reflect the natural influence of age growth on vocal cord vibration frequency. For the professional factor, the long-term high-frequency sound of teachers or sales professionals leads to high pronunciation intensity, and the daily pronunciation intensity of programmers or researchers is low. The daily average pronunciation intensity of each professional group is set; For the dialect factor, the northern dialect takes 0.04, the southern dialect takes 0.07, and the Cantonese takes 0.09. The influence of dialect pronunciation habits on amplitude is compensated according to the quantitative analysis results of the length difference between dialect and Mandarin finals; For the previous psychological factor, the previous anxiety history takes 0.09, the previous depression history takes 0.11, and the no previous history takes 0.02. The speech characteristics of patients with previous psychological diseases and healthy people are compared to obtain the results; α, β, γ and δ are adjustment factors. Through 10-fold cross-validation, α=0.3, β=0.25, γ=0.2 and δ=0.35 are determined from historical data to ensure that each factor contributes reasonably to the correction result;

[0073] The corrected corrected fundamental frequency mean F0 corr and the corrected amplitude A corr The fundamental frequency coefficient of variation CV F0 and the amplitude peak ratio P amp are recalculated, and the pronunciation habits are marked for final lengthening, initial fluctuation and unstable tone for comprehensive judgment, and the psychological characteristics are updated again;

[0074] For example, if the original fundamental frequency mean F0 m of the consultation object is 230 Hz, and the fundamental frequency coefficient of variation CV F0 is 0.32, it is marked as emotionally sensitive tendency; if the corrected corrected fundamental frequency mean F0 corr is 210 Hz, the fundamental frequency coefficient of variation CV F0 is 0.18, and there is no any pronunciation habit abnormality, the comprehensive update is emotionally stable tendency; if the corrected fundamental frequency coefficient of variation CV F0 is 0.22, and there is a final lengthening habit, the update is emotionally sensitive tendency, which takes into account the influence of objective characteristics and behavior habits;

[0075] Step three: detect stress position, key words, speech rate and pause characteristics through fundamental frequency mutation and amplitude mutation, construct a multi-dimensional emotion analysis model, and optimize emotion focus positioning, emotion type judgment and change trend tracking problems;

[0076] In specific embodiments, the stress position is detected by fundamental frequency mutation and amplitude mutation, and the core vocabulary of stress corresponding to emotional expression; a 3-second sliding window is set to extract the maximum fundamental frequency F0 max and the average fundamental frequency F0 avg , the maximum amplitude A max and the average amplitude A avg ;

[0077] The base frequency maximum value of the current window is subtracted from the base frequency maximum value of the previous window, and then divided by the base frequency maximum value of the previous window to obtain a base frequency mutation value F0 change , which reflects the base frequency feature difference;

[0078] The amplitude maximum value of the current window is subtracted from the amplitude maximum value of the previous window, and then divided by the amplitude maximum value of the previous window to obtain an amplitude mutation value A change , which reflects the amplitude feature difference;

[0079] The joint distribution entropy H(F0 change ,A change ) of the base frequency mutation and the amplitude mutation is calculated, and if H(F0 change ,A change ) is greater than a preset distribution entropy threshold and the maximum value max(F0 change ,A change ) of the base frequency mutation value and the amplitude mutation value is greater than a preset mutation threshold, it is determined as a stress position, and the lower the entropy value, the stronger the coordination of the base frequency mutation and the amplitude mutation;

[0080] The stress position corresponding vocabulary is matched with the psychological counseling special keyword library, the "Psychological Counseling Common Emotion Vocabulary Scale" and the clinical counseling corpus are collected, the high-frequency emotion words of "sad", "fear", "happy", "relaxed", "stress", "family", "work" and "childhood" are screened out by the TF-IDF algorithm, and the keyword library is constructed by manual annotation of the psychological counselors to determine the type of each keyword positive / negative and the emotion correlation degree, and the marked keywords are used for emotion determination;

[0081] The window size of the sliding window is preset to 3 seconds, and the step size is preset to 1 second for real-time feature calculation;

[0082] The number of words recognized by speech recognition in the window is divided by the total duration of the sliding window to obtain the real-time speech speed V speech ;

[0083] The average speech speed of the current window is subtracted from the speech speed of the previous window, and then divided by the average speech speed of the previous window to obtain the speech speed change trend T v ;

[0084] The duration of the pause voice energy less than-40dB in the window is taken as the mean value to obtain the single pause duration T pause ;

[0085] The number of pauses in the window is divided by the total duration of the sliding window to obtain the pause frequency F pause ;

[0086] The emotion judgment rule is constructed in combination with multi-dimensional features, the rule is optimized by a decision tree algorithm training, key words, real-time speech speed, single pause duration and speech speed change trend are combined, emotion intensity values are analyzed, and emotion classification of the consultation object is judged according to the emotion intensity values;

[0087] For example, if there is a positive emotion keyword, while the real-time speech speed V speech is greater than 2.5 words per second, the single pause duration T pause is greater than 0.2 seconds, and the speech speed change trend T v is greater than -0.2, the emotion is marked as positive, and the emotion intensity value S stress is greater than or equal to 8, corresponding to extreme positive emotion, the emotion intensity value is in [5, 7], corresponding to moderate positive emotion. speech If there is a negative emotion keyword, while the real-time speech speed V pause is less than 1.5 words per second, the single pause duration T v is greater than 0.5 seconds, and the speech speed change trend T stress is less than -0.2, the emotion is marked as negative, and the emotion intensity value S speech is greater than or equal to 8, corresponding to extreme negative emotion, the emotion intensity value is in [5, 7], corresponding to moderate negative emotion. pause If there is no obvious keyword, while the real-time speech speed V word is in [1.5 words per second, 2.5 words per second], the single pause duration T word is in [0.2 seconds, 0.5 seconds], and the absolute value of the speech speed change trend is less than or equal to 0.2, the emotion is marked as neutral, and the intensity is in [3, 5];

[0088] The sum of the difference values of the emotion intensity values of the continuous three windows is obtained, and the emotion change trend is analyzed according to the sum of the difference values of the emotion intensity values;

[0089] If the emotion change trend is greater than 5, the emotion is rapidly warming up and the positive or negative emotion is intensified; if the sum of the difference values of the emotion change trend is less than -5, the emotion is rapidly cooling down and the emotion is relieved or turned to neutral; otherwise, the emotion change is stable;

[0090] Step four: a comprehensive stability value calculation method is designed in combination with the key word and emotion change association rule, which reflects the emotion stability and the influence degree of the key word, and provides a quantitative index for psychological health assessment;

[0091] In specific embodiments, for each marked keyword, the maximum value of the emotion intensity value within 3 seconds after the appearance of the keyword is subtracted from the average value of the emotion intensity value within 3 seconds before the appearance of the keyword, to calculate the emotion intensity change amount ΔE within 3 seconds before and after the appearance of the corresponding keyword, and the influence coefficient K word quantifies the degree of influence of the keyword on the emotion, and K word = ΔE / T interval , T intervalThe time difference for the keyword to appear to the peak of the emotion change is the positive keyword influence coefficient K word Greater than 0, negative keyword influence coefficient K word Less than 0, the greater the absolute value, the stronger the influence;

[0092] Calculate the standard deviation SD of the emotion intensity throughout the conversation E , reflecting the dispersion degree of emotion fluctuation, calculate the number of emotion type jumps N switch , the emotion type of the two consecutive windows is different and the intensity difference is greater than 2 points, which is counted as an emotion type jump, avoiding false statistics caused by small fluctuations, and the emotion fluctuation degree F E is analyzed by the formula SD switch ×(1+0.2×N fluct ), where 0.2 is the preset correction coefficient of jump number, which amplifies the negative impact of frequent jumps on stability, and comprehensively reflects the stability of emotion;

[0093] Based on the obtained influence coefficient, combined with the clinical interval mapping, the sum of the keyword influence coefficient |ΣK word | and the distribution interval of the emotion fluctuation degree F fluct and the corresponding psychological health influence degree are established to establish a two-dimensional interval deduction item table;

[0094] According to the interval to which the sum of the current keyword influence coefficient belongs and the interval to which the emotion fluctuation degree belongs, the score is deducted, the corresponding deduction item of the sum of the keyword influence coefficient is determined through clinical emotion stability correlation analysis, and the corresponding deduction item of the emotion fluctuation degree is set according to the consultation effect tracking data; The total emotion deduction value is obtained by accumulating the deduction values of the two items; Combined with the preset full score standard, the psychological health stability value S calc is obtained by subtracting the total emotion deduction value;

[0095] For example, if the sum of the keyword influence coefficient belongs to [0, 1], the deduction is 5 points, if the sum of the keyword influence coefficient belongs to (1, 2], the deduction is 10 points, if the sum of the keyword influence coefficient is greater than 2, the deduction is 15 points, if the emotion fluctuation degree belongs to [0, 0.5], the deduction is 3 points, if the emotion fluctuation degree belongs to (0.5, 1], the deduction is 6 points, if the emotion fluctuation degree is greater than 1, the deduction is 9 points;

[0096] Based on the obtained psychological health stability value, the psychological health stability of the consultation object is analyzed and judged;

[0097] The psychological health stability value is standardized, and the psychological health stability value range is [0 points, 100 points], [80 points, 100 points] is marked as stable, [60 points, 80 points) is marked as relatively stable, [40 points, 60 points) is marked as unstable, and [0 points, 40 points) is marked as high risk;

[0098] Suspect marking and early warning are performed according to the number of emotion type jumps;

[0099] If the number of emotion type jumps is greater than the frequent threshold or the absolute value of the influence coefficient is greater than the conflict threshold, suspect marking is performed. The number of emotion type jumps greater than the frequent threshold indicates that the emotional state changes frequently and the psychological stability is poor. The absolute value of the influence coefficient greater than the conflict threshold indicates that the strong influence keyword corresponds to the core psychological conflict. The marked content includes the keyword text, the occurrence time and the corresponding emotion change curve;

[0100] If the psychological health stability value is less than 40 points, the psychological health early warning prompt is automatically triggered, and the prompt information is sent to the counselor terminal at the same time to assist the counselor to adjust the counseling strategy in time;

[0101] Fusion of emotion stability value and multi-modal speech features, construction of mapping relationship between emotion change trajectory and crisis degree, realization of graded early warning and targeted intervention, multi-modal speech features include: accent position, pause interval and speech rate change;

[0102] The emotion stability value is set as the core index, combined with the accent position change, the pause interval change and the speech rate change rate, the D-S fusion of evidence theory is used to construct the correlation degree model to quantify the contribution degree of each feature to the emotional crisis;

[0103] The difference between the current window accent position and the average accent position of the previous 3 windows is calculated to obtain the accent position offset;

[0104] The difference between the current window average pause duration and the average pause duration of the previous 3 windows is calculated to obtain the pause interval change;

[0105] The average speech rate of the previous 3 windows is subtracted from the current speech rate and then divided by the average speech rate of the previous 3 windows to obtain the speech rate change rate;

[0106] The emotion stability value, accent position offset, pause interval change and speech rate change rate are respectively taken as four evidence bodies E1, E2, E3 and E4;

[0107] For each evidence body, the BPA function is set according to the clinical data to analyze the probability of supporting the crisis and the probability of not supporting the crisis;

[0108] The BPA of the four evidence bodies is fused by the D-S synthesis rule to obtain the final fusion probability, and the crisis correlation degree C is obtained crisis ;

[0109] According to the psychological health stability value change value and the crisis correlation degree, the emotion change corresponds to different crisis levels;

[0110] It should be noted that the psychological health stability value change value is the difference between the current psychological health stability value and the previous psychological health stability value.

[0111] For example, if the mental health stability value changes by less than -10, the mental health stability value decreases from ≥60 to <60, the emotional stability changes to instability; Crisis determination: if 0.8 > crisis association degree ≥ 0.6, it is marked as moderate crisis; if the crisis association degree is ≥ 0.8, it is marked as severe crisis; if the mental health stability value changes by more than 10, the mental health stability value increases from <60 to ≥60, the emotional instability changes to emotional stability; if the absolute value of the mental health stability value change is less than or equal to 5, the mental health stability value is 60 ≤ mental health stability value ≤ 80, the emotion is continuously stable; if the mental health stability value changes by less than or equal to 5, the mental health stability value is <40, the emotion is continuously unstable; otherwise, the emotion belongs to normal fluctuation;

[0112] Step five: For the problems of language barriers and severe emotional disorders, a three-level processing mechanism is constructed to solve the problems of emotional feature extraction distortion and crisis misjudgment of special groups through exclusive baseline and multi-modal evidence fusion, feature adaptation, cross-validation and decision correction;

[0113] In specific embodiments, the speech samples of the stuttering consultation object are collected, the abnormal speech sequence is aligned through dynamic time warping algorithm, and the pathological features and emotional correlation features are extracted;

[0114] For pathological features, the syllable repetition rate, the proportion of prolonged sound, and the frequency of blocking pause of the consultation object are obtained, the stuttering type exclusive baseline is established, and the developmental stuttering and acquired stuttering are distinguished according to the syllable repetition rate and the blocking pause frequency;

[0115] It should be noted that the syllable repetition rate is the ratio of single syllable repetition times to total syllables, the proportion of prolonged sound is the ratio of prolonged time between syllables to total pronunciation time, and the blocking pause frequency is the number of non-sound blocking pauses per unit time;

[0116] For emotional correlation features, the ratio of net speech speed after excluding pathological repetition and prolongation, the ratio of non-blocking pause duration to total pause duration, and the fundamental frequency variation coefficient after excluding stuttering segments are analyzed;

[0117] Using pathological feature masking and emotional feature reconstruction methods, the stuttering segments identified in the speech signal are masked based on exclusive baseline judgment, the emotional correlation features of non-stuttering segments are retained, and the misjudgment of speech rate or pause caused by stuttering alone is corrected;

[0118] For example, in combination with fuzzy logic reasoning, the combination of net speech speed <1.2 words per second, emotional pause proportion >60%, and fundamental frequency fluctuation coefficient >0.25 is determined as emotional anxiety tendency;

[0119] The distorted speech is preprocessed and feature is recovered, and for the distortion types of burst noise, tremolo and air sound during emotional outburst, a layered enhancement algorithm is adopted, the speech signal is decomposed through a wavelet threshold denoising algorithm, the soft threshold value processing is performed on the high frequency burst component, and the trend characteristics of the fundamental frequency and amplitude are retained;

[0120] Adaptive Kalman filtering is adopted, the state equation is constructed based on the mean value of the fundamental frequency of the previous three windows, and the high frequency fluctuation caused by tremolo is smoothed;

[0121] The harmonic-to-noise ratio is improved by the spectral subtraction method, and the enhanced spectral centroid and mel cepstral coefficient are extracted;

[0122] If the speech distortion degree is greater than the preset distortion limit value, multi-modal cross-validation is started, the dynamic change rate of 68 facial key points is extracted, consistency determination is performed on the emotional features of the distorted speech, the consistency rate reaches the consistency standard, and then the emotional type is confirmed; the abnormality degree of heart rate variability and skin conductivity is calculated, when both indicators reach the preset high arousal threshold, the negative emotional judgment of the distorted speech segment is strengthened, and the emotional missed judgment caused by distortion is avoided.

[0123] Embodiment 2: as Figure 2 indicated, the psychological counseling real-time speech recognition system based on multi-modal data provided by the embodiment of the application specifically comprises the following modules:

[0124] The analysis and judgment module: by constructing a gender-specific voice-psychological feature mapping baseline, the emotional stability tendency degree is initially analyzed and judged through the mean value and the coefficient of variation of the fundamental frequency, and the expression tendency degree is judged in combination with the peak value proportion of the amplitude;

[0125] The dynamic correction module: in combination with pronunciation habit details and real data, the mean value and amplitude of the fundamental frequency are dynamically corrected, and the psychological feature misjudgment caused by individual pronunciation difference is solved;

[0126] The optimization analysis module: by measuring the accent position, key words, speech speed and pause features, a multi-dimensional emotional analysis model is constructed, and the problems of emotional focus positioning, emotional type judgment and change trend tracking are optimized;

[0127] The quantitative evaluation module: in combination with the key word and emotional change correlation law, a comprehensive stability value calculation method is designed, the emotional stability and the influence degree of the key word are reflected, and a psychological health evaluation quantitative index is provided;

[0128] The adaptive correction module: for the problems of language barriers and serious emotional disorders, a three-level processing mechanism is constructed, feature adaptation, cross-validation and decision correction are performed through exclusive baseline and multi-modal evidence fusion.

[0129] The above describes one embodiment of the present application in detail, but the content is only the preferred embodiment of the present application and cannot be considered to limit the scope of the present application; the above formulas are all dimensionless values, and the formulas are obtained by collecting a large amount of data to simulate a formula of the most recent real situation, and the preset parameters in the formula are set by the person skilled in the art according to the actual situation and historical experience, and can be adjusted according to the actual situation; the above is only the preferred embodiment of the present application and cannot be used to limit the present application, and all equivalent changes and improvements made according to the scope of the present application should still belong to the patent coverage range of the present application.

Claims

1. A psychological counseling real-time speech recognition method based on multi-modal data, characterized in that, The method comprises the following steps: By constructing a gender-specific voice psychological feature mapping baseline, the initial analysis of the mean fundamental frequency and the coefficient of variation of the fundamental frequency is used to judge the emotional stability tendency degree, and the amplitude peak value proportion is used to judge the expression tendency degree; The method for constructing the gender-specific voice psychological feature mapping baseline is: The initial dialogue voice of the consultation object is collected by using a double microphone array, the main microphone is placed in front of the visitor to focus on the target voice, and the reference microphone is placed at the back to collect environmental noise, the sampling rate is uniformly set, and the quantization bit number is set; The LMS adaptive filtering algorithm is used to reduce noise of the double microphone signals, non-steady-state noise in the consultation room is suppressed, and pure voice signals are retained; According to the physiological range of the fundamental frequency of men and women, the mean fundamental frequency and the coefficient of variation of the fundamental frequency are calculated, and the amplitude peak value proportion is calculated in combination with the amplitude dynamic range of men and women; Based on the pure voice signals of male and female visitors of different ages and occupations, the gender-specific mapping rules are trained and optimized by using a random forest algorithm, and initial judgments of emotional stability tendency, emotional sensitivity tendency and expression tendency are realized; In combination with pronunciation habit details and real data, the mean fundamental frequency amplitude is dynamically corrected, and the psychological feature misjudgment caused by individual pronunciation difference is solved; By detecting the accent position, key words, speech speed and pause features through the fundamental frequency mutation and the amplitude mutation, a multi-dimensional emotion analysis model is constructed, and problems of emotion focus positioning, emotion type judgment and change trend tracking are optimized; In combination with the correlation law of key words and emotion changes, a comprehensive stability value calculation method is designed, which simultaneously reflects the emotional stability and the influence degree of key words, and provides a quantitative index for psychological health assessment; For the problems of language barriers and serious emotional disorders, a three-level processing mechanism is constructed, feature adaptation, cross-validation and decision correction are performed through the fusion of exclusive baseline and multi-modal evidence.

2. The real-time voice recognition method for psychological counseling based on multi-modal data according to claim 1, characterized in that, The method for dynamically correcting the mean fundamental frequency amplitude is: The preprocessed voice is phoneme-level disassembled by using a Kaldi speech recognition toolkit, three types of pronunciation habit features are extracted, the deviation rate of the average duration of vowel pronunciation and the standard duration of Mandarin, the standard deviation of consonant initial pronunciation intensity, and the deviation rate of the frequency of four tones and the standard tone of Mandarin are calculated; The structured information of the age, occupation, dialect background and previous psychological state of the consultation object is obtained, a correction factor is constructed, and the mean fundamental frequency is calculated by combining the original mean fundamental frequency with the age factor and the occupation factor; The mean fundamental frequency amplitude is recalculated by using the corrected mean fundamental frequency and amplitude, and the psychological feature is updated again in combination with the pronunciation habit marker. 3.The real-time voice recognition method for psychological counseling based on multi-modal data according to claim 1, wherein, The method for constructing the multi-dimensional emotion analysis model is: A 3-second sliding window is set to extract the maximum fundamental frequency, the mean fundamental frequency, the maximum amplitude and the mean amplitude; The joint distribution entropy of the fundamental frequency mutation value and the amplitude mutation value and the maximum values of the two are used to determine the accent position; the accent position vocabulary is matched with the key word library for psychological counseling selected by the TF-IDF algorithm; the real-time speech speed, the speech speed change trend, the single pause duration and the pause frequency are calculated by using a 3-second window and a 1-second step; The decision tree algorithm is used to train and optimize the emotion judgment rules, and the model is constructed in combination with the key words and multi-dimensional features. The method for judging the emotion classification of the consultation object is:

4. The real-time voice recognition method for psychological counseling based on multi-modal data according to claim 3, characterized in that, The emotion judgment rules are constructed in combination with multi-dimensional features, the rules are optimized through decision tree algorithm training, the emotion intensity value is analyzed in combination with keywords, real-time speech speed, single pause duration and speech speed change trend, and the emotion classification of the consultation object is judged according to the emotion intensity value. The method for obtaining the emotion classification of the consultation object is:

5. The real-time voice recognition method for psychological counseling based on multi-modal data according to claim 4, characterized in that, The sum of the emotion intensity value differences of three consecutive windows is obtained, and the emotion change trend is analyzed according to the sum of the emotion intensity value differences; If the emotion change trend is greater than 5, the emotion is rapidly warming up positively or the negative emotion is intensified; If the emotion change trend is less than -5, the emotion is rapidly cooling down, the emotion is relieved or turned to neutral; otherwise, the emotion change is stable. The method for obtaining the comprehensive stability value is:

6. The real-time voice recognition method for psychological counseling based on multi-modal data according to claim 1, characterized in that, The difference between the maximum emotion intensity value within 3 seconds after the appearance of the keyword and the average value within 3 seconds before the appearance of the keyword is calculated to define the influence coefficient; the standard deviation of the emotion intensity and the emotion type jump frequency are calculated, and the emotion fluctuation degree is obtained by correcting the emotion intensity standard deviation in combination with the jump frequency; A double-dimensional interval deduction item table is established in combination with mapping, the sum of the keyword influence coefficients and the deduction value of the emotion fluctuation degree are accumulated, the psychological health stability value is obtained by subtracting the total deduction value from the preset full score. The method for quantifying the psychological health evaluation is:

7. The real-time voice recognition method for psychological counseling based on multi-modal data according to claim 1, characterized in that, The emotion stability value is standardized, and is divided into stable, relatively stable, unstable and high-risk types; If the emotion type jump frequency is greater than the frequent threshold or the absolute value of the keyword influence coefficient is greater than the conflict threshold, the suspicious mark of the text containing the keyword, the appearance time and the emotion change curve is performed; if the emotion stability value is less than 40 points, an early warning prompt is automatically triggered and is synchronized to the counselor terminal; a correlation degree model is constructed by using the D-S evidence theory to fuse the emotion stability value, the pitch position offset amount, the pause interval change amount and the speech speed change rate, the danger correlation degree is obtained through the BPA function and the synthesis rule, and the danger level is divided in combination with the emotion stability value change value. The method for feature adaptation is: 8.The real-time voice recognition method for psychological counseling based on multi-modal data according to claim 1, wherein, The speech samples of the stuttering consultation object are collected, the abnormal speech sequence is aligned through the dynamic time warping algorithm, the pathological features are extracted to establish the stuttering type exclusive baseline, and the developmental and acquired stuttering is divided; the emotion related features are analyzed; The stuttering segments in the speech are processed by using the pathological feature mask to retain the emotion related features of the non-stuttering segments, and the misjudgment is corrected by using the fuzzy logic reasoning. The method for constructing a three-level processing mechanism is: 9.The real-time voice recognition method for psychological counseling based on multi-modal data according to claim 1, wherein, For the explosive sound, tremolo and air sound distortion in the emotion burst, a hierarchical enhancement algorithm is used, the speech signal is decomposed by wavelet threshold denoising and the high-frequency burst component is processed by soft threshold, the adaptive Kalman filter is used to smooth the high-frequency fluctuation of tremolo, the spectral subtraction is used to improve the harmonic-to-noise ratio and extract the spectral centroid and mel cepstral coefficient; If the speech distortion degree is greater than a preset limit value, a multi-modal cross-validation is started, the dynamic change rate of 68 facial key points and the emotion features of distorted speech are determined for consistency; The heart rate variability and the abnormality degree of skin conductance are calculated, and the negative emotion judgment is intensified when both reach a high awakening threshold. ​

Citation Information

Patent Citations

  • Voice quality inspection method, device and equipment

    CN116206593A

  • Method, system and device for intelligently generating outbound call skill by using LLM (Logical Link Model) and medium

    CN120123472A