Psychological counseling automatic recording method and system based on voice recognition
By using speech recognition and voiceprint feature clustering technology, multi-role voice streams are separated and emotional attribution is corrected in psychological counseling. This solves the problems of identity confusion and misjudgment of emotional state in multi-role voice streams, and achieves accuracy and reliability in psychological counseling records.
Patent Information
- Application Number
- CN202511839408.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-08
- Publication Date
- 2026-02-10
AI Technical Summary
Existing technologies struggle to accurately separate and label the voice streams of multiple characters in psychological counseling, leading to misjudgments in emotional state analysis and affecting the objectivity of records and the relevance of intervention recommendations.
The audio signals of psychological counseling are collected by a recording capture device, and speech recognition algorithms are applied to process noise and speech rate changes. The voiceprint feature clustering method is combined to separate the voice stream of the roles, and the emotion attribution judgment is corrected by the emotion analysis model to generate a structured counseling log.
It achieves precise separation of multi-role voice streams and captures dynamic changes in emotional states, improving the accuracy of emotion analysis and the reliability of psychological counseling records, and ensuring the accuracy of intervention recommendations.
Smart Images

Figure CN121506201A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of speech recognition technology, and in particular relates to a method and system for automatically recording psychological counseling based on speech recognition. Background Technology
[0002] In the field of psychological counseling, intelligent technology is gradually becoming an important means to improve service quality and efficiency. This field not only concerns the support and intervention for individual mental health, but also plays an indispensable role in social welfare and public health systems. Especially in the generation and management of counseling records, intelligent methods can help counselors organize information more efficiently, analyze emotional dynamics, and provide a scientific basis for subsequent interventions. However, how to accurately capture and process information in complex dialogue scenarios remains a challenge that urgently needs to be overcome.
[0003] Currently, while some solutions have attempted to assist in recording psychological counseling sessions using speech recognition and text analysis, most suffer from deep-seated limitations. These methods often struggle to adapt to the complexities of multi-role and multi-emotional interactions in counseling scenarios, particularly in the dynamic shifts in role identities and subtle fluctuations in emotional states during dialogue, lacking the ability to deeply analyze and accurately differentiate between them. Furthermore, existing technologies often fall short when faced with semantic understanding within professional fields and the constraints of ethical norms, potentially leading to recorded content that deviates from actual needs or crosses privacy boundaries.
[0004] A deeper technical challenge lies in accurately separating the speech streams of different roles in dynamic dialogues and labeling their identities. The accuracy of identity labeling directly affects subsequent attribution of emotional states and semantic information. Deviations in this step will lead to misjudgments in sentiment analysis and intervention strategies. For example, during a consultation, the voices of the counselor and the client may be confused due to environmental noise or changes in speech rate. If the system cannot clearly distinguish between the two parties, it may incorrectly attribute the client's emotional fluctuations to the counselor, thus affecting the objectivity of the recording and the relevance of the intervention recommendations. More complexly, the challenge of identity labeling extends further to the real-time capture of emotional states, because emotional expression is often closely related to context and role identity. A lack of accurate identity differentiation will significantly reduce the reliability of emotional labels.
[0005] Therefore, achieving accurate separation and identity labeling of voice streams in multi-role dialogues, and on this basis, accurately capturing dynamic changes in emotional states, has become a key issue in improving the intelligence level of psychological counseling records. Solving this problem not only involves technological breakthroughs but also directly affects whether counseling records can truly reflect the essence of the dialogue, thereby providing reliable support for psychological intervention. Summary of the Invention
[0006] To address the aforementioned technical issues, this invention proposes an automatic recording method and system for psychological counseling based on speech recognition. This method can achieve accurate separation and identity labeling of speech streams in multi-role dialogues, and on this basis, accurately capture the dynamic changes in emotional states.
[0007] To achieve the above objectives, the present invention provides an automatic recording method for psychological counseling based on speech recognition, comprising:
[0008] The audio signal of the psychological counseling dialogue was collected by a recording capture device to obtain the raw voice data stream;
[0009] Based on the original speech data stream, a speech recognition algorithm is applied to process noise and speech rate changes to determine the separated multi-role speech stream;
[0010] If the separated multi-role speech streams show confusion in role identities, then the speaker clustering method is used to compare voiceprint features to obtain the identity-labeled speech streams;
[0011] Semantic content and tone features are extracted from the identity-labeled speech stream to obtain emotional state labels;
[0012] Based on the emotional state tags and the fusion of dialogue context information, a real-time emotional dynamic change sequence is obtained;
[0013] If there is a deviation in the real-time emotional dynamic change sequence, the attribution judgment is corrected by the emotional analysis model to determine the target emotional attribution record;
[0014] A structured counseling log is generated based on the target's emotional affiliation record.
[0015] Optionally, the audio signal of the psychological counseling dialogue can be acquired using a recording capture device to obtain the raw speech data stream, including:
[0016] The initial audio signal is obtained by capturing the dialogue audio in the psychological counseling scene in real time using a recording device;
[0017] Based on the initial audio signal, a preset noise filtering tool is used to perform preliminary cleaning of the signal to obtain a denoised audio stream;
[0018] The denoised audio stream is divided into multiple time segments using segmentation processing technology, and the speech content range of each segment is determined.
[0019] If the audio quality of a certain time segment is lower than a preset threshold, signal enhancement processing is performed on it to obtain an optimized audio segment.
[0020] Based on the optimized audio segment, speech recognition technology is used to convert it into text data to obtain the original speech data stream.
[0021] Optionally, based on the original speech data stream, a speech recognition algorithm is applied to process noise and speech rate variations to determine the separated multi-role speech stream, including:
[0022] The initial audio signal is obtained from the original speech data stream, and the signal is initially cleaned using a pre-established filtering technique to obtain the denoised basic speech data.
[0023] For the denoised basic speech data, a speech recognition algorithm is applied to detect and correct speech rate fluctuations, and to determine the speech segments after speech rate adjustment.
[0024] Based on the speech segments with adjusted speech rate, the voice features of multiple characters are obtained, and the speech signals of different characters are separated by hierarchical analysis technology to obtain the preliminary separated character speech streams;
[0025] For the initially separated character speech streams, a time series comparison method is used to determine the continuity of the speech flow. If an interruption or overlap is detected, the segments are spliced together to obtain continuous character speech data.
[0026] Based on continuous character voice data, the independent voice features of each character are obtained, and clustering methods are applied to group the features to determine the final independent voice stream;
[0027] For the final independent speech stream, structured speech flow data is generated through format standardization processing, resulting in a separated multi-role speech stream.
[0028] Optionally, if the separated multi-role speech streams show confusion regarding role identities, then a speaker clustering method is used to compare voiceprint features to obtain the identity-labeled speech streams, including:
[0029] Based on the separated multi-role speech stream, a pre-established voiceprint feature extraction model is used to obtain the voiceprint feature data of each speech segment;
[0030] By extracting voiceprint feature data, the speaker clustering method is used to group different speech segments and determine the preliminary role identity classification results.
[0031] If there are overlapping roles or unclear groupings in the preliminary classification results, a second comparison is performed on the voiceprint features of the overlapping segments to determine whether they belong to the same role, thus obtaining a more accurate classification and grouping.
[0032] Based on the classification and grouping after the second comparison, corresponding identity labeling information is generated for each role, and a set of voice segments with identity labels is obtained;
[0033] By integrating the annotated audio segments over time, a continuous stream of identity-annotated audio is generated, and the final role-identity correspondence is determined.
[0034] If there are still some segments with inconsistent identity labels in the integrated speech stream, logical correction is performed using the identity information of the context speech segments to obtain the final corrected identity-labeled speech stream.
[0035] Optionally, extracting semantic content and tone features from the identity-labeled speech stream to obtain emotional state labels includes:
[0036] The semantic information of the identity-tagged speech stream is extracted using signal processing techniques to form preliminary semantic text content;
[0037] Based on the extracted semantic text content, a pre-established semantic analysis model is used to analyze the key expressions in the text and obtain a preliminary judgment result on sentiment tendency.
[0038] For voice segments of the same identity, acoustic analysis tools are used to extract tone features, analyze changes in pitch and speech rate, and determine the intensity of emotion reflected by the tone.
[0039] If the tone characteristics are consistent with the semantic sentiment tendency judgment, the two are combined to obtain the sentiment state category of the identity; if they are inconsistent, the tone characteristics are given priority and the sentiment tendency judgment is readjusted.
[0040] Based on the determined emotional state category, generate corresponding emotional state classification labels.
[0041] Optionally, the real-time emotional dynamic change sequence obtained by fusing dialogue context information based on the emotional state tags includes:
[0042] Based on the emotional state tags, combined with the time dimension and interaction process in the dialogue content, a pre-established emotional analysis model is used to analyze the changing trend of emotional state at different points in time and determine the continuity characteristics of emotional dynamics.
[0043] Based on the continuous characteristics of the emotional dynamics, the interaction process information in the context and dialogue content is integrated, and the trend of emotional state change is matched with real-time monitoring data through a logical comparison method to determine the significance of emotional changes.
[0044] If the significance of the emotional change exceeds a preset threshold, the change trend is associated with the label information to generate a corresponding dynamic change sequence fragment, thus obtaining real-time updated emotional sequence data.
[0045] Optionally, if there is a deviation in the real-time emotional dynamic change sequence, the attribution judgment is corrected through an emotional analysis model, and the target emotional attribution record is determined to include:
[0046] If there is a deviation in the real-time emotional dynamic change sequence, the model correction process is triggered, and the emotional analysis model is used to perform in-depth analysis of the data to determine the corrected emotional classification.
[0047] Key features are extracted from the corrected sentiment classification, and data processing techniques are used to clean and organize the dynamically changing data to output structured sentiment attribution information.
[0048] Based on the structured sentiment attribution information, analytical tools are used to track the fluctuations of the sentiment sequence, determine whether there is a persistent deviation, and obtain deviation tracking records.
[0049] By using the deviation tracking record and combining it with the correction mechanism to dynamically adjust the parameters of the sentiment analysis model, the target sentiment attribution record is determined.
[0050] The present invention also provides an automatic recording system for psychological counseling based on speech recognition, including: a data acquisition module, an identity labeling module, an emotion sequence acquisition module, an emotion attribution recording module, and a counseling log generation module;
[0051] The data acquisition module is used to acquire audio signals of psychological counseling dialogues through a recording capture device to obtain raw voice data streams;
[0052] The identity labeling module is used to process noise and speech rate changes by applying a speech recognition algorithm to the original speech data stream to determine the separated multi-role speech stream; if the separated multi-role speech stream shows role identity confusion, then the speaker clustering method is used to compare voiceprint features to obtain the identity-labeled speech stream.
[0053] The emotion sequence acquisition module is used to extract semantic content and tone features from the identity-labeled speech stream to obtain emotion state labels; and to fuse dialogue context information based on the emotion state labels to obtain a real-time emotion dynamic change sequence.
[0054] The emotion attribution recording module is used to determine whether there is a deviation in the real-time emotion dynamic change sequence. If there is a deviation in the real-time emotion dynamic change sequence, the attribution judgment is corrected by the emotion analysis model to determine the target emotion attribution record.
[0055] The consultation log generation module is used to generate structured consultation logs based on the target's emotional affiliation record.
[0056] Compared with the prior art, the present invention has the following advantages and technical effects:
[0057] This invention addresses the complex business challenges of multi-role speech recognition, capturing dynamic emotional changes, and accurately attributing emotions in psychological counseling scenarios, proposing an integrated solution. This problem involves challenges such as role confusion in speech data streams, real-time fluctuations in emotional states, and contextual bias. This invention uses speech recognition algorithms to handle noise and speech rate variations, combines voiceprint feature clustering to resolve role confusion, extracts semantic and tone features to determine emotional states, integrates contextual information to generate dynamic change sequences, and corrects biases through an emotion analysis model, ultimately forming accurate emotion attribution records and structured counseling logs. The core innovation of this invention lies in its multimodal data fusion and real-time emotion correction mechanism, ensuring the accuracy of emotion analysis and the reliability of intervention criteria, significantly improving the effectiveness of emotion insight and decision support in psychological counseling. Attached Figure Description
[0058] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:
[0059] Figure 1 This is a flowchart of an automatic recording method for psychological counseling based on speech recognition, according to an embodiment of the present invention. Detailed Implementation
[0060] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0061] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0062] This embodiment proposes an automatic recording method for psychological counseling based on speech recognition, such as... Figure 1 As shown, the specific steps include:
[0063] The audio signal of the psychological counseling dialogue was collected by a recording capture device to obtain the raw voice data stream;
[0064] Based on the original speech data stream, a speech recognition algorithm is applied to process noise and speech rate variations to determine the separated multi-role speech stream;
[0065] If the separated multi-role speech streams show confusion in role identities, then the speaker clustering method is used to compare voiceprint features to obtain the identity-labeled speech streams;
[0066] Semantic content and tone features are extracted from identity-labeled speech streams to obtain emotional state labels;
[0067] By fusing dialogue context information with emotional state tags, a real-time sequence of emotional dynamic changes can be obtained.
[0068] If there are deviations in the real-time emotional dynamic change sequence, the attribution judgment is corrected by the emotional analysis model to determine the target emotional attribution record;
[0069] A structured counseling log is generated based on the target's emotional affiliation record.
[0070] Furthermore, the audio signal of the psychological counseling dialogue is acquired through a recording capture device, resulting in the raw voice data stream, including:
[0071] The initial audio signal is obtained by capturing the dialogue audio in the psychological counseling scene in real time using a recording device;
[0072] Based on the initial audio signal, a preset noise filtering tool is used to perform preliminary cleaning of the signal to obtain a denoised audio stream;
[0073] For the denoised audio stream, segmentation processing technology is used to divide it into multiple time segments and determine the range of speech content in each segment;
[0074] If the audio quality of a certain time segment is lower than a preset threshold, signal enhancement processing is performed on it to obtain an optimized audio segment.
[0075] Based on the optimized audio clips, speech recognition technology is used to convert them into text data to obtain the original speech data stream.
[0076] Specifically, in psychological counseling scenarios, capturing real-time audio of conversations using recording equipment is the starting point of the entire process. Suppose a counselor and client are discussing emotion management issues; the recording equipment captures their conversation, forming the initial audio signal. This initial signal may contain environmental noise, such as air conditioning hum or traffic noise, which can interfere with subsequent analysis. Therefore, using pre-set noise filtering tools for initial cleanup is crucial. A specific implementation could be using spectral analysis to distinguish the frequency range of background noise from the human voice, thus preserving the main speech content and obtaining a denoised audio stream. The advantage of this method is improved audio clarity, laying the foundation for subsequent processing.
[0077] For the denoised audio stream, segmentation processing technology breaks it down into multiple time segments. Assuming a 30-minute dialogue is divided into 30-second segments, the speech content range of each segment is determined using endpoint detection technology to ensure that each segment contains complete sentences. If the audio quality of a segment falls below a preset threshold due to the visitor's low volume (e.g., below 20 decibels), signal enhancement processing is required. Specific enhancement methods could include increasing the volume or adjusting the dynamic range to obtain an optimized audio segment. This processing effectively avoids information loss and ensures the accuracy of speech recognition.
[0078] In the speech recognition stage, the optimized audio segments are converted into text data. Suppose an audio segment contains the phrase "I've been feeling very anxious lately," speech recognition technology can generate corresponding dialogue text. This can be achieved using a deep learning-based speech model to map the audio waveform to text, significantly improving conversion efficiency. Next, the dialogue text is segmented to extract key semantic units, such as words like "anxiety" and "lately," and semantic analysis is used to determine the text's topic distribution, for example, identifying the topic as "emotional distress." This helps to quickly focus on the core issue.
[0079] Furthermore, based on the original speech data stream, a speech recognition algorithm is applied to process noise and speech rate variations, and the separated multi-role speech stream is determined to include:
[0080] The initial audio signal is obtained from the raw speech data stream, and the signal is initially cleaned using a pre-established filtering technique to obtain the denoised basic speech data.
[0081] For the denoised basic speech data, a speech recognition algorithm is applied to detect and correct speech rate fluctuations, and to determine the speech segments after speech rate adjustment.
[0082] Based on the speech segments with adjusted speech rate, the voice features of multiple characters are obtained, and the speech signals of different characters are separated by hierarchical analysis technology to obtain the preliminary separated character speech streams;
[0083] For the initially separated character speech streams, a time series comparison method is used to determine the continuity of the speech flow. If an interruption or overlap is detected, the segments are spliced together to obtain continuous character speech data.
[0084] Based on continuous character voice data, the independent voice features of each character are obtained, and clustering methods are applied to group the features to determine the final independent voice stream;
[0085] For the final independent speech stream, structured speech flow data is generated through format standardization processing, resulting in a separated multi-role speech stream.
[0086] Specifically, in the field of audio processing for psychological counseling dialogues, extracting the initial signal from the raw speech data stream and cleaning it is a crucial first step.
[0087] In one possible implementation, the acquired audio signal can be processed using pre-defined filtering techniques to remove low-frequency interference from the background, such as the continuous hum of an indoor fan or the honking of vehicles outside the window, ensuring relatively clean basic speech data. This processing provides a clear starting point for subsequent analysis.
[0088] For the denoised base speech data, speech rate fluctuation detection and correction are particularly important. Speech rate fluctuations may occur due to the visitor's emotional excitement or hesitation, affecting the accuracy of speech recognition.
[0089] In one possible implementation, a speech recognition algorithm can analyze the number of syllables per second. Assuming a normal speaking speed is 150 syllables per minute, if a segment of speech is detected to suddenly drop to 80 syllables per minute, the system will automatically stretch or compress the timeline of that audio segment to bring it closer to the standard speaking speed range, creating an adjusted speech fragment. This helps to more accurately capture speech features during subsequent character separation.
[0090] In the character voice feature separation stage, distinguishing the voice signals of multiple characters is the core challenge.
[0091] In one possible implementation, hierarchical analysis technology can be used to extract the voices of counselors and clients based on differences in timbre and pitch. Assuming the counselor's voice frequency range is 100-200 Hz, while the client's voice frequency range is 200-300 Hz, the system will separate an initial character-based speech stream based on this characteristic. This separation method can effectively distinguish the audio content of different speakers.
[0092] For the initially separated voice streams, continuity judgment is a key step in ensuring the integrity of the voice stream.
[0093] One possible implementation involves using a time-series comparison method. If a segment of the speech stream is found to have an interruption of more than 5 seconds on the timeline, possibly due to a pause by the speaker or loss of audio capture, the system will automatically search for the semantic connections between the preceding and following segments and splice them together to form continuous, role-based speech data. This avoids information fragmentation and improves the integrity of the speech stream.
[0094] The application of clustering methods is crucial when acquiring the unique voice features of each character.
[0095] In one possible implementation, the system groups the speech data based on features such as volume and tone. For example, if the counselor's voice volume is consistently around 50 decibels, while the client's voice volume fluctuates between 30 and 60 decibels, the system would cluster these features into two independent speech streams. This grouping method helps to more clearly present the speech characteristics of each role.
[0096] Furthermore, if the separated multi-role speech streams show confusion regarding role identities, then a speaker clustering method is used to compare voiceprint features to obtain the identity-labeled speech streams, including:
[0097] Based on the separated multi-role speech streams, a pre-established voiceprint feature extraction model is used to obtain the voiceprint feature data for each speech segment;
[0098] By extracting voiceprint feature data, the speaker clustering method is used to group different speech segments and determine the preliminary role identity classification results.
[0099] If there are overlapping roles or unclear groupings in the preliminary classification results, a second comparison is performed on the voiceprint features of the overlapping segments to determine whether they belong to the same role, thus obtaining a more accurate classification and grouping.
[0100] Based on the classification and grouping after the second comparison, corresponding identity labeling information is generated for each role, and a set of voice segments with identity labels is obtained;
[0101] By integrating the annotated audio segments over time, a continuous stream of identity-annotated audio is generated, and the final role-identity correspondence is determined.
[0102] If there are still some segments with inconsistent identity labels in the integrated speech stream, logical correction is performed using the identity information of the context speech segments to obtain the final corrected identity-labeled speech stream.
[0103] Specifically, when processing raw multi-role speech stream data, to address the issue of identity confusion, a voiceprint feature extraction model can be used to distinguish different speakers. The principle of voiceprint feature extraction lies in capturing the unique acoustic characteristics of speech, such as timbre and pitch variations, which are unique to each individual, just like fingerprints.
[0104] In one possible implementation, suppose a meeting recording scenario is used to record a conversation between three participants for a total duration of 30 minutes. The system first divides the audio into multiple short segments, each 2 seconds long, and then extracts voiceprint features from each segment to form a feature vector for subsequent analysis.
[0105] For speaker clustering methods, distance-based clustering can be used to group segments with similar voiceprint features into the same group. Suppose that in the aforementioned conference recording, the system initially divides all segments into 3 groups, but finds that one group has significantly fewer segments, accounting for only 5% of the total, potentially indicating classification bias. In this case, a secondary comparison can be introduced. By calculating the similarity between feature vectors, it can further confirm whether these segments belong to other groups, thereby optimizing the classification results.
[0106] When generating identity labeling information, a unique identifier can be assigned to each role, such as role A, role B, and role C, and these identifiers can be bound to corresponding audio segments. Assuming that in a meeting recording, role A's audio segments account for 40%, role B's for 35%, and role C's for 25%, the system will label each segment with its corresponding identity based on the clustering results, forming a preliminary label set. This approach helps to quickly identify the audio content of different roles during subsequent integration.
[0107] For time-series integration, labeled segments can be rearranged chronologically to form a continuous audio stream. For example, if during integration, a 2-second segment is labeled as role A, but the preceding and following 10-second segments are labeled as role B, the system will infer from the context that the segment may be incorrectly labeled and correct it. This method effectively improves the coherence of identity labeling in audio streams.
[0108] If inconsistencies in identity labeling persist after integration, they can be resolved through contextual logic correction. For example, suppose a segment of speech in a meeting recording has blurred voiceprint features due to background noise, preventing the system from directly identifying the speaker. However, by analyzing the preceding and following 5-second segments, both labeled as "role C," it can be inferred that the segment also belongs to "role C." This correction method further improves the accuracy of identity labeling, providing a reliable foundation for subsequent speech analysis.
[0109] Furthermore, semantic content and tone features are extracted from the identity-labeled speech stream to obtain emotional state labels, including:
[0110] Signal processing techniques are used to extract semantic information from identity-tagged speech streams to form preliminary semantic text content;
[0111] Based on the extracted semantic text content, a pre-established semantic analysis model is used to analyze the key expressions in the text and obtain a preliminary judgment result on sentiment tendency.
[0112] For voice segments of the same identity, acoustic analysis tools are used to extract tone features, analyze changes in pitch and speech rate, and determine the intensity of emotion reflected by the tone.
[0113] If the tone characteristics and the semantic sentiment tendency judgment results are consistent, the two are combined to obtain the sentiment state category of the identity; if they are inconsistent, the tone characteristics are given priority and the sentiment tendency judgment is readjusted.
[0114] Based on the determined emotional state category, generate corresponding emotional state classification labels.
[0115] Specifically, in the business of sentiment analysis of voice streams, the initially collected raw audio data often contains multiple types of sound information, requiring preprocessing to separate the voice segments of different speakers. Suppose a 10-minute audio segment containing three speakers is captured in a meeting recording scenario. Using signal separation technology, the audio is divided into multiple independent segments, each corresponding to the voice data of one speaker. This separation lays the foundation for subsequent analysis.
[0116] When extracting semantic information, speech-to-text tools can be used to generate corresponding text content for each speaker's audio segment. For example, if the first speaker's segment contains the expression "Thank you very much for your efforts," a semantic analysis model can initially determine that their emotional tendency is positive. Subsequently, analyzing key expressions such as "thank you" and "efforts" further confirms the emotion as gratitude.
[0117] When analyzing tone characteristics, acoustic tools can be used to detect pitch and speech rate for audio segments from the same speaker. For example, if a second speaker's pitch is high and their speech rate is fast, lasting approximately 15 seconds, this, combined with contextual analysis, might reflect an excited or anxious emotional state. If the semantic text indicates "time is tight, we must speed things up," then the tone is consistent with the semantics, and the emotional intensity can be classified as high.
[0118] If the tone and semantic judgment are inconsistent—for example, if the semantic text of the third speaker is "things are going well," leaning towards a positive tone, but the tone is low and the speaking speed is slow, lasting about 20 seconds—it may imply fatigue or uncertainty. In this case, prioritize the tone and adjust the emotional tendency to a neutral to slightly negative one. This adjustment helps to more closely reflect the speaker's true emotional state.
[0119] When generating emotional state classification labels, each speaker can be labeled with a specific category. For example, the first label might be "grateful," the second "eager," and the third "exhausted." These labels clearly reflect their respective emotional states, facilitating subsequent integrated analysis.
[0120] Furthermore, by fusing dialogue context information with emotion state labels, a real-time emotion dynamic change sequence is obtained, including:
[0121] Based on sentiment state tags, combined with the time dimension and interaction process in the dialogue content, a pre-established sentiment analysis model is used to analyze the changing trend of sentiment state at different points in time and determine the continuity characteristics of sentiment dynamics.
[0122] Based on the continuous nature of emotional dynamics, this study integrates information from the context and dialogue content regarding the interaction process. Through logical comparison, it matches the changing trends of emotional states with real-time monitoring data to determine the significance of emotional changes.
[0123] If the significance of the emotional change exceeds a preset threshold, the change trend is associated with the label information to generate a corresponding dynamic change sequence fragment, thus obtaining real-time updated emotional sequence data.
[0124] Specifically, when analyzing a dialogue, sentiment and context can be extracted from the text data. For example, in a dialogue about customer service feedback, a customer expresses dissatisfaction with the service. The text data can be processed by segmenting the dialogue into individual words using word segmentation tools, and then semantic analysis techniques can be used to identify negative sentiment keywords such as "poor service" and "long waiting time," initially labeling the sentiment as "dissatisfaction." This initial labeling relies on natural language processing tools to identify keywords and sentence structures, enabling rapid identification of sentiment tendencies.
[0125] When analyzing sentiment trends by combining the time dimension and the interaction process, it's important to pay attention to the customer's emotional fluctuations from the beginning to the end of the conversation. Assuming a conversation lasts 10 minutes, with the customer's tone remaining calm for the first 5 minutes and gradually showing impatience in the last 5 minutes, a sentiment analysis model can reveal that the emotional state shifts from "neutral" to "negative," exhibiting a continuous characteristic. This analytical approach helps capture subtle changes in emotional dynamics, providing a basis for subsequent judgments.
[0126] By fusing the dynamic continuity features of emotions, the information in the context where the customer mentions "multiple unsuccessful attempts to contact them" can be combined with the trend of emotional changes to determine the significance of these changes. Assuming a preset threshold of emotional intensity fluctuations exceeding 50%, if the customer's emotional intensity rises from 30% to 80%, it is considered a significant change, generating a corresponding dynamic change sequence segment. This matching method can more accurately reflect the authenticity of emotional fluctuations.
[0127] Furthermore, if there are deviations in the real-time emotional dynamic change sequence, the attribution judgment is corrected through the sentiment analysis model to determine the target sentiment attribution record, which includes:
[0128] If there are deviations in the real-time emotional dynamics sequence, the model correction process is triggered, and the emotional analysis model is used to perform in-depth analysis of the data to determine the corrected emotional classification.
[0129] Key features are extracted from the corrected sentiment classification, and data processing techniques are used to clean and organize the dynamically changing data to output structured sentiment attribution information.
[0130] Based on structured sentiment attribution information, analytical tools are used to track the fluctuations in sentiment sequences, determine whether there are persistent biases, and obtain bias tracking records.
[0131] By tracking deviations and using a correction mechanism, the parameters of the sentiment analysis model are dynamically adjusted to determine the target sentiment attribution record.
[0132] Specifically, in the process of real-time monitoring of dynamic changes in emotional sequences, sensors or voice input devices can capture changes in the user's tone and speaking speed during a conversation. Assuming a 5-minute conversation, the system collects data every 10 seconds, recording the fluctuations in the user's emotions. This method helps to initially capture emotional fluctuations, laying the foundation for subsequent analysis.
[0133] In one possible implementation, the initial processing of a pre-established sentiment analysis model can be performed. The model can be designed based on natural language processing techniques, combining the sentiment orientation of words to make judgments. For example, when keywords such as "happy" or "disappointed" are detected in a conversation, the model will provide an initial sentiment attribution based on the context, which could be "positive" or "negative." This helps to quickly form a preliminary classification, providing a reference for subsequent corrections.
[0134] A correction process is triggered by a bias in the initial sentiment attribution result. This can be achieved by setting a sentiment fluctuation threshold. Assuming the threshold is set at a change in sentiment intensity exceeding 30%, if the system detects a sudden increase in sentiment intensity from 20% to 60% in a dialogue, the correction process is triggered. Deep analysis technology is used to re-analyze semantic and tone features, ultimately determining the corrected sentiment classification as "excitement" rather than "calm." This correction mechanism improves classification accuracy.
[0135] One possible implementation involves extracting key features from the corrected sentiment classification and then cleaning the data by filtering out irrelevant noise. For example, non-emotionally relevant background noise or irrelevant words in the dialogue can be removed, retaining only emotionally relevant interjections and keywords to output structured sentiment attribution information. This cleaning process helps reduce interference and ensure data quality.
[0136] When using analytical tools to track fluctuations in emotional sequences and determine persistent biases, timeline analysis can identify users whose emotional state remains consistently "anxious" for three consecutive minutes with fluctuations not exceeding 10%, thus recording this as a persistent bias. This tracking record provides crucial information for subsequent model adjustments and helps optimize the analysis results.
[0137] In one possible implementation, the parameters of the sentiment analysis model can be dynamically adjusted using a correction mechanism. This allows for adjustments to the model's weighting based on deviation tracking records. For example, the weight of tone features could be increased from 40% to 50% to more accurately reflect user emotions. This dynamic adjustment can continuously optimize model performance and improve the accuracy of sentiment attribution judgments.
[0138] This embodiment also provides an automatic recording system for psychological counseling based on speech recognition, including: a data acquisition module, an identity labeling module, an emotion sequence acquisition module, an emotion attribution recording module, and a counseling log generation module;
[0139] The data acquisition module is used to acquire audio signals of psychological counseling dialogues through a recording capture device to obtain raw voice data streams;
[0140] The identity labeling module is used to process noise and speech rate variations based on the original speech data stream using speech recognition algorithms to determine the separated multi-role speech streams. If the separated multi-role speech streams show role identity confusion, then the speaker clustering method is used to compare voiceprint features to obtain the identity-labeled speech stream.
[0141] The emotion sequence acquisition module is used to extract semantic content and tone features from the identity-labeled speech stream to obtain emotion state labels; and to obtain a real-time emotion dynamic change sequence by fusing dialogue context information based on the emotion state labels.
[0142] The emotion attribution recording module is used to determine whether there is a deviation in the real-time emotion dynamic change sequence. If there is a deviation in the real-time emotion dynamic change sequence, the attribution judgment is corrected through the emotion analysis model to determine the target emotion attribution record.
[0143] The consultation log generation module is used to generate structured consultation logs based on the target's emotional affiliation records.
[0144] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for automatically recording psychological counseling sessions based on speech recognition, characterized in that, include: The audio signal of the psychological counseling dialogue was collected by a recording capture device to obtain the raw voice data stream; Based on the original speech data stream, a speech recognition algorithm is applied to process noise and speech rate changes to determine the separated multi-role speech stream; If the separated multi-role speech streams show confusion in role identities, then the speaker clustering method is used to compare voiceprint features to obtain the identity-labeled speech streams; Semantic content and tone features are extracted from the identity-labeled speech stream to obtain emotional state labels; Based on the emotional state tags and the fusion of dialogue context information, a real-time emotional dynamic change sequence is obtained; If there is a deviation in the real-time emotional dynamic change sequence, the attribution judgment is corrected by the emotional analysis model to determine the target emotional attribution record; A structured counseling log is generated based on the target's emotional affiliation record.
2. The method for automatically recording psychological counseling based on speech recognition according to claim 1, characterized in that, The audio signal of the psychological counseling dialogue was acquired by a recording capture device, and the raw voice data stream included: The initial audio signal is obtained by capturing the dialogue audio in the psychological counseling scene in real time using a recording device; Based on the initial audio signal, a preset noise filtering tool is used to perform preliminary cleaning of the signal to obtain a denoised audio stream; The denoised audio stream is divided into multiple time segments using segmentation processing technology, and the speech content range of each segment is determined. If the audio quality of a certain time segment is lower than a preset threshold, signal enhancement processing is performed on it to obtain an optimized audio segment. Based on the optimized audio segment, speech recognition technology is used to convert it into text data to obtain the original speech data stream.
3. The method for automatically recording psychological counseling based on speech recognition according to claim 1, characterized in that, Based on the original speech data stream, a speech recognition algorithm is applied to process noise and speech rate variations to determine the separated multi-role speech stream, which includes: The initial audio signal is obtained from the original speech data stream, and the signal is initially cleaned using a pre-established filtering technique to obtain the denoised basic speech data. For the denoised basic speech data, a speech recognition algorithm is applied to detect and correct speech rate fluctuations, and to determine the speech segments after speech rate adjustment. Based on the speech segments with adjusted speech rate, the voice features of multiple characters are obtained, and the speech signals of different characters are separated by hierarchical analysis technology to obtain the preliminary separated character speech streams; For the initially separated character speech streams, a time series comparison method is used to determine the continuity of the speech flow. If an interruption or overlap is detected, the segments are spliced together to obtain continuous character speech data. Based on continuous character voice data, the independent voice features of each character are obtained, and clustering methods are applied to group the features to determine the final independent voice stream; For the final independent speech stream, structured speech flow data is generated through format standardization processing, resulting in a separated multi-role speech stream.
4. The method for automatically recording psychological counseling based on speech recognition according to claim 1, characterized in that, If the separated multi-role speech streams show confusion regarding role identities, then a speaker clustering method is used to compare voiceprint features to obtain the identity-labeled speech streams, including: Based on the separated multi-role speech stream, a pre-established voiceprint feature extraction model is used to obtain the voiceprint feature data of each speech segment; By extracting voiceprint feature data, the speaker clustering method is used to group different speech segments and determine the preliminary role identity classification results. If there are overlapping roles or unclear groupings in the preliminary classification results, a second comparison is performed on the voiceprint features of the overlapping segments to determine whether they belong to the same role, thus obtaining a more accurate classification and grouping. Based on the classification and grouping after the second comparison, corresponding identity labeling information is generated for each role, and a set of voice segments with identity labels is obtained; By integrating the annotated audio segments over time, a continuous stream of identity-annotated audio is generated, and the final role-identity correspondence is determined. If there are still some segments with inconsistent identity labels in the integrated speech stream, logical correction is performed using the identity information of the context speech segments to obtain the final corrected identity-labeled speech stream.
5. The method for automatically recording psychological counseling based on speech recognition according to claim 1, characterized in that, Semantic content and tone features are extracted from the identity-labeled speech stream to obtain emotional state labels, including: The semantic information of the identity-tagged speech stream is extracted using signal processing techniques to form preliminary semantic text content; Based on the extracted semantic text content, a pre-established semantic analysis model is used to analyze the key expressions in the text and obtain a preliminary judgment result on sentiment tendency. For voice segments of the same identity, acoustic analysis tools are used to extract tone features, analyze changes in pitch and speech rate, and determine the intensity of emotion reflected by the tone. If the tone characteristics are consistent with the semantic sentiment tendency judgment, the two are combined to obtain the sentiment state category of the identity; if they are inconsistent, the tone characteristics are given priority and the sentiment tendency judgment is readjusted. Based on the determined emotional state category, generate corresponding emotional state classification labels.
6. The method for automatically recording psychological counseling based on speech recognition according to claim 1, characterized in that, Based on the emotional state tags and the fusion of dialogue context information, the real-time emotional dynamic change sequence is obtained, including: Based on the emotional state tags, combined with the time dimension and interaction process in the dialogue content, a pre-established emotional analysis model is used to analyze the changing trend of emotional state at different points in time and determine the continuity characteristics of emotional dynamics. Based on the continuous characteristics of the emotional dynamics, the interaction process information in the context and dialogue content is integrated, and the trend of emotional state change is matched with real-time monitoring data through a logical comparison method to determine the significance of emotional changes. If the significance of the emotional change exceeds a preset threshold, the change trend is associated with the label information to generate a corresponding dynamic change sequence fragment, thus obtaining real-time updated emotional sequence data.
7. The method for automatically recording psychological counseling based on speech recognition according to claim 1, characterized in that, If there is a deviation in the real-time emotional dynamic change sequence, the attribution judgment is corrected through the emotional analysis model, and the target emotional attribution record is determined to include: If there is a deviation in the real-time emotional dynamic change sequence, the model correction process is triggered, and the emotional analysis model is used to perform in-depth analysis of the data to determine the corrected emotional classification. Key features are extracted from the corrected sentiment classification, and data processing techniques are used to clean and organize the dynamically changing data to output structured sentiment attribution information. Based on the structured sentiment attribution information, analytical tools are used to track the fluctuations of the sentiment sequence, determine whether there is a persistent deviation, and obtain deviation tracking records. By using the deviation tracking record and combining it with the correction mechanism to dynamically adjust the parameters of the sentiment analysis model, the target sentiment attribution record is determined.
8. An automatic recording system for psychological counseling based on speech recognition, used to implement the method as described in any one of claims 1-7, characterized in that, include: The module includes a data acquisition module, an identity labeling module, an emotion sequence acquisition module, an emotion attribution recording module, and a counseling log generation module. The data acquisition module is used to acquire audio signals of psychological counseling dialogues through a recording capture device to obtain raw voice data streams; The identity labeling module is used to process noise and speech rate changes by applying a speech recognition algorithm to the original speech data stream to determine the separated multi-role speech stream; if the separated multi-role speech stream shows role identity confusion, then the speaker clustering method is used to compare voiceprint features to obtain the identity-labeled speech stream. The emotion sequence acquisition module is used to extract semantic content and tone features from the identity-labeled speech stream to obtain emotion state labels; Based on the emotional state tags and the fusion of dialogue context information, a real-time emotional dynamic change sequence is obtained; The emotion attribution recording module is used to determine whether there is a deviation in the real-time emotion dynamic change sequence. If there is a deviation in the real-time emotion dynamic change sequence, the attribution judgment is corrected by the emotion analysis model to determine the target emotion attribution record. The consultation log generation module is used to generate structured consultation logs based on the target's emotional affiliation record.