Psychological counseling speech transposition method and system based on consulting role constraint
By using non-voice emotion event preliminary detection and multi-dimensional role constraint determination, the problem of inaccurate discourse boundaries in psychological counseling scenarios has been solved, achieving accuracy and clinical usability of speech transcription in psychological counseling scenarios and meeting the needs of clinical analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU FIRST PEOPLES HOSPITAL
- Filing Date
- 2026-07-03
- Publication Date
- 2026-07-31
AI Technical Summary
Existing speech recognition technology cannot accurately identify speech boundaries in psychological counseling scenarios, resulting in inaccurate transcription results, incompatibility with complex speech phenomena, and affecting the accuracy and completeness of clinical analysis.
By conducting preliminary screening of non-verbal emotional events, determining multi-dimensional counselor role constraints, and reconstructing discourse unit boundaries, combined with role constraints in psychological counseling, we can achieve precise separation of counselors and clients and division of intention-level units, ensuring the accuracy and clinical usability of discourse unit boundaries.
It improves the accuracy and clinical usability of speech transcription in psychological counseling scenarios, ensures the integrity of output data and the precision of clinical analysis, and can accurately reflect "who said what, when, and in what way", meeting the needs of applications such as clinical supervision and risk screening.
Smart Images

Figure CN122493833A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech processing technology, specifically to a method and system for transposing psychological conversation speech based on counselor role constraints. Background Technology
[0002] Compared to regular meetings and customer service calls, psychological counseling has a stronger procedural, interactive, and clinically interpretable nature. While the speaking roles in these counseling sessions are typically relatively fixed, complex phenomena such as frequent rotations, interruptions, overlapping speech, prolonged silences, hesitation, self-correction, and emotional fluctuations can occur. Therefore, the follow-up analysis of psychological counseling sessions goes beyond simply obtaining a continuous text; it requires a structured sequence of discourse units that accurately reflects "who said what, when, and how," providing a reliable foundation for clinical supervision, risk screening, counseling review, automated summarization, audit trails, and intelligent auxiliary analysis.
[0003] In existing technologies, common solutions mainly rely on automatic speech recognition, speaker separation, and ordinary sentence segmentation to complete audio transcription. However, most of these solutions are designed for general dialogue scenarios and can usually only output coarse-grained text fragments or simple speaker tags. In particular, the segmentation boundaries formed by automatic speech recognition often serve the convenience of decoding and reading, but this is not the same as the discourse unit boundaries required for clinical analysis. This can easily lead to the incorrect splitting of complete expressions or the improper merging of multiple expressions with different intentions, thereby affecting the accuracy, completeness, and interpretability of subsequent analysis results.
[0004] The patent "Method, Apparatus, Electronic Device and Storage Medium for Audio Annotation" (publication number CN121811860A) determines the start time point of the dialogue timeline associated with the audio to be processed based on the start time information of a multi-person dialogue; starting from the start time point, it extracts dialogue fragments carrying the first start and end time information from the audio to be processed along the dialogue timeline; and it annotates the audio to be processed in units of dialogue fragments. Although this solution achieves speaker separation, text transcription and paragraph division in multi-role dialogue, it only solves the general problem of "pure text transcription without timestamps and speaker identifiers". It does not perform role priors, and is still prone to confusion when role switching is frequent or voiceprints are similar. In addition, the lack of discourse unit boundary reconstruction leads to the problem of complete expression being split and multiple intentions being merged.
[0005] The patent, "A Sales Voice Dialogue Speaker Segmentation and Labeling Method Based on Role Separation," publication number CN121034335B, introduces a "sales-customer" business role prior for sales scenarios (binding roles through sales feature words), solving the problem of "general speaker separation without business role mapping." It also handles complex speech phenomena such as silence and overlapping speech, improving the coherence of role labeling. However, the questioning, clarification, and restatement behaviors of psychological counselors are completely different from sales scripts and cannot be directly adapted. Furthermore, the speech segment boundaries are still based on Voice Activity Detection (VAD) and voiceprint clustering, which cannot meet the boundary requirements of "complete expressive intent" in clinical analysis. Summary of the Invention
[0006] This application addresses the issues of inaccurate speech boundaries and incompatibility with complex speech phenomena in general speech recognition technology when applied to the clinical context of psychological counseling, resulting in poor speech conversion quality. It proposes a speech transposition method and system for psychological counseling based on counselor role constraints. Through initial detection of non-vocal emotional events, multi-dimensional counselor role constraint determination, role label correction, and speech unit boundary reconstruction, this method solves the problem of low transcription accuracy caused by role confusion and arbitrary boundary segmentation in general speech recognition technology in psychological counseling scenarios. It achieves precise separation of counselor and client speech and intention-level unit segmentation, ensuring the clinical usability of the output data.
[0007] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows: In a first aspect, embodiments of this application provide a method for transposing psychological conversation dialogue based on counselor role constraints, the method comprising: The process involves: performing a preliminary non-verbal emotional event detection on the original audio of psychological counseling sessions to obtain candidate speech segments; determining the role affiliation of the candidate speech segments based on counselor role constraints to obtain initial role labels; modifying the initial role labels based on the counselor role constraints to obtain target role labels; reconstructing the discourse unit boundaries of the candidate speech segments identified by the target role labels to obtain time-ordered discourse unit sequences; and detecting and labeling the clinical process behavior data corresponding to each discourse unit sequence to obtain structured discourse data.
[0008] This solution addresses the issues of inaccurate speech boundaries and incompatibility with complex speech phenomena in psychological counseling scenarios by performing preliminary detection of non-vocal emotion events, role attribution determination and constraint correction, speech unit boundary reconstruction, and clinical process behavior data annotation. It achieves accurate speaker separation, reconstruction of speech unit boundaries that meet clinical analysis needs, and structured annotation of behavior data. This solution resolves the shortcomings of general speech transcription results, such as incomplete expression, incorrect intent merging, and loss of clinical interaction information, at the source. It significantly improves the ability of transcribed text to reflect "who said what, when, and in what way" for clinical analysis, thereby improving the clinical usability, effectiveness, and accuracy of the output data. This further meets the needs of clinical supervision, risk screening, and other applications for accurate restoration of complete expression intent and interaction details.
[0009] Optionally, the step of performing a preliminary non-voice emotional event detection on the original audio of the psychological interview to obtain candidate speech segments includes: performing basic preprocessing on the acquired original audio of the psychological interview, including audio noise reduction, echo suppression, and gain normalization; performing non-voice noise detection and filtering on the preprocessed audio information based on dual-channel VAD to obtain initial speech segments; performing privacy desensitization on the initial speech segments to obscure user privacy information to obtain valid speech segments with timestamps; extracting basic acoustic features based on the valid speech segments; comparing the basic acoustic features with typical acoustic features of non-voice emotional events based on feature threshold rules to identify non-voice emotional events; and performing deduplication and time correction on non-voice emotional events within the same time period to obtain candidate speech segments.
[0010] Optionally, the step of determining the role affiliation of the candidate speech segments based on counselor role constraints to obtain initial role labels includes: performing speech recognition on the candidate speech segments to generate text recognition results with timestamps; extracting the voiceprint features of the candidate speech segments, combining them with the text recognition results, and determining the initial role labels of the candidate speech segments based on counselor role constraints; wherein, the multi-dimensional counselor role constraints include general constraints, language behavior constraints, cross-conversation template constraints, and school of thought adaptation constraints; the general constraints are used to filter voiceprint clustering results that do not meet the conditions based on the basic attributes of the role; the language behavior constraints are used to calculate the language behavior matching degree between the current speech segment and the language pattern of the target counselor role based on the text recognition results; the cross-conversation template constraints are used to calculate the matching degree between the voiceprint features or language pattern of the current speech segment and the exclusive conversation template of the target counselor role; the school of thought adaptation constraints are used to adjust the calculation result of the language behavior matching degree according to the counseling school to which the target counselor role belongs.
[0011] Optionally, determining the initial role label of the candidate speech segment based on the consultant role constraint includes: normalizing the voiceprint features of the candidate speech segment and then performing initial voiceprint clustering using a clustering algorithm; filtering the initial voiceprint clustering results according to the general constraint to obtain a set of role clusters for the candidate speech segment; retrieving the corresponding exclusive voiceprint template according to the current consultant role ID; calculating the similarity between each role cluster and the exclusive voiceprint template; if the similarity is greater than or equal to a first similarity threshold, comparing the similarity of role clusters in the same set to determine the initial role label of the role cluster; if the similarity is less than a second similarity threshold, obtaining the average language behavior matching degree of each role cluster based on the language behavior constraint, and determining the initial role label based on the average language behavior matching degree; wherein, the first similarity threshold is greater than the second similarity threshold.
[0012] Optionally, the step of modifying the initial role labels based on the consultation role constraints to obtain target role labels includes: establishing a comprehensive role evaluation function based on the language behavior constraints, cross-conversation template constraints, and genre adaptation constraints; using the comprehensive role evaluation function to calculate the comprehensive score of the initial role labels of the candidate speech segments; and modifying the initial role labels based on the comprehensive score under the general constraints to obtain target role labels.
[0013] Optionally, the step of reconstructing discourse unit boundaries for the candidate speech segments identified by the target role label to obtain a discourse unit sequence based on time order includes: extracting the speech segment boundary and word-level semantic boundary of each candidate speech segment based on time order; performing invalid boundary filtering and initial reconstruction on the speech segment boundary and word-level semantic boundary to obtain candidate boundaries; determining the clinical unit type to which the candidate speech segment belongs based on the target role label matching a clinical discourse unit type library, and retrieving the corresponding boundary score weight to generate a weight vector for the candidate boundary; obtaining an initial boundary score based on the weight vector combined with the basic feature information of the candidate boundary, and if the initial boundary score is greater than or equal to a preset threshold, then the candidate boundary is used as a reserved boundary; and traversing all reserved boundaries in time order to generate a discourse unit sequence, where each discourse unit contains consecutive candidate speech segments.
[0014] Optionally, the step of extracting the speech segment boundaries and word-level semantic boundaries of each candidate speech segment based on time sequence includes: extracting two basic speech boundaries for each candidate speech segment; if the role labels of two adjacent segments are different, then marking the basic speech boundary at the last time point of the previous segment as the role switching boundary; based on the text recognition results of the candidate speech segments, extracting the sentence-end punctuation boundary and the ordinary word end boundary where there is no sentence-end punctuation and the time interval between the sentence-end punctuation and the next word exceeds a preset threshold; wherein, the basic speech boundaries and the role switching boundary are used as speech segment boundaries, and the sentence-end punctuation boundary and the ordinary word end boundary are used as word-level semantic boundaries.
[0015] Optionally, the invalid boundary filtering and initial reconstruction of speech segment boundaries and word-level semantic boundaries to obtain candidate boundaries includes: integrating all speech segment boundaries and word-level semantic boundaries and sorting them by time; calculating the time difference between two adjacent boundaries; if the time difference is less than or equal to a preset time difference threshold, it is determined to be a duplicate boundary, and the boundaries are merged based on the boundary priority; if the time difference is greater than the preset time difference threshold, two independent boundaries are retained; the merged or / and independent boundaries are used as temporary boundaries, and the temporary boundaries are subjected to invalid filtering for time overlap and exceeding the start and end time of the session to obtain candidate boundaries; wherein, ordinary word end boundaries are low priority, basic speech boundaries and sentence end punctuation boundaries are medium priority, and role switching boundaries are high priority.
[0016] Optionally, the step of detecting and labeling the clinical process behavior data corresponding to each discourse unit sequence to obtain structured discourse data includes: establishing a global time-energy mapping table and a unit-semantic mapping table; determining silence events based on the time interval between two temporally adjacent discourse units; if the time interval is greater than a threshold, writing the silence duration into the metadata of the discourse unit at the later time to obtain a list of silent discourse units; traversing the list of silent discourse units, identifying adjacent discourse units with time overlap based on the global time-energy mapping table, calculating the energy ratio of the two discourse units within the overlapping time period based on the overlap duration and the overlap duration ratio to obtain a list of overlapping events; based on the list of overlapping events, determining the speaking order according to the start time of two corresponding discourse units, determining the interruption and interrupted type based on the energy ratio of the discourse units and marking the interruption type of the relevant discourse units, determining the interruption nature based on keywords and sentence features in the text recognition results, and obtaining a list of interrupted discourse units; integrating the metadata of the list of silent discourse units and the list of interrupted discourse units to obtain structured discourse data.
[0017] Secondly, embodiments of this application provide a psychological conversation speech transposition system based on counselor role constraints, comprising: The audio processing module performs a preliminary non-vocal emotional event detection on the original audio of the psychological interview to obtain candidate speech segments; the role determination module determines the role affiliation of the candidate speech segments based on counselor role constraints to obtain initial role labels; and performs constraint correction on the initial role labels based on the counselor role constraints to obtain target role labels; the boundary reconstruction module reconstructs the discourse unit boundaries of the candidate speech segments identified by the target role labels to obtain a discourse unit sequence based on time order; and the behavior detection module detects and labels the clinical process behavior data corresponding to each discourse unit sequence to obtain structured discourse data.
[0018] The beneficial effects of this application are: 1. This application uses a discourse unit boundary reconstruction mechanism to reconstruct the boundaries of speech segments with clearly defined roles, guided by the complete expressive intent required for clinical analysis. This ensures that each discourse unit carries a relatively complete clinical action performed by a specific role, thus guaranteeing the clinical usability and effectiveness of the transposed speech text. 2. This application, by detecting and labeling clinical process behavioral data, enables more accurate calculation of the silence duration between units, identification of overlap between units, and determination of interruption relationships on the basis of a previously constructed complete speech unit sequence delivered by a clearly defined role at a precise time. This transforms a raw consultation audio full of complex interactions and emotional signals into a structured data that is not only accurate in text and clear in roles, but also retains complete information on the interaction process such as silence, overlap, and interruption. This overcomes the core problems of poor speech transcription quality and loss of clinical interaction information caused by fragmented processing in general technologies. 3. Initial screening of non-verbal emotional events allows for the separation and labeling of non-lexical but clinically informative emotional signals such as sighs and crying from valid speech at the source. This eliminates interference for subsequent processing, providing pure and information-free speech fragments and preventing the confusion caused by these complex phenomena directly entering the transcription and role determination stages. Based on this, utilizing prior knowledge of relatively fixed roles in psychological counseling, accurate role assignment is performed on pure candidate speech fragments. This process is no longer an isolated voiceprint matching task but is placed within a stable counseling interaction framework, ensuring a high degree of coherence and accuracy in identifying "who is speaking," thus laying an accurate foundation for role attribution of discourse units. Attached Figure Description
[0019] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings. The drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings.
[0020] Figure 1 Flowchart of the psychological conversation speech transposition method based on counselor role constraints provided in this application embodiment; Figure 2 This is a schematic diagram of the initial role tag acquisition process provided in an embodiment of this application; Figure 3 This is a schematic diagram of the clinical process behavior data detection workflow provided in the embodiments of this application; Figure 4 A schematic diagram of a psychological conversation speech transposition system module based on counselor role constraints provided in this application embodiment. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description of this application is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely one preferred embodiment of this application and are only used to explain this application. They do not limit the scope of protection of this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0022] Some of the terms or terms that appear in the description of the embodiments of this application shall be interpreted as follows: Discourse unit: refers to the smallest analytical unit that is completed by the same person within a relatively continuous period of time and has a relatively complete clinical expression intention; Consultant role constraints: These refer to the a priori restrictions imposed on the role categories, behavioral patterns, and stability within a psychological counseling setting. VAD stands for Voice Activity Detection, whose core task is to distinguish speech segments from non-speech segments (such as silence, background noise, etc.) in a continuous audio stream. Clinical process behavior data: refers to additional information beyond the text content used to describe the characteristics of the interview process, such as interruptions, silences, overlaps, speech rate, and volume. Boundary reconstruction: refers to the process of merging, splitting, or adjusting the boundaries of candidate speech segments generated by automatic speech recognition to form speech unit boundaries that meet the needs of clinical analysis; Clinical unit type: refers to the clinical functions and expressive attributes of a single discourse unit, used to dynamically adjust the scoring weights for discourse boundary reconstruction. Among them, the counselor's general functions include questioning, clarification, restatement, summarization, and intervention, while the client's general functions include narration, emotional expression, resistance, and self-exploration.
[0023] Example 1: As Figure 1As shown, a method for transposing psychological conversation discourse based on counselor role constraints includes steps S1-S5, wherein: S1. Perform a preliminary non-voice emotional event detection on the original audio of the psychological interview to obtain candidate voice segments.
[0024] In an optional embodiment, step S1 includes: Basic preprocessing was performed on the collected raw audio of psychological interviews, including audio noise reduction, echo suppression and gain normalization. Based on the dual-channel VAD, non-speech noise is detected and filtered from the pre-processed audio information to obtain the initial speech segment; The initial voice segment is desensitized to obscure user privacy information, and a valid voice segment with a timestamp is obtained. Based on the effective speech segments, basic acoustic features are extracted. Based on the feature threshold rule, the basic acoustic features are compared with the typical acoustic features of non-speech emotional events to identify non-speech emotional events. Then, non-speech emotional events within the same time period are deduplicated and time-corrected to obtain candidate speech segments.
[0025] Specifically, the basic acoustic features include frame-level energy, frame-level fundamental frequency, speech activity markers (including start frame time and end frame time), and the original PCM segment. Frame-level energy is calculated based on the root mean square energy of audio frames within a time period and is used to detect the energy ratio of overlapping segments and prosodic abrupt changes to help identify the interruptor. Frame-level fundamental frequency is the fundamental frequency trajectory of the original audio and is used to detect changes in speech rate and identify emotional pauses (distinguishing between clinical silence and technical pauses) to assist in subsequent boundary reconstruction. Speech activity markers refer to the VAD tagging based on whether each frame of the VAD output is speech, used to detect silence events. The original PCM segment is used for subsequent extraction of voiceprint features.
[0026] Specifically, during non-verbal emotion event detection, statistical features are obtained based on the fundamental acoustic features of each valid speech segment: average energy (arithmetic mean of the energy of all frames within the valid speech segment), energy variance (variance of the energy of all frames), average fundamental frequency (average of the remaining valid fundamental frequency frames excluding those with a fundamental frequency of 0 within the valid speech segment), fundamental frequency variance (variance of all valid fundamental frequency frames), and number of energy peaks (number of peak points where the energy exceeds the average energy fluctuation range). Furthermore, based on the typical acoustic features of various non-verbal emotion events in psychological interviews, threshold comparison rules are used to identify non-verbal emotion events. Types of non-verbal emotion events include, but are not limited to, sighing events, sobbing events, and crying events. For example, for a sighing event, the speech duration is 0.2-0.5 seconds, the average energy is less than 50% of the average energy of the same speech segment in the same conversation, the energy peak is a single peak appearing in the first half of the segment, and the average fundamental frequency is less than 30% of the average fundamental frequency of the same speech segment in the same conversation. These statistical characteristics of the effective speech segment are compared one by one with the typical characteristics of the event. When all conditions are met, the confidence level for emotional event detection is high; if one condition is missing, the confidence level is medium; if two or more conditions are missing, the event is not labeled. It should be understood that the specific values of the typical characteristics of the above events can be set according to the actual clinical presentation of the client, as long as non-vocal emotional events within the effective speech segment can be accurately detected.
[0027] Furthermore, for results marked as multiple events within the same time period, the event type with the highest confidence is retained to mark non-speech events in the valid speech segments; adjacent events of the same type with an interval of a time threshold (such as 0.1 seconds) are merged into a continuous event to correct the occurrence and end times of the event, and the marking position of the frame number corresponding to the non-speech emotion event is adjusted according to the corrected time.
[0028] In this embodiment, the dual-channel VAD (Voice-Aware Detection) performs a binary classification detection of "speech / non-speech," accurately distinguishing only valid human speech from all non-speech signals, thus reducing the false negative rate of speech. For example, all non-speech signals such as breathing sounds, keyboard sounds, and background noise are uniformly marked as "invalid," and the initial speech segments obtained are only speech fragments. However, VAD cannot capture the fine-grained acoustic features of emotional events, such as the fundamental frequency fluctuation pattern of crying and the spectral attenuation features of sighing. Expanding the classification categories of the VAD channels would lead to a surge in model parameters, increased inference latency, and a significant decrease in the original "speech / non-speech" binary classification accuracy. Therefore, non-speech emotional event detection is performed after VAD to further reduce the false positive rate of similar events, such as distinguishing between crying and sobbing. This avoids directly performing detection on all audio fragments in the subsequent clinical process behavior detection stage, reducing computational overhead and improving the accuracy of clinical process behavior detection.
[0029] S2. Based on the consultant role constraint, determine the role affiliation of the candidate voice segments to obtain initial role labels.
[0030] In an optional embodiment, combined with Figure 2 As shown, step S2 includes: The candidate speech segments are subjected to speech recognition to generate text recognition results with timestamps; Extract the voiceprint features of the candidate speech segments, combine them with the text recognition results, and determine the initial role label of the candidate speech segments according to the consultation role constraint; The multi-dimensional consulting role constraints include general constraints, language behavior constraints, cross-conversation template constraints, and genre adaptation constraints. The general constraints are used to filter voiceprint clustering results that do not meet the conditions based on the basic attributes of the role; The language behavior constraint is used to calculate the language behavior matching degree between the current speech segment and the language pattern of the target counselor role based on the text recognition results; The cross-conversation template constraint is used to calculate the degree of matching between the voiceprint features or language patterns of the current speech segment and the exclusive conversation template of the target consultation role. The school of thought adaptation constraint is used to adjust the calculation results of the language behavior matching degree according to the school of thought to which the target counselor belongs.
[0031] Specifically, speech recognition of candidate speech segments can be performed using ASR (Automatic Speech Recognition), a technology that automatically converts human speech signals into editable text.
[0032] Specifically, voiceprint features (hereinafter referred to as voiceprints) are extracted using a voiceprint model based on PCM segments in candidate speech fragments. Voiceprints are then clustered to distinguish the roles of the counselor and the client, avoiding errors in speech boundary reconstruction (i.e., merging or splitting) caused by role confusion, and preventing the merging of speeches from different roles. There are three role types: counselor, client, and a third role representing the companion / observer / used to mark background sounds.
[0033] It should be noted that the constraints on each counseling role are essentially rules for processing the voiceprint and text semantic features of candidate speech segments. The role attribution of candidate speech segments is determined through the execution of these constraint rules. Specifically, the general constraint rules define inviolable hard constraint judgment logic by fixing the core set of psychological counseling roles. Basic role attributes include the number of roles, global role uniqueness, minimum turn duration, and role switching frequency. A hard constraint verification interface is generated based on the constraints corresponding to the basic role attributes to filter the voiceprint clustering results. Specifically, the role number constraint means that in a two-person counseling session, the number of role clusters (i.e., speaker clusters) can only be 2. If the clustering result is 1 or greater than or equal to 3, it is considered abnormal. The global role uniqueness constraint means that in the same session, the counselor role can only correspond to 1 voiceprint, and the client role can only correspond to 1 voiceprint.
[0034] Language behavior constraint rules: Based on historical language pattern-annotated corpora, specific language behavior features of counselors / clients are extracted to train a clinical language behavior binary classifier. Language behavior features are extracted from the text recognition results of real-time candidate speech segments, and the cosine similarity with these specific language behavior features is calculated to output the language behavior matching degree (i.e., the probability that the current candidate speech segment is likely to be a counselor's role). Language behavior features include, but are not limited to, typical sentence structure features (such as "What do you think?", "Can you tell me?"), keywords, and question frequency. The clinical language behavior binary classifier adopts an industry-standard single-output binary classifier, outputting only two categories of results: counselor tendency / client tendency. It can be fine-tuned based on the BERT-base language model. The last layer of the model uses a sigmoid activation function to output a 0-1 probability value (language behavior matching degree) for a single neuron. This value directly represents the probability that the sample belongs to the positive class (counselor). The probability of the negative class (client) is 1 minus the positive class probability. In this embodiment, only the positive class probability is considered. For example, an output of 0.9 indicates that there is a 90% probability that the segment was spoken by the counselor, resulting in a language behavior matching score of 0.9. An output of 0.5 indicates that it is indistinguishable, with a language behavior matching score of 0.5 (neutral). It is understandable that the language behavior matching score is only used for the comprehensive evaluation of initial role labels and is not used to separately determine role labels, thus avoiding distortion in role identification due to semantic pattern boundary interference.
[0035] Cross-conversation template constraints: Output conversation template matching score. Specifically, this involves: extracting voiceprints from speech segments based on the consultant's historical consultation data, calculating the mean vector of all voiceprints, and using this as the consultant's unique voiceprint template; simultaneously calculating the average turn length, questioning frequency, summary frequency, and proportion of commonly used phrases from the historical consultation data to generate a unique language behavior template; defining cross-conversation template matching score items, including: voiceprint template matching score, language behavior template matching score, and total cross-conversation template matching score (conversation template matching score). Where: Voiceprint template matching degree = 1 - cosine distance between the current segment voiceprint and the exclusive voiceprint template of the corresponding target consultant role; Language behavior template matching degree = 1 - cosine distance between the language behavior characteristics of the current segment and the exclusive language behavior template of the corresponding target counselor role; Total cross-session template matching score = Voiceprint template matching score w1+ Language Behavior Template Matching w2, w1, and w2 are the weight coefficients for each template matching degree, set based on the actual acquisition of voiceprint and language behavior features, with w1+w2=1.
[0036] School of thought adaptation constraint: This involves statistically analyzing the differences in language behavior characteristics among counselors from different schools of thought, generating a school of thought adaptation parameter table. The table lists a one-to-one correspondence between school type, sentence structure features, and school of thought adaptation weight coefficients. Sentence structure features are extracted based on the text recognition results of the current segment, including but not limited to interrogative, declarative, and exclamatory sentences. Corresponding weight coefficients are then matched to these sentence structure features to adjust the language behavior matching degree, thus obtaining the degree to which the language behavior matches the current counseling school's characteristics. The weights only apply to the language behavior matching degree sub-item, improving the accuracy of role recognition. Here, "school of thought" refers to the entire set of therapeutic theories, technical systems, and interaction styles followed by the counselor. The final output of the school of thought adaptation constraint is the school of thought adaptation score, where the school of thought score equals the original language behavior matching degree. Genre adaptation weight coefficient.
[0037] In some examples, the rules for adjusting the weight of school of thought include: CBT school: high frequency of questioning, high frequency of summarizing, and high proportion of directive language, increasing the weight of question sentence features; Humanistic school: high proportion of empathic language, low frequency of questioning, and high proportion of listening time, decreasing the weight of summary sentence features; Psychoanalytic school: high proportion of explanatory language, high tolerance for silence, and mostly open-ended questions, increasing the weight of explanatory language matching.
[0038] Furthermore, comprehensive evaluation refers to the weighted summation of the confidence levels of multiple independent pieces of evidence for the proposition that "the fragment belongs to the counselor," which is based on the fundamental idea that a more reliable posterior probability can be obtained by weighted combination of evidence from multiple independent sources.
[0039] In this embodiment, by extracting typical language behavior features of counselors and clients and calculating language behavior matching degree, role tendencies are distinguished from a semantic level, making up for the judgment defects of pure voiceprint clustering when voiceprints are similar; by matching counselors' historical voiceprints with language behavior exclusive templates, counselor identity is prioritized, greatly improving the accuracy and global consistency of initial role mapping in new sessions; by dynamically adjusting the corresponding scoring weights according to the differences in language behavior features of different counseling schools, the role behavior characteristics of schools such as CBT and humanistic counseling are adapted, further improving the accuracy of role determination.
[0040] In an optional embodiment, determining the initial role label of the candidate speech segment based on the consultant role constraint includes: After normalizing the voiceprint features of the candidate speech segments, a clustering algorithm is used to perform initial voiceprint clustering. The initial voiceprint clustering results are then filtered according to the general constraints to obtain the role cluster set of the candidate speech segments. Based on the current consultant role ID, retrieve the corresponding exclusive voiceprint template, calculate the similarity between each role cluster and the exclusive voiceprint template, and if the similarity is greater than or equal to the first similarity threshold, compare the similarity of role clusters in the same set to determine the initial role label of the role cluster. If the similarity is less than the second similarity threshold, the average language behavior matching degree of each role cluster is obtained based on the language behavior constraints, and the initial role label is determined based on the average language behavior matching degree. Among them, the first similarity threshold is greater than the second similarity threshold.
[0041] In this embodiment, the general constraints, by forcibly limiting the number of roles, global uniqueness, minimum turn duration, and role switching frequency, eliminate obviously erroneous clustering and role mapping results, thus laying a solid foundation for the underlying logical boundaries of subsequent soft constraint optimization and ensuring the basic stability and rationality of role attribution determination.
[0042] In some embodiments, a unique voiceprint template is obtained based on the current counselor role ID. The cosine similarity between the center vector of each role cluster and the counselor's voiceprint template is calculated for initial role label mapping. The initial role mapping rules are as follows: if the similarity between role cluster a and its unique voiceprint template is ≥ a first similarity threshold (e.g., 0.7) and higher than cluster b (e.g., 0.2 or more higher), then cluster a is the counselor and cluster b is the client; if the similarity between cluster b and the counselor template is ≥ 0.7 and higher than cluster a (e.g., 0.2 or more higher), then cluster b is the counselor and cluster a is the client; if the difference in similarity is < a second similarity threshold (e.g., 0.1), then the language behavior-assisted mapping process begins. Language behavior-assisted mapping includes: calculating the average language behavior matching degree for all segments of each role cluster; the cluster with the higher average matching degree is labeled as the counselor, and the other cluster as the client. It is understood that the threshold values given above are examples, and the actual threshold settings can be adjusted based on the voiceprint extraction results.
[0043] In this embodiment, when the voiceprint matching degree is high, the role is quickly locked using a dedicated template. When the voiceprint matching degree is low, the system automatically downgrades and relies on the language behavior matching degree for a fallback judgment. This dynamic adaptation mechanism avoids the risk of misjudgment caused by a single threshold setting and ensures that a stable and accurate initial role label can be obtained under different signal-to-noise ratio environments, whether the voiceprint is clear or interfered with.
[0044] S3. Based on the consulting role constraints, modify the initial role labels to obtain the target role labels.
[0045] In an optional embodiment, step S3 includes: A comprehensive role evaluation function is established based on the aforementioned language behavior constraints, cross-conversation template constraints, and genre adaptation constraints. The overall score of the initial role label for the candidate voice segment is calculated using the aforementioned role comprehensive evaluation function; The initial role labels are corrected based on the comprehensive score under the general constraints to obtain the target role labels.
[0046] Specifically, after extracting and normalizing the corresponding dimensions of the execution process of language behavior constraints, cross-conversation template constraints, and genre adaptation constraints, a single-constraint score is obtained. These single-constraint scores are then weighted and summed to obtain a comprehensive score for role attribution (i.e., the confidence level that the target role label of a candidate speech segment belongs to the consultant). The single-constraint score includes voiceprint score, language behavior matching degree, conversation template matching degree, adjacent role consistency score, and genre score. The voiceprint score is obtained through voiceprint clustering similarity: voiceprint score = 1 - (similarity between the role cluster and the exclusive voiceprint template / 2). The adjacent role consistency score refers to the consistency of candidate speech segments arranged in chronological order. If two adjacent segments have the same initial role label, the adjacent role consistency score is 1; if the two adjacent segments have different initial role labels, the score is 0. It is understood that the weight of each single constraint scoring item in this embodiment is based on industry practice and academic research in the field of speaker separation, and is set and fine-tuned in combination with the special characteristics of the fixed dual roles of "counselor-client" in the psychological counseling scenario. It belongs to the industry's general benchmark weight, and will not be explained here.
[0047] In some possible embodiments, the initial role labels are modified based on the general constraints of the comprehensive score, including: For each candidate audio segment, calculate the combined score of its assigned counselor and client. If the score of the assigned counselor is higher than that of the assigned client, retain the initial role label. If the score of the assigned client is higher than that of the assigned counselor, correct the initial role label. If the difference between the two scores is less than a threshold, mark it as a low-confidence segment, retain the initial label, and add a label to be corrected. All corrected results must pass the general hard constraint rule; otherwise, restore the initial label.
[0048] In other possible embodiments, the initial role labels are modified based on the general constraints of the comprehensive score, including: Arrange all candidate audio segments with initial role labels in chronological order. Use a sliding window to iterate through the segments and correct the roles of the intermediate audio segments based on the overall score. The correction conditions include: the duration of the intermediate audio segment is less than the preset duration, the overall score of the intermediate audio segment is less than the preset score, and the initial role labels of the preceding and following audio segments are the same but different from those of the intermediate audio segment. When all conditions are met, correct the initial role label of the intermediate audio segment to be consistent with those of the preceding and following audio segments.
[0049] In this embodiment, by fusing multi-dimensional quantitative scoring based on voiceprint clustering similarity, language behavior matching, cross-conversation template matching, and genre adaptation correction, the limitations of general speaker separation relying solely on a single voiceprint feature are overcome. This effectively avoids role mapping errors in scenarios such as similar voiceprints, environmental noise interference, and voiceprint fatigue during long conversations. Simultaneously, utilizing role behavior priors specific to psychological counseling, the undifferentiated "speaker A / B" labels output by general technologies are accurately mapped to clinically meaningful counselor / client labels, improving the accuracy and clinical relevance of role recognition. Furthermore, the sliding window smoothing correction leverages the inherent characteristics of relatively orderly turn transitions and a high proportion of continuous speaking in psychological counseling. The sliding window mechanism corrects isolated low-confidence erroneous segments, eliminating common problems in general technologies such as frequent role jumps and mislabeling of ultra-short segments. This ensures the continuity and consistency of the role sequence, providing a reliable role foundation for subsequent core steps such as clinical semantic-driven discourse boundary reconstruction and interruption and overlap event detection.
[0050] S4. Reconstruct the speech unit boundaries of the candidate speech segments identified by the target role label to obtain a speech unit sequence based on time order.
[0051] In an optional embodiment, step S4 includes: S41. Extract the speech segment boundaries and word-level semantic boundaries of each candidate speech segment based on time sequence; S42. Perform invalid boundary filtering and initial reconstruction on speech segment boundaries and word-level semantic boundaries to obtain candidate boundaries; S43. Based on the target role label, match the clinical discourse unit type library to determine the clinical unit type to which the candidate speech segment belongs, and retrieve the corresponding boundary score weights to generate the weight vector of the candidate boundary; S44. Based on the weight vector and the basic feature information of the candidate boundary, obtain the initial boundary score. If the initial boundary score is greater than or equal to a preset threshold, then the candidate boundary is used as the reserved boundary. S45. Traverse all reserved boundaries in chronological order to generate a sequence of speech units, each speech unit containing consecutive candidate speech segments.
[0052] Specifically, each candidate speech segment is input into a pre-trained clinical discourse unit classifier to obtain its corresponding clinical unit type. Counselor unit types include questioning, clarification, restatement, summarization, and intervention units; client unit types include narrative, emotional expression, resistance, and self-exploration units. Based on the current segment's clinical unit type, the corresponding boundary scoring dimension weights are retrieved from a pre-defined weight configuration table. In the weight configuration table, each clinical unit type corresponds to a weight coefficient for a scoring dimension, which includes, but is not limited to, pauses, role switching, syntactic completeness, semantic breaks, prosodic breaks, and semantic continuity.
[0053] Specifically, the basic feature information of candidate boundaries includes at least temporal features, role label features, semantic features, syntactic features, and prosodic features. Temporal features are obtained based on the pause duration of the segments before and after the boundary; role label features are obtained based on whether the role labels of the segments before and after the boundary are consistent (i.e., whether a role switch has occurred). If they are inconsistent, it indicates a role switch and is recorded as 1; otherwise, it is recorded as 0; semantic features are obtained by calculating the semantic similarity of the text before and after the boundary to determine whether there is a semantic break; syntactic features are obtained by judging the syntactic integrity of the preceding segment text. The syntactic integrity judgment process is: based on dependency parsing, whether there is a complete subject-verb-object structure and whether it ends with a sentence punctuation mark; prosodic features are obtained based on the average speech rate of the segments before and after the boundary, and the energy difference or fundamental frequency difference of the time period before and after the boundary. Prosodic features are different from voiceprints. Prosodic features are the core basis for distinguishing clinical discourse boundaries from general ASR decoding boundaries. In psychological counseling, the client's emotional pauses (such as the pause after crying) are not the end of expression, but general ASR will misjudge them as boundaries.
[0054] Furthermore, for each candidate boundary, the scores for each dimension are calculated based on the extracted basic feature information and the corresponding matched weights, and the scores for each dimension are added together to obtain the initial boundary score.
[0055] For example, the scores for each dimension include: pause score = min(pause duration / 2, 1), where 1 is taken when the pause exceeds the threshold; role switching score = 1 (different roles before and after), role switching score = 0 (same roles before and after); syntactic integrity score = 1 (syntactic integrity of the previous segment), syntactic integrity score = 0.5 (partially complete syntactic structure), syntactic integrity score = 0 (incomplete syntactic structure); semantic fragmentation score = 1 - semantic similarity, where the lower the similarity, the higher the degree of fragmentation; prosodic fragmentation score = min((speech rate difference + energy difference + fundamental frequency difference) / 3, 1); semantic continuity score = semantic similarity (the higher the similarity, the higher the degree of continuity).
[0056] In this embodiment, an initial boundary score is obtained based on the basic feature information of the candidate boundary to determine the retained boundary, providing an operable initial sequence for subsequent clinically guided fine processing and avoiding the problem that the whole text without boundaries cannot be analyzed in subsequent steps.
[0057] In some possible embodiments, after generating the speech unit sequence, a clinical turn-taking pattern library is invoked to perform pattern matching on the initial sequence. This library predefines various core conversation turn-taking patterns, and the matching rules are as follows: using multiple consecutive speech units as windows (e.g., 3-5), the role labels and clinical type combinations of the units are matched. For example, when matching the "Counselor-Question Unit → Client-Answer Unit" pattern, if a "Boundary to be Merged" exists within the client's answer unit, the boundary is forcibly merged to ensure the integrity of the answer. When matching the "Client-Narrative Unit → Counselor-Feedback Unit" pattern, a boundary is forcibly retained before the counselor's feedback unit, even if the boundary score is below a threshold. When matching the "Interruption Pattern (Previous Unit Not Finished + Next Unit Begins)," a boundary is forcibly inserted at the interruption point and marked as an "Interruption Boundary." For turn breaks caused by overlapping speech, if two consecutive units have the same role and a semantic similarity greater than a preset threshold, they are merged into one unit, and an overlap / merge marker is added.
[0058] Furthermore, the modified discourse unit sequence is traversed to filter out ultra-long units whose duration exceeds the threshold time (e.g., 45 seconds). Clinical intent clustering is performed on these ultra-long units. The clustering steps include: segmenting the text of the ultra-long unit into sentences and extracting the semantic embedding of each sentence; using a clustering algorithm to cluster the sentence embeddings to generate different clinical intent clusters; each clinical intent cluster corresponds to a potential segmentation point located at the boundary of sentences between clusters; for each potential segmentation point, it is detected whether there are significant abrupt changes in speech rate, energy, and fundamental frequency before and after it. If the change amplitude exceeds the threshold, it is determined to be a significant abrupt change; if there is a prosodic change, the point is confirmed as a valid segmentation point; the ultra-long unit is divided into multiple sub-units according to the valid segmentation points, and the same homologous unit ID is added to all sub-units to record that they come from the same original ultra-long unit; the boundary score is recalculated for each sub-unit based on the initial boundary calculation method to determine whether to retain it.
[0059] The final discourse unit sequence is generated based on ultra-long unit segmentation processing. All discourse units are arranged in chronological order. Each unit contains: unit ID, role label, start time, end time, text content, boundary confidence, clinical unit type, merging / segmentation marker, and homologous unit ID.
[0060] In this embodiment, by using boundary correction based on clinical turn-taking patterns and segmenting ultra-long discourse units based on clinical intent, the boundaries can be corrected based on the initial sequence and the fixed turn-taking interaction patterns between the counselor and the client. This ensures that complete turns such as questioning and answering, feedback and response are not incorrectly split. At the same time, ultra-long discourses are segmented from the perspective of clinical expression intent, avoiding multiple clinical intents mixed in a single unit. This effectively solves the defects of rigid boundaries, fragmented semantics, inapplicability of ultra-long expressions, and inaccurate boundaries of complex interactions in the initial sequence. As a result, the final discourse units are more in line with the actual needs of clinical analysis in psychological counseling, significantly improving boundary accuracy, semantic integrity and clinical usability.
[0061] In an optional embodiment, step S41 includes: Extract two basic speech boundaries for each candidate speech segment. If the role labels of two adjacent segments are different, mark the basic speech boundary of the last time point of the previous segment as the role switching boundary. Based on the text recognition results of the candidate speech segments, extract the sentence-ending punctuation boundaries and the ordinary word ending boundaries that have no sentence-ending punctuation and whose time interval with the next word exceeds a preset threshold. Specifically, the boundaries of basic speech and role switching are used as the boundaries of speech segments, and the boundaries of sentence-end punctuation and ordinary word endings are used as the boundaries of word-level semantics.
[0062] In this embodiment, boundaries refer to time points. Speech segment boundaries are coarse-grained basic boundaries representing audio activity based on candidate speech segments; word-level semantic boundaries are fine-grained internal boundary points of segments based on the text recognition results of the audio, representing semantic and syntactic levels. If the text recognition result corresponding to the current candidate speech segment contains a period, question mark, exclamation mark, or semicolon at the end of the word, the time point corresponding to the punctuation mark is extracted and marked as the sentence-end punctuation boundary; if the current word does not have a sentence-end punctuation mark and the time interval between it and the next word is greater than a preset threshold (e.g., 0.1 seconds), the time point corresponding to the current word is marked as the ordinary word end boundary.
[0063] In an optional embodiment, step S42 includes: Integrate all speech segment boundaries and word-level semantic boundaries and sort them by time. Calculate the time difference between two adjacent boundaries. If the time difference is less than or equal to a preset time difference threshold, it is determined to be a duplicate boundary, and the boundaries are merged based on the boundary priority. If the time difference is greater than the preset time difference threshold, two independent boundaries are retained. Merged or / and independent boundaries are used as temporary boundaries. These temporary boundaries are then filtered for time overlap and invalidity exceeding the session start and end time to obtain candidate boundaries. Among them, the boundary of ordinary word ending is low priority, the boundary of basic speech and the boundary of sentence end punctuation are medium priority, and the boundary of role switching is high priority.
[0064] In some possible embodiments, invalid filtering specifically includes time overlap filtering, extremely short interval boundary filtering, and out-of-session boundary filtering. Time overlap filtering involves filtering a boundary if its timestamp is earlier than the previous boundary. Extremely short interval boundary filtering involves calculating the time interval between the current boundary point and the previous boundary point; if the interval is less than a threshold (e.g., 0.1 seconds), the current boundary is filtered as an extremely short interval boundary. Out-of-session boundary filtering involves filtering a boundary if its timestamp is earlier than the session start time or later than the session end time. The final candidate boundaries include the boundary ID and the IDs of the nearest candidate speech segments before and after the boundary point.
[0065] In this embodiment, candidate speech segment boundaries cover coarse-grained boundaries of speech activities and role switching, while word-level semantic boundaries cover fine-grained semantic and syntactic boundaries. Combining the two ensures no potential boundaries are missed. Role switching boundaries have the highest priority, ensuring that turn-taking transitions inevitably become candidate boundaries. Sentence-ending punctuation boundaries have the next highest priority, conforming to the logic of natural language expression. Ordinary word spacing boundaries have the lowest priority and are only used as a supplement. Temporal overlap detection is performed again on merged or / and independent temporary boundaries to avoid the same physical boundary being split into multiple candidate boundaries. The merging threshold setting needs to be adapted to the timestamp errors of the speech recognition task and the VAD task.
[0066] Furthermore, by employing a multi-level priority mechanism, the system can prioritize retaining role transition points that reflect the essence of dialogue interaction when dealing with overlapping speech or rapid switching. At the same time, it uses sentence-end punctuation to assist in verifying semantic integrity. This priority-based boundary filtering strategy effectively filters out redundant boundaries caused by breathing pauses or filler words, thereby improving the signal-to-noise ratio of the candidate boundary set.
[0067] S5. Detect and label the clinical process behavior data corresponding to each discourse unit sequence to obtain structured discourse data.
[0068] In an optional embodiment, combined with Figure 3 As shown, step S5 includes: Establish a global time-energy mapping table and a unit-semantic mapping table; The silence event is determined based on the time interval between two temporally adjacent speech units. If the time interval is greater than a threshold, the silence duration is written into the metadata of the speech unit at the next time step to obtain a list of silent speech units. Traverse the list of silent speech units, identify adjacent speech units with time overlap based on the global time-energy mapping table, calculate the energy ratio of the two speech units within the overlapping time period based on the overlap duration and the overlap duration ratio, and obtain the list of overlapping events; Based on the list of overlapping events, the speaking order is determined according to the start time of the two corresponding discourse units. The interruption and interrupted types are determined according to the energy ratio of the discourse units, and the interruption type is marked for the relevant discourse units. The interruption nature is determined based on the keywords and sentence features in the text recognition results, and a list of interrupted discourse units is obtained. The metadata of the silent discourse unit list and the interrupted discourse unit list is integrated to obtain structured discourse data.
[0069] Specifically, establishing the global time-energy mapping table and the unit-semantic mapping table includes: The frame-level energy of the original audio is converted into energy values corresponding to the same unit timestamp, forming a global time-energy mapping table; Based on the timestamps of the text recognition results, word-level texts are associated with their respective discourse units to obtain a unit-lexical meaning mapping table. It should be noted that the essence of boundary reconstruction is to merge multiple consecutive word-level text recognition results into a clinically meaningful complete unit. It only retains the overall relationship of these words forming this unit, but does not retain the precise timestamp of each word within the unit. However, clinical process behavior detection (such as interruption) requires fine-grained word, time, and unit mapping relationships. Therefore, it is necessary to associate word-level texts with their respective discourse units and construct a unit-lexical meaning mapping table.
[0070] In some embodiments, the step of detecting overlapping events specifically includes: The list of silent speech units is sorted in ascending order of start time. A double loop iterates through all pairs of silent speech units in the list. The outer loop iterates through speech unit A, which started earlier, and the inner loop iterates through speech unit B, which started later than speech unit A. If the start time of speech unit B is earlier than the end time of speech unit A, then the two units are considered to have time overlap. For each pair of speech units A and B with time overlap, the following calculation is performed: The overlap start time is the larger of the start time of discourse unit A and the start time of discourse unit B. The overlap end time is the smaller value between the end time of discourse unit A and the end time of discourse unit B; Overlap duration: Overlap end time minus overlap start time; The overlap duration percentage (reflecting the degree of impact of overlap on the shorter unit) is calculated by dividing the overlap duration by the total duration of the shorter of the two discourse units.
[0071] Furthermore, from the global time-energy mapping table, all frame-level energy data from the overlap start time to the overlap end time are extracted, and the total energy of discourse unit A and discourse unit B during this overlap time period is calculated respectively: the energy values of all frames within the corresponding time period are summed. Then, the energy ratio of each party is calculated: the energy ratio of discourse unit A = the total energy of the overlap segment of discourse unit A ÷ (the total energy of the overlap segment of discourse unit A + the total energy of the overlap segment of discourse unit B), and the energy ratio of discourse unit B = 1 - the energy ratio of discourse unit A.
[0072] Furthermore, if the overlap duration is less than a preset threshold, the overlap duration is not marked as an overlap event, and a list of overlap events is finally generated. The list includes the overlap identifier, overlap duration, overlap duration percentage, and energy percentage of the two overlapping discourse units.
[0073] In some embodiments, the steps of interruption detection and interruption property classification include: For each valid overlap event, the start times of the two units are compared to determine the first and second speakers. If the energy percentage of the second speaker exceeds a first energy threshold and is greater than the energy percentage of the first speaker, the second speaker is determined to be the interruptor, and the first speaker is the interrupted party. If the difference in energy percentage between the two parties is less than a second energy threshold, it is determined that both parties spoke simultaneously and are not marked as an interruption. The first energy threshold is greater than the second energy threshold. Based on the interruption type, the relevant discourse units are identified, resulting in a list of discourse units with the interruptor and basic overlap metadata.
[0074] Furthermore, all discourse units marked as interruption initiators are traversed, and their word-level text recognition results are extracted (only the semantics of text within overlapping time periods are extracted). The interruption nature is then classified based on keywords and sentence structure rules. The interruption nature includes at least: supportive interruption, where the text contains empathy / confirmation keywords, classified as a supportive interruption used for empathy or clarification, such as keywords like "I understand," "I get it," "I know," "Yes," etc.; corrective interruption, where the text contains corrective keywords, classified as a corrective interruption used to correct misunderstandings, such as keywords like "Incorrect," "Actually," "To be precise," "No," etc.; and competitive interruption, which does not meet the above two categories and is a statement, question, or rebuttal, classified as a competitive interruption intended to seize the discourse power. The interruption nature (supportive / corrective / competitive) is written into the metadata of the interruption initiator unit, resulting in a list of discourse units classified by interruption nature.
[0075] In this embodiment, basic discourse units that originally only contained text, time, and role labels are upgraded to structured data that carries details of clinical interaction: the annotation of long silence duration quantifies nonverbal pauses with significant clinical importance in the conversation, providing objective evidence for identifying the client's thinking, resistance, and emotional blockage; the determination of the interruption initiator based on time and energy, as well as the classification of interruptions as supportive, corrective, and competitive, clearly presents the interaction patterns and power dynamics between the counselor and the client, which can be directly used to assess the quality of the counseling relationship and the counselor's intervention skills; and the standardized metadata encapsulation makes this process information statistically quantifiable, traceable, and comparable, transforming the output data from a simple transcript of the conversation into a core clinical resource that can support clinical supervision, counseling quality assessment, and client status tracking, thereby enhancing the clinical value of the speech transcription results.
[0076] Example 2, as Figure 4 As shown, the psychological conversation language transposition system based on counselor role constraints includes: The audio processing module is used to perform a preliminary non-voice emotional event detection on the original audio of the psychological interview and obtain candidate voice segments. The role determination module is used to determine the role affiliation of the candidate speech segments based on the consultation role constraints to obtain initial role labels; and to perform constraint correction on the initial role labels based on the consultation role constraints to obtain target role labels. The boundary reconstruction module is used to reconstruct the discourse unit boundaries of the candidate speech segments identified by the target role label and obtain a discourse unit sequence based on time order. The behavior detection module is used to detect and label the clinical process behavior data corresponding to each discourse unit sequence, and obtain structured discourse data.
[0077] In this embodiment, by having each module work collaboratively based on a unified timeline and role constraints, data silos or temporal misalignments caused by independent processing by each module are avoided. This ensures the traceability and consistency of the entire chain from raw audio input to structured speech data output, significantly improving the automation and analysis efficiency of psychointerview data processing, thereby enhancing the clinical usability of interview speech transposition data.
[0078] The above-described embodiments are preferred embodiments of this application and are not intended to limit the specific scope of this application. The scope of this application includes but is not limited to the specific embodiments described above. All equivalent changes made in accordance with the shape, structure, and method of this application are within the protection scope of this application.
Claims
1. A counseling role constraint-based psychological counseling speech transposition method, characterized in that: Includes the following steps: Perform a preliminary non-voice emotional event detection on the original audio of psychological interviews to obtain candidate voice segments; Based on the consultant role constraint, the candidate voice segments are assigned a role to obtain initial role labels; Based on the aforementioned consultant role constraints, the initial role labels are constrained and modified to obtain the target role labels; The candidate speech segments identified by the target role tags are reconstructed into speech unit boundaries to obtain a speech unit sequence sorted by time. Detect and label the clinical process behavior data corresponding to each discourse unit sequence to obtain structured discourse data.
2. The counseling role constraint based psychological counseling speech transposition method according to claim 1, characterized in that: The initial non-voice emotional event detection of the original audio of the psychological interview to obtain candidate voice segments includes: Basic preprocessing was performed on the collected raw audio of psychological interviews, including audio noise reduction, echo suppression and gain normalization. Based on the dual-channel VAD, non-speech noise is detected and filtered from the pre-processed audio information to obtain the initial speech segment; The initial voice segment is desensitized to obscure user privacy information, and a valid voice segment with a timestamp is obtained. Based on the effective speech segments, basic acoustic features are extracted. Based on the feature threshold rule, the basic acoustic features are compared with the typical acoustic features of non-speech emotional events to identify non-speech emotional events. Then, non-speech emotional events within the same time period are deduplicated and time-corrected to obtain candidate speech segments.
3. The counseling role constraint based psychological interview speech transposition method according to claim 1, characterized in that: The step of determining the role affiliation of candidate voice segments based on consultation role constraints to obtain initial role labels includes: The candidate speech segments are subjected to speech recognition to generate text recognition results with timestamps; Extract the voiceprint features of the candidate speech segments, combine them with the text recognition results, and determine the initial role label of the candidate speech segments according to the consultation role constraint; The constraints on the consultant role include general constraints, language behavior constraints, cross-conversation template constraints, and genre adaptation constraints. The general constraints are used to filter voiceprint clustering results that do not meet the conditions based on the basic attributes of the role; The language behavior constraint is used to calculate the language behavior matching degree between the current speech segment and the language pattern of the target counselor role based on the text recognition results; The cross-conversation template constraint is used to calculate the degree of matching between the voiceprint features or language patterns of the current speech segment and the exclusive conversation template of the target consultation role. The school of thought adaptation constraint is used to adjust the calculation results of the language behavior matching degree according to the school of thought to which the target counselor belongs.
4. The counseling role constraint based psychological interview speech transposition method according to claim 3, characterized in that: The step of determining the initial role label of the candidate speech segment based on the consultant role constraint includes: After normalizing the voiceprint features of the candidate speech segments, a clustering algorithm is used to perform initial voiceprint clustering. The initial voiceprint clustering results are then filtered according to the general constraints to obtain the role cluster set of the candidate speech segments. Based on the current consultant role ID, retrieve the corresponding exclusive voiceprint template, calculate the similarity between each role cluster and the exclusive voiceprint template, and if the similarity is greater than or equal to the first similarity threshold, compare the similarity of role clusters in the same set to determine the initial role label of the role cluster. If the similarity is less than the second similarity threshold, the average language behavior matching degree of each role cluster is obtained based on the language behavior constraints, and the initial role label is determined based on the average language behavior matching degree. Among them, the first similarity threshold is greater than the second similarity threshold.
5. The counseling role constraint based psychological interview speech transposition method according to claim 4, characterized in that: The step of modifying the initial role labels based on the consultant role constraints to obtain the target role labels includes: A comprehensive role evaluation function is established based on the aforementioned language behavior constraints, cross-conversation template constraints, and genre adaptation constraints. The overall score of the initial role label for the candidate voice segment is calculated using the aforementioned role comprehensive evaluation function; The initial role labels are corrected based on the comprehensive score under the general constraints to obtain the target role labels.
6. The counseling role constraint based psychological interview speech transposition method according to claim 3, characterized in that: The step of reconstructing discourse unit boundaries for the candidate speech segments identified by the target role tags to obtain a time-ordered discourse unit sequence includes: The speech segment boundaries and word-level semantic boundaries of each candidate speech segment are extracted based on the time sequence. Invalid boundary filtering and initial reconstruction are performed on speech segment boundaries and word-level semantic boundaries to obtain candidate boundaries; Based on the target role label matching clinical discourse unit type library, the clinical unit type to which the candidate speech segment belongs is determined, and the corresponding boundary score weight is retrieved to generate the weight vector of the candidate boundary; An initial boundary score is obtained based on the weight vector and the basic feature information of the candidate boundary. If the initial boundary score is greater than or equal to a preset threshold, the candidate boundary is used as the reserved boundary. All reserved boundaries are traversed in chronological order to generate a sequence of discourse units, each discourse unit containing consecutive candidate speech segments.
7. The counseling role constraint based psychological interview speech transposition method according to claim 6, characterized in that: The extraction of speech segment boundaries and word-level semantic boundaries for each candidate speech segment based on time sequence includes: Extract two basic speech boundaries for each candidate speech segment. If the role labels of two adjacent segments are different, mark the basic speech boundary of the last time point of the previous segment as the role switching boundary. Based on the text recognition results of the candidate speech segments, extract the sentence-ending punctuation boundaries and the ordinary word ending boundaries that have no sentence-ending punctuation and whose time interval with the next word exceeds a preset threshold. Specifically, the boundaries of basic speech and role switching are used as the boundaries of speech segments, and the boundaries of sentence-end punctuation and ordinary word endings are used as the boundaries of word-level semantics.
8. The counseling role constraint based psychological interview speech transposition method according to claim 7, characterized in that: The process of filtering invalid boundaries and performing initial reconstruction on speech segment boundaries and word-level semantic boundaries to obtain candidate boundaries includes: Integrate all speech segment boundaries and word-level semantic boundaries and sort them by time. Calculate the time difference between two adjacent boundaries. If the time difference is less than or equal to a preset time difference threshold, it is determined to be a duplicate boundary, and the boundaries are merged based on the boundary priority. If the time difference is greater than the preset time difference threshold, two independent boundaries are retained. Merged or / and independent boundaries are used as temporary boundaries. These temporary boundaries are then filtered for time overlap and invalidity exceeding the session start and end time to obtain candidate boundaries. Among them, the boundary of ordinary word ending is low priority, the boundary of basic speech and the boundary of sentence end punctuation are medium priority, and the boundary of role switching is high priority.
9. The psychological conversation speech transposition method based on counselor role constraints according to claim 6, characterized in that: The process involves detecting and labeling the clinical process behavior data corresponding to each discourse unit sequence to obtain structured discourse data, including: Establish a global time-energy mapping table and a unit-semantic mapping table; The silence event is determined based on the time interval between two temporally adjacent speech units. If the time interval is greater than a threshold, the silence duration is written into the metadata of the speech unit at the next time step to obtain a list of silent speech units. Traverse the list of silent speech units, identify adjacent speech units with time overlap based on the global time-energy mapping table, calculate the energy ratio of the two speech units within the overlapping time period based on the overlap duration and the overlap duration ratio, and obtain the list of overlapping events; Based on the list of overlapping events, the speaking order is determined according to the start time of the two corresponding discourse units. The interruption and interrupted types are determined according to the energy ratio of the discourse units, and the interruption type is marked for the relevant discourse units. The interruption nature is determined based on the keywords and sentence features in the text recognition results, and a list of interrupted discourse units is obtained. By integrating the metadata of the silent discourse unit list, the overlapping event list, and the interrupted discourse unit list, structured discourse data is obtained.
10. A psychological conversation speech transposition system based on counselor role constraints, characterized in that: The method for transposing psychological conversation discourse based on counselor role constraints as described in any one of claims 1-9 includes: The audio processing module is used to perform a preliminary non-voice emotional event detection on the original audio of the psychological interview and obtain candidate voice segments. The role determination module is used to determine the role affiliation of the candidate speech segments based on the consultation role constraints to obtain initial role labels; and to perform constraint correction on the initial role labels based on the consultation role constraints to obtain target role labels. The boundary reconstruction module is used to reconstruct the discourse unit boundaries of the candidate speech segments identified by the target role label and obtain a discourse unit sequence based on time order. The behavior detection module is used to detect and label the clinical process behavior data corresponding to each discourse unit sequence, and obtain structured discourse data.