Psychological interview language interaction structure analysis method and system

By constructing dialogue state sequences through speech recognition and semantic intent recognition, identifying key interaction nodes, and generating structured interaction features, the problem of difficult-to-structure representation of language interaction in psychological interview training is solved, and objective evaluation is achieved.

CN121528217BActive Publication Date: 2026-04-21CHENGDU IND VOCATIONAL TECHN COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHENGDU IND VOCATIONAL TECHN COLLEGE
Filing Date
2026-01-15
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing psychological interview training systems struggle to transform unstructured interview language interactions into structured data while allowing trainees to express themselves freely, resulting in inconsistent and unobjective assessment results.

Method used

By acquiring language interaction data for speech recognition and semantic intent recognition, a dialogue state sequence is constructed, dialogue feature information is extracted, key interaction nodes are identified and structured interaction features are generated, and these features are mapped to skill dimension feature data for scoring.

Benefits of technology

It achieves systematic processing of the interview process, transforming it into quantifiable and comparable structured interactive features, supporting objective analysis and evaluation, and solving the problem of inconsistent evaluation results in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121528217B_ABST
    Figure CN121528217B_ABST
Patent Text Reader

Abstract

This invention provides a method and system for structured analysis of language interaction in psychological interviews, belonging to the field of natural language processing technology. The method includes: performing speech recognition processing on language interaction data to generate corresponding text data; extracting semantic intent information of trainees' language interactions and constructing a dialogue state sequence during the interview based on the semantic intent information; extracting dialogue feature information representing the evolutionary relationship of trainees' interactive behaviors during the interview; performing structured interaction analysis on the language interactions during the interview, identifying key interaction nodes and classifying trainees' language interaction behaviors to generate a set of structured interaction features; generating skill dimension feature data representing trainees' interview skill performance according to preset skill dimension mapping rules, and outputting corresponding skill dimension scoring results. This invention achieves structured representation and comparable evaluation of trainees' language interaction behaviors in psychological interview training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural language processing technology, and more specifically to a method and system for structured analysis of language interaction in psychological interviews. Background Technology

[0002] In psychological interview training and assessment scenarios, trainees typically need to engage in open-ended dialogues with interviewees through verbal interaction to train interview skills such as questioning, responding, guiding, and controlling the interaction. This type of interview process is highly flexible; the trainee's verbal expression, interaction rhythm, and response strategies all dynamically change as the interview progresses, resulting in a distinctly unstructured interview process.

[0003] Existing psychological interview training systems or language interaction training platforms typically focus on providing a dialogue environment or pre-set interaction scripts to guide trainees through a predetermined training process. However, in practical applications, interview training often requires allowing trainees to express themselves freely. Pre-set scripts struggle to cover the complex and varied interaction scenarios in real interviews, leading to discrepancies between the training process and actual application scenarios. Furthermore, the large amount of language interaction data generated in free dialogue modes lacks a unified structured representation, making subsequent analysis and evaluation of the training process reliant on manual playback or subjective judgment, hindering scalable and objective processing.

[0004] On the other hand, some existing technologies attempt to perform sentiment analysis or keyword statistics on interview language to aid in assessing trainees' performance. However, these methods typically focus only on single statements or local semantic features, failing to consider the contextual relationships and interactive evolution patterns during the interview process. Consequently, they struggle to reflect the dynamic changes in trainees' interactive behaviors. Furthermore, the interviewees, dialogue situations, and interactive stimuli faced by different trainees often vary, making the analysis results based on language content lack alignability and comparability, further limiting the objectivity and consistency of the assessment results.

[0005] Therefore, in the context of psychological interview training, there is an urgent need for a technical solution that can systematically process the language interactions formed during the interview process without restricting the trainees' freedom of expression. This solution can transform the originally unstructured interview language interaction process into an interactive feature representation with clear structural relationships, and based on the structured interactive features, achieve objective analysis and evaluation of the trainees' interview skills performance, so as to meet the actual needs of large-scale training and standardized assessment. Summary of the Invention

[0006] The purpose of this invention is to provide a method and system for structured analysis of language interaction in psychological interviews, so as to at least solve the problem that the interview process under free language interaction conditions is difficult to transform into structured and comparable data.

[0007] To achieve the above objectives, the first aspect of this invention provides a method and system for structured analysis of language interaction in psychological interviews. The method involves acquiring language interaction data between trainees and preset virtual interviewees during psychological interview training, performing speech recognition processing on the language interaction data to generate corresponding text data; performing semantic intent recognition processing on the text data to extract semantic intent information of the trainees' language interactions, and constructing a dialogue state sequence during the interview process based on the semantic intent information; performing dialogue state tracking and dialogue analysis processing on the dialogue state sequence to extract dialogue feature information representing the evolution of trainees' interactive behaviors during the interview; performing structured interaction analysis processing on the language interactions during the interview based on the dialogue feature information to identify key interaction nodes and classify the trainees' language interaction behaviors to generate a set of structured interaction features; and generating skill dimension feature data representing the trainees' interview skill performance according to preset skill dimension mapping rules based on the set of structured interaction features, and outputting corresponding skill dimension scoring results.

[0008] Optionally, the process involves acquiring language interaction data between trainees and preset virtual interviewees during psychological interview training, and performing speech recognition processing on the language interaction data to generate corresponding text data. This includes: during psychological interview training, collecting trainees' voice input signals in real time, and generating raw voice data containing timestamp information based on the voice input signals; performing voice activity detection processing on the raw voice data to identify valid voice segments and eliminate non-voice segments, generating valid voice data for speech recognition processing; performing speech recognition processing based on the valid voice data to generate text data corresponding to the trainees' voice input content, and associating the text data with the corresponding timestamp information.

[0009] Optionally, semantic intent recognition processing is performed based on the text data to extract semantic intent information of the trainees' language interactions, and a dialogue state sequence during the interview process is constructed based on the semantic intent information. This includes: performing semantic parsing processing on the text data to extract semantic intent features that characterize the trainees' speaking intent type and interaction mode; mapping the trainees' speaking content to preset semantic intent labels based on the semantic intent features to form a semantic intent label sequence corresponding to each speaking moment during the interview; constructing a dialogue state sequence describing the evolution of language interaction during the interview based on the semantic intent label sequence and the temporal sequence corresponding to the text data; and using the dialogue state sequence as input data for subsequent dialogue state tracking and dialogue analysis processing.

[0010] Optionally, the text data is subjected to semantic parsing processing to extract semantic intent features that characterize the trainee's speaking intention type and interaction mode. This includes: segmenting the text data into speaking units based on a preset interview interaction unit division rule to obtain multiple interactive statement units divided according to semantic integrity; extracting the pragmatic position features of each interactive statement unit in the interview interaction process, wherein the pragmatic position features are used to characterize the interaction stage of the interactive statement unit in the current interview round; and constructing a semantic intent feature vector to describe the trainee's speaking interaction mode based on the pragmatic position features of the interactive statement units, combined with the sentence structure features and response orientation features of the corresponding statement units.

[0011] Optionally, based on the dialogue state sequence, dialogue state tracking and dialogue analysis are performed to extract dialogue feature information representing the evolution of trainees' interactive behaviors during the interview. This includes: constructing a state transition structure based on the dialogue state sequence to represent the transition relationship between adjacent interactive states during the interview, thereby describing the evolution path of trainees' language interactive behaviors during the interview; performing continuous state update processing on the state transition structure to generate state evolution features representing the changing trends of trainees' interactive behaviors in different interview stages; and extracting behavioral evolution features based on the state evolution features to represent the stability of trainees' interactions, the frequency of interaction adjustments, and the changes in interaction strategies during the interview, as dialogue feature information.

[0012] Optionally, based on the state evolution features, behavioral evolution features are extracted to characterize the stability of the trainees' interactions, the frequency of interaction adjustments, and the changes in their interaction strategies during the interview, as dialogue feature information. This includes: calculating the distribution of state retention duration between adjacent interview states based on the state evolution features to generate stability feature parameters characterizing the stability of the trainees' interactions; statistically analyzing the frequency of state transitions during the interview and their distribution on the interview timeline based on the state evolution features to generate adjustment frequency feature parameters characterizing the frequency of the trainees' interaction adjustments; identifying interaction pattern change segments formed by multiple consecutive state transitions during the interview based on the state evolution features, and generating strategy change feature parameters characterizing the changes in the trainees' interaction strategies based on the interaction pattern change segments; and combining the stability feature parameters, adjustment frequency feature parameters, and strategy change feature parameters as behavioral evolution features.

[0013] Optionally, based on the dialogue feature information, structured interaction analysis is performed on the language interactions during the interview process to identify key interaction nodes and classify the students' language interaction behaviors to generate a set of structured interaction features. This includes: constructing an interaction structure representation based on the dialogue feature information to describe the evolution of language interactions during the interview, wherein the interaction structure representation includes at least state evolution relationships, speaking order relationships, and interaction response relationships; based on the interaction structure representation, determining the interaction positions that meet the conditions of state transition mutation, interaction strategy change, or significant change in interaction stability during the interview process according to preset node determination rules, and identifying the corresponding key interaction nodes; for each key interaction node, extracting node behavior features that characterize the students' language interaction methods in the corresponding interview stage by combining the dialogue feature information before and after the key interaction node; classifying the students' language interaction behaviors during the interview process according to preset behavior classification rules based on the node behavior features, and generating corresponding behavior category identifiers; associating the key interaction nodes, the node behavior features, and the behavior category identifiers to form a set of structured interaction features for describing the language interaction process during the interview.

[0014] Optionally, based on the structured interaction feature set, skill dimension feature data representing the interviewer's interview skill performance is generated according to a preset skill dimension mapping rule. This includes: extracting interaction feature subsets corresponding to different interview skill dimensions based on the structured interaction feature set, wherein the interview skill dimensions include at least questioning guidance ability, response coherence, and interaction control ability; for each interview skill dimension, calculating dimensional feature parameters reflecting the trainee's interaction performance characteristics under the corresponding interaction feature subset according to a preset dimensional mapping relationship; and performing intra-dimensional normalization processing on each dimensional feature parameter to generate skill dimension feature data representing the trainee's relative performance level under different interview skill dimensions.

[0015] Optionally, the output rules for the corresponding skill dimension scoring results are as follows: Based on the skill dimension feature data, calculate the original score value of each interview skill dimension according to the preset scoring mapping relationship corresponding to each interview skill dimension; perform cross-dimensional alignment processing on each of the original score values ​​to eliminate the difference in scoring scale between different interview skill dimensions and generate standardized scoring results for comparability between dimensions; based on the standardized scoring results, construct a skill scoring vector to describe the relative performance distribution of trainees under different interview skill dimensions; associate the skill scoring vector with the corresponding interview skill dimension identifier to form a structured skill dimension scoring result, and output it according to the preset output format.

[0016] A second aspect of the present invention provides a structured analysis system for language interaction in psychological interviews. The system includes: a data acquisition unit, configured to acquire language interaction data formed between a trainee and a preset virtual interviewee during psychological interview training, and to perform speech recognition processing on the language interaction data to generate corresponding text data; an intent recognition unit, configured to perform semantic intent recognition processing based on the text data, extract semantic intent information of the trainee's language interaction, and construct a dialogue state sequence during the interview based on the semantic intent information; a feature extraction unit, configured to perform dialogue state tracking and dialogue analysis processing based on the dialogue state sequence, and extract dialogue feature information representing the evolution of trainee interaction behavior during the interview; a structured processing unit, configured to perform structured interaction analysis processing on the language interaction during the interview based on the dialogue feature information, identify key interaction nodes, classify the trainee's language interaction behavior, and generate a set of structured interaction features; and an output unit, configured to generate skill dimension feature data representing the trainee's interview skill performance according to preset skill dimension mapping rules based on the set of structured interaction features, and output corresponding skill dimension scoring results.

[0017] Through the above technical solution, this invention achieves systematic processing of free language interaction data during psychological interview training, transforming the originally loosely sequenced and poorly structured interview language process into structured interaction features with clear state evolution relationships and behavioral category identifiers. On the one hand, through semantic intent recognition and dialogue state sequence construction, the trainee's language expression during the interview process is mapped into continuously trackable dialogue states, thereby characterizing the dynamic evolution process of the interview interaction. On the other hand, through dialogue state tracking and structured interaction analysis, key interaction nodes in the interview process are identified and language interaction behaviors are classified, transforming the trainee's interview behavior from raw text form into a quantifiable and comparable set of structured features. Based on this, through preset skill dimension mapping rules, the structured interaction features are further mapped into skill dimension feature data and corresponding scoring results are output. Thus, without restricting the trainee's free expression, objective analysis and unified evaluation of interview skills performance are achieved, solving the problems of difficult structured representation of language interaction and lack of consistency in evaluation results in existing psychological interview training.

[0018] Other features and advantages of the embodiments of the present invention will be described in detail in the following detailed description section. Attached Figure Description

[0019] The accompanying drawings are provided to further illustrate embodiments of the present invention and form part of the specification. They are used together with the following detailed description to explain the embodiments of the present invention, but do not constitute a limitation thereof. In the drawings:

[0020] Figure 1 This is a flowchart of the steps of a structured analysis method for language interaction in psychological interviews provided by one embodiment of the present invention;

[0021] Figure 2 This is a detailed flowchart of step S20 of the structured analysis method for language interaction in psychological interviews provided by one embodiment of the present invention.

[0022] Figure 3 This is a schematic diagram of the training system architecture for structured analysis of psychological interview language interaction provided by one embodiment of the present invention;

[0023] Figure 4 This is a system structure diagram of a psychological interview language interaction structured analysis system provided in one embodiment of the present invention;

[0024] Figure 5 This is an internal structural diagram of a computer device provided in one embodiment of the present invention. Detailed Implementation

[0025] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0026] like Figure 1 As shown, embodiments of the present invention provide a method for structured analysis of language interaction in psychological interviews, the method comprising:

[0027] Step S10: Obtain the language interaction data formed between the trainee and the preset virtual interviewee during the psychological interview training process, and perform speech recognition processing on the language interaction data to generate corresponding text data.

[0028] Specifically, during the psychological interview training process, the trainees' voice input signals are collected in real time, and raw voice data containing timestamp information is formed based on the voice input signals; voice activity detection processing is performed on the raw voice data to identify valid voice segments and eliminate non-voice segments, generating valid voice data for voice recognition processing; voice recognition processing is performed based on the valid voice data to generate text data corresponding to the trainees' voice input content, and the text data is associated with the corresponding timestamp information.

[0029] In this embodiment of the invention, after the psychological interview training begins, the training system continuously collects the trainees' voice input signals. The voice input signals can be continuous voice stream data acquired in real time by a microphone or other voice acquisition device. To ensure the correspondence between the voice content and the interview process during subsequent processing, time stamp information is added to the voice input signals while they are being collected, forming raw voice data containing timestamp information. The timestamp information is used to characterize the positional relationship of each voice segment on the interview timeline.

[0030] Furthermore, speech activity detection (VAD) processing is performed on the raw speech data to distinguish between valid speech segments and non-speech segments in the speech signal. Specifically, by analyzing the energy changes, spectral distribution, or other speech activity determination features of the raw speech data, speech segments containing the student's actual speech content are identified, and background noise, silent segments, or non-speech signals are removed. After this step, valid speech data containing only valid speech content is generated, thereby reducing the interference of invalid data on subsequent speech recognition processing.

[0031] After obtaining the valid speech data, speech recognition processing is performed on it. Speech recognition processing converts continuous speech signals into corresponding text representations according to a preset recognition model or rules. During this process, each speech segment in the valid speech data maintains a correlation with its corresponding timestamp information, ensuring that the generated text data not only contains semantic content but also retains time-location information consistent with the interview's chronological order.

[0032] Specifically, the output of speech recognition processing is a set of text data fragments, each corresponding to a segment of a trainee's speech during the interview, and each text data fragment corresponds one-to-one with the timestamp information in the original speech data. This method forms a "text content-timestamp" associated data structure, used to characterize the distribution of trainees' language interaction content over time during psychological interview training. Finally, the timestamped text data is used as input data for subsequent semantic intent recognition processing and dialogue state sequence construction processing.

[0033] In another possible implementation, to further enhance the alignment and traceability of interview language interaction data in the time dimension, this embodiment introduces a speech boundary self-calibration mechanism to refine and adjust the time stamp of text data during the speech acquisition and speech recognition processing.

[0034] Specifically, after real-time acquisition of student voice input signals and formation of raw voice data containing timestamp information, in addition to performing routine voice activity detection processing, continuous monitoring of voice energy changes and speech rate changes within effective voice segments is also conducted. Based on the monitoring results, the speech start boundary point and speech end boundary point are further identified within each effective voice segment to distinguish multiple consecutive speech units that may be contained within the same voice segment.

[0035] Based on this, fine-grained timestamp information is reassigned to each identified speech unit, ensuring that each speech unit corresponds to an independent time interval. Subsequently, during speech recognition processing, the valid speech data is segmented according to the boundaries of the speech units, generating text data segments that correspond one-to-one with each speech unit. These text data segments maintain association with their corresponding fine-grained timestamp information during generation, thus forming a text-time mapping relationship accurate to the speech level.

[0036] This implementation method allows for a more refined distribution of text data along the timeline without altering the overall speech recognition process. This facilitates the differentiation of continuous speech in subsequent semantic intent recognition and provides a more accurate temporal sequence for constructing dialogue state sequences. This approach is particularly suitable for interactive scenarios during interviews where trainees exhibit significant changes in speaking speed and frequent, continuous responses, thus improving the accuracy of locating key interaction nodes in subsequent structured interaction analysis.

[0037] Step S20: Perform semantic intent recognition processing based on the text data, extract semantic intent information of the trainees' language interaction, and construct a dialogue state sequence during the interview process based on the semantic intent information.

[0038] Specifically, based on the aforementioned text data, this embodiment performs semantic intent recognition processing on the language interaction content formed by trainees during psychological interview training to extract semantic intent information that reflects the trainees' speaking purpose and interaction mode. The semantic intent recognition processing parses the sentence structure, semantic orientation, and contextual relationships in the text data, mapping the trainees' speaking content into a preset semantic intent representation. Based on this, according to the sequential relationship of the text data on the interview timeline, the semantic intent information corresponding to consecutive statements is organized and associated to construct a dialogue state sequence describing the evolution of language interaction during the interview. This dialogue state sequence is used to characterize the changes in the trainees' interaction states at different interview stages and serves as the basic input data for subsequent dialogue state tracking, dialogue analysis, and structured interaction analysis, thereby providing a unified data expression form for further structured processing of the interview language interaction process. Specifically, as... Figure 2 Step S20 includes the following steps:

[0039] Step S201: Perform semantic parsing on the text data to extract semantic intent features that characterize the type of student's speaking intent and interaction method.

[0040] Specifically, based on preset interview interaction unit segmentation rules, the text data is segmented into speaking units to obtain multiple interactive statement units divided according to semantic integrity; for each interactive statement unit, its pragmatic position features in the interview interaction process are extracted, and the pragmatic position features are used to characterize the interaction stage of the interactive statement unit in the current interview round; based on the pragmatic position features of the interactive statement units, combined with the sentence structure features and response orientation features of the corresponding statement units, a semantic intent feature vector is constructed to describe the trainees' speaking interaction methods.

[0041] Specifically, after completing the segmentation of the speech units, this implementation method denotes each interactive statement unit as u. i and retain u i Round index r in the interview rounds i And the sequential index k within the round i Semantic parsing does not encode the entire text at once, but first processes the u... i Pragmatic position features, sentence structure features, and response orientation features are extracted, and then these three types of features are fused in the same feature space to form a semantic intent feature vector f. i The reason is that in interview training scenarios, the same statement may correspond to different interaction methods in different rounds, and the sentence structure and response direction can provide reusable clues to determine whether the speech is asking, clarifying, restating, empathizing, or advancing.

[0042] In a specific implementation, pragmatic positional feature p i Can be indexed by round r i Sequential index k i and the unit u of the previous statement i-1 The time interval Δt i Common construction, for example, (r) i ,k i ,Δt i Mapped to vector p i ∈R dp ; Sentence structure features i A vector s can be constructed from syntactic dependency relation statistics, interrogative word trigger markers, negation structure markers, subject-verb-object skeleton lengths, etc. i ∈R ds ; response pointing to feature g i Used to describe u iThe connection between the statements made by the interviewees in the previous virtual interview round or the reference to a certain topic in the preceding text can be resolved through the results of referential resolution, the location of keyword backreferences, and the reference to the statements made by the interviewees in the previous round. ri The matching scores are constructed as a vector. Then, the three types of features are fused to obtain the semantic intent feature vector:

[0043] ;

[0044] Among them, f i Represents the interactive statement unit u i The semantic intent feature vector; p i Represents the pragmatic position feature vector; s i Represents the feature vector of sentence structure; g i W indicates that the response points to the feature vector; p W s W g p respectively i s i g i A parameter matrix mapped to a uniform dimension; b f For the bias vector; [p i s i g i ] indicates vector concatenation; W c b c Here, σ represents the parameter matrix and bias vector used to generate the gated components; σ(·) represents the Sigmoid nonlinear function; tanh(·) represents the hyperbolic tangent function; and ⊙ represents element-wise multiplication. Through this fusion method, pragmatic location, sentence structure, and response orientation can form mutually constraining relationships within the same vector, avoiding intent drift caused by relying solely on content vocabulary. The final output f... i This will serve as input for subsequent semantic intent label mapping.

[0045] Step S202: Based on the semantic intent features, map the trainee's speech content to preset semantic intent tags to form a semantic intent tag sequence corresponding to each speaking moment during the interview.

[0046] Specifically, the preset semantic intent tag set is denoted as L={l1,l2,…,l M} where M is the preset number of labels. The label content can cover common interaction intent types and interaction methods in interview training, such as open-ended questions, closed-ended questions, clarification and follow-up questions, restatement and confirmation, emotional transition, summary and advancement, boundary setting, etc. The mapping process does not use a hard threshold for single judgment, but first calculates the posterior probability distribution of each label, and then obtains the final label according to the preset selection rules, thereby providing traceable uncertainty information for the subsequent construction of dialogue state sequences.

[0047] In one specific implementation, the f output in step S201 is... i The input is fed into the intent classifier to obtain a score z for each label. i ∈R M Furthermore, prior category weights and a temperature coefficient are introduced to accommodate the differences in the frequency of different labels and the differences in discrimination difficulty in the training corpus. Subsequently, the posterior probability π is obtained through Softmax. i :

[0048]

[0049] Where, π{ i,m} represents the interactive statement unit u i Corresponding tag l m The posterior probability of z{ i,m} indicates the label l m Unnormalized scoring; W o b is the output layer parameter matrix; o α is the output layer bias vector; m For tag l m The prior bias term is used to reflect the preset label frequency or importance constraints; τ is a temperature coefficient used to adjust the sharpness of the probability distribution; M is the total number of labels. Then, labels y are obtained according to the preset label selection rules. i The preset label selection rules may include, but are not limited to: selecting the label corresponding to the maximum posterior probability; or, when the difference between the maximum posterior probability and the second-largest posterior probability is less than a preset difference threshold, retaining the candidate label set and further judging it in the subsequent state sequence construction. Finally, each interactive statement unit u i The tag y i Arranged chronologically, forming a semantic intent tag sequence Y=[y1,y2,…,y N ], where N is the number of interactive statement units. The semantic intent tag sequence Y and each u i The timestamp information is kept one-to-one, providing discrete observation sequences for subsequent modeling of state transitions and evolutionary relationships.

[0050] Step S203: Based on the semantic intent tag sequence and combined with the temporal order relationship corresponding to the text data, construct a dialogue state sequence to describe the evolution of language interaction during the interview.

[0051] Specifically, in this implementation, the dialogue state sequence is defined as S=[s1,s2,…,s…]. N ], where s i Used to characterize the interactive statement unit u i The interview interaction status at the location. Compared to only putting y iAs different from other states, this implementation constructs states as a combination of "label-driven + time-constrained + contextual memory," so that states can both reflect the current intent and inherit from previous evolutions. To this end, the label y is first... i Mapped to embedding vector e i ∈R de And introduce the time increment Δt corresponding to the time sequence relationship. i and round boundary marker q i This is used to distinguish state transitions across rounds. Subsequently, a recursive method is used to generate the state vector h. i and with h i As a dialogue state i The expression .

[0052] In one specific implementation, state recursion can employ a gated update mechanism, explicitly introducing a transition constraint matrix A to limit unreasonable intentional jumps, thereby ensuring that state evolution conforms to the basic sequential logic of the interview interaction. The state update can be represented as:

[0053]

[0054] Among them, h i This represents the state vector corresponding to the i-th interactive statement unit; e i The semantic intent label y i The embedding vector; φ(Δt) i ) represents the time increment Δt i The time encoding function mapped to a vector; q i Represents the round boundary marker vector, used to indicate u i Is this the start of a new round? i-1} represents the previous state vector; U e U t U q U h b is the parameter matrix; h ρ is the bias vector; ρ(·) represents the preset nonlinear combination function, which can be tanh or a gated combination form; A y{i-1},yi This indicates that the transition constraint matrix A is formed by the previous label y. i-1 To the current tag y i The transition term is used to express the preset permissible transition relationship; λ is a preset coefficient used to adjust the strength of the transition constraint influence. Through the above construction method, the state vector h... i Simultaneously, it absorbs current tag information, time sequence information, round boundary information, and historical memory information, and applies structural constraints to tag jumps through A. Finally, it will... i Defined as h iWe obtain the dialogue state sequence S and maintain the correspondence between S, the semantic intent label sequence Y, and the timestamp information so as to perform dialogue state tracking and dialogue analysis in subsequent steps.

[0055] Step S204: Use the dialogue state sequence as input data for subsequent dialogue state tracking and dialogue analysis.

[0056] Specifically, the dialogue state sequence S, the semantic intent label sequence Y, and the timestamp information T=[t1,t2,…,t] of the text data are combined. N They are bound together to form a unified data object D. Where, (s i ,y i ,t i The data object D is represented by a triplet across the state, label, and time layers of the same interactive statement unit (UI). Subsequently, based on the windowing requirements of subsequent processing, the data object D is fragmented. For example, if a sliding window approach is used for subsequent dialogue state tracking, then window fragment D is constructed with a preset window length K. k Consistency in continuous tracking is achieved by using overlapping sections shared between adjacent windows. If subsequent dialogue analysis requires aggregation by round, it is based on the round boundary marker q. i The process involves dividing D into multiple round sets, with each round set maintaining a monotonically increasing timestamp to avoid state drift caused by mixing across rounds. Finally, the window segment D is... k Alternatively, the set of rounds can be used as input data for subsequent dialogue state tracking and dialogue analysis, and y can be retained in the input data. i With s i This establishes a correspondence so that when identifying key interaction nodes and classifying behaviors in subsequent processes, it is possible to trace back to the specific intent label and corresponding text fragment, thus achieving a data loop throughout the entire process.

[0057] Step S30: Based on the dialogue state sequence, perform dialogue state tracking and dialogue analysis processing to extract dialogue feature information that represents the evolution of trainees' interactive behaviors during the interview.

[0058] Specifically, based on the dialogue state sequence, a state transition structure is constructed to represent the transition relationship between adjacent interactive states during the interview, thereby describing the evolution path of the learner's language interaction behavior during the interview. The state transition structure is continuously updated to generate state evolution features that characterize the changing trends of the learner's interaction behavior in different interview stages. Based on the state evolution features, behavioral evolution features that characterize the learner's interaction stability, interaction adjustment frequency, and interaction strategy changes during the interview are extracted as dialogue feature information.

[0059] Furthermore, based on the state evolution features, behavioral evolution features are extracted to characterize the stability of the trainees' interactions, the frequency of interaction adjustments, and the changes in their interaction strategies during the interview, serving as dialogue feature information. This includes: calculating the distribution of state duration between adjacent interview states based on the state evolution features to generate stability feature parameters characterizing the stability of the trainees' interactions; statistically analyzing the frequency of state transitions during the interview and their distribution along the interview timeline based on the state evolution features to generate adjustment frequency feature parameters characterizing the frequency of the trainees' interaction adjustments; identifying interaction pattern change segments formed by multiple consecutive state transitions during the interview based on the state evolution features, and generating strategy change feature parameters characterizing the changes in the trainees' interaction strategies based on these interaction pattern change segments; and combining the stability feature parameters, adjustment frequency feature parameters, and strategy change feature parameters as behavioral evolution features.

[0060] In this embodiment of the invention, based on the dialogue state sequence, dialogue state tracking and dialogue analysis are performed to extract dialogue feature information characterizing the evolution of trainees' interactive behaviors during the interview. In this embodiment, the dialogue state sequence is represented as...

[0061]

[0062] in, Indicates the relationship with the first Each interactive statement unit corresponds to a dialogue state vector, and each dialogue state vector All are associated with the corresponding timestamp Maintaining relevance. By analyzing the continuous changes in adjacent dialogue states, the evolutionary path of trainees' language interaction during the interview process can be depicted.

[0063] Therefore, a state transition structure is first constructed based on the differences between adjacent dialogue states. For any two adjacent dialogue states... and By calculating the magnitude of its change in the state space and combining it with the corresponding time interval information, state transition weights are defined. , used to characterize from state Evolution to state The transfer intensity is calculated as follows:

[0064]

[0065] in, This represents the difference vector between adjacent dialogue state vectors. This represents a pre-defined positive semi-definite weight matrix, used to assign different weights to the changes in each dimension of the state difference vector. This indicates the time interval between adjacent interactive statement units. This represents a pre-defined non-zero stable term. Using the above method, a state transition weight sequence reflecting the continuous changes in language interaction during the interview can be obtained. It is used to describe the evolution trajectory of student interaction behavior on the timeline.

[0066] Building upon this, to characterize the overall interaction trends at different stages of the interview process, the state transition weight sequence is continuously aggregated to generate state evolution features. Specifically, a sliding window mechanism is introduced to weight and aggregate multiple consecutive state transition weights to obtain a state evolution intensity sequence. The calculation method is as follows:

[0067]

[0068] in, This indicates the preset window length. This represents the normalized coefficient of the state transition weights within the window. This represents the decay parameter, used to enhance attention to changes in the state near the current location. The state evolution intensity sequence calculated above can reflect the overall trend of language interaction during the interview process, from stable to changing.

[0069] Furthermore, based on the state evolution intensity sequence, behavioral evolution features are extracted to characterize the stability of the trainees' interactions, the frequency of interaction adjustments, and the changes in their interaction strategies during the interview. To this end, the state evolution intensity sequence is compared with a preset threshold. Compare and generate evolutionary event indicator sequences. And based on this, behavioral evolution characteristic parameters are calculated. Specifically, stability characteristic parameters... Adjusting frequency characteristic parameters and strategy change characteristic parameters They are defined as follows:

[0070]

[0071]

[0072] in, Indicates the first The length of a continuous stable segment Indicates the number of stable segments; Indicates the first The time position corresponding to the significant change in the state of the next interaction. Indicates the number of events that show significant changes; Indicates the first The number of consecutive change events in a change cluster This indicates the number of variable clusters. Finally, the... , and The combined data form a behavioral evolution feature vector, which is used to characterize the overall evolutionary characteristics of trainees' language interaction behavior during the interview process, and serves as input data for subsequent structured interaction analysis.

[0073] Step S40: Based on the dialogue feature information, perform structured interaction analysis processing on the language interaction during the interview, identify key interaction nodes, classify the language interaction behavior of trainees, and generate a set of structured interaction features.

[0074] Specifically, based on the dialogue feature information, an interaction structure representation is constructed to describe the evolution of language interaction during the interview. This interaction structure representation includes at least state evolution relationships, speaking order relationships, and interaction response relationships. Based on the interaction structure representation, according to preset node determination rules, interaction positions that satisfy conditions such as abrupt state transitions, changes in interaction strategies, or significant changes in interaction stability are determined, and corresponding key interaction nodes are identified. For each key interaction node, combined with the dialogue feature information before and after the key interaction node, node behavioral features are extracted to characterize the trainee's language interaction style in the corresponding interview stage. Based on the node behavioral features, according to preset behavior classification rules, the trainee's language interaction behavior during the interview is classified, generating corresponding behavior category identifiers. The key interaction nodes, the node behavioral features, and the behavior category identifiers are associated to form a structured interaction feature set describing the interview language interaction process.

[0075] In this embodiment of the invention, based on the dialogue feature information, structured interaction analysis is performed on the language interactions during the interview process to identify key interaction nodes and classify the trainees' language interaction behaviors, generating a set of structured interaction features. In this embodiment, the dialogue feature information is obtained from the aforementioned dialogue state tracking and dialogue analysis processing, and includes at least a dialogue state sequence. State transition weight sequence State evolution intensity sequence and behavioral evolution feature vectors Furthermore, all of the above features correspond to the timestamp sequence. Maintain a one-to-one correspondence.

[0076] Based on this, we first construct an interaction structure representation to describe the evolution of language interaction during the interview process. This interaction structure representation is defined as a weighted ordered structure:

[0077]

[0078] Among them, the node set Each node in Corresponding to the first point in the interview process One interactive statement unit; edge set Used to describe the interaction relationships between nodes. For any adjacent nodes and In the interactive structure representation, directed edges are established. and the corresponding state transition weights The weight attribute of this edge reflects the intensity of state evolution between adjacent interaction positions. Thus, the interaction structure representation simultaneously includes state evolution relationships, speaking order relationships, and interaction response relationships implied by time order, thereby forming a structured representation of the interview language interaction process.

[0079] After obtaining the interaction structure representation, based on preset node determination rules, interaction positions that meet specific evolutionary conditions during the interview process are determined to identify key interaction nodes. Specifically, for each node... Taking into account its corresponding state transition weights and State evolution intensity and behavioral evolution feature vectors Construct a node saliency determination function based on the relevant components in the data. When node A node is identified as a key interaction node if it meets at least one of the following criteria:

[0080] 1) Adjacent state transition weights at nodes A mutation occurs at that point, that is The change exceeds the preset threshold.

[0081] 2) Node The corresponding state evolution intensity A significant shift compared to the mean of the preceding and following intervals indicates a change in interaction stability.

[0082] 3) Nodes The behavioral evolution feature vectors within the given segment exhibit a clustering trend along the strategy change dimension, reflecting the phased adjustments of the interaction strategy. Using the above determination method, a set of key interaction nodes is obtained on the timeline:

[0083]

[0084] This is used to characterize the locations where significant changes occur in language interaction during the interview. For each key interaction node, further combining the dialogue feature information before and after the key interaction node, node behavioral features are extracted to represent the trainee's language interaction style in the corresponding interview stage. Specifically, for any key interaction node... Centered on this, select a local interval on the timeline containing several preceding and following interactive statement units to form a node context set:

[0085]

[0086] in, The radius of the preset context window is used. Within this context set, dialogue features such as state evolution strength, state transition weights, and semantic intent distribution are aggregated to form a node behavior feature vector. This is used to describe the characteristics of a learner's language interaction style near this key interaction node.

[0087] After obtaining the node behavior feature vectors, the language interaction behaviors of trainees during the interview process are classified based on preset behavior classification rules. The behavior classification rules are used to classify the node behavior feature vectors...

[0088] Mapped to discrete behavior category identifiers The behavior categories distinguish at least different interaction types, such as stable and continuous interaction, strategy-adjusting interaction, or mutation-response interaction. This approach assigns a unique behavior category identifier to each key interaction node, thereby providing a structured characterization of learners' language interaction behavior at the node level.

[0089] Finally, the set of key interaction nodes will be... The corresponding set of node behavior feature vectors and behavioral category identifier set By associating these features, a structured set of interactive characteristics can be formed to describe the language interaction process during the interviews.

[0090]

[0091] The structured interaction feature set maintains a consistent data correspondence across the time, state, and behavior dimensions, and can serve as input data for subsequent interview skill dimension mapping and scoring processing, thereby enabling structured analysis and aligned representation of the interview language interaction process.

[0092] Step S50: Based on the structured interaction feature set, generate skill dimension feature data representing the trainees' interview skill performance according to the preset skill dimension mapping rules, and output the corresponding skill dimension score results.

[0093] Specifically, based on the structured interaction feature set, a subset of interaction features corresponding to different interview skill dimensions is extracted. The interview skill dimensions include at least the ability to guide questions, the coherence of responses, and the ability to control interactions. For each interview skill dimension, based on the corresponding subset of interaction features and according to a preset dimension mapping relationship, dimensional feature parameters reflecting the interactive performance characteristics of trainees under the interview skill dimension are calculated. The dimensional feature parameters are then normalized within the dimension to generate skill dimension feature data that characterizes the relative performance level of trainees under different interview skill dimensions.

[0094] Furthermore, the output rules for the corresponding skill dimension scoring results are as follows: Based on the skill dimension feature data, calculate the original score value of each interview skill dimension according to the preset scoring mapping relationship corresponding to each interview skill dimension; perform cross-dimensional alignment processing on each of the original score values ​​to eliminate the difference in scoring scale between different interview skill dimensions and generate standardized scoring results for comparability between dimensions; based on the standardized scoring results, construct a skill scoring vector to describe the relative performance distribution of trainees under different interview skill dimensions; associate the skill scoring vector with the corresponding interview skill dimension identifier to form a structured skill dimension scoring result, and output it according to the preset output format.

[0095] In this embodiment of the invention, based on the structured interaction feature set, skill dimension feature data representing trainees' interview skill performance is generated according to a preset skill dimension mapping rule, and the corresponding skill dimension score results are output. In this embodiment, the structured interaction feature set is represented as...

[0096]

[0097] in, Indicates the first Key interaction nodes, This represents the node behavior feature vector corresponding to the key interaction node. This represents the corresponding behavioral category identifier. The structured interaction feature set maintains a consistent correspondence between the time and behavioral dimensions, supporting analysis at the interview skills level.

[0098] Based on this, firstly, based on a pre-defined set of interview skills dimensions:

[0099]

[0100] The structured interaction feature set is decomposed, wherein the interview skill dimensions include at least the ability to guide questioning, the coherence of responses, and the ability to control interaction. For any interview skill dimension... Based on the preset skill dimension mapping rules, from the structured interaction feature set Extract a subset of interactive features related to the skill dimension, and in this way, achieve the ordered aggregation of structured interactive features at the skill dimension level.

[0101] For each interview skill dimension Based on the corresponding interaction feature subset The dimensional feature parameters reflecting the trainees' interactive performance characteristics under the interview skill dimension are calculated according to the preset dimensional mapping relationship. Specifically, the behavioral feature vectors of each node in the interactive feature subset are... We perform weighted aggregation based on the behavior category weights to obtain the original feature vector of the skill dimension. The calculation method is as follows:

[0102]

[0103] in, Indicates the size of the subset of interactive features. Indicator and behavior category identifier The corresponding preset weight coefficients are used to reflect the differences in importance of different interactive behaviors within this skill dimension. Through the above aggregation method, the skill dimension feature parameters can comprehensively reflect the performance of multiple key interactive nodes within this dimension.

[0104] After obtaining the original dimensional feature vectors corresponding to each interview skill dimension, intra-dimensional normalization is performed on the dimensional feature parameters to eliminate the influence of differences in interview sample length or interaction frequency on the results, generating skill dimension feature data. Specifically, for the skill dimensions... The original feature vector Perform the following normalization process:

[0105]

[0106] in, and These represent the skill dimensions. The mean vector and standard deviation vector within a preset sample range, This represents a pre-defined stable term. Through this processing, skill dimension feature data is obtained to characterize the relative performance levels of trainees across different interview skill dimensions. .

[0107] Furthermore, based on the skill dimension feature data, and according to the preset scoring mapping relationship corresponding to each interview skill dimension, the original score value of each interview skill dimension is calculated. Specifically, for each skill dimension... The corresponding skill dimension feature data Mapped to raw rating values The calculation method is as follows:

[0108]

[0109] in, Indicating skill dimensions The corresponding rating mapping weight vector, This indicates the rating bias. The original rating values ​​are used to reflect the trainee's absolute performance level in a single skill dimension.

[0110] To eliminate rating scale discrepancies across different interview skill dimensions, cross-dimensional alignment was performed on each raw rating value to generate standardized rating results for inter-dimensional comparability. Specifically, the set of raw rating values... Mapped to a standardized set of ratings The calculation method is as follows:

[0111]

[0112] in, Indicating skill dimensions The corresponding standardized scoring results. Based on the standardized scoring results, a skill scoring vector is constructed:

[0113]

[0114] This is used to describe the relative performance distribution of trainees across different interview skill dimensions. Finally, the skill score vectors are associated with the corresponding interview skill dimension identifiers to form structured skill dimension score results, which are then output according to a preset output format for subsequent teaching assessments or training feedback.

[0115] In one specific implementation, the structured analysis method for language interaction in psychological interviews proposed in this invention is applied to a psychological interview training system, such as... Figure 3 The psychological interview training system is used to organize the language interaction process between trainees and virtual interviewees in a controlled training environment, and to perform structured processing and analysis of the language data during the interaction process.

[0116] In this embodiment, trainees participate in psychological interview training via voice input, and their voice signals are first collected and processed by a voice recognition module. The voice recognition module performs voice activity detection and voice recognition processing on the trainee's continuous voice input during the interview, generating text data corresponding to the trainee's speech content, and transmitting this text data to the psychological interview subject system as the basis for subsequent analysis. The psychological interview subject system includes a semantic intent recognition module and a role and scenario model module. The semantic intent recognition module performs semantic parsing on the text data to extract semantic intent information contained in the trainee's speech and converts this semantic intent information into a dialogue feature representation.

[0117] The role and scenario model module provides a unified background setting and state parameters for virtual interviewees. These state parameters include at least the virtual character's psychological state, emotional intensity, and interview stage markers. Based on this role and scenario model, virtual interviewees maintain consistent response characteristics during training, ensuring that different trainees face consistent interviewee conditions during testing. The virtual interviewees generate corresponding interview responses based on dialogue feature information from the psychological interviewee system and output corresponding voice or text feedback through the natural language response generation module.

[0118] During the interview, the psychological interview training system simultaneously tracks and analyzes the dialogue state between the trainee and the virtual interviewee. Specifically, the dialogue state tracking module constructs a dialogue state sequence based on semantic intent recognition results and continuously updates the transition relationships between adjacent states to form dialogue analysis results that reflect the evolution of the interview process. These dialogue analysis results, along with the dialogue logs from the interview process, are transmitted to the structured interactive analysis module.

[0119] The structured interaction analysis module identifies key interaction nodes during the interview process based on dialogue analysis results, and classifies the trainees' language interaction behaviors in different interview stages to generate a structured interaction feature set. This structured interaction feature set is further used by the skills dimension assessment module to quantitatively analyze the trainees' interview performance according to preset skills dimension mapping rules, outputting corresponding skills dimension scores, performance summaries, and improvement suggestions.

[0120] Through the above system architecture, this embodiment realizes the unified collection, analysis and structured expression of language interaction data during psychological interview training, enabling trainees' interview behavior to be evaluated comparably while maintaining consistency with the interviewees, and providing basic data support for subsequent training feedback and ability analysis.

[0121] like Figure 4As shown, this invention provides a structured analysis system for language interaction in psychological interviews. The system includes: a data acquisition unit, used to acquire language interaction data formed between trainees and preset virtual interviewees during psychological interview training, and to perform speech recognition processing on the language interaction data to generate corresponding text data; an intent recognition unit, used to perform semantic intent recognition processing based on the text data, extract semantic intent information of trainees' language interaction, and construct a dialogue state sequence during the interview process based on the semantic intent information; a feature extraction unit, used to perform dialogue state tracking and dialogue analysis processing based on the dialogue state sequence, and extract dialogue feature information representing the evolution of trainees' interactive behaviors during the interview; a structured processing unit, used to perform structured interaction analysis processing on the language interaction during the interview process based on the dialogue feature information, identify key interaction nodes and classify trainees' language interaction behaviors to generate a set of structured interaction features; and an output unit, used to generate skill dimension feature data representing trainees' interview skill performance according to preset skill dimension mapping rules based on the set of structured interaction features, and output the corresponding skill dimension score results.

[0122] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor A01, a network interface A02, memory (not shown), and a database (not shown) connected via a system bus. The processor A01 provides computing and control capabilities. The memory includes internal memory A03 and a non-volatile storage medium A04. The non-volatile storage medium A04 stores an operating system B01, a computer program B02, and a database (not shown). The internal memory A03 provides an environment for the operation of the operating system B01 and the computer program B02 stored in the non-volatile storage medium A04. The network interface A02 is used for communication with external terminals via a network connection. When the computer program B02 is executed by the processor A01, it implements a structured analysis method for psychological interview language interaction.

[0123] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a microcontroller, chip, or processor to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0124] The optional embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the embodiments of the present invention are not limited to the specific details described above. Within the scope of the technical concept of the embodiments of the present invention, various simple modifications can be made to the technical solutions of the embodiments of the present invention, and these simple modifications all fall within the protection scope of the embodiments of the present invention. It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the embodiments of the present invention will not further describe the various possible combinations.

[0125] Furthermore, various different embodiments of the present invention can be combined in any way, as long as they do not violate the spirit of the embodiments of the present invention, they should also be regarded as the content disclosed by the embodiments of the present invention.

Claims

1. A structured analysis method for language interaction in psychological interviews, characterized in that, The method includes: Acquire language interaction data formed between trainees and preset virtual interviewees during psychological interview training, and perform speech recognition processing on the language interaction data to generate corresponding text data; Based on the text data, semantic intent recognition processing is performed to extract semantic intent information of the trainees' language interactions, and a dialogue state sequence during the interview process is constructed based on the semantic intent information. Based on the dialogue state sequence, dialogue state tracking and dialogue analysis are performed to extract dialogue feature information that represents the evolution of trainees' interactive behaviors during the interview process. Based on the dialogue feature information, structured interaction analysis is performed on the language interaction during the interview process to identify key interaction nodes and classify the language interaction behavior of trainees to generate a set of structured interaction features. Based on the structured interaction feature set, skill dimension feature data representing the trainees' interview skill performance is generated according to the preset skill dimension mapping rules, and the corresponding skill dimension score results are output.

2. The structured analysis method for language interaction in psychological interviews according to claim 1, characterized in that, Acquire language interaction data between trainees and preset virtual interviewees during psychological interview training, and perform speech recognition processing on the language interaction data to generate corresponding text data, including: During the psychological interview training process, the trainees' voice input signals are collected in real time, and raw voice data containing timestamp information is formed based on the voice input signals. The original speech data is subjected to speech activity detection processing to identify valid speech segments and eliminate non-speech segments, thereby generating valid speech data for speech recognition processing; Based on the valid speech data, speech recognition processing is performed to generate text data corresponding to the student's speech input, and the text data is associated with the corresponding timestamp information.

3. The structured analysis method for language interaction in psychological interviews according to claim 1, characterized in that, Based on the text data, semantic intent recognition processing is performed to extract semantic intent information from the learner's language interactions, and a dialogue state sequence during the interview process is constructed based on the semantic intent information, including: The text data is subjected to semantic parsing to extract semantic intent features that characterize the type of student's speaking intent and interaction method; Based on the semantic intent features, the trainees' speech content is mapped to preset semantic intent tags to form a semantic intent tag sequence corresponding to each speech moment during the interview. Based on the semantic intent tag sequence and combined with the temporal order relationship corresponding to the text data, a dialogue state sequence is constructed to describe the evolution of language interaction during the interview. The dialogue state sequence is used as input data for subsequent dialogue state tracking and dialogue analysis.

4. The structured analysis method for language interaction in psychological interviews according to claim 3, characterized in that, The text data is subjected to semantic parsing to extract semantic intent features that characterize the type of student's speaking intent and interaction method, including: Based on the preset interview interaction unit division rules, the text data is segmented into speech units to obtain multiple interactive statement units divided according to semantic integrity. For each interactive statement unit, its pragmatic position features in the interview interaction process are extracted. The pragmatic position features are used to characterize the interaction stage of the interactive statement unit in the current interview round. Based on the pragmatic position features of the interactive statement units, and combined with the sentence structure features and response orientation features of the corresponding statement units, a semantic intent feature vector is constructed to describe the interactive mode of student speech.

5. The structured analysis method for language interaction in psychological interviews according to claim 1, characterized in that, Based on the dialogue state sequence, dialogue state tracking and dialogue analysis are performed to extract dialogue feature information that characterizes the evolution of trainees' interactive behaviors during the interview, including: Based on the dialogue state sequence, a state transition structure is constructed to represent the transition relationship between adjacent interactive states during the interview process, so as to describe the evolution path of the learner's language interaction behavior during the interview process. The state transition structure is continuously updated to generate state evolution features that characterize the changing trends of trainees' interactive behaviors in different interview stages. Based on the state evolution features, behavioral evolution features are extracted to characterize the stability of the trainees' interactions, the frequency of interaction adjustments, and the changes in interaction strategies during the interview process, as dialogue feature information.

6. The structured analysis method for language interaction in psychological interviews according to claim 5, characterized in that, Based on the aforementioned state evolution features, behavioral evolution features are extracted to characterize the stability of the trainees' interactions, the frequency of interaction adjustments, and the changes in interaction strategies during the interview process, as dialogue feature information, including: Based on the state evolution characteristics, the distribution of state retention duration between adjacent interview states is calculated to generate stability feature parameters for characterizing the stability of trainees' interactions. Based on the state evolution characteristics, the frequency of state transitions during the interview process and their distribution on the interview time axis are statistically analyzed to generate adjustment frequency characteristic parameters to characterize the frequency of trainees' interaction adjustments. Based on the state evolution characteristics, the interaction pattern change segments formed by multiple consecutive state transitions during the interview are identified, and strategy change feature parameters are generated based on the interaction pattern change segments to characterize the changes in the trainees' interaction strategies. The stability feature parameter, adjustment frequency feature parameter, and strategy change feature parameter are combined as behavioral evolution features.

7. The structured analysis method for language interaction in psychological interviews according to claim 1, characterized in that, Based on the aforementioned dialogue feature information, structured interaction analysis is performed on the language interactions during the interview process to identify key interaction nodes and classify the learners' language interaction behaviors, generating a set of structured interaction features, including: Based on the dialogue feature information, an interaction structure representation is constructed to describe the evolution of language interaction during the interview, wherein, The interaction structure includes at least state evolution relationships, speaking order relationships, and interaction response relationships; Based on the interaction structure representation, according to the preset node determination rules, the interaction positions that meet the conditions of state transition change, interaction strategy change or significant change in interaction stability during the interview process are determined, and the corresponding key interaction nodes are identified. For each of the key interaction nodes, combined with the dialogue feature information before and after the key interaction node, node behavior features are extracted to characterize the language interaction mode of the trainees in the corresponding interview stage. Based on the node behavior characteristics, the language interaction behaviors of trainees during the interview process are classified according to the preset behavior classification rules, and corresponding behavior category labels are generated. The key interaction nodes, the node behavior features, and the behavior category identifiers are associated to form a structured set of interaction features for describing the interview language interaction process.

8. The structured analysis method for language interaction in psychological interviews according to claim 7, characterized in that, Based on the structured interaction feature set, and according to preset skill dimension mapping rules, skill dimension feature data representing trainees' interview skill performance is generated, including: Based on the structured interaction feature set, extract the interaction feature subsets corresponding to different interview skill dimensions. The interview skill dimensions include at least the ability to guide questions, the coherence of responses, and the ability to control interactions. For each of the interview skill dimensions, based on the corresponding subset of interaction features and according to the preset dimension mapping relationship, the dimension feature parameters reflecting the interaction performance characteristics of trainees under the interview skill dimensions are calculated. The feature parameters of each dimension are normalized within the dimension to generate skill dimension feature data that characterizes the relative performance level of trainees under different interview skill dimensions.

9. The structured analysis method for language interaction in psychological interviews according to claim 8, characterized in that, The output rules for the corresponding skill dimension scoring results are as follows: Based on the skill dimension feature data, the original score value of each interview skill dimension is calculated according to the preset score mapping relationship corresponding to each interview skill dimension. Cross-dimensional alignment is performed on each of the original rating values ​​to eliminate differences in rating scales between different interview skill dimensions and generate standardized rating results that are comparable across dimensions. Based on the standardized scoring results, a skill scoring vector is constructed to describe the relative performance distribution of trainees across different interview skill dimensions; The skill score vector is associated with the corresponding interview skill dimension identifier to form a structured skill dimension score result, which is then output according to a preset output format.

10. A structured analysis system for language interaction in psychological interviews, characterized in that, The system includes: The data acquisition unit is used to acquire language interaction data formed between trainees and preset virtual interviewees during psychological interview training, and to perform speech recognition processing on the language interaction data to generate corresponding text data. The intent recognition unit is used to perform semantic intent recognition processing based on the text data, extract semantic intent information of the trainee's language interaction, and construct a dialogue state sequence during the interview based on the semantic intent information. The feature extraction unit is used to perform dialogue state tracking and dialogue analysis processing based on the dialogue state sequence, and extract dialogue feature information that represents the evolution of trainees' interactive behavior during the interview process; The structured processing unit is used to perform structured interaction analysis processing on the language interaction during the interview based on the dialogue feature information, identify key interaction nodes and classify the language interaction behavior of trainees, and generate a set of structured interaction features. The output unit is used to generate skill dimension feature data representing the trainees' interview skill performance based on the structured interaction feature set and according to the preset skill dimension mapping rules, and output the corresponding skill dimension score results.

Citation Information

Patent Citations

  • Conversation state tracking model optimization system and method in electric power service scene

    CN120407725A

  • Dynamic intention chain modeling and causal cleaning method and system under multi-round dialogue scene

    CN121146044A