Psychological interview language interaction structured analysis method and system

By constructing a dialogue state sequence through speech recognition and semantic intent recognition, identifying key interaction nodes, and generating a set of structured interaction features, the problem of the difficulty in structurally representing language interaction in psychological interview training is solved, and objective evaluation and scoring are achieved.

CN121528217AActive Publication Date: 2026-02-13CHENGDU IND VOCATIONAL TECHN COLLEGE
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202610049383.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-15
Publication Date
2026-02-13
Estimated Expiration
2046-01-15

AI Technical Summary

Technical Problem

Existing psychological interview training systems struggle to structure the language interaction process under the premise of trainees' free expression, resulting in inconsistent and unobjective assessment results that fail to meet the needs of large-scale and standardized training.

Method used

By acquiring language interaction data, speech recognition and semantic intent recognition are performed to construct a dialogue state sequence, extract interaction feature information, identify key interaction nodes and generate a structured interaction feature set, and output the scoring results by combining preset skill dimension mapping rules.

Benefits of technology

It achieves systematic processing of the interview process, transforms language interaction into structured features, and can objectively analyze and evaluate trainees' interview skills performance, thus solving the problem of inconsistent evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121528217A_ABST
    Figure CN121528217A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a psychological interview language interaction structured analysis method and system, and belongs to the technical field of natural language processing. The method comprises the following steps: performing voice recognition processing on language interaction data to generate corresponding text data; semantic intention information of student language interaction is extracted, and a dialogue state sequence in the interview process is constructed based on the semantic intention information; dialogue feature information representing the trainee interaction behavior evolution relation in the interview process is extracted; performing structured interaction analysis processing on language interaction in the interview process, identifying key interaction nodes, performing behavior classification on language interaction behaviors of students, and generating a structured interaction feature set; and according to a preset skill dimension mapping rule, generating skill dimension feature data representing trainee interview skill performance, and outputting a corresponding skill dimension scoring result. According to the scheme of the invention, structured representation and comparability evaluation of language interaction behaviors of students in psychological interview training are realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of natural language processing, in particular to a psychological interview language interaction structured analysis method and system. BACKGROUND

[0002] In the psychological interview training and examination scene, the trainee usually needs to carry out open dialogue with the interviewee through language interaction to train the interview skills such as questioning, responding, guiding and interaction control. The interview process has high freedom, and the language expression content, interaction rhythm and response strategy of the trainee dynamically change with the interview process, so that the interview process presents obvious unstructured characteristics.

[0003] The existing psychological interview training system or language interaction training platform usually focuses on providing a dialogue environment or a preset interaction script to guide the trainee to complete the established training process. However, in actual application, the interview training often needs to allow the trainee to express freely, and the preset script is difficult to cover the complex and changeable interaction situation in the real interview process, resulting in deviation between the training process and the actual application scene. At the same time, a large amount of language interaction data generated in the free dialogue mode lacks a unified structured representation form, so that the subsequent analysis and evaluation of the training process rely on manual playback or subjective judgment, and it is difficult to realize the scale and objectivity.

[0004] On the other hand, part of the existing technology tries to analyze the emotion or keyword of the interview language to assist in evaluating the performance of the trainee, but this kind of method usually only focuses on single speech or local semantic features, and cannot combine the context relationship and interaction evolution law in the interview process, so it is difficult to reflect the dynamic change characteristics of the trainee's interaction behavior in the interview process. In addition, the interviewees, dialogue situations and interaction stimulation conditions faced by different trainees in the interview process often differ, so that the analysis results based on language content lack alignability and comparability, further limiting the objectivity and consistency of the evaluation results.

[0005] Therefore, in the psychological interview training scene, a technical scheme is needed to systematically process the language interaction formed in the interview process without limiting the free expression of the trainee, to convert the originally unstructured interview language interaction process into an interaction feature representation with clear structure relationship, and to realize the objective analysis and evaluation of the trainee's interview skill performance based on the structured interaction features, to meet the actual needs of large-scale training and standardized examination. SUMMARY

[0006] The purpose of the embodiments of the present application is to provide a psychological interview language interaction structured analysis method and system to at least solve the problem that the interview process under the condition of free language interaction is difficult to be converted into structured and comparable analysis data.

[0007] To achieve the above object, the present application provides a psychological interview language interaction structured analysis method and system. Language interaction data formed between a trainee and a preset virtual interviewee during psychological interview training is acquired, and speech recognition processing is performed on the language interaction data to generate corresponding text data. Semantic intention recognition processing is performed based on the text data to extract semantic intention information of the trainee's language interaction, and a dialogue state sequence in the interview process is constructed based on the semantic intention information. Dialogue state tracking and dialogue analysis processing are performed based on the dialogue state sequence to extract dialogue feature information representing the evolution relationship of the trainee's interactive behavior in the interview process. Structured interaction analysis processing is performed on the language interaction in the interview process based on the dialogue feature information to identify key interaction nodes and classify the trainee's language interaction behavior, and a structured interaction feature set is generated. Skill dimension feature data representing the trainee's interview skill performance is generated based on the structured interaction feature set according to a preset skill dimension mapping rule, and corresponding skill dimension score results are output.

[0008] Optionally, the language interaction data formed between the trainee and the preset virtual interviewee during the psychological interview training is acquired, and speech recognition processing is performed on the language interaction data to generate corresponding text data, including: during the psychological interview training, real-time collection of the trainee's speech input signal is performed, and original speech data containing timestamp information is formed based on the speech input signal; speech activity detection processing is performed on the original speech data to identify valid speech segments and eliminate non-speech segments, and valid speech data for speech recognition processing is generated; speech recognition processing is performed based on the valid speech data to generate text data corresponding to the trainee's speech input content, and the text data is associated with corresponding timestamp information.

[0009] Optionally, the semantic intention recognition processing is performed based on the text data to extract semantic intention information of the trainee's language interaction, and the dialogue state sequence in the interview process is constructed based on the semantic intention information, including: semantic analysis processing is performed on the text data to extract semantic intention features representing the trainee's speech intention type and interaction mode; based on the semantic intention features, the trainee's speech content is mapped to a preset semantic intention label to form a semantic intention label sequence corresponding to each speech time in the interview process; based on the semantic intention label sequence, the dialogue state sequence describing the evolution process of the language interaction in the interview process is constructed in combination with the time sequence relationship of the text data; the dialogue state sequence is taken as input data for subsequent dialogue state tracking processing and dialogue analysis processing.

[0010] Optionally, the text data is subjected to semantic analysis processing to extract semantic intention features for representing the intention type and interaction mode of the student's speech, including: based on a preset interview interaction unit division rule, the text data is subjected to speech unit segmentation to obtain a plurality of interactive sentence units divided according to semantic integrity; for each interactive sentence unit, its pragmatic position feature in the interview interaction process is extracted, the pragmatic position feature is used to represent the interaction stage of the interactive sentence unit in the current interview round; based on the pragmatic position feature of the interactive sentence unit, combined with the sentence structure feature and response pointing feature of the corresponding sentence unit, a semantic intention feature vector for describing the student's speech interaction mode is constructed.

[0011] Optionally, based on the dialogue state sequence, dialogue state tracking and dialogue analysis processing are performed to extract dialogue feature information representing the evolution relationship of the student's interactive behavior in the interview process, including: based on the dialogue state sequence, a state transition structure for representing the transition relationship between adjacent interaction states in the interview process is constructed to describe the evolution path of the student's language interactive behavior in the interview process; the state transition structure is subjected to continuous state updating processing to generate a state evolution feature for representing the change trend of the student's interactive behavior in different interview stages; based on the state evolution feature, behavior evolution features for representing the interaction stability, interaction adjustment frequency and interaction strategy change of the student in the interview process are extracted as dialogue feature information.

[0012] Optionally, based on the state evolution feature, behavior evolution features for representing the interaction stability, interaction adjustment frequency and interaction strategy change of the student in the interview process are extracted as dialogue feature information, including: based on the state evolution feature, the state holding time length distribution between adjacent interview states is calculated to generate a stability feature parameter for representing the interaction stability of the student; based on the state evolution feature, the frequency of state transition occurrence in the interview process and its distribution on the interview time axis are counted to generate an adjustment frequency feature parameter for representing the interaction adjustment frequency of the student; based on the state evolution feature, an interaction mode change segment formed by a plurality of continuous state transitions in the interview process is identified, and based on the interaction mode change segment, a strategy change feature parameter for representing the interaction strategy change of the student is generated; the stability feature parameter, adjustment frequency feature parameter and strategy change feature parameter are combined as the behavior evolution feature.

[0013] Optionally, based on the dialogue feature information, a structured interaction analysis processing is performed on the language interaction in the interview process, key interaction nodes are identified, and the language interaction behavior of the student is classified, and a structured interaction feature set is generated, including: based on the dialogue feature information, an interaction structure representation for describing the evolution process of the language interaction in the interview process is constructed, wherein the interaction structure representation at least includes state evolution relationship, speaking order relationship and interaction response relationship; based on the interaction structure representation, according to a preset node determination rule, the interaction positions that meet the state transition mutation, interaction strategy change or interaction stability significant change conditions in the interview process are determined, and the corresponding key interaction nodes are identified; for each key interaction node, the node behavior features for representing the language interaction mode of the student in the corresponding interview stage are extracted in combination with the dialogue feature information before and after the key interaction node; based on the node behavior features, according to a preset behavior classification rule, the language interaction behavior of the student in the interview process is classified, and a corresponding behavior category identifier is generated; the key interaction nodes, the node behavior features and the behavior category identifier are associated to form a structured interaction feature set for describing the interview language interaction process.

[0014] Optionally, based on the structured interaction feature set, according to a preset skill dimension mapping rule, skill dimension feature data representing the interview skill performance of the student is generated, including: based on the structured interaction feature set, an interaction feature subset corresponding to different interview skill dimensions is extracted, the interview skill dimensions at least including questioning guidance ability, response coherence and interaction regulation ability; for each interview skill dimension, based on the corresponding interaction feature subset, according to a preset dimension mapping relationship, a dimension feature parameter reflecting the interaction performance feature of the student in the interview skill dimension is calculated; each dimension feature parameter is processed by dimension internal normalization to generate skill dimension feature data for representing the relative performance level of the student in different interview skill dimensions.

[0015] Optionally, the output rule of the corresponding skill dimension score result is: based on the skill dimension feature data, according to the preset score mapping relationship corresponding to each interview skill dimension, the original score value of each interview skill dimension is calculated; cross-dimension alignment processing is performed on each original score value to eliminate the score scale difference between different interview skill dimensions, and a standardized score result for comparison between dimensions is generated; based on the standardized score result, a skill score vector for describing the relative performance distribution of the student in different interview skill dimensions is constructed; the skill score vector is associated with the corresponding interview skill dimension identifier to form a structured skill dimension score result, and is output according to a preset output format.

[0016] In a second aspect, the present application provides a psychological interview language interaction structured analysis system, comprising: an acquisition unit configured to acquire language interaction data formed between a trainee and a preset virtual interviewee during a psychological interview training process, and perform speech recognition processing on the language interaction data to generate corresponding text data; an intention recognition unit configured to perform semantic intention recognition processing based on the text data, extract semantic intention information of the trainee's language interaction, and construct a dialogue state sequence in the interview process based on the semantic intention information; a feature extraction unit configured to perform dialogue state tracking and dialogue analysis processing based on the dialogue state sequence, and extract dialogue feature information representing the evolution relationship of the trainee's interactive behavior in the interview process; a structured processing unit configured to perform structured interaction analysis processing on the language interaction in the interview process based on the dialogue feature information, identify key interaction nodes, and classify the trainee's language interaction behavior, to generate a structured interaction feature set; and an output unit configured to generate skill dimension feature data representing the trainee's interview skill performance according to a preset skill dimension mapping rule based on the structured interaction feature set, and output corresponding skill dimension score results.

[0017] Through the above technical solution, the present application realizes systematic processing of free language interaction data in the psychological interview training process, and converts the originally loose and unclear structured interview language process into structured interaction features with clear state evolution relationship and behavior category identification. On the one hand, through semantic intention recognition and dialogue state sequence construction, the trainee's language expression in the interview process is mapped into a continuously trackable dialogue state, so as to depict the dynamic evolution process of the interview interaction. On the other hand, through dialogue state tracking and structured interaction analysis, key interaction nodes in the interview process are identified and the language interaction behavior is classified, so that the trainee's interview behavior is converted from the original text form into a structured feature set that is quantifiable and comparable. On this basis, through a preset skill dimension mapping rule, the structured interaction features are further mapped into skill dimension feature data and the corresponding score results are output, so as to realize objective analysis and unified evaluation of the interview skill performance without limiting the trainee's free expression, and solve the problems of difficulty in structuring the language interaction and lack of consistency in the evaluation results in the existing psychological interview training.

[0018] Other features and advantages of the present application will be described in detail in the following detailed description. BRIEF DESCRIPTION OF DRAWINGS

[0019] The accompanying drawings are included to provide a further understanding of the present application and constitute a part of the specification, and are used together with the following detailed description to explain the present application, but do not constitute a limitation of the present application. In the drawings: Figure 1is a step flow chart of a psychological interview language interaction structured analysis method provided by an embodiment of the present application; Figure 2 is a detailed flow chart of step S20 of a psychological interview language interaction structured analysis method provided by an embodiment of the present application; Figure 3 is a training system architecture schematic diagram of a psychological interview language interaction structured analysis provided by an embodiment of the present application; Figure 4 is a system structure diagram of a psychological interview language interaction structured analysis system provided by an embodiment of the present application; Figure 5 is an internal structure diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0020] The specific embodiments of the present application are described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application.

[0021] As shown in Figure 1 , an embodiment of the present application provides a psychological interview language interaction structured analysis method, which comprises: Step S10: obtaining language interaction data formed between a trainee and a preset virtual interviewee in a psychological interview training process, and performing speech recognition processing on the language interaction data to generate corresponding text data.

[0022] Specifically, in the psychological interview training process, the speech input signal of the trainee is collected in real time, and original speech data containing timestamp information is formed based on the speech input signal; speech activity detection processing is performed on the original speech data to identify valid speech segments and eliminate non-speech segments, thereby generating valid speech data for speech recognition processing; speech recognition processing is performed based on the valid speech data to generate text data corresponding to the speech input content of the trainee, and the text data is associated with the corresponding timestamp information.

[0023] In the embodiment of the present application, after the psychological interview training starts, the speech input signal of the trainee is continuously collected by the training system. The speech input signal can be continuous speech stream data obtained in real time by a microphone or other speech collection device. In order to ensure the correspondence between the speech content and the interview process in the subsequent processing process, time mark information is added to the speech input signal at the same time as the speech input signal is collected, to form original speech data containing timestamp information. The timestamp information is used to represent the positional relationship of each speech segment on the interview time axis.

[0024] Further, voice activity detection (VAD) is performed on the original speech data to distinguish valid speech segments from non-speech segments in the speech signal. Specifically, by analyzing the energy variation, spectral distribution or other voice activity determination features of the original speech data, the speech segments containing the actual speech content of the interviewee are identified, and the background noise, silence segments or non-speech signals are removed. After this step, valid speech data containing only valid speech content is generated, thereby reducing the interference of invalid data on subsequent speech recognition processing.

[0025] After obtaining the valid speech data, speech recognition processing is performed on the valid speech data. The speech recognition processing can convert continuous speech signals into corresponding text representation forms according to a preset recognition model or recognition rule. In this process, each speech segment in the valid speech data is associated with its corresponding timestamp information, so that the generated text data not only contains semantic content, but also retains time positioning information consistent with the order of the interview.

[0026] Specifically, the output result of the speech recognition processing is a set of text data segments, each of which corresponds to a speech content of the interviewee in the interview process, and each of which corresponds to the timestamp information in the original speech data. In this way, a data structure associated with "text content—timestamp" can be formed to represent the distribution of language interaction content of the interviewee in the psychological interview training process over time. Finally, the text data associated with the timestamp information is used as input data for subsequent semantic intent recognition processing and dialogue state sequence construction processing.

[0027] In another possible implementation, to further enhance the alignability and traceability of the interview language interaction data in the time dimension, the embodiment introduces a speech boundary self-calibration mechanism to refine the time label of the text data during speech collection and speech recognition processing.

[0028] Specifically, after collecting the speech input signal of the interviewee in real time and forming the original speech data containing the timestamp information, in addition to performing the conventional voice activity detection processing, the speech energy variation and speech speed variation within the valid speech segment are also continuously monitored. Based on the monitoring result, the speech starting boundary point and the speech ending boundary point are further identified within each valid speech segment to distinguish multiple continuous speech units that may be contained in the same speech segment.

[0029] On this basis, the identified speech units are respectively re-assigned with fine-grained timestamp information, so that each speech unit corresponds to an independent time interval mark. Subsequently, when performing speech recognition processing, the valid speech data is segmented and recognized according to the speech unit boundaries, to generate text data segments corresponding to each speech unit. The text data segments are associated with their corresponding fine-grained timestamp information during the generation process, thereby forming a text-time mapping relationship accurate to the speech level.

[0030] Through this embodiment, the distribution of text data on the time axis can be made more fine-grained without changing the overall speech recognition process, which is conducive to the subsequent distinguishing processing of continuous speech in semantic intent recognition processing, and provides a more accurate time sequence basis for constructing a dialogue state sequence. This method is particularly suitable for interactive situations where the interviewee's speech speed changes significantly and continuous responses are frequent, and helps to improve the accuracy of positioning key interactive nodes in subsequent structured interactive analysis.

[0031] Step S20: performing semantic intent recognition processing based on the text data, extracting semantic intent information of the interviewee's language interaction, and constructing a dialogue state sequence in the interview process based on the semantic intent information.

[0032] Specifically, based on the above text data, in this embodiment, semantic intent recognition processing is performed on the language interaction content formed by the interviewee in the psychological interview training process to extract semantic intent information that can reflect the interviewee's speech purpose and interaction mode. The semantic intent recognition processing maps the interviewee's speech content into a preset semantic intent representation form by analyzing the sentence structure, semantic direction and context association relationship in the text data. On this basis, according to the sequential relationship of the text data on the interview time axis, the semantic intent information corresponding to the continuous speech is organized and associated to construct a dialogue state sequence for describing the evolution relationship of language interaction in the interview process. The dialogue state sequence is used to depict the interaction state change of the interviewee in different interview stages, and serves as the basic input data for subsequent dialogue state tracking, dialogue analysis and structured interactive analysis processing, thereby providing a unified data expression form for further structured processing of the interview language interaction process. Specifically, as Figure 2 Step S20 includes the following steps: Step S201: performing semantic analysis processing on the text data to extract semantic intent features representing the interviewee's speech intent type and interaction mode.

[0033] Specifically, based on preset interview interaction unit segmentation rules, the text data is segmented into speaking units to obtain multiple interactive statement units divided according to semantic integrity; for each interactive statement unit, its pragmatic position features in the interview interaction process are extracted, and the pragmatic position features are used to characterize the interaction stage of the interactive statement unit in the current interview round; based on the pragmatic position features of the interactive statement units, combined with the sentence structure features and response orientation features of the corresponding statement units, a semantic intent feature vector is constructed to describe the trainees' speaking interaction methods.

[0034] Specifically, after completing the segmentation of the speech units, this implementation method denotes each interactive statement unit as u. i and retain u i Round index r in the interview rounds i And the sequential index k within the round i Semantic parsing does not encode the entire text at once, but first processes the u... i Pragmatic position features, sentence structure features, and response orientation features are extracted, and then these three types of features are fused in the same feature space to form a semantic intent feature vector f. i The reason is that in interview training scenarios, the same statement may correspond to different interaction methods in different rounds, and the sentence structure and response direction can provide reusable clues to determine whether the speech is asking, clarifying, restating, empathizing, or advancing.

[0035] In a specific implementation, pragmatic positional feature p i Can be indexed by round r i Sequential index k i and the unit u of the previous statement i-1 The time interval Δt i Common construction, for example, (r) i ,k i ,Δt i Mapped to vector p i ∈R dp ; Sentence structure features i A vector s can be constructed from syntactic dependency relation statistics, interrogative word triggering markers, negation structure markers, subject-verb-object skeleton lengths, etc. i ∈R ds ; response pointing to feature g i Used to describe u i The connection between the statements made by the interviewees in the previous virtual interview round or the reference to a certain topic in the preceding text can be resolved through the results of referential analysis, the location of keyword references, and the relationship between the statements made by the interviewees in the previous round. ri The matching scores are constructed as a vector. Then, the three types of features are fused to obtain the semantic intent feature vector: ; Among them, f i Represents the interactive statement unit u i The semantic intent feature vector; p i Represents the pragmatic position feature vector; s i Represents the feature vector of sentence structure; g i W indicates that the response points to the feature vector; p W s W g p respectively i s i g i A parameter matrix mapped to a uniform dimension; b f For the bias vector; [p i s i g i ] indicates vector concatenation; W c b c Here, σ represents the parameter matrix and bias vector used to generate the gated components; σ(·) represents the Sigmoid nonlinear function; tanh(·) represents the hyperbolic tangent function; and ⊙ represents element-wise multiplication. Through this fusion method, pragmatic location, sentence structure, and response orientation can form mutually constraining relationships within the same vector, avoiding intent drift caused by relying solely on content vocabulary. The final output f... i This will serve as input for subsequent semantic intent label mapping.

[0036] Step S202: Based on the semantic intent features, map the trainee's speech content to preset semantic intent tags to form a semantic intent tag sequence corresponding to each speaking moment during the interview.

[0037] Specifically, the preset semantic intent tag set is denoted as L={l1,l2,…,l M} where M is the preset number of labels. The label content can cover common interaction intent types and interaction methods in interview training, such as open-ended questions, closed-ended questions, clarification and follow-up questions, restatement and confirmation, emotional transition, summary and advancement, boundary setting, etc. The mapping process does not use a hard threshold for single judgment, but first calculates the posterior probability distribution of each label, and then obtains the final label according to the preset selection rules, thereby providing traceable uncertainty information for the subsequent construction of dialogue state sequences.

[0038] In one specific implementation, the f output in step S201 is... i The input is fed into the intent classifier to obtain a score z for each label. i ∈R M Furthermore, prior category weights and a temperature coefficient are introduced to accommodate the differences in the frequency of different labels and the differences in discrimination difficulty in the training corpus. Subsequently, the posterior probability π is obtained through Softmax. i:

[0039] Where, π{ i,m} represents the interactive statement unit u i Corresponding tag l m The posterior probability of z{ i,m} indicates the label l m Unnormalized scoring; W o b is the output layer parameter matrix; o α is the output layer bias vector; m For tag l m The prior bias term is used to reflect the preset label frequency or importance constraints; τ is a temperature coefficient used to adjust the sharpness of the probability distribution; M is the total number of labels. Then, labels y are obtained according to the preset label selection rules. i The preset label selection rules may include, but are not limited to: selecting the label corresponding to the maximum posterior probability; or, when the difference between the maximum posterior probability and the second-largest posterior probability is less than a preset difference threshold, retaining the candidate label set and further judging it in the subsequent state sequence construction. Finally, each interactive statement unit u i The tag y i Arranged chronologically, forming a semantic intent tag sequence Y=[y1,y2,…,y N ], where N is the number of interactive statement units. The semantic intent tag sequence Y and each u i The timestamp information is kept one-to-one, providing discrete observation sequences for subsequent modeling of state transitions and evolutionary relationships.

[0040] Step S203: Based on the semantic intent tag sequence and combined with the temporal order relationship corresponding to the text data, construct a dialogue state sequence to describe the evolution of language interaction during the interview.

[0041] Specifically, in this implementation, the dialogue state sequence is defined as S=[s1,s2,…,s…]. N ], where s i Used to characterize the interactive statement unit u i The interview interaction status at the location. Compared to only putting y i As different from other states, this implementation constructs the state as a combination of "label-driven + time-constrained + contextual memory," so that the state can both reflect the current intention and inherit from previous evolutions. To this end, the label y is first... i Mapped to embedding vector e i ∈R de And introduce the time increment Δt corresponding to the time sequence relationship. i and round boundary marker q iThis is used to distinguish state transitions across rounds. Subsequently, a recursive method is used to generate the state vector h. i and with h i As a dialogue state i The expression .

[0042] In one specific implementation, state recursion can employ a gated update mechanism, explicitly introducing a transition constraint matrix A to limit unreasonable intentional jumps, thereby ensuring that state evolution conforms to the basic sequential logic of the interview interaction. The state update can be represented as:

[0043] Among them, h i This represents the state vector corresponding to the i-th interactive statement unit; e i The semantic intent label y i The embedding vector; φ(Δt) i ) represents the time increment Δt i The time encoding function mapped to a vector; q i Represents the round boundary marker vector, used to indicate u i Is this the start of a new round? i-1} represents the previous state vector; U e U t U q U h b is the parameter matrix; h ρ is the bias vector; ρ(·) represents the preset nonlinear combination function, which can be tanh or a gated combination form; A y{i-1},yi This indicates that the transition constraint matrix A is formed by the previous label y. i-1 To the current tag y i The transition term is used to express the preset permissible transition relationship; λ is a preset coefficient used to adjust the strength of the transition constraint influence. Through the above construction method, the state vector h... i Simultaneously, it absorbs current tag information, time sequence information, round boundary information, and historical memory information, and applies structural constraints to tag jumps through A. Finally, it will... i Defined as h i We obtain the dialogue state sequence S and maintain the correspondence between S, the semantic intent label sequence Y, and the timestamp information, so as to perform dialogue state tracking and dialogue analysis in subsequent steps.

[0044] Step S204: Use the dialogue state sequence as input data for subsequent dialogue state tracking and dialogue analysis.

[0045] Specifically, the dialogue state sequence S, the semantic intent label sequence Y, and the timestamp information T=[t1,t2,…,t] of the text data are combined. NThey are bound together to form a unified data object D. Where, (s i ,y i ,t i The data object D is represented by a triplet across the state, label, and time layers of the same interactive statement unit (UI). Subsequently, based on the windowing requirements of subsequent processing, the data object D is fragmented. For example, if a sliding window approach is used for subsequent dialogue state tracking, then window fragment D is constructed with a preset window length K. k Consistency in continuous tracking is achieved by using overlapping sections shared between adjacent windows. If subsequent dialogue analysis requires aggregation by round, it is based on the round boundary marker q. i The process involves dividing D into multiple round sets, with each round set maintaining a monotonically increasing timestamp to avoid state drift caused by mixing across rounds. Finally, the window segment D is... k Alternatively, the set of rounds can be used as input data for subsequent dialogue state tracking and dialogue analysis, and y can be retained in the input data. i With s i This establishes a correspondence so that when identifying key interaction nodes and classifying behaviors in subsequent processes, it is possible to trace back to the specific intent label and corresponding text fragment, thus achieving a data loop throughout the entire process.

[0046] Step S30: Based on the dialogue state sequence, perform dialogue state tracking and dialogue analysis processing to extract dialogue feature information that represents the evolution of trainees' interactive behaviors during the interview.

[0047] Specifically, based on the dialogue state sequence, a state transition structure is constructed to represent the transition relationship between adjacent interactive states during the interview, thereby describing the evolution path of the learner's language interaction behavior during the interview. The state transition structure is continuously updated to generate state evolution features that characterize the changing trends of the learner's interaction behavior in different interview stages. Based on the state evolution features, behavioral evolution features that characterize the learner's interaction stability, interaction adjustment frequency, and interaction strategy changes during the interview are extracted as dialogue feature information.

[0048] Furthermore, based on the state evolution features, behavioral evolution features are extracted to characterize the stability of the trainees' interactions, the frequency of interaction adjustments, and the changes in their interaction strategies during the interview, serving as dialogue feature information. This includes: calculating the distribution of state duration between adjacent interview states based on the state evolution features to generate stability feature parameters characterizing the stability of the trainees' interactions; statistically analyzing the frequency of state transitions during the interview and their distribution along the interview timeline based on the state evolution features to generate adjustment frequency feature parameters characterizing the frequency of the trainees' interaction adjustments; identifying interaction pattern change segments formed by multiple consecutive state transitions during the interview based on the state evolution features, and generating strategy change feature parameters characterizing the changes in the trainees' interaction strategies based on these interaction pattern change segments; and combining the stability feature parameters, adjustment frequency feature parameters, and strategy change feature parameters as behavioral evolution features.

[0049] In this embodiment of the invention, based on the dialogue state sequence, dialogue state tracking and dialogue analysis are performed to extract dialogue feature information characterizing the evolution of trainees' interactive behaviors during the interview. In this embodiment, the dialogue state sequence is represented as...

[0050] in, Indicates the relationship with the first Each interactive statement unit corresponds to a dialogue state vector, and each dialogue state vector All are associated with the corresponding timestamp Maintaining relevance. By analyzing the continuous changes in adjacent dialogue states, the evolutionary path of trainees' language interaction during the interview process can be depicted.

[0051] Therefore, a state transition structure is first constructed based on the differences between adjacent dialogue states. For any two adjacent dialogue states... and By calculating the magnitude of its change in the state space and combining it with the corresponding time interval information, the state transition weights are defined. , used to characterize from state Evolution to state The transfer intensity is calculated as follows:

[0052] in, This represents the difference vector between adjacent dialogue state vectors. This represents a pre-defined positive semi-definite weight matrix, used to assign different weights to the changes in each dimension of the state difference vector. This indicates the time interval between adjacent interactive statement units. This represents a pre-defined non-zero stable term. Using the above method, a state transition weight sequence reflecting the continuous changes in language interaction during the interview can be obtained. It is used to describe the evolution trajectory of student interaction behavior on the timeline.

[0053] Building upon this, to characterize the overall interaction trends at different stages of the interview process, the state transition weight sequence is continuously aggregated to generate state evolution features. Specifically, a sliding window mechanism is introduced to weight and aggregate multiple consecutive state transition weights to obtain a state evolution intensity sequence. The calculation method is as follows:

[0054] in, This indicates the preset window length. This represents the normalized coefficient of the state transition weights within the window. This represents the decay parameter, used to enhance attention to changes in the state near the current location. The state evolution intensity sequence calculated above can reflect the overall trend of language interaction during the interview process, from stable to changing.

[0055] Furthermore, based on the state evolution intensity sequence, behavioral evolution features are extracted to characterize the stability of the trainees' interactions, the frequency of interaction adjustments, and the changes in their interaction strategies during the interview. To this end, the state evolution intensity sequence is compared with a preset threshold. Compare and generate evolutionary event indicator sequences. And based on this, behavioral evolution characteristic parameters are calculated. Specifically, stability characteristic parameters... Adjusting frequency characteristic parameters and strategy change characteristic parameters They are defined as follows:

[0056]

[0057] in, Indicates the first The length of a continuous stable segment Indicates the number of stable segments; Indicates the first The time position corresponding to the significant change in the state of the next interaction. Indicates the number of events that show significant changes; Indicates the first The number of consecutive change events in a change cluster This indicates the number of variable clusters. Finally, the... , and The combined data form a behavioral evolution feature vector, which is used to characterize the overall evolutionary characteristics of trainees' language interaction behavior during the interview process, and serves as input data for subsequent structured interaction analysis.

[0058] Step S40: Based on the dialogue feature information, perform structured interaction analysis processing on the language interaction during the interview, identify key interaction nodes, classify the language interaction behavior of trainees, and generate a set of structured interaction features.

[0059] Specifically, based on the dialogue feature information, an interaction structure representation is constructed to describe the evolution of language interaction during the interview. This interaction structure representation includes at least state evolution relationships, speaking order relationships, and interaction response relationships. Based on the interaction structure representation, according to preset node determination rules, interaction positions that satisfy conditions such as abrupt state transitions, changes in interaction strategies, or significant changes in interaction stability are determined, and corresponding key interaction nodes are identified. For each key interaction node, combined with the dialogue feature information before and after the key interaction node, node behavioral features are extracted to characterize the trainee's language interaction style in the corresponding interview stage. Based on the node behavioral features, according to preset behavior classification rules, the trainee's language interaction behavior during the interview is classified, generating corresponding behavior category identifiers. The key interaction nodes, the node behavioral features, and the behavior category identifiers are associated to form a structured interaction feature set describing the interview language interaction process.

[0060] In this embodiment of the invention, based on the dialogue feature information, structured interaction analysis is performed on the language interactions during the interview process to identify key interaction nodes and classify the trainees' language interaction behaviors, generating a set of structured interaction features. In this embodiment, the dialogue feature information is obtained from the aforementioned dialogue state tracking and dialogue analysis processing, and includes at least a dialogue state sequence. State transition weight sequence State evolution intensity sequence and behavioral evolution feature vectors Furthermore, all of the above features correspond to the timestamp sequence. Maintain a one-to-one correspondence.

[0061] Based on this, we first construct an interaction structure representation to describe the evolution of language interaction during the interview process. This interaction structure representation is defined as a weighted ordered structure:

[0062] Among them, the node set Each node in Corresponding to the first point in the interview process One interactive statement unit; edge set Used to describe the interaction relationships between nodes. For any adjacent nodes and In the interactive structure representation, directed edges are established. and the corresponding state transition weights The weight attribute of this edge reflects the intensity of state evolution between adjacent interaction positions. Thus, the interaction structure representation simultaneously includes state evolution relationships, speaking order relationships, and interaction response relationships implied by time order, thereby forming a structured representation of the interview language interaction process.

[0063] After obtaining the interaction structure representation, based on preset node determination rules, interaction positions that meet specific evolutionary conditions during the interview process are determined to identify key interaction nodes. Specifically, for each node... Taking into account its corresponding state transition weights and State evolution intensity and behavioral evolution feature vectors Construct a node saliency determination function based on the relevant components in the data. When node A node is identified as a key interaction node if it meets at least one of the following criteria: 1) Adjacent state transition weights at nodes A mutation occurs at that point, that is The change exceeds the preset threshold.

[0064] 2) Node The corresponding state evolution intensity A significant shift compared to the mean of the preceding and following intervals indicates a change in interaction stability.

[0065] 3) Nodes The behavioral evolution feature vectors within the given segment exhibit a clustering trend along the strategy change dimension, reflecting the phased adjustments of the interaction strategy. Using the above determination method, a set of key interaction nodes is obtained on the timeline:

[0066] This is used to characterize the locations where significant changes occur in language interaction during the interview. For each key interaction node, further combining the dialogue feature information before and after the key interaction node, node behavioral features are extracted to represent the trainee's language interaction style in the corresponding interview stage. Specifically, for any key interaction node... Centered on this, select a local interval on the timeline containing several preceding and following interactive statement units to form a node context set:

[0067] in, The radius of the preset context window is used. Within this context set, dialogue features such as state evolution strength, state transition weights, and semantic intent distribution are aggregated to form a node behavior feature vector. This is used to describe the characteristics of a learner's language interaction style near this key interaction node.

[0068] After obtaining the node behavior feature vectors, the language interaction behaviors of trainees during the interview process are classified based on preset behavior classification rules. The behavior classification rules are used to classify the node behavior feature vectors...

[0069] Mapped to discrete behavior category identifiers The behavior categories distinguish at least different interaction types, such as stable and continuous interaction, strategy-adjusting interaction, or mutation-response interaction. This approach assigns a unique behavior category identifier to each key interaction node, thereby providing a structured characterization of learners' language interaction behavior at the node level.

[0070] Finally, the set of key interaction nodes will be... The corresponding set of node behavior feature vectors and behavioral category identifier set By associating these features, a structured set of interactive characteristics can be formed to describe the language interaction process during the interviews.

[0071] The structured interaction feature set maintains a consistent data correspondence across the time, state, and behavior dimensions, and can serve as input data for subsequent interview skill dimension mapping and scoring processing, thereby enabling structured analysis and aligned representation of the interview language interaction process.

[0072] Step S50: Based on the structured interaction feature set, generate skill dimension feature data representing the trainees' interview skill performance according to the preset skill dimension mapping rules, and output the corresponding skill dimension score results.

[0073] Specifically, based on the structured interaction feature set, a subset of interaction features corresponding to different interview skill dimensions is extracted. The interview skill dimensions include at least the ability to guide questions, the coherence of responses, and the ability to control interactions. For each interview skill dimension, based on the corresponding subset of interaction features and according to a preset dimension mapping relationship, dimensional feature parameters reflecting the interactive performance characteristics of trainees under the interview skill dimension are calculated. The dimensional feature parameters are then normalized within the dimension to generate skill dimension feature data that characterizes the relative performance level of trainees under different interview skill dimensions.

[0074] Furthermore, the output rules for the corresponding skill dimension scoring results are as follows: Based on the skill dimension feature data, calculate the original score value of each interview skill dimension according to the preset scoring mapping relationship corresponding to each interview skill dimension; perform cross-dimensional alignment processing on each of the original score values ​​to eliminate the difference in scoring scale between different interview skill dimensions and generate standardized scoring results for comparability between dimensions; based on the standardized scoring results, construct a skill scoring vector to describe the relative performance distribution of trainees under different interview skill dimensions; associate the skill scoring vector with the corresponding interview skill dimension identifier to form a structured skill dimension scoring result, and output it according to the preset output format.

[0075] In this embodiment of the invention, based on the structured interaction feature set, skill dimension feature data representing trainees' interview skill performance is generated according to a preset skill dimension mapping rule, and the corresponding skill dimension score results are output. In this embodiment, the structured interaction feature set is represented as...

[0076] in, Indicates the first Key interaction nodes, This represents the node behavior feature vector corresponding to the key interaction node. This represents the corresponding behavioral category identifier. The structured interaction feature set maintains a consistent correspondence between the time and behavioral dimensions, supporting analysis at the interview skills level.

[0077] Based on this, firstly, based on a pre-defined set of interview skills dimensions:

[0078] The structured interaction feature set is decomposed, wherein the interview skill dimensions include at least the ability to guide questioning, the coherence of responses, and the ability to control interaction. For any interview skill dimension... Based on the preset skill dimension mapping rules, from the structured interaction feature set Extract a subset of interactive features related to the skill dimension, thereby achieving the ordered aggregation of structured interactive features at the skill dimension level.

[0079] For each interview skill dimension Based on the corresponding interaction feature subset The dimensional feature parameters reflecting the trainees' interactive performance characteristics under the interview skill dimension are calculated according to the preset dimensional mapping relationship. Specifically, the behavioral feature vectors of each node in the interactive feature subset are... We perform weighted aggregation based on the behavior category weights to obtain the original feature vector of the skill dimension. The calculation method is as follows:

[0080] in, Indicates the size of the subset of interactive features. Indicator and behavior category identifier The corresponding preset weight coefficients are used to reflect the differences in importance of different interactive behaviors within this skill dimension. Through the above aggregation method, the skill dimension feature parameters can comprehensively reflect the performance of multiple key interactive nodes within this dimension.

[0081] After obtaining the original dimensional feature vectors corresponding to each interview skill dimension, intra-dimensional normalization is performed on the dimensional feature parameters to eliminate the influence of differences in interview sample length or interaction frequency on the results, generating skill dimension feature data. Specifically, for the skill dimensions... The original feature vector Perform the following normalization process:

[0082] in, and These represent the skill dimensions. The mean vector and standard deviation vector within a preset sample range, This represents a pre-defined stable term. Through this processing, skill dimension feature data is obtained to characterize the relative performance levels of trainees across different interview skill dimensions. .

[0083] Furthermore, based on the skill dimension feature data, and according to the preset scoring mapping relationship corresponding to each interview skill dimension, the original score value of each interview skill dimension is calculated. Specifically, for each skill dimension... The corresponding skill dimension feature data Mapped to raw rating values The calculation method is as follows:

[0084] in, Indicating skill dimensions The corresponding rating mapping weight vector, This indicates the rating bias. The original rating values ​​are used to reflect the trainee's absolute performance level in a single skill dimension.

[0085] To eliminate rating scale differences between different interview skill dimensions, cross-dimensional alignment was performed on each raw rating value to generate standardized rating results for comparability across dimensions. Specifically, the set of raw rating values ​​was... Mapped to a standardized set of ratings The calculation method is as follows:

[0086] in, Indicating skill dimensions The corresponding standardized scoring results. Based on the standardized scoring results, a skill scoring vector is constructed:

[0087] This is used to describe the relative performance distribution of trainees across different interview skill dimensions. Finally, the skill score vectors are associated with the corresponding interview skill dimension identifiers to form structured skill dimension score results, which are then output according to a preset output format for subsequent teaching evaluation or training feedback.

[0088] In one specific implementation, the structured analysis method for language interaction in psychological interviews proposed in this invention is applied to a psychological interview training system, such as... Figure 3 The psychological interview training system is used to organize the language interaction process between trainees and virtual interviewees in a controlled training environment, and to perform structured processing and analysis of the language data during the interaction process.

[0089] In this embodiment, trainees participate in psychological interview training via voice input, and their voice signals are first collected and processed by a voice recognition module. The voice recognition module performs voice activity detection and voice recognition processing on the trainee's continuous voice input during the interview, generating text data corresponding to the trainee's speech content, and transmitting this text data to the psychological interview subject system as the basis for subsequent analysis. The psychological interview subject system includes a semantic intent recognition module and a role and scenario model module. The semantic intent recognition module performs semantic parsing on the text data to extract semantic intent information contained in the trainee's speech and converts this semantic intent information into a dialogue feature representation.

[0090] The role and scenario model module provides a unified background setting and state parameters for virtual interviewees. These state parameters include at least the virtual character's psychological state, emotional intensity, and interview stage markers. Based on this role and scenario model, virtual interviewees maintain consistent response characteristics during training, ensuring that different trainees face consistent interviewee conditions during testing. The virtual interviewees generate corresponding interview responses based on dialogue feature information from the psychological interviewee system and output corresponding voice or text feedback through the natural language response generation module.

[0091] During the interview, the psychological interview training system simultaneously tracks and analyzes the dialogue state between the trainee and the virtual interviewee. Specifically, the dialogue state tracking module constructs a dialogue state sequence based on semantic intent recognition results and continuously updates the transition relationships between adjacent states to form dialogue analysis results that reflect the evolution of the interview process. These dialogue analysis results, along with the dialogue logs from the interview process, are transmitted to the structured interactive analysis module.

[0092] The structured interaction analysis module identifies key interaction nodes during the interview process based on dialogue analysis results, and classifies the trainees' language interaction behaviors in different interview stages to generate a structured interaction feature set. This structured interaction feature set is further used by the skills dimension assessment module to quantitatively analyze the trainees' interview performance according to preset skills dimension mapping rules, outputting corresponding skills dimension scores, performance summaries, and improvement suggestions.

[0093] Through the above system architecture, this embodiment realizes the unified collection, analysis and structured expression of language interaction data during psychological interview training, enabling trainees' interview behavior to be evaluated comparably while maintaining consistency with the interviewees, and providing basic data support for subsequent training feedback and ability analysis.

[0094] like Figure 4As shown, this invention provides a structured analysis system for language interaction in psychological interviews. The system includes: a data acquisition unit, used to acquire language interaction data formed between trainees and preset virtual interviewees during psychological interview training, and to perform speech recognition processing on the language interaction data to generate corresponding text data; an intent recognition unit, used to perform semantic intent recognition processing based on the text data, extract semantic intent information of trainees' language interaction, and construct a dialogue state sequence during the interview process based on the semantic intent information; a feature extraction unit, used to perform dialogue state tracking and dialogue analysis processing based on the dialogue state sequence, and extract dialogue feature information representing the evolution of trainees' interactive behaviors during the interview; a structured processing unit, used to perform structured interaction analysis processing on the language interaction during the interview process based on the dialogue feature information, identify key interaction nodes and classify trainees' language interaction behaviors to generate a set of structured interaction features; and an output unit, used to generate skill dimension feature data representing trainees' interview skill performance according to preset skill dimension mapping rules based on the set of structured interaction features, and output the corresponding skill dimension score results.

[0095] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor A01, a network interface A02, memory (not shown), and a database (not shown) connected via a system bus. The processor A01 provides computing and control capabilities. The memory includes internal memory A03 and a non-volatile storage medium A04. The non-volatile storage medium A04 stores an operating system B01, a computer program B02, and a database (not shown). The internal memory A03 provides an environment for the operation of the operating system B01 and the computer program B02 stored in the non-volatile storage medium A04. The network interface A02 is used for communication with external terminals via a network connection. When the computer program B02 is executed by the processor A01, it implements a structured analysis method for psychological interview language interaction.

[0096] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a microcontroller, chip, or processor to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0097] The optional embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the embodiments of the present invention are not limited to the specific details described above. Within the scope of the technical concept of the embodiments of the present invention, various simple modifications can be made to the technical solutions of the embodiments of the present invention, and these simple modifications all fall within the protection scope of the embodiments of the present invention. It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the embodiments of the present invention will not further describe the various possible combinations.

[0098] Furthermore, various different embodiments of the present invention can be combined in any way, as long as they do not violate the spirit of the embodiments of the present invention, they should also be regarded as the content disclosed by the embodiments of the present invention.

Claims

1. A structured analysis method for language interaction in psychological interviews, characterized in that, The method includes: Acquire language interaction data formed between trainees and preset virtual interviewees during psychological interview training, and perform speech recognition processing on the language interaction data to generate corresponding text data; Based on the text data, semantic intent recognition processing is performed to extract semantic intent information of the trainees' language interactions, and a dialogue state sequence during the interview process is constructed based on the semantic intent information. Based on the dialogue state sequence, dialogue state tracking and dialogue analysis are performed to extract dialogue feature information that represents the evolution of trainees' interactive behaviors during the interview process. Based on the dialogue feature information, structured interaction analysis is performed on the language interaction during the interview process to identify key interaction nodes and classify the language interaction behavior of trainees to generate a set of structured interaction features. Based on the structured interaction feature set, skill dimension feature data representing the trainees' interview skill performance is generated according to the preset skill dimension mapping rules, and the corresponding skill dimension score results are output.

2. The structured analysis method for language interaction in psychological interviews according to claim 1, characterized in that, Acquire language interaction data between trainees and preset virtual interviewees during psychological interview training, and perform speech recognition processing on the language interaction data to generate corresponding text data, including: During the psychological interview training process, the trainees' voice input signals are collected in real time, and raw voice data containing timestamp information is formed based on the voice input signals. The original speech data is subjected to speech activity detection processing to identify valid speech segments and eliminate non-speech segments, thereby generating valid speech data for speech recognition processing; Based on the valid speech data, speech recognition processing is performed to generate text data corresponding to the student's speech input, and the text data is associated with the corresponding timestamp information.

3. The structured analysis method for language interaction in psychological interviews according to claim 1, characterized in that, Based on the text data, semantic intent recognition processing is performed to extract semantic intent information from the learner's language interactions, and a dialogue state sequence during the interview process is constructed based on the semantic intent information, including: The text data is subjected to semantic parsing to extract semantic intent features that characterize the type of student's speaking intent and interaction method; Based on the semantic intent features, the trainees' speech content is mapped to preset semantic intent tags to form a semantic intent tag sequence corresponding to each speech moment during the interview. Based on the semantic intent tag sequence and combined with the temporal order relationship corresponding to the text data, a dialogue state sequence is constructed to describe the evolution of language interaction during the interview. The dialogue state sequence is used as input data for subsequent dialogue state tracking and dialogue analysis.

4. The structured analysis method for language interaction in psychological interviews according to claim 3, characterized in that, The text data is subjected to semantic parsing to extract semantic intent features that characterize the type of student's speaking intent and interaction method, including: Based on the preset interview interaction unit division rules, the text data is segmented into speech units to obtain multiple interactive statement units divided according to semantic integrity. For each interactive statement unit, its pragmatic position features in the interview interaction process are extracted. The pragmatic position features are used to characterize the interaction stage of the interactive statement unit in the current interview round. Based on the pragmatic position features of the interactive statement units, and combined with the sentence structure features and response orientation features of the corresponding statement units, a semantic intent feature vector is constructed to describe the interactive mode of student speech.

5. The structured analysis method for language interaction in psychological interviews according to claim 1, characterized in that, Based on the dialogue state sequence, dialogue state tracking and dialogue analysis are performed to extract dialogue feature information that characterizes the evolution of trainees' interactive behaviors during the interview, including: Based on the dialogue state sequence, a state transition structure is constructed to represent the transition relationship between adjacent interactive states during the interview process, so as to describe the evolution path of the learner's language interaction behavior during the interview process. The state transition structure is continuously updated to generate state evolution features that characterize the changing trends of trainees' interactive behaviors in different interview stages. Based on the state evolution features, behavioral evolution features are extracted to characterize the stability of the trainees' interactions, the frequency of interaction adjustments, and the changes in interaction strategies during the interview process, as dialogue feature information.

6. The structured analysis method for language interaction in psychological interviews according to claim 5, characterized in that, Based on the aforementioned state evolution features, behavioral evolution features are extracted to characterize the stability of the trainees' interactions, the frequency of interaction adjustments, and the changes in interaction strategies during the interview process, as dialogue feature information, including: Based on the state evolution characteristics, the distribution of state retention duration between adjacent interview states is calculated to generate stability feature parameters for characterizing the stability of trainees' interactions. Based on the state evolution characteristics, the frequency of state transitions during the interview process and their distribution on the interview time axis are statistically analyzed to generate adjustment frequency characteristic parameters to characterize the frequency of trainees' interaction adjustments. Based on the state evolution characteristics, the interaction pattern change segments formed by multiple consecutive state transitions during the interview are identified, and strategy change feature parameters are generated based on the interaction pattern change segments to characterize the changes in the trainees' interaction strategies. The stability feature parameter, adjustment frequency feature parameter, and strategy change feature parameter are combined as behavioral evolution features.

7. The structured analysis method for language interaction in psychological interviews according to claim 1, characterized in that, Based on the aforementioned dialogue feature information, structured interaction analysis is performed on the language interactions during the interview process to identify key interaction nodes and classify the learners' language interaction behaviors, generating a set of structured interaction features, including: Based on the dialogue feature information, an interaction structure representation is constructed to describe the evolution of language interaction during the interview, wherein, The interaction structure includes at least state evolution relationships, speaking order relationships, and interaction response relationships; Based on the interaction structure representation, according to the preset node determination rules, the interaction positions that meet the conditions of state transition abrupt change, interaction strategy change or significant change in interaction stability during the interview process are determined, and the corresponding key interaction nodes are identified. For each of the key interaction nodes, combined with the dialogue feature information before and after the key interaction node, node behavior features are extracted to characterize the language interaction mode of the trainees in the corresponding interview stage. Based on the node behavior characteristics, the language interaction behaviors of trainees during the interview process are classified according to the preset behavior classification rules, and corresponding behavior category labels are generated. The key interaction nodes, the node behavior features, and the behavior category identifiers are associated to form a structured set of interaction features for describing the interview language interaction process.

8. The structured analysis method for language interaction in psychological interviews according to claim 7, characterized in that, Based on the structured interaction feature set, and according to preset skill dimension mapping rules, skill dimension feature data representing trainees' interview skill performance is generated, including: Based on the structured interaction feature set, extract the interaction feature subsets corresponding to different interview skill dimensions. The interview skill dimensions include at least the ability to guide questions, the coherence of responses, and the ability to control interactions. For each of the interview skills dimensions, based on the corresponding subset of interaction features and according to the preset dimension mapping relationship, the dimension feature parameters reflecting the interaction performance characteristics of trainees under the interview skills dimension are calculated. The feature parameters of each dimension are normalized within the dimension to generate skill dimension feature data that characterizes the relative performance level of trainees under different interview skill dimensions.

9. The structured analysis method for language interaction in psychological interviews according to claim 8, characterized in that, The output rules for the corresponding skill dimension scoring results are as follows: Based on the skill dimension feature data, the original score value of each interview skill dimension is calculated according to the preset score mapping relationship corresponding to each interview skill dimension. Cross-dimensional alignment is performed on each of the original rating values ​​to eliminate differences in rating scales between different interview skill dimensions and generate standardized rating results that are comparable across dimensions. Based on the standardized scoring results, a skill scoring vector is constructed to describe the relative performance distribution of trainees across different interview skill dimensions; The skill score vector is associated with the corresponding interview skill dimension identifier to form a structured skill dimension score result, which is then output according to a preset output format.

10. A structured analysis system for language interaction in psychological interviews, characterized in that, The system includes: The data acquisition unit is used to acquire language interaction data formed between trainees and preset virtual interviewees during psychological interview training, and to perform speech recognition processing on the language interaction data to generate corresponding text data. The intent recognition unit is used to perform semantic intent recognition processing based on the text data, extract semantic intent information of the trainee's language interaction, and construct a dialogue state sequence during the interview based on the semantic intent information. The feature extraction unit is used to perform dialogue state tracking and dialogue analysis processing based on the dialogue state sequence, and extract dialogue feature information that represents the evolution of trainees' interactive behavior during the interview process; The structured processing unit is used to perform structured interaction analysis processing on the language interaction during the interview based on the dialogue feature information, identify key interaction nodes and classify the language interaction behavior of trainees, and generate a set of structured interaction features. The output unit is used to generate skill dimension feature data representing the trainees' interview skill performance based on the structured interaction feature set and according to the preset skill dimension mapping rules, and output the corresponding skill dimension score results.

Citation Information

Patent Citations

  • Semi-structured interview system based on fine-tuning large language model

    CN118886503A

  • Multi-mode psychological problem detection system based on psychological interviews

    CN120148866A

  • Conversation state tracking model optimization system and method in electric power service scene

    CN120407725A

  • Multi-round dialogue intention recognition method and system based on adaptive semantic understanding

    CN120805921A

  • Multi-modal interview automatic quality analysis and evaluation method and system based on large model

    CN120849791A