Emotional dialogue intelligent interview method and system
By using multimodal perception and heterogeneous graph transformation technology, semantic-emotion coupling is monitored in real time, and adaptive guiding prompts are generated. This solves the problems of accuracy and adaptability in emotional state recognition in human-computer interaction, and improves the quality and accuracy of interviews.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU YUNZHICHU TECHNOLOGY CO LTD
- Filing Date
- 2026-04-29
- Publication Date
- 2026-07-21
AI Technical Summary
Existing technologies struggle to accurately identify respondents' emotional states in human-computer interaction, especially when faced with semantic masking and emotional fluctuations. This makes it difficult to implement adaptive interview strategy adjustments, leading to rigid interview logic and decreased feedback quality.
By synchronously collecting speech physical characteristics, text semantic content, and nonverbal behavioral features through a multimodal perception unit, a semantic-emotion coupling feature model is constructed to monitor semantic logic and emotional fluctuations in real time. A heterogeneous graph converter is used to track interview history information, generate adaptive guiding prompts, and inject emotion modulation factors to achieve adaptive collection of in-depth interview information.
It improves the accuracy of identifying the true psychological state of respondents, enables adaptive reshaping of interview logic, enhances interview quality and the authenticity of feedback, and has cross-scenario robustness and sub-second-level emotional perception accuracy.
Smart Images

Figure CN122432295A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence technology, and specifically relates to an intelligent interview method and system for emotional dialogue. Background Technology
[0002] Intelligent interaction technology, as an important branch of artificial intelligence, has been widely applied in diverse scenarios such as psychological assessment, assisted medical treatment, and talent interviews by simulating human communication logic and feedback mechanisms. Among these applications, accurate identification of emotional states is a core parameter for ensuring interview quality and natural interaction. Through in-depth analysis of respondents' real-time feedback, the system can obtain objective and valuable behavioral data, providing high-granularity data support for subsequent decision-making. Currently, there are various intelligent interview solutions, but all of them exhibit significant bottlenecks in complex human-computer interaction applications. (1) Text semantic analysis method: This type of approach mainly relies on natural language processing technology to analyze the sentiment polarity of respondents' responses by constructing semantic feature models. However, since the text is a high-level cognitive output processed by the respondents' subjective consciousness, it is easily affected by defensive psychology or social expectations, resulting in emotional masking and causing a deviation between the textual appearance and the inner psychology.
[0003] (2) Pre-set path retrieval method: Based on the identified semantic tendencies, the system retrieves and matches subsequent interview paths from a pre-set question bank. This method has a single perception dimension and lacks a dynamic monitoring mechanism for the coupling relationship between semantic logic and emotional fluctuations. As a result, the system can only extract modified surface semantic information. When faced with the critical state of the interviewee's insincerity, the interview logic tends to be rigid and cannot trigger adaptive strategy reshaping for the contradiction points.
[0004] Emotion recognition algorithms based on traditional natural language processing mainly face the following three types of challenges: 1. Static Semantic Feature Matching Method: The Subjective Concealment Challenge: Interviewees often exhibit a conflict between semantic feedback and affective valence during interviews. Due to the lack of cross-validation between semantic logic and emotional fluctuations in current technology, the system cannot identify anxiety or rejection characteristics displayed by interviewees while providing positive semantic content, making it difficult to detect the underlying avoidance or hesitation. Limitations of Perceptual Dimension: Relying solely on text-level feature mapping ignores the physical consistency of speech characteristics (such as pitch and speech rate variations) and nonverbal behavioral characteristics (such as micro-expression changes), resulting in interview data lacking authenticity and failing to address the underlying motivations of interviewees.
[0005] 2. Fixed Logic Jump Scheme: Insufficient Time Quantification Precision: In high-frequency dynamic interactions like interviews, traditional systems often rely on simple round counting to track contextual information. Lacking the ability to capture non-linear semantic-emotional evolution trends, the system cannot identify the micro-evolution of the user's emotional state at the sub-second level, making it difficult to generate targeted follow-up questions. Lack of Adaptability: Such methods tend to establish fixed mapping relationships. When increased defensiveness is detected in the interviewee, the guidance strategy cannot be dynamically adjusted. Mechanical, repetitive questioning easily leads to interviewee resistance, significantly reducing feedback quality.
[0006] 3. Machine Learning-Based Sentiment Fitting Methods (AI Models): Black Box Effect and Logical Inexplicability: While some nonlinear models can fit sentiment labels to existing data, they struggle to clearly explain the logical connection between semantic evolution and respondents' psychological motivations. Model-generated responses often lack empathy, failing to balance user psychological safety while deeply mining implicit information. Generalization Ability and Sample Dependence: AI models are highly dependent on the distribution of the training set. Due to the highly specific nature of interview events (such as the logical differences between talent interviews and psychological assessments), models trained for specific scenarios are difficult to directly transfer, and without a large number of labeled samples, they struggle to identify complex cross-round logical contradictions. Summary of the Invention
[0007] To address the aforementioned technical issues, this application provides an intelligent interview method for emotional dialogue and its interactive system.
[0008] The first aspect of this application provides an intelligent interview method for emotional dialogue, characterized by comprising: The multimodal perception unit simultaneously collects the physical characteristics of the interviewee's speech, the semantic content of the text, and the nonverbal behavior features, and extracts the corresponding acoustic physical parameters, semantic feature vectors, and micro-expression change features. A semantic-emotion coupling feature model is constructed. By extracting dialogue state features, the mapping relationship between semantic logic and emotional fluctuations is monitored in real time, and the conflict coefficient between semantic feedback and emotional valence is calculated. Based on the conflict coefficient, the mental state of the respondents is monitored. When a negative correlation is detected between semantic feedback and emotional valence, the conflict exploration sub-logic is triggered. An improved heterogeneous graph converter is used to track historical rounds of the interview process. Semantic nodes, emotional state nodes, and timestamp features in the interview are constructed into a heterogeneous graph. An attention mechanism is used to capture the non-linear semantic-emotional evolution trend and generate context vectors. Based on the context vector, adaptive guiding prompts with empathy-relief components are generated. By injecting emotion regulation factors into the guiding prompts, respondents are guided to gradually release their true psychological motivations, thus completing the adaptive collection of in-depth interview information.
[0009] Preferably, the multimodal perception unit includes an array of microphones and a high frame rate camera arranged within a range of 0.5m to 1.5m directly in front of the interviewee, for the purpose of synchronously capturing acoustic signals and facial key point data.
[0010] Preferably, an acoustic feature sequence, including pitch, energy, and speech rate, is extracted by performing a short-time Fourier transform on the physical properties of speech; at the same time, a facial key point detection algorithm is used to align the micro-expression transition features with the semantic feature vector within a time accuracy of 200ms.
[0011] Preferably, the cross-correlation calculation using the polarity score of the semantic feedback and the negativeness score of the emotional valence includes: ;in, This represents the semantic-sentiment conflict coefficient at time t; This represents the semantic affirmative polarity score calculated from the semantic feature vector at time t; This represents the emotional valence score obtained at time t by mapping acoustic physical parameters to micro-expression change features.
[0012] Preferably, the process of injecting sentiment modulators into the adaptive guiding prompts includes: ;in, This indicates the strength of the generated guiding prompts; α represents the preset empathy weight adjustment coefficient. This indicates the respondents' emotional anxiety index; This represents the original bootstrapping baseline generated based on the context vector.
[0013] Preferably, the determination step size for the negative correlation state is the time step size of one round of interview interaction.
[0014] According to a second aspect of this application, an intelligent interview system for emotional dialogue is provided, characterized in that it includes: The multimodal perception unit is used to simultaneously collect the physical characteristics of the interviewee's speech, the semantic content of the text, and the non-verbal behavior features, and extract the corresponding acoustic physical parameters, semantic feature vectors, and micro-expression change features. The intelligent analysis module is used to build a semantic-emotion coupling feature model. By extracting dialogue state features, it monitors the mapping relationship between semantic logic and emotional fluctuations in real time and calculates the conflict coefficient between semantic feedback and emotional valence. The logic scheduling unit is used to monitor the mental state of the respondents based on the conflict coefficient, and activates the conflict exploration sub-logic when it detects that semantic feedback and emotional valence are negatively correlated. A heterogeneous graph processing device is used to track historical rounds of the interview process using an improved heterogeneous graph converter, constructing a heterogeneous graph from semantic nodes, emotional state nodes and timestamp features in the interview, and generating context vectors through an attention mechanism. The strategy generation module is used to generate adaptive guiding prompts with empathy-relief components based on the context vector, thereby completing the adaptive collection of in-depth interview information.
[0015] Preferably, the camera in the multimodal perception unit supports a sampling frequency of 60fps or higher to ensure the continuity of micro-expression transition feature extraction.
[0016] According to a third aspect of this application, a computing device is provided, characterized in that it includes: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, wherein when the computer-executable instructions are executed by the processor, they implement the steps of any one of the emotional dialogue intelligent interview methods.
[0017] According to a fourth aspect of this application, a computer-readable storage medium is provided, characterized in that it stores computer-executable instructions, which, when executed by a processor, implement the steps of any one of the emotional dialogue intelligent interview methods.
[0018] The technical principles and beneficial effects of the above-mentioned solution of the present invention are as follows: (1) Solving the problem of identifying true intentions under semantic masking: Compared with traditional text analysis methods, the conflict coefficient extracted in this application is an "endogenous psychological indicator." Regardless of how the respondents subjectively process the textual representation (such as a 180° semantic reversal), their acoustic anomalies and micro-expression envelope features remain highly correlated with the underlying motivation. Real-time monitoring of indicators ensures the accuracy of judging insincere statements.
[0019] (2) Automatic compensation for differences in interview environment and individual expression: Even if different respondents have different speaking speed and tone benchmarks when responding, the evolution of their semantic-emotion coupling feature model in the logical space is physically homogeneous. By capturing the nonlinear evolution trend through the heterogeneous graph converter, the influence of emotional expression distortion caused by individual personality differences is eliminated, and the adaptive reshaping of interview logic is realized.
[0020] (3) Balancing accuracy and empathy injection: Feature alignment at the 200ms level is employed to improve the response resolution of emotion perception to sub-second levels. Under the same interaction conditions, through... The formula dynamically injects emotion regulation factors, controlling the proportion of comforting components in the prompts within a reasonable threshold, thus resolving the contradiction between in-depth information mining and the triggering of defensive psychology in respondents in principle.
[0021] (4) Robustness across scenarios under small sample size: This application uses an improved heterogeneous graph converter to model node associations, which has natural logical interpretability. Occasional semantic deviations that occur in individual rounds will be smoothed by the attention mechanism of the global heterogeneous graph, thereby ensuring the robustness of information collection in different small sample scenarios such as talent interviews and psychological assessments. Attached Figure Description
[0022] Figure 1 This is a flowchart of an intelligent emotional dialogue interview method provided in this application; Figure 2 This is a schematic diagram of the acoustic feature sequence of the physical characteristics of the respondent's speech provided in this application; Figure 3 This is a flowchart illustrating the micro-expression transition features of facial key points provided in this application. Detailed Implementation
[0023] It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. In the following description, suffixes such as "module," "part," or "unit" used to denote elements are used solely for illustrative purposes and have no inherent meaning. Therefore, "module," "part," or "unit" may be used interchangeably.
[0024] Figure 1 This is a flowchart of an intelligent emotional dialogue interview method provided in this application, such as... Figure 1 As shown, it includes: Step S101: Simultaneously collect the physical characteristics of the interviewee's speech, the semantic content of the text, and the non-verbal behavior features through the multimodal perception unit, and extract the corresponding acoustic physical parameters, semantic feature vectors, and micro-expression change features. Step S102: Construct a semantic-emotion coupling feature model, extract dialogue state features, monitor the mapping relationship between semantic logic and emotional fluctuations in real time, and calculate the conflict coefficient between semantic feedback and emotional valence. Step S103: Monitor the mental state of the interviewee according to the conflict coefficient. When a negative correlation is detected between semantic feedback and emotional valence, trigger the contradiction exploration sub-logic. Use the improved heterogeneous graph converter to track the historical round information of the interview process. Construct a heterogeneous graph by combining the semantic nodes, emotional state nodes and timestamp features in the interview. Capture the non-linear semantic-emotional evolution trend through the attention mechanism and generate a context vector. Step S104: Generate adaptive guiding prompts with empathy-relief components based on the context vector. By injecting emotion modulation factors into the guiding prompts, the interviewee is guided to gradually release their true psychological motivations, completing the adaptive collection of in-depth interview information. The multimodal perception unit includes an array of microphones and a high-frame-rate camera positioned within a range of 0.5m to 1.5m directly in front of the interviewee to achieve simultaneous capture of acoustic signals and facial key point data.
[0025] In some possible implementations, acoustic feature sequences, including pitch, energy, and speech rate, are extracted by performing a short-time Fourier transform on the physical properties of speech; at the same time, a facial key point detection algorithm is used to align micro-expression transition features with semantic feature vectors within a 200ms time precision.
[0026] In some possible implementations, cross-correlation calculations using the polarity score of the semantic feedback and the negativeness score of the emotional valence include: ;in, This represents the semantic-sentiment conflict coefficient at time t; This represents the semantic affirmative polarity score calculated from the semantic feature vector at time t; This represents the emotional valence score obtained at time t through mapping acoustic physical parameters to micro-expression change features. In some possible implementations, the injection process of emotional modulatory factors into the adaptive guiding prompt includes: ;in, This indicates the strength of the generated guiding prompts; α represents the preset empathy weight adjustment coefficient. This indicates the respondents' emotional anxiety index; This represents the original bootstrapping baseline generated based on the context vector.
[0027] In some possible implementations, the determination step size for the negatively correlated state is the time step size of one round of interview interaction. This application addresses complex psychological concealment behaviors of respondents by using a semantic-affective coupling feature model to extract the evolutionary trends of psychological states with physical consistency across different modalities. The specific steps are as follows: Step 1: Multimodal synchronous high-frequency signal acquisition; an array microphone and a 60fps high frame rate camera are deployed 1.0m in front of the interviewee. The system acquires speech sequences at a sampling rate of 16kHz. Simultaneously acquire facial image sequences at a frequency of 60fps .
[0028] Step 2: Feature Extraction and Temporal Alignment; The speech signal is processed by frame segmentation, and Mel-frequency cepstral coefficients (MFCCs) are extracted using a 25ms window; Simultaneously, 68 facial feature points are identified using a facial keypoint algorithm. The extracted acoustic features and micro-expression features are mapped to a unified time axis, ensuring alignment accuracy within 200ms.
[0029] Step 3: Semantic-sentiment conflict calculation; using a natural language processing model to perform semantic polarity analysis on the speech-to-text content, and obtain... .
[0030] 1. Regarding acoustic parameters (such as fundamental frequency) The system fuses features (such as tremors) and facial features (such as simulated electromyography signals from the brow area) to calculate the true emotional valence score. .
[0031] 2. Construct a conflict assessment function to eliminate the influence of subjective defense on the signal appearance: in, and Semantic terms and physical feature terms. The textual representation of the respondents after conscious processing; This represents uncontrolled physical feedback controlled by the vegetative nervous system. Conflict coefficient. Defined as follows: and The normalized deviation. It is a scalar that reflects the degree to which respondents are consistent in their words and actions. When When the threshold is exceeded (e.g., 0.65), the respondent is determined to be in a state of emotional concealment.
[0032] Step 4: Heterogeneous graph modeling and trend capture; transforming the interview process into a non-linear dynamic graph, with node types including semantic nodes. Emotional nodes and time nodes : 1. Establish edge relationships .
[0033] 2. Utilize an improved heterogeneous graph converter for multi-head attention aggregation to generate context vectors reflecting the depth of the interview. Technical principle: Heterogeneous graph architecture can record the offset trajectory of semantic nodes over time, and is highly sensitive to sudden emotional shifts (such as a sudden increase in speech rate when discussing specific sensitive topics).
[0034] Step 5: Adaptive guide word generation; utilizing the generated guide words... Drive the generative model and inject sentiment modulators. : —α: Empathy weighting coefficient; Meaning: A proportional factor measuring the intensity of system intervention. Physical significance: In this application, it determines whether the system intervenes through "gentle reassurance" or "direct questioning" after discovering a contradiction. Emotional Anxiety Index; Meaning: Anxiety level extracted based on acoustic energy and micro-expression transition frequency. Physical meaning: When... At higher levels, the formula increases the buffering component in the guiding words through a multiplicative effect, reducing the respondents' defensiveness.
[0035] Step Six: Adaptive Data Acquisition and Strategy Closure. Follow up with questions based on the generated prompts. A decrease indicates the guidance is effective; continue in-depth data collection. If the α coefficient is maintained at a high level, the α coefficient is adjusted to achieve adaptive reshaping of the interview logic.
[0036] Figure 2 This is a schematic diagram of an intelligent emotional dialogue interview system provided in this application, such as... Figure 2 As shown, it includes: The multimodal perception unit is used to simultaneously collect the physical characteristics of the interviewee's speech, the semantic content of the text, and the non-verbal behavior features, and extract the corresponding acoustic physical parameters, semantic feature vectors, and micro-expression change features. The intelligent analysis module is used to build a semantic-emotion coupling feature model, monitor the mapping relationship between semantic logic and emotional fluctuations in real time, and calculate the conflict coefficient. The logic scheduling unit is used to activate the contradiction-finding sub-logic when a negative correlation is detected between semantic feedback and emotional valence. A heterogeneous graph processing device is used to construct heterogeneous graphs from interview information and generate context vectors through an attention mechanism; The strategy generation module is used to generate adaptive guiding prompts with empathy-relief components based on the context vector, thereby completing the adaptive collection of in-depth interview information.
[0037] The camera in the multimodal perception unit supports a sampling frequency of over 60fps to ensure the continuity of micro-expression transition feature extraction.
[0038] This application also provides a computing device, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the steps of any one of the emotional dialogue intelligent interview methods.
[0039] This application also provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of any of the described emotional dialogue intelligent interview methods.
[0040] Example: This example uses a talent selection interview at a large company as the application scenario. When respondents answer sensitive questions such as "reasons for leaving the company," there is a high risk of semantic and emotional disconnect.
[0041] Signal preprocessing: Sampling and adjustment: The microphone array performs real-time gain control on the speech signal, filtering out 50Hz power frequency noise. The camera locates the interviewee's face and extracts the motion trajectory of key facial points.
[0042] This example compares the depth of information collection (measured by the key motivation identification rate) using the "traditional semantic analysis method", the "fixed question bank jump method", and the "intelligent interview method of this application", as shown in Table 1: Table 1. Comparison of the effects of different interview methods In-depth analysis of results 1. Solving the problem of identifying true intentions under semantic masking: When respondents verbally express "very satisfied with the previous company culture" (semantic polarity) =0.92) but accompanied by downturned corners of the mouth and limited speech rate (affective valence) When =0.21), if Figure 3 As shown, traditional algorithms fail to recognize text due to their over-reliance on it. However, the conflict coefficient calculated in this application... =0.77, successfully triggering the contradiction exploration logic and revealing its true defensive psychology.
[0043] 2. Eliminating the Quantification Trap of Individual Expression Differences: Different respondents have different personality baselines. This application captures "dynamic evolution trends" rather than "absolute values" through a heterogeneous graph converter. Compared with traditional AI models, this application uses an attention mechanism to find logical connections between nodes in a global heterogeneous graph. Even if a respondent's tone is naturally low, the system can accurately capture their psychological fluctuations through their relative changes on sensitive topics.
[0044] 3. Balancing Accuracy and Empathy Injection: This application employs 200ms feature alignment to improve the response resolution of emotion perception to sub-second levels. (The formula is used to...) Dynamically generated prompts (such as "I noticed you hesitated a bit when you talked about this part. Could you please share your specific feelings at that time?") can effectively alleviate the pressure on respondents compared to mechanical follow-up questions, while ensuring a response accuracy of 0.01 seconds and having stronger interactive resilience.
[0045] In summary, this application has the following technical advantages: 1. Solving the problem of identifying true intent under semantic masking: Compared with traditional text methods, the conflict coefficient extracted in this application is an "endogenous psychological indicator," which is obtained through... Real-time monitoring of indicators ensures the accuracy of judging insincere statements.
[0046] 2. Automatic compensation for individual expression differences: By capturing non-linear evolution trends through a heterogeneous graph converter, the effects of expression distortion caused by individual personality and background differences are eliminated.
[0047] 3. Balance between speed measurement accuracy and feature preservation: Sub-second feature alignment was achieved, preserving transient emotional fluctuations during the interview process and avoiding information smoothing caused by long window analysis.
[0048] 4. Robustness across scenarios with small sample sizes: This application uses a heterogeneous graph converter to model node associations, which has natural logical interpretability and can achieve robust information collection without the need for pre-training with massive amounts of specific scenario data.
[0049] The preferred embodiments of this application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and spirit of this application shall be within the scope of the claims.
[0050] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. An intelligent interview method for emotional dialogue, characterized in that, include: The multimodal perception unit simultaneously collects the physical characteristics of the interviewee's speech, the semantic content of the text, and the nonverbal behavior features, and extracts the corresponding acoustic physical parameters, semantic feature vectors, and micro-expression change features. A semantic-emotion coupling feature model is constructed. By extracting dialogue state features, the mapping relationship between semantic logic and emotional fluctuations is monitored in real time, and the conflict coefficient between semantic feedback and emotional valence is calculated. Based on the conflict coefficient, the mental state of the respondents is monitored. When a negative correlation is detected between semantic feedback and emotional valence, the conflict exploration sub-logic is triggered. An improved heterogeneous graph converter is used to track historical rounds of the interview process. Semantic nodes, emotional state nodes, and timestamp features in the interview are constructed into a heterogeneous graph. An attention mechanism is used to capture the non-linear semantic-emotional evolution trend and generate context vectors. Based on the context vector, adaptive guiding prompts with empathy-relief components are generated. By injecting emotion regulation factors into the guiding prompts, respondents are guided to gradually release their true psychological motivations, thus completing the adaptive collection of in-depth interview information.
2. The intelligent interview method for emotional dialogue according to claim 1, characterized in that, The multimodal perception unit includes an array of microphones and a high frame rate camera arranged in a range of 0.5m to 1.5m directly in front of the interviewee, for the synchronous capture of acoustic signals and facial key point data.
3. The intelligent interview method for emotional dialogue according to claim 1, characterized in that, By performing a short-time Fourier transform on the physical properties of speech, an acoustic feature sequence including pitch, energy, and speech rate is extracted; at the same time, a facial key point detection algorithm is used to align micro-expression transition features with semantic feature vectors within a time accuracy of 200ms.
4. The intelligent interview method for emotional dialogue according to claim 3, characterized in that, The calculation of the cross-correlation between the polarity score and the negativeness score of the emotional valence includes: ;in, This represents the semantic-sentiment conflict coefficient at time t; This represents the semantic affirmative polarity score calculated from the semantic feature vector at time t; This represents the emotional valence score obtained at time t by mapping acoustic physical parameters to micro-expression change features.
5. The intelligent interview method for emotional dialogue according to claim 4, characterized in that, The process of injecting sentiment modulators into the adaptive guiding prompts includes: ;in, This indicates the strength of the generated guiding prompts; α represents the preset empathy weight adjustment coefficient. This indicates the respondents' emotional anxiety index; This represents the original bootstrapping baseline generated based on the context vector.
6. The intelligent interview method for emotional dialogue according to claim 1, characterized in that, The determination step size for the negative correlation state is the time step size of one round of interview interaction.
7. An intelligent interview system for emotional dialogue, characterized in that, include: The multimodal perception unit is used to simultaneously collect the physical characteristics of the interviewee's speech, the semantic content of the text, and the non-verbal behavior features, and extract the corresponding acoustic physical parameters, semantic feature vectors, and micro-expression change features. The intelligent analysis module is used to build a semantic-emotion coupling feature model. By extracting dialogue state features, it monitors the mapping relationship between semantic logic and emotional fluctuations in real time and calculates the conflict coefficient between semantic feedback and emotional valence. The logic scheduling unit is used to monitor the mental state of the respondents based on the conflict coefficient, and activates the conflict exploration sub-logic when it detects that semantic feedback and emotional valence are negatively correlated. A heterogeneous graph processing device is used to track historical rounds of the interview process using an improved heterogeneous graph converter, constructing a heterogeneous graph from semantic nodes, emotional state nodes and timestamp features in the interview, and generating context vectors through an attention mechanism. The strategy generation module is used to generate adaptive guiding prompts with empathy-relief components based on the context vector, thereby completing the adaptive collection of in-depth interview information.
8. The intelligent interview system for emotional dialogue according to claim 7, characterized in that, The camera in the multimodal perception unit supports a sampling frequency of over 60fps to ensure the continuity of micro-expression transition feature extraction.
9. A computing device, characterized in that, include: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the emotional dialogue intelligent interview method according to any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that, It stores computer-executable instructions that, when executed by a processor, implement the steps of the emotional dialogue intelligent interview method according to any one of claims 1 to 6.